I study how visual generative models should represent information and allocate computation.
My recent work explores adaptive representations for diffusion models across space, scale, and time (Foveated Diffusion, Spectral Progressive Diffusion),
with broader interests in video, world models, and the interface between reasoning and visual generation.
I've also worked on 4D world modeling at Waymo Research.
Previously, I've worked on efficient representations and rendering for 3D scenes (Textured Gaussians, GWS)
and computational imaging systems, including work at Meta Reality Labs.
Invited talks:
NVIDIA Research, Rhoda AI, Black Forest Labs, Cohere, and Decart.
A biologically-inspired diffusion framework that employs spatially adaptive tokenization
to concentrate compute on selected regions, achieving up to 4× speedups in image and video synthesis.
A fundamentally new alpha blending algorithm for Gaussian splats generates random-phase, full-bandwidth light field holograms for VR displays,
enabling natural defocus blur, physically-accurate parallax, and occlusion.
When state-of-the-art neural rendering meets next-generation holographic VR displays: converting
optimized Gaussian splats to holograms that support natural focus cues.
The novel combination of a multisource laser array and a dynamic Fourier amplitude modulator significantly improves
holographic VR display étendue and light field hologram image quality.
A near-eye display design that pairs inverse-designed metasurface waveguides with AI-driven holographic displays
to enable full-colour 3D augmented reality from a compact glasses-like form factor.
The inclusion of parallax cues in CGH rendering plays a crucial role in enhancing perceptual realism,
and we show this through a live demonstration of 4D light field holograms.
A novel light-efficiency loss function, AI-driven
CGH techniques, and camera-in-the-loop calibration greatly improves holographic projector
brightness and image quality.
We propose an image-to-image translation algorithm based on generative adversarial networks
that rectifies fisheye images without the need of paired training data.