Ang Cao

I am a Research Scientist at Google DeepMind. My research focuses on pushing the frontier of foundation model intelligence across digital worlds and, ultimately, the physical world.

Currently, I am particularly interested in discovering scalable, self-sustaining sources of learning signals that enable intelligent systems to improve beyond the limits of curated human supervision, including reinforcement learning, AI for AI (AI4AI), recursive self- and joint improvement, and learning from past experience.

Previously, I completed my Ph.D. at the University of Michigan, advised by Justin Johnson, and worked closely with Andrew Owens and JJ Park. I was also fortunate to intern at Meta FAIR and Meta GenAI.

profile photo

News

  • Grateful to be recognized as a CVPR 2026 Outstanding Reviewer.

  • Joined Google DeepMind as a Research Scientist.

  • Completed my Ph.D. at the University of Michigan.

  • LiftGS, LOCATE 3D, and ViLP were accepted to ICML 2025.

Research interests

My research aims to advance foundation-model intelligence in digital worlds and, ultimately, the physical world. I approach this goal from three complementary perspectives: Environments, Intelligent systems, and Learning and continual improvement.

Environments. I am interested in creating open-ended, interactive environments that agents can explore and interact with. My work involves developing highly scalable and efficient representations of the world, particularly in 3D and 4D, together with methods for reconstructing and generating such environments.

Intelligent systems. I study how to push the capability frontier of foundation models, particularly their ability to perceive, reason about, and act within their environments. I am especially interested in extending these capabilities from digital spaces to the physical world.

Learning and continual improvement. I am currently focused on how intelligent systems can continually improve beyond curated human supervision. I explore scalable, self-sustaining sources of learning signals, including interactions with environments, feedback from the system itself and from other models, and accumulated experience. My work spans reinforcement learning, AI for AI (AI4AI), and automated research (autoresearch), with an emphasis on recursive self-improvement and the joint improvement of agents and their environments.

ViLP More work coming soon
Work Experience

Research Scientist, Google DeepMind

Dec 2025 – Present


Research Scientist Intern, Meta FAIR, MPK, USA

Worked with Sasha Sax and Franziska Meier

May 2024 – Dec 2024


Research Scientist Intern, Meta GenAI, London, UK

Worked with David Novotny, Andrea Vedaldi, and Natalia Neverova

May 2023 – Nov 2023


Publications

* indicates equal contribution; indicates a mentored project.

Gen-Points tracked subject
Generative Point Tracking and Forecasting
Xuanchen Lu, Ang Cao, Chao Feng, Andrew Owens
CVPR 2026
paper project
point tracking motion forecasting generative modeling

We unify point tracking and future motion forecasting with an omni generative model.

LiftGS 3D grounding result
From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMs
ICML 2025
paper project
3D VLG 2D supervision VLM distillation

We train a 3D visual-language grounding model using only 2D supervision distilled through differentiable rendering.

LOCATE 3D object localization result
LOCATE 3D: Real-World Object Localization via Self-Supervised Learning in 3D
Sergio Arnaud, Paul McVay, Ada Martin, Arjun Majumdar, ... Ang Cao, ... Franziska Meier
ICML 2025
paper project
3D localization self-supervised learning language grounding

We localize language-referred objects directly in real-world RGB-D observation streams using self-supervised 3D representations.

ViLP visual grounding examples
Probing Visual Language Priors in VLMs
Tiange Luo*, Ang Cao*, Gunhee Lee, Justin Johnson, Honglak Lee
ICML 2025
paper
VLMs visual priors diffusion data

We identify visual ignorance in VLMs and improve their visual grounding through self-improving DPO.

Fast3R reconstruction result
Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass
CVPR 2025
paper project code demo
3D reconstruction large-scale views pose estimation

A single transformer reconstructs scenes and estimates cameras from more than 1,000 images in one forward pass.

Lightplane rendered objects
Lightplane: Highly-Scalable Components for Neural 3D Fields
3DV 2025
paper project docs code
neural fields memory efficiency rendering

Lightplane provides highly memory-efficient splatting and rendering components for scalable neural 3D fields.

Meta 3D Gen assets
Meta 3D Gen
Meta GenAI 3DGen Team
Tech report · 2024
paper project
text-to-3D asset generation

A fast, high-quality pipeline for text-to-3D asset generation.

DreamGaussian4D animated assets
DreamGaussian4D: Generative 4D Gaussian Splatting
arXiv 2023
paper project code
4D generation Gaussian splatting video diffusion

We do 4D generation with Gaussian Splatting by distilling motions from video diffusion models.

Text2Room generated room
Text2Room: Extracting Textured 3D Meshes from 2D Text-to-Image Models
ICCV 2023 · Oral
paper project code video
text-to-3D mesh generation scene synthesis

We generate meshes of full 3D rooms using text-to-image models.

HexPlane dynamic scene
HexPlane: A Fast Representation for Dynamic Scenes
Ang Cao, Justin Johnson
CVPR 2023
paper project code
dynamic scenes 4D representation NeRF

An elegant representation for dynamic 3D scenes using six feature planes.

FWD method overview
FWD: Real-time Novel View Synthesis with Forward Warping and Depth
CVPR 2022
paper project code video patent
novel view synthesis depth forward warping

We show point rasterization can be really fast for sparse view novel view synthesis.

Object detector inversion method overview
Inverting and Understanding Object Detectors
Ang Cao, Justin Johnson
Tech report · 2021
paper code
object detection model inversion

Revealing intriguing properties of detectors by applying our layout inversion technique.

Unified signal compression model overview
Unified Signal Compression Using Generative Adversarial Networks
Bowen Liu*, Ang Cao*, Hun-Seok Kim
ICASSP 2020
paper
signal compression generative models ADMM

We compress image and speech signals into quantized latent vectors using generative models and ADMM optimization.

Teaching Experience
Computer graphics course project scene
EECS 498/598: Computer Graphics and Generative Models (Fall 2024)

Teaching assistant (GSI/Head TA), working with JJ Park