Generalizable 3D Understanding and Generation

Chao Xu
Ph.D., 2025
WU, YINGNIAN
How should we model the physical world for embodied agents? To carry out mundane tasks, humans exhibit remarkable capabilities in world modeling, mentally constructing 3D scenes from limited viewpoints, leveraging strong perceptual and generative priors to infer occluded regions and understand environmental affordances for interaction. Crucially, these abilities generalize across diverse scenes and objects. While modern foundation models display a similar breadth in language and in 2D domains, transferring such generalizability to 3D representations remains a key challenge for creating truly capable embodied agents.

This dissertation explores methods that enhance the generalization capacity of 3D vision systems for both understanding and generation. Key contributions include: devising efficient feed-forward pipelines that leverage 2D pre-trained priors to generate high-quality 3D shapes from challenging inputs such as a single image or sparse, unposed views; developing techniques to enhance the multi-view consistency and fidelity of generated 3D assets; and introducing methods for functional, part-level 3D understanding that discover cross-category affordances.

Collectively, these efforts significantly advance the state-of-the-art in robust, generalizable 3D generation from limited views and in achieving more fine-grained, functional object understanding. This dissertation systematically rethinks and refines the pipelines to unite the strengths of 2D generative priors with 3D representation learning. The research presented provides foundational steps towards more scalable, generalizable, and functionally aware 3D vision systems. The dissertation concludes by summarizing these contributions and outlining promising directions for future research in pursuit of physically grounded artificial intelligence.

2025