Generalizable 3D Understanding and Generation
This dissertation explores methods that enhance the generalization capacity of 3D vision systems for both understanding and generation. Key contributions include: devising efficient feed-forward pipelines that leverage 2D pre-trained priors to generate high-quality 3D shapes from challenging inputs such as a single image or sparse, unposed views; developing techniques to enhance the multi-view consistency and fidelity of generated 3D assets; and introducing methods for functional, part-level 3D understanding that discover cross-category affordances.
Collectively, these efforts significantly advance the state-of-the-art in robust, generalizable 3D generation from limited views and in achieving more fine-grained, functional object understanding. This dissertation systematically rethinks and refines the pipelines to unite the strengths of 2D generative priors with 3D representation learning. The research presented provides foundational steps towards more scalable, generalizable, and functionally aware 3D vision systems. The dissertation concludes by summarizing these contributions and outlining promising directions for future research in pursuit of physically grounded artificial intelligence.

