Learning to Create 3D Content
Modeling highly realistic virtual worlds in 3D is the holy grail for visual content creation, and obtaining high quality 3D models is an indispensable part of it. 3D models are used extensively in the film and gaming industry, and we are foreseeing growing demand from novel applications such as social media and mixed reality. Unfortunately, the creation of 3D models is a complicated process that is tedious for the artists and inaccessible for ordinary users. By streamlining this 3D content creation process using machine learning methods, we will not only allow artists to create better quality artworks with less effort, but also enable ordinary users to create 3D effortlessly. Recently, with the advances in generative modeling techniques and the increasing availability of 2D and 3D datasets, we are starting to see significant improvements on the quality of synthesized 3D shapes. However, a useful 3D generative model requires more than good quality. First, we need controllability. The model should offer an intuitive interface for the users to express their ideas, customize the results, and achieve their goals. Moreover, we need scalability. The model should be ready to handle larger datasets that will be available in the future, and produce higher output resolution that will meet future demands. This dissertation centers around improving the quality, control and scalability of 3D generative models. First, we show that by coupling 3D shapes with their abstract, primitive-based representation in a data-driven manner, we can enable a user with no 3D modeling experience to interactively manipulate a 3D shape while preserving the identity of the object. We then show that we can go beyond shape manipulation. By incorporating unstructured 2D images into the 3D generation process, we are able to directly synthesize photorealistic 3D scenes from coarse voxels created intuitively by the user.Lastly, we propose a general framework for scaling up implicit neural representations, an important component for modeling continuous signals including 3D shapes. We conclude the dissertation with future avenues of research, where I will also show how the ideas and technique I developed during my PhD can help shape the future of 3D synthesis.