Hybrid Generative Models for 2D and 3D Computer Vision
Deep Learning has made tremendous progress in the last decade, making breakthroughs in visual perception and imagination. Discriminative models have achieved human-level performance on several tasks. Generative models have also shown great promise for 2D and 3D synthesis; however, their performance still lags behind discriminative models. Several approaches have been proposed to enhance the performance of generative models. One promising direction is leveraging multiple representations to guide the synthesis process. In this dissertation, we explore how generative models can benefit from hybrid representations, in which a stronger representation guides a weaker one. We propose models for 3D synthesis conditioned on images or point clouds, label-conditioned 2D generation, and 2.5D motion generation. A major drawback of current generative models is that they require gigantic datasets for training. We present a few-shot synthesis method for decreasing the model’s dependence on labeled data. Finally, we discuss adversarial applications of generative models in which adversaries abuse visual realism to deceive humans and machines.