Neural Representation for Camera Models and Dynamic Scenes
The advent of digital and smartphone cameras has revolutionized photography by making it accessible to almost everyone. Computational photography offers exciting opportunities to democratize creative photography, allowing beginners to create amazing visual content effortlessly. This is achieved by leveraging computer vision algorithms and machine learning to enable users to create high-quality visual content effortlessly. In my thesis dissertation, I will discuss three novel approaches we proposed for representing camera models and 3D scenes using deep neural networks. The first part of this dissertation focuses on the development of a neural lens model, which simplifies the process of pre-capture camera calibration and refines camera parameters during 3D reconstruction. The proposed method accurately models the effects of the camera lens distortion using a neural network, which can be easily integrated into downstream neural rendering systems. My second work focuses on camera orientation estimation from single images by incorporating 3D geometry. Our proposed system can be trained end-to-end with camera poses and intermediate representations of surface geometry, outperforming a black-box regression from image to camera parameters. Third, I will present a compelling application of my research on creating 3D videography from input videos casually captured by cellphones. By using such algorithms, we can create immersive experiences that were previously only possible with specialized equipment and expertise. Lastly, I will discuss my current and upcoming research, which seek to expand the possibilities of artistic expression and content creation. I will highlight some exciting directions in the future, through enhancing scene rendering quality, implementing novel features like AR/VR and visual effects, and ultimately providing users with a more effortless and creative photography experience.