Cornell University
Library
Cornell UniversityLibrary

eCommons

Help
Log In(current)
  1. Home
  2. Cornell University Graduate School
  3. Cornell Theses and Dissertations
  4. Learning to Represent and Recognize Multimodal Videos

Learning to Represent and Recognize Multimodal Videos

File(s)
Qian_cornellgrad_0058F_13689.pdf (3.04 MB)
Permanent Link(s)
https://doi.org/10.7298/7bje-bs37
https://hdl.handle.net/1813/114738
Collections
Cornell Theses and Dissertations
Author
Qian, Rui
Abstract

In today's digital landscape, the staggering growth of video resources has resulted in a wealth of visual, auditory, and textual information readily available on the internet. To fully harness the potential of learning from multimodal videos, it is crucial to develop efficient techniques for processing and analyzing this information. This dissertation delves into two primary areas: label-efficient representation learning from videos and multimodal video recognition. We leverage the advances in large-scale deep learning to fully exploit the potential of internet videos. Label-efficient representation learning from videos is motivated by the fact that internet videos often lack high-quality labels. We first propose a contrastive learning-based framework to learn features in an unsupervised manner. This approach pulls together feature representations from the same video and pushes apart those from different videos. We then investigate how to learn more refined temporal features for time-sensitive tasks, such as temporal event classification and detection, and more precise spatial features for location-aware tasks, including tracking and detection. Lastly, we explore the use of stronger transformer backbones to integrate multimodal data and form a unified, robust representation. Regarding multimodal recognition for videos, we initially examine the challenging case of fine-grained video recognition to determine if different modalities can assist each other in ambiguous scenarios. We create a new expert-curated audiovisual benchmark and discover that multimodal fusion significantly benefits performance. We subsequently investigate open-vocabulary multimodal video recognition and propose a framework that can effectively classify videos from any given category.

Description
153 pages
Date Issued
2023-08
Keywords
Audiovisual
•
Multimodal Learning
•
Open-vocabulary recognition
•
Representation Learning
•
Self-supervised Learning
•
Video Understanding
Committee Chair
Belongie, Serge
Committee Member
Hariharan, Bharath
Lee, Clarence
Degree Discipline
Computer Science
Degree Name
Ph. D., Computer Science
Degree Level
Doctor of Philosophy
Rights
Attribution 4.0 International
Rights URI
https://creativecommons.org/licenses/by/4.0/
Type
dissertation or thesis
Link(s) to Catalog Record
https://newcatalog.library.cornell.edu/catalog/16219401

Site Statistics | Help

About eCommons | Policies | Terms of use | Contact Us

copyright © 2002-2026 Cornell University Library | Privacy | Web Accessibility Assistance