Towards Comprehensive Visual Understanding
An image is worth a thousand words, conveying information that goes beyond the visual content therein. Traditional computer vision tasks focus on the recognition of tangible properties of images, such as objects and scenes. Relatively little attention has been paid to tasks that involve private states where subjectivity analysis is relevant. This area includes detecting cyberbullying and hate speech, identifying emotions, and understanding rhetoric and intentions. This dissertation presents our work in exploring new challenges and approaches towards comprehensive visual understanding, with both subjectivity and objectivity in images in mind. Specifically, on the challenge side, we focus on a specific aspect of subjectivity: the intent behind social media images. We introduce an intent dataset, Intentonomy, annotated with 28 intent categories derived from a social psychology taxonomy. We then systematically study whether, and to what extent, commonly used visual information, i.e., object and context, contribute to human intent understanding. On the approach side, we present three approaches: (1) an intent classifier that attends to object and context classes in images as well as textual information in the form of hashtags; (2) a streamlined pre-training method that uses pseudo labels derived from human responses to social media posts. (3) a parameter-efficient transfer learning method for adapting ever-increasing pre-trained vision models. We find our dataset to be very challenging for visual recognition systems and our approaches to be empirically effective on representative visual understanding tasks.