Generative AI for Digital Pathology: Advancing Diffusion, Vision-Language, and Safety-Aware Models
This dissertation investigates the transformative intersection of Artificial Intelligence (AI) and Digital Pathology, presenting both a comprehensive review and a series of methodological contributions that highlight the synergistic potential of Generative AI (GenAI) in this domain. It spans a broad range of tasks—from synthesizing histopathology images and modeling pathologists’ decision-making processes to generating synthetic diagnostic text for enhancing vision-language model (VLM) representation and alignment. The work further explores the integration of safety and context awareness into generative frameworks to significantly improve model performance and trustworthiness. Chapter 1 provides an in-depth review of recent advancements in AI for histopathology, encompassing both foundational techniques and state-of-the-art developments. It also surveys publicly available datasets that have been instrumental in driving progress in the field. Chapter 2 addresses the challenge of uncertainty in mitotic figure detection—a clinically significant yet inherently subjective task. To model this uncertainty, a probabilistic diffusion model is introduced to synthesize cell nuclei undergoing mitosis. This generative framework identifies key visual features that inform diagnostic decisions and offers a novel interpretability tool to support pathologists. Chapter 3 explores the application of VLMs in digital pathology, focusing on limitations related to the scarcity of large-scale image-caption datasets and the sensitivity of zero-shot classification to prompt design. To address these challenges, language rewrites generated by a large language model (LLM) are used to enrich an existing dataset, resulting in substantial performance gains. Additionally, a novel context modulation layer is proposed to enhance image-text alignment. Building on this foundation, Chapter 4 introduces an extended training paradigm for digital pathology VLMs that leverages synthetic captions. Alongside the standard CLIP loss, a new text-only loss is formulated using stochastic cosine similarity between original captions and various LLM-derived concepts (cell, organ, disease). The training also incorporates “safety-aware” incorrect captions that are LLM-generated texts that simulate clinically relevant misinterpretations to improve model robustness. The similarity between original and incorrect captions is explicitly penalized, guiding the model to better distinguish between meaningful and misleading descriptions. This strategy leads to significant improvements in both zero-shot classification and image-text retrieval tasks. Collectively, this dissertation advances the role of GenAI in digital pathology by tackling key challenges related to data scarcity, diagnostic uncertainty, and model reliability. The proposed methods contribute to the development of more robust, interpretable, and clinically trustworthy AI-assisted diagnostic systems.