Cornell University
Library
Cornell UniversityLibrary

eCommons

Help
Log In(current)
  1. Home
  2. Cornell University Graduate School
  3. Cornell Theses and Dissertations
  4. AI-Driven Scientific Discovery

AI-Driven Scientific Discovery

File(s)
Du_cornellgrad_0058F_15455.pdf (34.58 MB)
Permanent Link(s)
https://doi.org/10.7298/dtfm-7e56
https://hdl.handle.net/1813/126513
Collections
Cornell Theses and Dissertations
Author
Du, Yuanqi
Abstract

Scientific discovery has long been constrained by two rate-limiting factors: the exponential expansion of hypothesis space and the prohibitive cost to validate hypothesis through high-fidelity simulation or experiment. Recent advances in Artificial Intelligence present an unprecedented opportunity to accelerate the whole process. This thesis presents a unified computational perspective on accelerating scientific discovery through probabilistic machine learning with chemistry as a driving application domain. The first part reframes hypothesis search as a sampling problem. Generative models learn from massive observed hypotheses to provide a strong prior over the vast hypothesis space. Sampling algorithms can be naturally applied at inference time to target distributions of interest, from annealed, reward-tilted, posterior, product distributions and any combination of them for flexible use. We will present both the recipe to build generative models that respect physical symmetry and a new unifying framework to scale them at inference time. The second part bridges the de facto tool to understand chemical systems at molecular level, molecular dynamics simulation, and recent advances in probabilistic inference. Despite the development of both theory and computation in physical chemistry through decades, non-equilibrium behaviors are notoriously challenging to understand, due to their statistical rare and dynamical fleeting nature, even from simulation. We will present how we can tackle this grand challenge that often arises in simulating non-equilibrium behaviors---high dissipation or equivalently high variance, by learning non-equilibrium processes that actively minimize dissipation. The third part bridges the first two parts into an emerging paradigm of automated scientific discovery workflow. Modern large language models are unprecedented agents to assist in controlling the discovery process. Nevertheless, discovery in natural sciences does not come with fixed discovery goal (thus not a well-defined search or optimization problem) but necessitates empirical evidence to navigate the goal space along with accessible computational tools or experiments to validate the hypothesis. We will present our work, autonomous goal-evolving agent that pushes the first step along this new horizon, an agent initialized with a discovery goal and iteratively steers the discovery direction, translates the goal to executable programs (i.e. scoring functions), and search through the hypothesis space with them. All together, the pace of scientific discovery is being largely accelerated by the new era of probabilistic machine learning approaches, from search, simulation to automation.

Description
230 pages
Date Issued
2026-05
Committee Chair
Gomes, Carla
Committee Member
Bindel, David
Selman, Bart
Degree Discipline
Computer Science
Degree Name
Ph. D., Computer Science
Degree Level
Doctor of Philosophy
Rights
Attribution 4.0 International
Rights URI
https://creativecommons.org/licenses/by/4.0/
Type
dissertation or thesis

Site Statistics | Help

About eCommons | Policies | Terms of use | Contact Us

copyright © 2002-2026 Cornell University Library | Privacy | Web Accessibility Assistance