Cornell University
Library
Cornell UniversityLibrary

eCommons

Help
Log In(current)
  1. Home
  2. Cornell University Graduate School
  3. Cornell Theses and Dissertations
  4. Between Randomness and Arbitrariness: Some Lessons for Reliable Machine Learning at Scale

Between Randomness and Arbitrariness: Some Lessons for Reliable Machine Learning at Scale

File(s)
Cooper_cornellgrad_0058F_14358.pdf (37.46 MB)
Permanent Link(s)
https://doi.org/10.7298/fpbk-nk85
https://hdl.handle.net/1813/116423
Collections
Cornell Theses and Dissertations
Author
Cooper, A.
Abstract

To develop rigorous knowledge about ML models -- and the systems in which they are embedded -- we need reliable measurements. But reliable measurement is fundamentally challenging, and touches on issues of reproducibility, scalability, uncertainty quantification, epistemology, and more. This dissertation addresses criteria needed to take reliability seriously: both criteria for designing meaningful metrics, and for methodologies that ensure that we can dependably and efficiently measure these metrics at scale and in practice. In doing so, this dissertation articulates a research vision for a new field of scholarship at the intersection of machine learning, law, and policy. Within this frame, we cover topics that fit under three different themes. First, we quantify and mitigate sources of arbitrariness in machine learning, with respect to hyperparameter optimization and social prediction contexts. We clarify important connections between machine-learning arbitrariness, rooted in non-determinism, with legal notions of arbitrariness that implicate legal rules and due process. Second, we tame randomness in uncertainty estimation and optimization algorithms, in order to achieve scalability without sacrificing reliability. We discuss how across computing, and particularly in machine learning, scalability and reliability are typically in trade-off. Analogous trade-offs in law and policy make this type of trade-off a useful abstraction for communicating about machine-learning capabilities and risks to policymakers and other non-expert stakeholders. Third, we provide methods for evaluating generative-AI systems, with specific focuses on quantifying memorization in language models and training latent diffusion models on open-licensed data. These contributions have urgent and significant connections to U.S. copyright law. We provide an abridged discussion of landmark legal scholarship that details the complicated relationships between generative-AI supply chain and copyright. By making contributions in these three themes, this dissertation serves as an empirical proof by example that research on reliable measurement for machine learning is intimately and inescapably bound up with research in law and policy. These different disciplines pose similar research questions about reliable measurement in machine learning. They are, in fact, two complementary sides of the same research vision, which, broadly construed, aims to construct machine-learning systems that cohere with broader societal values.

Description
655 pages
Date Issued
2024-08
Keywords
artificial intelligence
•
law
•
machine learning
•
technology policy
•
uncertainty estimation
Committee Chair
De Sa, Christopher
Committee Member
Sampson, Adrian
Grimmelmann, James
Kleinberg, Jon
Degree Discipline
Computer Science
Degree Name
Ph. D., Computer Science
Degree Level
Doctor of Philosophy
Type
dissertation or thesis
Link(s) to Catalog Record
https://newcatalog.library.cornell.edu/catalog/16611881

Site Statistics | Help

About eCommons | Policies | Terms of use | Contact Us

copyright © 2002-2026 Cornell University Library | Privacy | Web Accessibility Assistance