Cornell University
Library
Cornell UniversityLibrary

eCommons

Help
Log In(current)
  1. Home
  2. Cornell University Graduate School
  3. Cornell Theses and Dissertations
  4. Data-Aware Algorithms for Secure and Efficient Machine Learning Systems

Data-Aware Algorithms for Secure and Efficient Machine Learning Systems

File(s)
Tiwari_cornellgrad_0058F_15495.pdf (1.46 MB)
Permanent Link(s)
https://doi.org/10.7298/4w7p-sp77
https://hdl.handle.net/1813/126505
Collections
Cornell Theses and Dissertations
Author
Tiwari, Trishita
Abstract

Machine learning models are increasingly deployed in real-world systems where they must operate under strict inference-time constraints, including data access control, privacy guarantees, and efficiency requirements. Despite this, most modern ML pipelines are designed and evaluated in a largely data-agnostic manner: training data is freely entangled in model parameters, memorization is measured without accounting for statistical generalization, and representations are optimized without regard to how they are queried at inference time. This mismatch leads to systems that are difficult to control, difficult to interpret, and inefficient in practice. This dissertation argues for an data-aware approach to machine learning system design. It presents three complementary contributions that incorporate data structure and constraints directly into model architectures, evaluation metrics, and optimization objectives. First, it introduces information flow control for machine learning, formalizing a notion of non-interference and proposing a modular Transformer architecture that enforces access policies at inference time through secure gating and aggregation. Second, it develops Prior-Aware Memorization, an efficient, training-free metric that distinguishes genuine memorization of training data from statistical generalization, enabling more accurate measurement of privacy and copyright leakage in large language models. Third, it proposes Query-Aware Compression for semantic search, deriving an analytic, training-free compression method that minimizes similarity distortion under real query distributions. Across language modeling and large-scale retrieval tasks, empirical results demonstrate that data-aware designs can substantially improve security and efficiency with minimal overhead. Collectively, this work advances the design of machine learning systems that are aligned with real-world data constraints and provides a foundation for building models that are not only accurate, but also controllable and trustworthy.

Description
151 pages
Date Issued
2026-05
Keywords
Machine Learning
•
Memorization
•
Privacy
•
Security
•
Vector Compression
Committee Chair
Suh, Gookwon Edward
Committee Member
Agarwal, Rachit
Weinberger, Kilian
Brann, Ross
Degree Discipline
Computer Science
Degree Name
Ph. D., Computer Science
Degree Level
Doctor of Philosophy
Rights
Attribution 4.0 International
Rights URI
https://creativecommons.org/licenses/by/4.0/
Type
dissertation or thesis

Site Statistics | Help

About eCommons | Policies | Terms of use | Contact Us

copyright © 2002-2026 Cornell University Library | Privacy | Web Accessibility Assistance