Cornell University
Library
Cornell UniversityLibrary

eCommons

Help
Log In(current)
  1. Home
  2. Cornell University Graduate School
  3. Cornell Theses and Dissertations
  4. Code Generation with Large Language Models: Inductive Reasoning and Calibration

Code Generation with Large Language Models: Inductive Reasoning and Calibration

File(s)
Li_cornellgrad_0058F_15000.pdf (14.13 MB)
Permanent Link(s)
https://doi.org/10.7298/xbdg-dx50
https://hdl.handle.net/1813/120777
Collections
Cornell Theses and Dissertations
Author
Li, Wen-Ding
Abstract

Large Language Models (LLMs) have demonstrated remarkable capabilities across various domains, including code generation. However, complex inductive reasoning, deriving general rules from limited observations, remains a significant challenge. Programming-by-Examples (PBE) aims to synthesize programs from input-output examples, representing an important inductive reasoning task in programming languages with practical applications. We propose an approach to enhance LLMs on PBE using code-grounded synthetic data generation to provide high-quality training data for finetuning LLMs and address the scarcity of domain-specific data. Furthermore, we demonstrate how scaling test-time computation significantly improves inference results in this PBE setting. Our approach achieves state-of-the-art results on common PBE benchmarks including string, number sequence, and logo graphics domains. We further extend our methods to ARC-AGI, a very challenging benchmark requiring visual inductive reasoning from a few examples involving concepts such as physics, objects and symmetry. By applying our synthetic data and test-time scaling method, and then combining with transduction, we can approach human-level performance on ARC-AGI, demonstrating the framework's effectiveness even in highly challenging, visually-grounded domains. Unlike PBE and ARC-AGI tasks where examples enable direct validation, real-world code generation often begins with ambiguous natural language specifications. This inherent ambiguity creates uncertainty about code correctness. We develop an approach that samples both code and tests from LLMs and uses execution results to build a classifier that estimates correctness probabilities. The method produces human-interpretable predicates explaining code behavior, a feature that users preferred in the user study, and helps create more trustworthy program synthesis while maintaining state-of-the-art accuracy.

Description
255 pages
Date Issued
2025-08
Keywords
Code Generation
•
Inductive Reasoning
•
Large Language Models
•
Program Synthesis
•
Programming-by-Examples
•
Synthetic Data Generation
Committee Chair
Ellis, Kevin
Committee Member
Legunsen, Owolabi
Sampson, Adrian
Degree Discipline
Computer Science
Degree Name
Ph. D., Computer Science
Degree Level
Doctor of Philosophy
Type
dissertation or thesis

Site Statistics | Help

About eCommons | Policies | Terms of use | Contact Us

copyright © 2002-2026 Cornell University Library | Privacy | Web Accessibility Assistance