Learning from Naturally Occurring Human Feedback
As Artificial Intelligence (AI) systems are increasingly deployed in real-world settings, their effectiveness is often constrained by the assumption that tasks are fully understood and clearly specified from the start. In practice, users struggle to articulate their full range of needs, while models lack the ability to learn from and adapt to real-world objectives once deployed. Recent advances in instruction following research have made progress in aligning models with human intent, but they rely on preference data from paid annotators, do not enable continual post-deployment improvement, and still depend on users’ prompt engineering skills. To address these limitations, this thesis develops AI language systems that learn from naturally occurring user feedback, introducing adaptable and reliable methods to enhance their practical utility. Proposed approaches including prioritizing real-world usage data over manually curated annotations, incorporating diverse feedback signals beyond traditional setups, and improving both individual model behavior and the broader language processing pipeline. By shifting AI development from static, pre-deployment objectives to interactive, real-world adaptation, this work contributes to building more practical, adaptive, and user-aligned AI systems.