Cornell University
Library
Cornell UniversityLibrary

eCommons

Help
Log In(current)
  1. Home
  2. Cornell University Graduate School
  3. Cornell Theses and Dissertations
  4. DP-CP-RAG: DUAL-PROJECTION CONSTRAINT-PROVIDER RETRIEVAL-AUGMENTED GENERATION FOR EVIDENCE-GROUNDED PORT-DOMAIN QUESTION ANSWERING

DP-CP-RAG: DUAL-PROJECTION CONSTRAINT-PROVIDER RETRIEVAL-AUGMENTED GENERATION FOR EVIDENCE-GROUNDED PORT-DOMAIN QUESTION ANSWERING

Access Restricted

Access to this document is restricted. Some items have been embargoed at the request of the author, but will be made publicly available after the "No Access Until" date.

During the embargo period, you may request access to the item by clicking the link to the restricted file(s) and completing the request form. If we have contact information for a Cornell author, we will contact the author and request permission to provide access. If we do not have contact information for a Cornell author, or the author denies or does not respond to our inquiry, we will not be able to provide access. For more information, review our policies for restricted content.

File(s)
Li_cornell_0058O_12718.pdf (2.51 MB)
No Access Until
2028-06-22
Permanent Link(s)
https://doi.org/10.7298/masp-5v76
https://hdl.handle.net/1813/126244
Collections
Cornell Theses and Dissertations
Author
Li, Yichen
Abstract

Question answering in port operational analytics requires more than plain text retrieval or direct large language model generation. In this setting, evidence is distributed across technical documents, structured operational tables, and domain-specific KPI definitions. The core challenge is not only how to retrieve information, but also when retrieval is necessary, how natural-language questions should be grounded to database schema, and how structured evidence should be reconciled with documentary semantics. This thesis proposes DP-CP-RAG, a dual-projection constraint-provider retrieval-augmented generation framework for evidence-grounded question answering in the port domain. The framework integrates adaptive routing, rule-based identifier extraction, intent extraction, schema grounding, SQL synthesis and execution, constraint gating, claim verification, and fusion-based answer generation within a unified end-to-end pipeline. It supports three processing depths for non-domain, knowledge-only, and data-dependent queries, thereby improving both efficiency and answer reliability. Experiments are conducted on three benchmark groups: a 25-question multiple-choice benchmark, a 12-question open benchmark covering integration, distractor, and refusal settings, and a 15-question user-query benchmark. The strongest gains are observed on the multiple-choice benchmark, where the full system improves over direct answering by up to 32 percentage points for Mixtral and Llama and 28 percentage points for Gemini under best-of-3 evaluation. On the user-query benchmark, the full system also substantially improves performance on operational questions. Overall, the results show that structured evidence grounding, constraint-aware filtering, and execution-based verification are more effective than direct answering or naive vector-only retrieval for definition-sensitive and cross-table reasoning in port operations. These findings suggest that combining documentary constraints with executable relational evidence is a practical path toward trustworthy domain-specific question answering.

Description
91 pages
Date Issued
2026-05
Keywords
Large Language Models
•
Retrieval-Augmented Generation
•
Text-to-SQL
Committee Chair
Gao, Huaizhu
Committee Member
Dean, Sarah
Degree Discipline
Systems Engineering
Degree Name
M.S., Systems Engineering
Degree Level
Master of Science
Type
dissertation or thesis

Site Statistics | Help

About eCommons | Policies | Terms of use | Contact Us

copyright © 2002-2026 Cornell University Library | Privacy | Web Accessibility Assistance