On the Limitations of Data: Mismatches between Neural Models of Language and Humans
The majority of work at the intersection of computational linguistics and natural language processing aims to show, process by process, that human linguistic behavior (and knowledge) is reducible to a simple learning objective (e.g., predicting the next word) applied to unstructured linguistic data (e.g., written data). This dissertation uses three test cases to show concrete instances where current reductionist approaches fall short of human linguistic knowledge. In the first case study, implicit causality, competition among multiple linguistic processes is shown to obscure human-like behavior in models. This challenges existing methodologies that rely on the investigation of individual linguistic processes in isolation and points to a mismatch between human linguistic systems and those built solely on the basis of linguistic data. In the second case study, ambiguous relative clause attachment, models of Spanish and English are compared to show that, while models appear to mimic humans in English, they fail to do so in Spanish. The failure of computational models of Spanish follows from a mismatch between data produced by speakers and speakers' interpretation preferences, and it is argued that this reflects fundamental limitations of text data. In the third case study, Principle B and incremental processing, it is demonstrated that, while humans use hard constraints to restrict their online processing of pronouns, computational models do not. The inability of models to process language incrementally like humans indicates a mismatch between linguistic data and the human parser. This dissertation argues that data are not sufficient to instruct models about fundamental aspects of human language. Ultimately, in using techniques from psycholinguistics and careful cross-linguistic comparison, it is argued that neural models can reveal specific areas of linguistic knowledge where data is not enough, suggesting in turn what the human mind itself must contribute.