Machine Reading Comprehension: Challenges and Approaches
Machine reading comprehension (MRC) tasks have attracted substantial attention from both academia and industry. These tasks require a machine reader to answer questions relevant to a given document provided as input. In this dissertation, we mainly focus on non-extractive MRC, in which a significant percentage of candidate answers are not restricted to text spans from the reference document or corpus. In comparison to extractive MRC tasks, non-extractive MRC tasks contain a significant percentage of questions focusing on the implicitly expressed facts, events, opinions, or emotions in the given text, requiring diverse types of world knowledge (e.g., commonsense, paraphrase, and arithmetic knowledge) and advanced reading skills (e.g., logical reasoning, summarization, and sentiment analysis). This dissertation presents our work in exploring new challenges and approaches for non-extractive MRC. Specifically, on the challenge side, we create the first MRC dataset that focuses on in-depth multi-turn multi-party dialogue understanding and the first free-form multiple-choice Chinese MRC dataset that requires various kinds of prior knowledge. On the approach side, we propose three general reading strategies and a method of utilizing contextualized knowledge to improve non-extractive MRC. We find our datasets to be very challenging for reading comprehension systems and our approaches to be empirically effective on representative non-extractive MRC tasks.