Practical Data Quality for Modern Data & Modern Uses, With Applications To America's COVID-19 Data
Modern data is often assumed to be of high quality. We dismiss this assumption using examples from America’s COVID-19 data. We explore the types and origins of these issues by diving into the data production process, noting that the root causes are not particular to this dataset but are typical of modern data in general. Data quality issues are frequently surprising, sometimes baffling, and often overwritten. We recover these issues by creating data releases which enable us to replay history. Using novel visualizations, we are able to surface quality issues within and across releases. Two such issues are of particular concern: major restatements, which rewrite history, and non-retroactive changes, which restart history. Data quality is defined by use case. We explore this through two applications: allocation of finite resources and surge prediction. While doing so, we propose κ-accuracy, a practical solution to a common obstacle, and argue for quality-based data selection. Research into data quality is urgently needed. The field is young, the theory uncoordinated, and the metrics all but nonexistent. Yet, assessing data quality is essential to using data properly. Researchers have developed so many ways to use data, and so few ways to assess whether or not we should. We hope this work will inspire research and investment into data quality.