Cornell University
Library
Cornell UniversityLibrary

eCommons

Help
Log In(current)
DigitalCollections@ILR
ILR School
  1. Home
  2. ILR School
  3. Centers, Institutes, Programs
  4. Labor Dynamics Institute
  5. NSF Census Research Network
  6. Presentations by Cornell University NCRN node
  7. Managing Confidentiality and Provenance across Mixed Private and Publicly-Accessed Data and Metadata

Managing Confidentiality and Provenance across Mixed Private and Publicly-Accessed Data and Metadata

File(s)
Presentation-FCSM2013-subdoc.tex (21.47 KB)
LaTeX main content
Presentation-FCSM2013.tex (3.62 KB)
LaTeX shell for presentation
Presentation-FCSM2013.pdf (2.17 MB)
Main presentation file
Presentation-FCSM2013-appendix.tex (3.93 KB)
LaTeX appendix to main document
Permanent Link(s)
https://hdl.handle.net/1813/34534
Collections
Presentations by Cornell University NCRN node
Author
Vilhuber, Lars
Abowd, John
Block, William
Lagoze, Carl
Williams, Jeremy
Abstract

Social science researchers are increasingly interested in making use of confidential micro-data that contains linkages to the identities of people, corporations, etc. The value of this linking lies in the potential to join these identifiable entities with external data such as genome data, geospatial information, and the like. Leveraging these linkages is an essential aspect of “big data” scholarship. However, the utility of these confidential data for scholarship is compromised by the complex nature of their management and curation. This makes it difficult to fulfill US federal data management mandates and interferes with basic scholarly practices such as validation and reuse of existing results.

We describe in this paper our work on the CED2AR prototype, a first step in providing researchers with a tool that spans the confidential/publicly-accessible divide, making it possible for researchers to identify, search, access, and cite those data. The particular points of interest in our work are the cloaking of metadata fields and the expression of provenance chains. For the former, we make use of existing fields in the DDI (Data Description Initiative) specification and suggest some minor changes to the specification. For the latter problem, we investigate the integration of DDI with recent work by the W3C PROV working group that has developed a generalizable and extensible model for expressing data provenance.

Sponsorship
NSF Grant #1131848
Date Issued
2013-11
Publisher
2013 Federal Committee on Statistical Methodology Research Conference
Keywords
PROV
•
DDI
•
Provenance
•
Confidentiality
Type
presentation

Site Statistics | Help

About eCommons | Policies | Terms of use | Contact Us

copyright © 2002-2026 Cornell University Library | Privacy | Web Accessibility Assistance