LLMpediaThe first transparent, open encyclopedia generated by LLMs

DCC Curation Lifecycle Model

Note: This article was automatically generated by a large language model (LLM) from purely parametric knowledge (no retrieval). It may contain inaccuracies or hallucinations. This encyclopedia is part of a research project currently under review.
Article Genealogy
Parent: National Digital Stewardship Alliance Hop 6 terminal

This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.

DCC Curation Lifecycle Model
NameDCC Curation Lifecycle Model
AbbreviationDCC
TypeData curation model
Developed byDigital Curation Centre

DCC Curation Lifecycle Model The DCC Curation Lifecycle Model is a conceptual framework for managing research data through stages of planning, appraisal, ingest, preservation, access, and reuse. It provides guidance for institutions, archives, and researchers to ensure long-term value of datasets across scholarly domains and organizational contexts. The model is referenced by libraries, archives, and funding agencies as part of data management and stewardship practices.

Overview

The model articulates discrete yet interconnected activities including policy planning and project management that inform preservation actions used by United Kingdom Research and Innovation, National Institutes of Health, European Commission, Wellcome Trust, and other funders. It emphasizes active curation during research workflows encountered at institutions such as University of Oxford, University of Cambridge, Harvard University, Stanford University, and Massachusetts Institute of Technology. The framework complements standards promulgated by bodies like International Organization for Standardization, Research Data Alliance, Digital Preservation Coalition, OpenAIRE, and Committee on Data for Science and Technology.

History and Development

The model was developed by the Digital Curation Centre in response to evolving data policies and practices influenced by case studies from repositories such as ICPSR, UK Data Service, Dryad Digital Repository, Zenodo, and Figshare. Its iterations reflect dialogues with stakeholders including the British Library, Library of Congress, National Archives (United Kingdom), European Research Council, and national consortia in Australia, Canada, and the United States National Science Foundation. The evolution of the model paralleled initiatives like the Open Science Framework, the FAIR Guiding Principles, and the adoption of mandates by funders such as Horizon 2020 and NIH Data Sharing Policy.

Core Components and Processes

Core activities in the model cover appraisal, ingest, preservation planning, metadata creation, storage, access, and reuse, aligning with practices at repositories such as ArXiv, PubMed Central, PLOS, Springer Nature, and Elsevier. It maps responsibilities among stakeholders including university libraries like Columbia University Libraries, consortia such as California Digital Library, and standards organizations including Dublin Core Metadata Initiative and ISO/IEC JTC 1. The model integrates policy influences from bodies such as Jisc, Research Councils UK, and European Commission Directorate-General for Research and Innovation while recognizing practical workflows at archives like The National Archives (UK) and museums including the British Museum.

Implementation and Use Cases

Organizations implement the model via institutional repositories, data management plans, and preservation services used by projects at Wellcome Sanger Institute, CERN, European Space Agency, NASA, and clinical trials networks tied to World Health Organization protocols. Use cases span disciplines represented at Max Planck Society, Chinese Academy of Sciences, Australian National University, University of Cape Town, and multidisciplinary initiatives like the Human Genome Project and Large Hadron Collider. Implementation often involves partnerships with commercial platforms offered by Amazon Web Services, Google Cloud, and Microsoft Azure as well as open infrastructure such as CKAN and Dataverse.

Tools, Standards, and Interoperability

The model interoperates with metadata schemas and tools endorsed by Dublin Core Metadata Initiative, PREMIS, DataCite, and OGC. Software ecosystems include Archivematica, Preservica, Fedora Commons, Invenio, BitCurator, and workflow engines used in projects at European Bioinformatics Institute and National Center for Biotechnology Information. Alignment with identifiers from ORCID, DOI, Handle System, and vocabularies from Library of Congress Subject Headings and Getty Vocabulary Program supports discoverability and reuse. Crosswalks with standards by W3C, OAI-PMH, and SPARQL endpoints enable integration with scholarly infrastructures like Crossref, OpenAIRE, and Scholix.

Impact and Evaluation

Adoption of the model has informed policy and infrastructure choices at universities including Yale University, Princeton University, University of Melbourne, and University of Toronto, and influenced national strategies in Canada, France, Germany, Japan, and South Africa. Evaluations often reference metrics used by CORE, citation data tracked by Scopus and Web of Science, and repository performance studies citing Jisc and Digital Preservation Coalition reports. Case studies from initiatives such as H2020 projects and national research infrastructure programs illustrate benefits in data discoverability, reuse, and compliance with mandates from NIH, European Research Council, and philanthropic funders like the Gates Foundation.

Criticisms and Limitations

Critics argue that the model is conceptual and may lack prescriptive operational detail needed by small organizations, as noted in reviews by Scholarly Communications units at University of Edinburgh and University College London. Concerns include the resource intensity highlighted by audits from National Audit Office (United Kingdom) and challenges integrating legacy systems at institutions such as Smithsonian Institution and National Institutes of Health Clinical Center. Interoperability issues remain when mapping to domain-specific standards used by communities around GenBank, PDB, and disciplinary repositories like SSRN.

Category:Data curation