This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.
| DataVerse Project | |
|---|---|
| Name | DataVerse Project |
| Developer | Harvard University, IQSS, DataONE, NCSA, MIT |
| Released | 2006 |
| Programming language | Java (programming language), Ruby (programming language), Python (programming language) |
| Operating system | Linux, Windows, macOS |
| License | BSD license, MIT License |
DataVerse Project The DataVerse Project is an open-source repository platform and ecosystem designed to publish, cite, and preserve research data across academic and cultural institutions. It integrates persistent identifiers, metadata schemas, and access controls to support reproducible research workflows used by universities, libraries, and funding agencies. The platform interoperates with institutional repositories, scholarly publishers, and national data infrastructures to enable discovery and reuse of datasets across disciplines.
The project provides a software stack for data publication, combining features found in Dataverse (software), Zenodo, Figshare, Dryad (repository), and ICPSR while interoperating with registries such as DataCite and Crossref. It supports persistent identifiers like DOI and integrates with identity providers including ORCID and Shibboleth. Institutions deploy the platform to connect with infrastructure projects such as European Open Science Cloud, Science Gateways Community Institute, NARA, and Duraspace while aligning with standards from W3C, ISO, and FAIR principles advocacy groups. The project’s governance model draws on consortial arrangements similar to Apache Software Foundation and DuraSpace partnerships.
Initial development began in the mid-2000s with contributions from Harvard University, IQSS, and collaborators associated with ICPSR and ODI (open data institute). Early funding and pilot deployments involved grants from agencies such as the National Science Foundation, Wellcome Trust, and European Commission. Subsequent development cycles incorporated integration with services from DataCite, Crossref, ORCID, and national infrastructures like JISC and SSHRC. Major milestones included interoperability work with CKAN-based portals, adoption by consortia such as Canadian Data Service and Australian Research Data Commons, and forks or integrations by projects at NCSA, MIT Libraries, and Stanford University.
The architecture uses a modular server-side core written in Java (programming language) with RESTful APIs inspired by Representational State Transfer patterns and client components in JavaScript frameworks. Key components mirror services like Elasticsearch for indexing, PostgreSQL for relational storage, and Amazon S3-style object stores for file preservation operated by providers such as AWS or Chronopolis. Metadata handling relies on schemas and frameworks from DataCite Metadata Schema, Dublin Core, and schema.org, while authentication integrates ORCID and Shibboleth implementations. The platform supports plugin architectures akin to Apache Tomcat modules and containerized deployment patterns using Docker and Kubernetes.
Data ingestion and curation workflows implement metadata profiles compatible with DataCite, Dublin Core, ISO 19115 for geospatial metadata, and domain standards used by repositories like PANGAEA and GenBank. The system enforces identifiers such as DOI and granular accessions comparable to BioSample and GEO (Gene Expression Omnibus), and it maps to discovery services like Crossref and ORCID records. Preservation strategies reference models from OAIS and standards from NARA and ISO 14721, while validation and provenance tracking adopt practices advocated by W3C PROV and registries such as ORCID and DataCite Commons.
Academic libraries at institutions such as Harvard University, Princeton University, Yale University, and University of Cambridge deploy the platform for institutional repositories, integrating with learning platforms like Moodle and research information systems such as Pure (software). Funders including NSF, NIH, and ERC recommend or mandate deposition workflows that the platform supports, enabling citation in journals like Nature, Science (journal), and PLOS. Domain repositories for social science, genomics, and climate science interoperate with tools from R Project for Statistical Computing, Python (programming language), and MATLAB to enable reproducible computation pipelines connected to workflow systems such as Galaxy (platform) and Jupyter Notebook.
Community governance follows a consortial model with steering committees and working groups reminiscent of Apache Software Foundation and Linux Foundation practices, drawing membership from universities, libraries, and national data centers including DANS, DataONE, and Australian Research Data Commons. The contributor community collaborates through issue trackers and code hosting platforms like GitHub and coordinates roadmaps at conferences such as Open Repositories, IDCC (International Digital Curation Conference), and Force11. Training and outreach leverage organizations like Society of American Archivists, Research Data Alliance, and Association of Research Libraries.
Operational security implements TLS and authentication stacks aligned with OAuth 2.0 and SAML profiles, integrates logging compatible with SIEM systems, and supports encryption-at-rest strategies available from AWS, Google Cloud Platform, and Azure (Microsoft Azure). Privacy and data handling practices align with regulatory frameworks such as GDPR and funder policies from NIH and NSF, while controlled-access workflows interoperate with data access committees modeled after dbGaP and European Genome-phenome Archive. Audit, retention, and legal deposit obligations reference standards from NARA and regional laws such as UK Data Protection Act.
Category:Open-source software projects