This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.
| PDBML | |
|---|---|
| Name | PDBML |
| Developer | Worldwide Protein Data Bank Consortium |
| Released | 2004 |
| Operating system | Cross-platform |
| License | Open data |
PDBML
PDBML is an XML-based archival format for macromolecular structural data used by the Worldwide Protein Data Bank Consortium. It provides a structured representation of atomic coordinates, experimental metadata, and annotation that interoperates with tools and archives associated with the Protein Data Bank, enabling exchange among researchers, databases, and software projects.
PDBML encodes structural biology records in an XML schema derived from efforts by the Worldwide Protein Data Bank Consortium, the Research Collaboratory for Structural Bioinformatics, and related initiatives to modernize molecular archive representations. The format is intended to parallel legacy formats such as the original Protein Data Bank flat-file while leveraging technologies promoted by the World Wide Web Consortium and the Extensible Markup Language community. PDBML records describe macromolecules studied at facilities like the European Synchrotron Radiation Facility, the Advanced Photon Source, and synchrotrons referenced in structural work associated with awards such as the Nobel Prize in Chemistry.
PDBML emerged during standardization efforts following projects led by the Research Collaboratory for Structural Bioinformatics, the Worldwide Protein Data Bank, and national centers including the Protein Data Bank Japan and the Electron Microscopy Data Bank. Early development drew on practices from the Brookhaven National Laboratory era of the Protein Data Bank and responded to increasing volumes of structures deposited by laboratories using methods acknowledged by prizes like the Lasker Award and institutions such as the National Institutes of Health. Committees and working groups involving contributors from universities like Stanford, Cambridge, and Kyoto coordinated schema drafts, with community input from software vendors and archives such as the Cambridge Crystallographic Data Centre and the Biological Magnetic Resonance Data Bank.
The PDBML schema maps archival concepts—authors, citations, experimental data, atomic coordinates—onto XML elements and types formalized with XML Schema constructs. The schema enables validation with tools championed by the World Wide Web Consortium and integrates controlled vocabularies maintained by committees with ties to organizations like the International Union of Crystallography and the International Nucleotide Sequence Database Collaboration. Records include identifiers that cross-reference collections such as UniProt, Gene Ontology annotations curated by the European Bioinformatics Institute, and journal articles indexed in PubMed and listed by publishers like Nature Publishing Group and Elsevier.
PDBML was developed alongside and in dialogue with PDBx/mmCIF, which serves as the master format in current deposition pipelines run by the Worldwide Protein Data Bank. Translators and mapping rules relate PDBML elements to PDBx/mmCIF data items maintained by working groups associated with the International Union of Crystallography and standards efforts linked to databases such as UniProt and RefSeq. Legacy Protein Data Bank flat-file records produced by centers like Rutgers University and archival systems at Brookhaven were considered in conversion strategies, while community adoption involved stakeholders including the European Bioinformatics Institute and the RCSB Protein Data Bank.
Support for PDBML exists in parsers and libraries implemented in programming environments used at academic centers such as Massachusetts Institute of Technology, California Institute of Technology, and University of Oxford. Software projects and visualization tools developed by teams at institutions like the European Molecular Biology Laboratory, the Howard Hughes Medical Institute, and commercial vendors implement import/export facilities. Toolchains often integrate with annotation services provided by UniProt, interaction databases like STRING, and visualization suites used in structural publications appearing in journals such as Science, Cell, and Proceedings of the National Academy of Sciences.
PDBML serves archival exchange among data centers including the RCSB Protein Data Bank, Protein Data Bank Japan, and the European Bioinformatics Institute, facilitating deposition workflows used by researchers at institutions such as Harvard Medical School and Max Planck Institutes. It is used in pipeline systems for validating crystallographic and cryo-EM models from facilities like the National Center for Electron Microscopy and synchrotrons tied to projects recognized by awards such as the Breakthrough Prize. Downstream applications include integration with sequence resources like UniProt, structural classification in SCOP and CATH, and cross-references in pathway databases curated by institutes like Kyoto University and Johns Hopkins University.
Limitations of PDBML include verbosity inherent to XML, challenges in rapid parsing compared with binary or compact text formats favored in high-throughput pipelines at centers like the European Synchrotron Radiation Facility, and the need to maintain parity with evolving PDBx/mmCIF dictionaries overseen by the International Union of Crystallography. Future directions discussed within the Worldwide Protein Data Bank community involve improved tooling from software groups at institutions such as Stanford, enhancements to validation workflows supported by the National Institutes of Health, and tighter integration with ontology efforts led by organizations like the Gene Ontology Consortium and the Open Biological and Biomedical Ontology Foundry.
Category:Data formats