LLMpediaThe first transparent, open encyclopedia generated by LLMs

PDB (Protein Data Bank)

Note: This article was automatically generated by a large language model (LLM) from purely parametric knowledge (no retrieval). It may contain inaccuracies or hallucinations. This encyclopedia is part of a research project currently under review.
Article Genealogy
Parent: Open Babel Hop 5 terminal

This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.

PDB (Protein Data Bank)
NameProtein Data Bank
CaptionCrystal structure of hemoglobin (representative)
Established1971
FounderBrookhaven National Laboratory, Walter Hamilton
TypeScientific data repository
LocationRutgers University, Research Collaboratory for Structural Bioinformatics
DisciplineStructural biology, Crystallography, Nuclear magnetic resonance

PDB (Protein Data Bank) The Protein Data Bank is a global archive for three-dimensional structural data of biological macromolecules, founded to serve crystallographers and structural biologists. It supports deposition, validation, curation, and dissemination of macromolecular structures to researchers associated with institutions like Brookhaven National Laboratory, Rutgers University, Research Collaboratory for Structural Bioinformatics, and agencies such as the National Institutes of Health. The resource underpins work in fields ranging from X-ray crystallography and nuclear magnetic resonance spectroscopy to cryo-electron microscopy and computational modeling.

History

The archive began in 1971 at Brookhaven National Laboratory with leaders including Walter Hamilton and contributors from laboratories such as Cold Spring Harbor Laboratory and European Molecular Biology Laboratory. During the 1980s and 1990s stewardship shifted through collaborations involving Research Collaboratory for Structural Bioinformatics, Protein Data Bank Japan, and Protein Data Bank Europe, reflecting growth driven by advances at facilities like Stanford Synchrotron Radiation Lightsource and Diamond Light Source. Milestones include the adoption of standardized file formats and community efforts following meetings at venues like Cold Spring Harbor Laboratory and Gordon Research Conferences, with influence from organizations such as the National Science Foundation and the Wellcome Trust shaping policy and funding.

Organization and governance

Governance has involved academic centers and funding agencies including Rutgers University, Research Collaboratory for Structural Bioinformatics, Protein Data Bank Europe, Protein Data Bank Japan, the National Institutes of Health, and the Wellcome Trust. Advisory boards have drawn experts from institutions such as Harvard University, Massachusetts Institute of Technology, University of Cambridge, Max Planck Society, and European Molecular Biology Laboratory to set deposition policies and validation standards. Partnerships with synchrotron facilities like Argonne National Laboratory and community stakeholders including societies such as the International Union of Crystallography govern technical and ethical frameworks.

Data content and structure

The archive stores atomic coordinates, experimental data, and metadata for proteins, nucleic acids, and complexes from groups at Harvard Medical School, University of California, Berkeley, Yale University, University of Oxford, and many others. Data types include coordinates from X-ray crystallography, restraints from nuclear magnetic resonance spectroscopy, maps from cryo-electron microscopy produced by instruments at European Synchrotron Radiation Facility and EMBL-EBI, and annotations using ontologies from UniProt, Gene Ontology, and Chemical Entities of Biological Interest. File formats evolved from legacy coordinate files to standardized formats compatible with software from vendors like Schrödinger, CCP4, and open-source projects such as PyMOL (software), UCSF Chimera, and MODELLER.

Data deposition and validation

Researchers from universities like Stanford University, Columbia University, Johns Hopkins University, and laboratories including Los Alamos National Laboratory deposit structures following policies influenced by journals such as Nature, Science, and Proceedings of the National Academy of Sciences. Validation pipelines integrate tools developed by projects at European Bioinformatics Institute, RCSB PDB, and consortia linked to International Union of Crystallography to check geometry, chemistry, and experimental fit. Community standards and workshops at centers like EMBO and Gordon Research Conferences guide improvements in validation metrics and deposition workflows.

Access, distribution, and tools

Public access is provided through portals maintained by organizations including RCSB PDB, PDBe, and PDBj, with mirrors and services at institutions like European Bioinformatics Institute and Rutgers University. Distribution channels support bulk download via repositories and APIs used by software from groups at Broad Institute, Rosetta Commons, and OpenMM (software project), and support visualization through clients such as PyMOL (software), UCSF ChimeraX, VMD (software), and web tools developed by Google and academic partners. Training and outreach coordinated with societies like American Chemical Society and Biophysical Society promote FAIR principles championed by initiatives like GO FAIR.

Usage and impact in research

Data underpin structural studies and drug discovery in sectors interacting with Pfizer, Merck & Co., Novartis, and academic groups at University of California, San Francisco and Imperial College London, informing computational design using platforms such as Rosetta (software), AutoDock, and machine learning models from teams at DeepMind and Google DeepMind. Structural entries support research published in journals including Cell, Nature Structural & Molecular Biology, and Journal of Molecular Biology, and enable analyses in projects linked to Human Genome Project, Enzyme Commission classifications, and systems biology consortia. The archive facilitates education and reproducibility for courses at MIT, Stanford University School of Medicine, and University of Cambridge.

Interoperability is maintained with resources like UniProt, Pfam, InterPro, SCOP, CATH, BRENDA, BioGRID, Reactome, KEGG, ChEMBL, and DrugBank, enabling cross-references to sequence, functional, and chemical data curated at institutions such as European Bioinformatics Institute, Swiss Institute of Bioinformatics, and National Center for Biotechnology Information. Collaborative standards with organizations including International Nucleotide Sequence Database Collaboration and projects like ELIXIR ensure metadata harmonization and programmatic access for integrative research across structural and systems biology.

Category:Biological databases