This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.
| CATH (protein structure classification) | |
|---|---|
| Name | CATH |
| Title | CATH (protein structure classification) |
| Description | Hierarchical protein domain classification resource |
| Scope | Structural bioinformatics |
CATH (protein structure classification) is a hierarchical resource that organizes protein domain structures into a consistent ontology for comparative analysis, annotation, and discovery. It integrates experimentally determined structures, computationally derived domain assignments, and evolutionary relationships to facilitate research spanning structural biology, molecular evolution, and bioinformatics. The resource is widely used by researchers at institutions such as European Bioinformatics Institute, Wellcome Trust Sanger Institute, and in collaborations with groups at University of Cambridge, University of Oxford, and Massachusetts Institute of Technology.
CATH classifies protein domains by combining automated algorithms and manual curation to assign structures into discrete levels: Class, Architecture, Topology, and Homologous superfamily. The resource complements other structural resources and initiatives including Protein Data Bank, SCOP, Pfam, InterPro, and efforts by consortia such as the Protein Structure Initiative and projects at European Molecular Biology Laboratory. Its outputs support pipelines at organizations like UniProt, Ensembl, NCBI, European Nucleotide Archive, and research groups in computational biology.
CATH emerged from collaborations between research groups led by investigators at Glasgow University and the University College London structural bioinformatics teams, with early influences from classification efforts at Brookhaven National Laboratory and the European Bioinformatics Institute. Over successive releases the project integrated methods from groups including those at MRC Laboratory of Molecular Biology and Cambridge University while interacting with initiatives such as the Human Genome Project and analyses published in journals linked to the Royal Society. Funding and institutional support have involved agencies like the Wellcome Trust and national research councils tied to United Kingdom science infrastructure.
At its core CATH employs a hierarchy: Class (secondary structure composition), Architecture (overall shape), Topology (fold family), and Homologous superfamily (evolutionary relationships). This hierarchical model interfaces conceptually with schemas used by databases such as Protein Data Bank and classification schemes from SCOP and annotations propagated into resources like UniProt. Decisions at the Homologous superfamily level reflect evolutionary judgments comparable to those considered by groups at Max Planck Institute and by authors publishing in venues associated with the European Molecular Biology Organization.
CATH ingests three-dimensional coordinates primarily from the Protein Data Bank and augments them with sequence data cross-referenced to databases such as UniProt, Pfam, and Ensembl. The methodology integrates structural alignment algorithms, hidden Markov models influenced by work at European Bioinformatics Institute, and manual curation steps by domain experts with affiliations including University of York and Imperial College London. Quality control and release practices have drawn on standards used by repositories like GenBank and tools developed in collaboration with groups at University of California, San Diego and Stanford University.
CATH provides downloadable datasets, searchable web interfaces, and programmatic access that researchers at centers such as Wellcome Trust Sanger Institute, European Bioinformatics Institute, and Max Planck Institute for Biophysical Chemistry use for comparative modeling, fold recognition, and genome annotation. Integration points include pipelines driven by HMMER implementations and workflows used by teams at Swiss Institute of Bioinformatics and software ecosystems developed alongside Rosetta (software), MODELLER, and visualization suites comparable to those from UCSF and groups at University of California, San Francisco.
CATH classifications underpin studies in structural genomics, functional annotation, and evolutionary analysis undertaken by consortia like the Protein Structure Initiative and academic labs at Imperial College London, University of Cambridge, and Harvard University. Outputs from CATH inform annotation in UniProt and comparative studies published in journals associated with the Nature Publishing Group, Cell Press, and the Proceedings of the National Academy of Sciences. The resource has influenced machine learning efforts at institutions such as Google DeepMind and academic collaborations at Massachusetts Institute of Technology focusing on fold prediction and structure–function relationships.
Critiques of CATH mirror broader debates in structural classification: boundary definitions, subjective manual curation, and differences with alternative schemes such as SCOP and domain definitions used by Pfam. Concerns raised in community discussions involving researchers from European Bioinformatics Institute, Wellcome Trust Sanger Institute, and university groups at University of Cambridge address reproducibility, interoperability with resources like UniProt and Protein Data Bank, and the challenge of scaling to predicted models emerging from projects at Google DeepMind and structural prediction efforts at DeepMind Technologies Limited. Ongoing development seeks to balance automated scalability with expert oversight practiced at institutions such as Imperial College London and MRC Laboratory of Molecular Biology.
Category:Protein structure databases