LLMpediaThe first transparent, open encyclopedia generated by LLMs

Carbohydrate-Active enZYmes database

Note: This article was automatically generated by a large language model (LLM) from purely parametric knowledge (no retrieval). It may contain inaccuracies or hallucinations. This encyclopedia is part of a research project currently under review.
Article Genealogy
Parent: glycosyltransferase family 2 Hop 5 terminal

This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.

Carbohydrate-Active enZYmes database
NameCarbohydrate-Active enZYmes database
AbbreviationCAZy
Established1998
FocusEnzyme classification, glycoscience, bioinformatics
CountryFrance
HostAix-Marseille University
Website(omitted)

Carbohydrate-Active enZYmes database The Carbohydrate-Active enZYmes database is a curated resource for enzymes that build, modify, or break down carbohydrates and glycoconjugates, serving researchers across molecular biology, structural biology, and biotechnology. It links experimental enzymology, sequence analysis, and structural data to ongoing work in genomics and metagenomics, supporting studies from academic laboratories to industrial research groups. The resource is widely used by investigators working with pathogens, plants, and microbial communities.

Overview

The resource organizes enzyme families by sequence-based homology and catalytic mechanism, enabling integration with databases and projects such as UniProt, Protein Data Bank, Ensembl, NCBI, and European Nucleotide Archive. It connects biochemical nomenclature used by International Union of Biochemistry and Molecular Biology and structural classifications used by SCOP and CATH while providing cross-references to model organism resources like Saccharomyces cerevisiae, Arabidopsis thaliana, and Escherichia coli. The platform supports comparative analyses relevant to initiatives including the Human Genome Project, Microbiome Project, and translational efforts linked to Bill & Melinda Gates Foundation-funded programs.

History and Development

The database was initiated by researchers at French institutions with ties to CNRS and Aix-Marseille University, emerging amid the expansion of sequence databases during the late 1990s similar to efforts by European Bioinformatics Institute and National Center for Biotechnology Information. Early development paralleled community-driven resources such as Pfam, PROSITE, and InterPro, and later incorporated methods developed in collaborations with groups at Massachusetts Institute of Technology, European Molecular Biology Laboratory, and industrial partners like Novozymes and Genentech. Key milestones included adding structural cross-links to entries from RCSB PDB and adapting to large-scale metagenomic datasets generated by consortia such as the TerraGenome Project and the Earth Microbiome Project.

Data Content and Classification

Content is organized into major enzyme classes—glycoside hydrolases, glycosyltransferases, polysaccharide lyases, carbohydrate esterases, and auxiliary activity enzymes—aligned with classifications used by the International Union of Biochemistry and Molecular Biology and semantic frameworks similar to Gene Ontology. Each family entry integrates sequence exemplars from GenBank and curated annotations from UniProtKB/Swiss-Prot, while structural exemplars are linked to entries in the Protein Data Bank and functional assays cited in journals such as Nature, Science, Cell, Journal of Biological Chemistry, and Glycobiology. Taxonomic breadth spans organisms studied at institutions like Harvard University, Stanford University, and Max Planck Society, and metagenomic samples collected by projects including Global Ocean Sampling Expedition.

Annotation and Curation Methods

Curation employs manual expert review combined with automated sequence-similarity searches using tools developed in concert with groups at European Bioinformatics Institute and Broad Institute, leveraging algorithms like BLAST and HMMER originally described by researchers at National Institutes of Health and Wellcome Trust Sanger Institute. Annotations reference experimental literature from publishers including Oxford University Press, Elsevier, and Wiley-Blackwell and draw on community standards promulgated by organizations such as the International Society for Computational Biology. Quality control protocols reflect best practices established in collaborations with Swiss Institute of Bioinformatics and benchmark datasets produced by consortia including Critical Assessment of protein Structure Prediction.

Database Access and Tools

The platform provides web-based browsing and downloadable datasets compatible with pipelines used at laboratories from Cold Spring Harbor Laboratory to industrial bioinformatics teams at DSM-Firmenich. Tools include family pages, sequence search interfaces, and links to external visualization tools developed by groups at EMBL-EBI, RCSB, and UCSF. Integration points include programmatic access modalities similar to APIs offered by Ensembl and data exchange formats used by BioMart and Galaxy Project workflows, enabling incorporation into computational projects led by researchers at Imperial College London and University of California, Berkeley.

Applications and Impact

The resource underpins applied research in biofuel development pursued at DOE-funded centers, agricultural biotechnology programs at USDA, and pharmaceutical enzyme discovery in companies such as Merck and Pfizer. It informs structural enzymology studies at facilities like Diamond Light Source and Advanced Photon Source and supports translational glycoengineering carried out at ETH Zurich and Delft University of Technology. By enabling annotation in microbiome and pathogen research, it contributes to public health efforts coordinated by World Health Organization and to sustainability initiatives supported by organizations such as European Commission research programs.

Community and Collaborations

Sustained development relies on collaborations with academic groups at University of Oxford, Johns Hopkins University, and University of Tokyo, as well as contributions from industrial partners and international consortia including ELIXIR and Global Biodata Coalition. Training and outreach activities align with workshops held at conferences like Gordon Research Conferences, Glyco25, and meetings organized by the American Society for Microbiology, fostering community standards and shared resources across the glycoscience and bioinformatics communities.

Category:Bioinformatics databases