LLMpediaThe first transparent, open encyclopedia generated by LLMs

IUPAC InChI

Note: This article was automatically generated by a large language model (LLM) from purely parametric knowledge (no retrieval). It may contain inaccuracies or hallucinations. This encyclopedia is part of a research project currently under review.
Article Genealogy
Parent: RDKit Hop 5 terminal

This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.

IUPAC InChI
NameIUPAC International Chemical Identifier
AcronymInChI
DeveloperInternational Union of Pure and Applied Chemistry
Introduced2005
Latest releaseIUPAC InChI Trust releases
Formattextual identifier

IUPAC InChI is a textual identifier standard created under the auspices of the International Union of Pure and Applied Chemistry to represent chemical substances in a machine-readable string. It was developed to enable interoperability among databases such as PubChem, ChemSpider, ChEBI, and DrugBank and to support cheminformatics tools used at institutions like European Bioinformatics Institute and National Center for Biotechnology Information. The standard complements identifier systems employed by organizations including CAS Registry, WIPO, World Health Organization, and projects such as Human Genome Project and Protein Data Bank that rely on precise cross-referencing.

Overview

InChI provides a layered, canonical text string that encodes structural information for small molecules, facilitating searches across resources like Google Scholar, Scopus, Web of Science, arXiv, Springer Nature, and repositories such as Zenodo and Figshare. It was designed by committees associated with IUPAC and the InChI Trust to interoperate with standards from bodies including ISO and software from vendors such as OpenEye Scientific, ChemAxon, RDKit, Accelrys, and PerkinElmer. InChI strings are often used alongside hashed forms like the InChIKey to enable fast lookup in services run by organizations like Wikidata, Wikipedia, EuroPMC, and national archives like the Library of Congress.

History and development

Work on InChI began in the early 2000s through collaborations that involved representatives from IUPAC, the Royal Society of Chemistry, and the International Union of Crystallography, and drew on algorithms from groups at NIST, Scripps Research, University of Cambridge, Massachusetts Institute of Technology, Stanford University, and University of California, Berkeley. Early milestones included pilot implementations showcased at conferences such as American Chemical Society national meetings and presentations at Gordon Research Conferences. The InChI Trust, established with support from benefactors including foundations and agencies like Wellcome Trust and National Institutes of Health, coordinates maintenance, releases, and community governance in partnership with publishers such as Elsevier and Wiley-Blackwell.

Chemical structure encoding and layers

InChI encodes molecular information through ordered layers: connectivity, hydrogen atoms, tautomeric information, isotopic composition, stereochemistry, and electronic charge states, enabling mapping to registry systems like CAS Registry, European Nucleotide Archive, UniProt, and KEGG. The layered approach supports representation of stereoisomers discussed in literature from Linus Pauling to Robert Burns Woodward and enables compatibility with cheminformatics toolkits from Open Babel, RDKit, ChemAxon, and commercial suites in use at companies such as Pfizer, Novartis, GlaxoSmithKline, and Roche. For computational workflows used by labs affiliated with CERN or projects like Human Cell Atlas, layers permit selective matching for searches in databases including Reaxys and SciFinder.

Versions and implementations

InChI has evolved through versions maintained by the InChI Trust and implemented in open-source libraries employed by projects like OpenWetWare and infrastructures run by European Molecular Biology Laboratory and NCBI. Software implementations include command-line tools, libraries embedded in Java, Python, and C++ ecosystems, and integrations with platforms such as KNIME, Galaxy (platform), Jupyter Notebook, and commercial ELN systems produced by Labguru and Benchling. Version updates have been discussed at meetings of organizations like IUPAC divisions, American Chemical Society committees, and working groups convened at events such as AChemS and GCC.

Applications and adoption

InChI is used widely in cheminformatics, drug discovery pipelines at firms like AstraZeneca and Bayer, regulatory submissions to agencies including the US Food and Drug Administration and European Medicines Agency, and data integration efforts at consortia like OpenPHACTS and ELIXIR. It supports indexing in literature platforms such as PubMed Central and data linkage in community resources such as Wikidata, enabling cross-references with projects like DBpedia and institutional repositories at universities including Harvard University and University of Oxford. InChIKey values facilitate web-scale discovery via search engines operated by Google, Microsoft Bing, and domain-specific services run by Chemical Abstracts Service.

Limitations and criticisms

Criticisms of InChI include difficulties representing polymers, mixtures, macromolecules, and materials studied at facilities like Diamond Light Source and European Synchrotron Radiation Facility, which has led researchers at institutions such as Max Planck Society and Lawrence Berkeley National Laboratory to seek complementary standards. Concerns raised by stakeholders including publishers like Nature Publishing Group and consortia such as CrossRef involve canonicalization ambiguities, tautomer handling, and stereochemical edge cases that affect interoperability with identifiers like CAS Registry Number and sequence identifiers used by GenBank. Debates at conferences hosted by Royal Society and panels convened by IUPAC have addressed governance, versioning, and extension mechanisms.

InChI interoperates with and complements identifiers such as the CAS Registry Number, InChIKey, SMILES developed at institutions including Daylight Chemical Information Systems and fields represented by FAIR Principles advocates, and registry keys used in databases like PubChem, ChemSpider, ChEBI, and DrugBank. Crosswalks and mapping efforts involve standards bodies such as ISO, data initiatives like FAIRsharing, and infrastructure projects funded by agencies like European Commission and National Science Foundation, facilitating integration with semantic resources including Wikidata and linked-data platforms managed by organizations such as DBpedia.

Category:Chemical identifiers