LLMpediaThe first transparent, open encyclopedia generated by LLMs

TEI header (teiHeader)

Note: This article was automatically generated by a large language model (LLM) from purely parametric knowledge (no retrieval). It may contain inaccuracies or hallucinations. This encyclopedia is part of a research project currently under review.
Article Genealogy
Parent: TEI Guidelines Hop 6 terminal

This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.

TEI header (teiHeader)
NameTEI header (teiHeader)
AbbreviationteiHeader
DomainDigital humanities, Text encoding

TEI header (teiHeader) The TEI header (teiHeader) is the metadata container defined by the Text Encoding Initiative used to describe digital texts and their editorial contexts. It interoperates with XML standards to provide descriptive, administrative, and technical metadata for scholarly editions, corpora, and archival resources. Implementations commonly appear in projects involving libraries, museums, and research infrastructures.

Overview

The teiHeader functions as the central metadata block in TEI P5 XML documents and connects editorial practice from projects such as Oxford University, British Library, Library of Congress, Bodleian Library, Harvard University, Yale University, Princeton University, Stanford University, Cambridge University, Max Planck Institute for the History of Science, Getty Research Institute, Smithsonian Institution, British Museum, Vatican Library, Bibliothèque nationale de France, National Archives (United Kingdom), National Archives and Records Administration, German National Library, Chinese Academy of Social Sciences, Kunsthistorisches Institut in Florenz, Columbia University, University of Toronto, University of Pennsylvania, King's College London, University of California, Berkeley with preservation frameworks like Dublin Core Metadata Initiative, PREMIS, MODS, and standards from International Organization for Standardization.

Structure and Elements

A teiHeader typically contains subcomponents such as fileDesc, profileDesc, and revisionDesc, aligning with XML element conventions used by World Wide Web Consortium, Unicode Consortium, ISO 8601 conventions, and bibliographic traditions from Anglo-American Cataloguing Rules and Library of Congress Subject Headings. Elements inside fileDesc include titleStmt, publicationStmt, and sourceDesc; profileDesc often carries language, creation, and creation-related markup that curators from British Library and Bibliothèque nationale de France use in editorial workflows. revisionDesc records provenance akin to practices at National Archives (United Kingdom) and National Archives and Records Administration; encodingDesc documents the technical environment with references to Extensible Markup Language and XSLT processors like those from Saxonica or Apache Software Foundation.

Usage and Purpose

Editors at institutions such as Perseus Project, Project Gutenberg, Early English Books Online, EEBO-TCP project, Gallica, Trove (National Library of Australia), and Europeana use the teiHeader to capture bibliographic description, editorial statements, rights, and machine-actionable provenance. It supports digital scholarly communication in contexts involving Humanities Commons, ACL Anthology, JSTOR, Project MUSE, and research infrastructures like CLARIN and DARIAH. The teiHeader enables interoperability with repositories run by Zenodo, Figshare, and national aggregators such as Europeana.

Creation and Editing Tools

Many tools support building and editing teiHeaders, for example oXygen XML Editor, SIL FieldWorks, Emacs, Notepad++, Sublime Text, Visual Studio Code, and project-specific environments like Tropy, Transkribus, eScriptorium, SCRIPTORIUM, TEI Boilerplate, and the OxGarage converter developed by teams at Max Planck Institute for the History of Science and Oxford University. Collaborative platforms such as those from GitHub, GitLab, Bitbucket, and institutional repositories at Digital Public Library of America integrate with continuous integration systems using Jenkins and Travis CI for automated checks.

Validation and Serialization

Validation of teiHeader content uses XML Schema, RELAX NG, and Schematron rules promulgated by the Text Encoding Initiative and validated by processors like jing, Saxon, and xmllint. Serialization workflows export TEI documents to formats used by PDF, HTML5, RDF, JSON-LD, and IIIF manifests for presentation in viewers such as Mirador and Universal Viewer; integration with linked data platforms like Wikidata and Geonames often requires mapping from teiHeader elements to Europeana Data Model or Schema.org.

Best Practices and Guidelines

Best practice recommendations come from communities and institutions such as Text Encoding Initiative, Oxford Text Archive, Center for Digital Scholarship at Princeton University, British Library, Library of Congress, DARIAH, CLARIN, and project handbooks like those from Perseus Project and TEI Consortium. Guidelines emphasize consistent use of identifiers (e.g., ISBN, ISNI, ORCID), clear rights statements referencing mechanisms like Creative Commons, and provenance capture compatible with PREMIS and archival standards used by National Archives (United Kingdom).

Examples and Case Studies

Case studies demonstrating teiHeader usage include digital editions by Perseus Project, archival digitization at the British Library, manuscript cataloguing at the Vatican Library, and scholarly corpora assembled by Project Gutenberg partners. Large-scale implementations with complex teiHeader metadata appear in initiatives such as EEBO-TCP project, HathiTrust, Gallica, Europeana aggregations, and specialized projects hosted by Max Planck Institute for the History of Science and Getty Research Institute.

Category:Text Encoding Initiative