This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.
| SAM (Sequence Alignment/Map) format | |
|---|---|
| Name | SAM (Sequence Alignment/Map) format |
| Introduced | 2009 |
| Creators | Li Heng, Broad Institute |
| File extension | .sam, .bam |
| Latest version | 1.6 |
SAM (Sequence Alignment/Map) format SAM (Sequence Alignment/Map) format is a plain-text, tab-delimited format for storing biological sequence alignments to a reference, designed for high-throughput sequencing workflows. It is widely adopted across bioinformatics projects and integrated into pipelines used by institutions such as the Broad Institute, Wellcome Sanger Institute, European Bioinformatics Institute, National Institutes of Health, and consortia like the 1000 Genomes Project, ENCODE Project, and The Cancer Genome Atlas. The format facilitates interoperability between aligners, variant callers, and visualization tools developed by groups including Genome Research Ltd., Illumina, Pacific Biosciences, Oxford Nanopore Technologies, and research labs at universities such as Harvard University, Stanford University, and Massachusetts Institute of Technology.
SAM was introduced to standardize representation of mapped short reads from platforms such as Illumina HiSeq, Applied Biosystems SOLiD, and later long-read platforms from Pacific Biosciences and Oxford Nanopore Technologies. It supports large-scale projects undertaken by entities like the Human Genome Project, 1000 Genomes Project, and clinical efforts at institutions such as Mayo Clinic and Johns Hopkins University. The design was motivated by the need for compatibility with downstream tools from groups like the Broad Institute and databases hosted by European Bioinformatics Institute. SAM serves as the textual counterpart to the binary BAM representation used by utilities associated with SAMtools and related software developed at the European Molecular Biology Laboratory and the National Center for Biotechnology Information.
A SAM file consists of a header section and an alignment section; each alignment row contains mandatory fields including QNAME, FLAG, RNAME, POS, MAPQ, CIGAR, RNEXT, PNEXT, TLEN, and SEQ and QUAL columns used by aligners like BWA, Bowtie, STAR, HISAT2, and Minimap2. Implementations used by groups such as University of California, Santa Cruz and projects like the Human Cell Atlas rely on MAPQ and CIGAR semantics for variant calling with tools such as GATK from the Broad Institute, FreeBayes from BayesHammer contributors, and VarScan implementations. FLAG bitwise encoding aligns with conventions used in software from organizations including Picard at the Broad Institute and utilities maintained by GitHub repositories from labs at European Bioinformatics Institute and Wellcome Sanger Institute.
Header lines beginning with '@' provide metadata tags (e.g., @HD, @SQ, @RG, @PG, @CO) used to record reference sequences and read-group provenance required by consortia such as ENCODE Project and clinical pipelines at Genomics England. The @SQ lines reference contigs often named according to assemblies from projects like GRCh38 created by the Genome Reference Consortium, GRCh37 archives at the National Center for Biotechnology Information, or specialized references used in studies from Broad Institute collaborators. @RG and @PG metadata support provenance tracking across workflows using workflow managers created by groups such as Galaxy (project), Nextflow, and Snakemake in institutional environments including European Bioinformatics Institute and Wellcome Sanger Institute.
SAM permits optional TAG:TYPE:VALUE fields to encode platform-specific or analysis-specific attributes, enabling integration with variant annotation tools from Ensembl, UCSC Genome Browser, and clinical annotation systems at institutions like Broad Institute and Stanford University School of Medicine. Common tags such as NM, MD, AS, and XS are used by aligners including BWA-MEM and Bowtie2 while allowing custom tags adopted in projects at Cold Spring Harbor Laboratory and by consortia like GTEx. This extensibility facilitates downstream processing by toolchains from groups such as Picard, SAMtools, GATK, and visualization in platforms like Integrative Genomics Viewer developed at Broad Institute and collaborators from Wadsworth Center.
To optimize storage, SAM is commonly converted to its binary counterpart BAM and further to CRAM for reference-based compression; conversions are supported by utilities from SAMtools originally developed at the Wellcome Trust Sanger Institute and the European Bioinformatics Institute. CRAM compression leverages reference sequences curated by the Genome Reference Consortium and encoded in archives maintained by NCBI and Ensembl. Cloud-scale projects run by Amazon Web Services, Google Cloud Platform, and Microsoft Azure integrate BAM/CRAM workflows for large datasets such as those produced by the 1000 Genomes Project, UK Biobank, and clinical genomics programs at Mayo Clinic.
A broad ecosystem supports SAM/BAM/CRAM including SAMtools, Picard, GATK, Bedtools, BCFtools, IGV from Broad Institute, aligners like BWA, Bowtie2, Minimap2, and platforms such as Galaxy (project), Nextflow, Snakemake, and commercial offerings from Illumina and Thermo Fisher Scientific. These tools are used in pipelines at institutions including Broad Institute, Wellcome Sanger Institute, European Bioinformatics Institute, Harvard Medical School, and projects like ENCODE Project and The Cancer Genome Atlas for variant discovery, RNA-seq processing, and structural variant detection.
SAM's plain-text nature leads to large file sizes and potential portability issues; hence BAM and CRAM are recommended for storage and transmission in infrastructures run by Amazon Web Services and research centers such as European Bioinformatics Institute. Ambiguities in optional TAG semantics and differing MAPQ interpretations across aligners like BWA-MEM and Bowtie2 have caused interoperability challenges noted in reviews by groups such as Broad Institute and research labs at Stanford University. Reference-dependent CRAM workflows require careful management of assemblies like GRCh38 and GRCh37 to ensure reproducible analyses in consortia such as 1000 Genomes Project and clinical programs at Genomics England.