LLMpediaThe first transparent, open encyclopedia generated by LLMs

FASTQ

Note: This article was automatically generated by a large language model (LLM) from purely parametric knowledge (no retrieval). It may contain inaccuracies or hallucinations. This encyclopedia is part of a research project currently under review.
Article Genealogy
Parent: BAM Hop 5 terminal

This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.

FASTQ
NameFASTQ
Extension.fastq, .fq
OwnerInternational Nucleotide Sequence Database Collaboration
GenreBioinformatics, sequencing data
Released2009 (de facto standardized)

FASTQ

FASTQ is a text-based file format used to represent nucleotide sequences and their corresponding quality scores for high-throughput sequencing data. It serves as a core interchange format between sequencing instruments from companies such as Illumina, Pacific Biosciences, and Oxford Nanopore Technologies, and analysis software developed by groups at Broad Institute, European Bioinformatics Institute, and National Center for Biotechnology Information. FASTQ underpins pipelines built with tools from Genome Analysis Toolkit, SAMtools, BBMap, and FastQC, and is integral to repositories like Sequence Read Archive and projects such as the 1000 Genomes Project and ENCODE Project.

Introduction

FASTQ combines sequence data with per-base quality information in a compact four-line record, enabling downstream workflows in genomics studies by laboratories at institutions including Wellcome Sanger Institute, Broad Institute, University of California, Berkeley and consortia like the International Human Genome Sequencing Consortium. It interfaces with standards and formats such as SAM/BAM/CRAM and enables analyses in software ecosystems exemplified by Bioconductor, Galaxy (platform), and Nextflow. FASTQ is essential for research performed at hospitals and research centers like Mayo Clinic, Johns Hopkins University, Massachusetts General Hospital and initiatives such as The Cancer Genome Atlas.

Format and Structure

A FASTQ record consists of four lines: an identifier line beginning with '@', a nucleotide sequence line, a '+' separator line (optionally repeating the identifier), and a quality score line using ASCII-encoded symbols. The format is used by sequencing platforms like Illumina HiSeq, Illumina NovaSeq, Roche 454, ABI SOLiD, and Ion Torrent instruments, and is parsed by software libraries such as HTSlib, BioPython, SeqIO and BioPerl. Typical workflows convert FASTQ into SAM format via alignment tools such as BWA, Bowtie 2, STAR (aligner), and HISAT2 for downstream analyses at centers like European Molecular Biology Laboratory and Cold Spring Harbor Laboratory.

Quality Scores and Encoding

Quality scores in FASTQ encode base-calling confidence using Phred-like scales originating from projects like the Phred (base-calling program), and have been encoded historically in variants such as Sanger, Illumina 1.3+, and Illumina 1.8+ encodings. Discussions of encoding impact pipelines used by GATK, FreeBayes, Platypus (variant caller), and variant databases like dbSNP and ClinVar influence best practices at organizations including National Institutes of Health and European Genome-phenome Archive. Tools including seqtk, Prinseq, and Trim Galore! adjust quality trimming based on Phred scores, and visualization tools such as FastQC and plotting packages in R (programming language) and Python (programming language) are commonly applied by groups at Stanford University and Massachusetts Institute of Technology.

Variants and Extensions

Multiple variants and extensions of FASTQ support paired-end reads, color-space encodings from ABI SOLiD, long reads from Pacific Biosciences and Oxford Nanopore Technologies, and auxiliary annotations used by consortia like 1000 Genomes Project and Genome in a Bottle. Compressed alternatives and containers such as gzip, bzip2, CRAM format, and specialized compressors like fqzcomp and DSRC are used in projects at European Nucleotide Archive and National Center for Biotechnology Information to reduce storage footprints. Metadata-rich formats and standards such as MIxS and integration with provenance systems like PROV (W3C) are applied in large-scale facilities including European Genome-phenome Archive and national infrastructures like ELIXIR.

Tools and Software Support

A broad ecosystem supports FASTQ: command-line utilities (e.g., cutadapt, Trimmomatic, fastp), alignment tools (BWA, Bowtie 2, Minimap2), quality control tools (FastQC, MultiQC), and conversion/manipulation libraries (HTSlib, BioPython, BioPerl, SeqKit). Workflow managers such as Snakemake, Nextflow, and Cromwell orchestrate FASTQ-centric pipelines in cloud environments provided by Amazon Web Services, Google Cloud Platform, and Microsoft Azure used by institutes like Broad Institute and EMBL-EBI. Visualization and analysis platforms including IGV, UCSC Genome Browser and Ensembl consume downstream products derived from FASTQ-originated alignments.

Applications and Data Management

FASTQ data feed applications spanning whole-genome sequencing, exome sequencing, RNA-Seq, ChIP-Seq, metagenomics, and single-cell sequencing used in studies by ENCODE Project, Human Microbiome Project, Earth Microbiome Project, and clinical programs at Mayo Clinic and Dana-Farber Cancer Institute. Data management strategies leverage archival repositories such as Sequence Read Archive, European Nucleotide Archive, and systems at National Center for Biotechnology Information and EMBL-EBI, often integrating access control and consent frameworks like those from dbGaP and GA4GH. Large consortia including TCGA, ICGC, and All of Us Research Program rely on standardized FASTQ workflows to ensure reproducibility across sites such as Wellcome Sanger Institute and Broad Institute.

History and Development

FASTQ arose from practical needs in the early 2000s as sequencing throughput increased with instruments from Solexa (later acquired by Illumina) and the development of base-callers like Phred (base-calling program). Standardization and community discussion involved organizations including EMBL-EBI, NCBI, and developer communities around BioPerl and BioPython, with influential projects such as 1000 Genomes Project and ENCODE Project driving best practices. Continued evolution has been shaped by technology advances at Illumina, Pacific Biosciences, and Oxford Nanopore Technologies, and by software projects like HTSlib, SAMtools, and workflow initiatives such as GA4GH standards and ELIXIR infrastructure.

Category:Bioinformatics file formats