LLMpediaThe first transparent, open encyclopedia generated by LLMs

BLAST (biological sequence)

Note: This article was automatically generated by a large language model (LLM) from purely parametric knowledge (no retrieval). It may contain inaccuracies or hallucinations. This encyclopedia is part of a research project currently under review.
Article Genealogy
Parent: CP2K Hop 5 terminal

This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.

BLAST (biological sequence)
NameBLAST
CaptionSequence similarity search example
Developed byAltschul, National Center for Biotechnology Information
Initial release1990
Operating systemCross-platform

BLAST (biological sequence) is a family of heuristic algorithms and software for comparing nucleotide or protein sequences to sequence databases and calculating the statistical significance of matches. It is widely used in computational biology, molecular biology, genomics, and bioinformatics for tasks such as annotation, evolutionary inference, and identification of homologs. The method accelerated sequence analysis workflows established during the rise of high-throughput sequencing and genomic projects.

Introduction

BLAST emerged as a practical alternative to exhaustive alignment methods during the expansion of databases such as GenBank, Swiss-Prot, RefSeq, and Protein Data Bank, enabling researchers at institutions like the National Institutes of Health, European Bioinformatics Institute, and Broad Institute to rapidly find regions of local similarity. It complements global alignment tools developed in contexts like work by Needleman–Wunsch algorithm and Smith–Waterman algorithm while integrating statistical frameworks influenced by studies from Karlin–Altschul and collaborations with groups at Columbia University and University of California, Berkeley.

History and development

The algorithm family was introduced in 1990 by a team associated with the National Center for Biotechnology Information and later refined through interactions with projects including Human Genome Project, Mouse Genome Project, and consortia such as Ensembl and International Nucleotide Sequence Database Collaboration. Key developments paralleled computational advances at companies and institutions like IBM, Intel, and Lawrence Berkeley National Laboratory and software paradigms seen in systems from Unix and GNU Project. Evolution of BLAST reflected broader shifts exemplified by initiatives such as 1000 Genomes Project, ENCODE Project, and the rise of cloud platforms like Amazon Web Services for scalable computation.

Algorithm and methodology

BLAST applies a seed-and-extend heuristic: it identifies short exact matches (seeds) between a query and database sequences, extends seeds to high-scoring segment pairs (HSPs), and evaluates significance with statistics developed by scholars including those behind the Karlin–Altschul equation. The approach contrasts with dynamic programming methods used in Smith–Waterman algorithm and Needleman–Wunsch algorithm by trading guaranteed optimality for speed, leveraging substitution matrices such as BLOSUM62 and gap penalties formalized in scoring systems used in resources like PAM matrices and standards from groups including International Union of Biochemistry and Molecular Biology. Implementation details draw from practices at computational centers like Los Alamos National Laboratory and algorithms research from ACM conferences.

Variants and tools (e.g., BLASTn, BLASTp, PSI-BLAST)

The BLAST family includes program variants tailored for specific sequence types and objectives: BLASTn for nucleotide–nucleotide comparisons, BLASTp for protein–protein comparisons, BLASTx for translated nucleotide to protein, tBLASTn and tBLASTx for translated searches, and iterative methods such as PSI-BLAST for profile-based detection of distant homologs. These tools are packaged in suites maintained by organizations like NCBI and incorporated into platforms such as UCSC Genome Browser, Galaxy Project, Bioconductor, and commercial systems from Illumina and Thermo Fisher Scientific.

Applications and use cases

BLAST underpins workflows in gene annotation for projects like RefSeq and Ensembl, pathogen identification in public health efforts by agencies such as Centers for Disease Control and Prevention and World Health Organization, metagenomics studies associated with Human Microbiome Project and Earth Microbiome Project, and phylogenetics used by researchers from institutions like Smithsonian Institution and The Sanger Institute. Clinical and translational use cases appear in diagnostics programs at Mayo Clinic, Johns Hopkins Hospital, and biotech firms including Genentech and Roche. Educational adoption is widespread in curricula at universities such as Harvard University, Massachusetts Institute of Technology, and Stanford University.

Performance, limitations, and improvements

BLAST performance scales with database size and hardware; enhancements include parallelized versions optimized for multicore CPUs, GPUs studied in collaborations with NVIDIA, and cluster deployments using middleware from Open Grid Forum. Limitations include sensitivity loss for remote homologs relative to exhaustive methods, reduced performance on repetitive regions observed in large genomes such as Homo sapiens assemblies, and challenges in handling ultra-large metagenomic datasets from initiatives like Global Ocean Sampling Expedition. Improvements have arrived via heuristics refinement, adoption of profile HMMs from HMMER algorithms, and hybrid pipelines integrating machine learning models developed in contexts like DeepMind and academic labs at University of California, San Diego.

Implementation and web/services

BLAST binaries and source code are distributed by the National Center for Biotechnology Information and integrated into web services hosted by organizations including NCBI, European Bioinformatics Institute, DNA Data Bank of Japan, and commercial providers like Qiagen. Community platforms such as Galaxy Project, Cytoscape, and KBase include BLAST wrappers; workflow systems like Snakemake and Nextflow orchestrate large-scale BLAST analyses on infrastructure providers such as Google Cloud Platform and Microsoft Azure.

References and resources

Key resources include documentation and executables from the National Center for Biotechnology Information, tutorials at training programs run by EMBL-EBI and workshops at conferences like ISMB and RECOMB, and textbooks used in courses at Cold Spring Harbor Laboratory and graduate programs at University of Cambridge. For community standards and database access, consult archives maintained by GenBank, UniProt, RefSeq, and consortium pages for projects like 1000 Genomes Project.

Category:Bioinformatics