LLMpediaThe first transparent, open encyclopedia generated by LLMs

STAR (aligner)

Note: This article was automatically generated by a large language model (LLM) from purely parametric knowledge (no retrieval). It may contain inaccuracies or hallucinations. This encyclopedia is part of a research project currently under review.
Article Genealogy
Parent: GTEx Project Hop 5 terminal

This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.

STAR (aligner)
STAR (aligner)
AI-generated (Stable Diffusion 3.5) · CC BY 4.0 · source
NameSTAR
DeveloperAlexander Dobin et al.
Released2012
Programming languageC++
PlatformLinux, macOS
LicenseGPL

STAR (aligner)

STAR is a fast RNA-seq read aligner designed for mapping high-throughput sequencing reads to reference genomes. Developed to handle spliced alignments for transcriptomics, it emphasizes speed, sensitivity, and support for complex junctions. It is widely used in pipelines for transcriptome analysis, single-cell profiling, and large consortia projects.

Introduction

STAR was introduced by a research group led by Alexander Dobin and quickly adopted by projects and institutions such as the ENCODE Project, GTEx Consortium, The Cancer Genome Atlas, and laboratories in universities like Harvard University, Stanford University, Broad Institute, and University of California, Berkeley. The aligner addresses challenges encountered in projects that require mapping reads across exon–exon junctions, a problem also studied by teams at Wellcome Sanger Institute, Max Planck Society, European Bioinformatics Institute, and Cold Spring Harbor Laboratory. STAR’s original description appeared alongside methodological work from groups including Illumina and research themes investigated in journals associated with Nature Publishing Group and Cold Spring Harbor Laboratory Press.

Algorithm and features

STAR implements a seed-and-extend approach leveraging a Suffix Array/SA and uncompressed suffix array-like tables inspired by bioinformatics methods used in tools from groups such as BLAST-related research and algorithmic advances from labs at MIT, University of Cambridge, and Stanford University. Key features include two-pass mapping for improved junction discovery, support for gapped alignments across splice junctions, handling of chimeric and fusion reads relevant to studies from Memorial Sloan Kettering Cancer Center and Dana-Farber Cancer Institute, and options for soft-clipping and multimapping read reporting that parallel concerns addressed by teams at European Molecular Biology Laboratory and Max Delbrück Center for Molecular Medicine. It integrates ideas comparable to alignment strategies used in projects at National Center for Biotechnology Information and algorithms developed at California Institute of Technology.

Performance and benchmarking

Benchmarking studies comparing STAR to other aligners such as tools originating from Johns Hopkins University and research groups behind aligners in publications from Genome Research and Bioinformatics (journal) show that STAR offers high throughput on multicore servers at institutions like Oak Ridge National Laboratory and Lawrence Berkeley National Laboratory. Comparative analyses often include alternatives developed in labs at Weizmann Institute of Science, University of Washington, and ETH Zurich. Performance evaluations by consortia such as 1000 Genomes Project and groups involved with Human Genome Project-related initiatives highlight STAR’s trade-offs between memory footprint and alignment speed, with many large-scale centers such as European Genome-phenome Archive data processors favoring STAR for speed-sensitive workflows.

Usage and parameters

Common usage patterns for STAR appear in pipelines developed by computational groups at Fred Hutchinson Cancer Research Center, Sanger Institute, and bioinformatics cores at Yale University and University of Michigan. Typical parameters include genome generation indices, read length settings, and annotation-guided options compatible with gene models produced by GENCODE, RefSeq, and projects like UCSC Genome Browser. Users integrate STAR into workflow managers such as Snakemake, Nextflow, and Cromwell for reproducible analyses in environments at European Molecular Biology Laboratory, Argonne National Laboratory, and commercial providers such as Amazon Web Services and Google Cloud Platform.

Output formats

STAR emits standard formats used across bioinformatics infrastructure maintained by groups like National Institutes of Health, including SAM/BAM files compatible with tools from Picard (software), SAMtools, and workflows used by Galaxy (platform). It also produces splice junction files and chimeric output consumed by downstream programs like quantifiers developed by teams at RSEM project and differential expression tools common in labs at Johns Hopkins University School of Medicine and Fred Hutchinson Cancer Research Center.

Applications and case studies

Applications span transcriptome profiling in studies led by centers such as Dana-Farber Cancer Institute and Memorial Sloan Kettering Cancer Center, single-cell RNA-seq protocols used in work from Broad Institute and groups collaborating with 10x Genomics, fusion detection in cancer research at MD Anderson Cancer Center, and population transcriptomics in consortia like GTEx Consortium. Case studies include tumor transcriptome characterization in projects affiliated with The Cancer Genome Atlas, developmental transcriptomics in collaborations with Wellcome Trust-funded institutes, and pathogen transcriptomics in studies conducted at Centers for Disease Control and Prevention and World Health Organization partner labs.

Limitations and future development

Limitations include high memory requirements noted by computational groups at Lawrence Livermore National Laboratory and alignment sensitivity trade-offs discussed in forums associated with International Society for Computational Biology and conferences such as RECOMB and ISMB. Future development priorities cited by communities at European Bioinformatics Institute and academic centers like University of California, San Diego include reduced memory footprint, improved handling of long reads generated by technologies from Oxford Nanopore Technologies and Pacific Biosciences, and tighter integration with transcript quantifiers used in projects at EMBL-EBI and Sanger Institute.

Category:Bioinformatics software