This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.
| Illumina–PacBio | |
|---|---|
| Name | Illumina–PacBio |
| Type | Hybrid sequencing strategy |
| Industry | Biotechnology |
| Introduced | 2010s |
Illumina–PacBio
Illumina–PacBio is a hybrid sequencing approach combining short-read Illumina sequencing and long-read Pacific Biosciences sequencing to improve genome assembly, structural variant detection, and transcript resolution. The method integrates data from instruments such as the Illumina HiSeq, Illumina NovaSeq, PacBio Sequel, and PacBio Sequel II to leverage high accuracy and read depth from Illumina MiSeq and high contiguity from PacBio RS II, enabling projects across model systems like Escherichia coli, Homo sapiens, Saccharomyces cerevisiae, and Arabidopsis thaliana. Researchers from institutions including the Broad Institute, Wellcome Sanger Institute, Cold Spring Harbor Laboratory, and J. Craig Venter Institute frequently apply hybrid strategies alongside bioinformatics tools developed at groups such as Genome Institute at Washington University, European Bioinformatics Institute, and National Center for Biotechnology Information.
Illumina–PacBio couples the Illumina short-read platforms known for low per-base error rates and high throughput with Pacific Biosciences single-molecule, real-time long reads noted for resolving repetitive and structural regions, an approach adopted in large consortia like the 1000 Genomes Project, Human Genome Project-Write, Earth BioGenome Project, and Genome in a Bottle. The hybrid paradigm supports assemblies used by projects at the Wellcome Trust Sanger Institute, comparative genomics at UC Berkeley, pathogen surveillance at Centers for Disease Control and Prevention and World Health Organization, and agricultural genomics at CIMMYT and CGIAR centers. Funding and standards from organizations such as the National Institutes of Health, European Commission, and Bill & Melinda Gates Foundation have shaped adoption, while standards bodies like the Global Alliance for Genomics and Health influence data sharing.
Typical Illumina–PacBio workflows begin with sample collection from sources like Human Microbiome Project cohorts, clinical samples processed at Mayo Clinic or Johns Hopkins Hospital, or ecological samples from Smithsonian Institution collections, followed by DNA extraction protocols refined at Cold Spring Harbor Laboratory and library prep kits from vendors including Illumina and Pacific Biosciences. Sequencing runs occur on instruments such as Illumina NovaSeq 6000 for short reads and PacBio Sequel IIe for HiFi reads; computational pipelines often employ assemblers and polishers developed by teams at Broad Institute and MIT, integrating software like SPAdes, Canu, Flye, Pilon, Arrow, and error correction tools from National Center for Biotechnology Information groups. Hybrid assembly strategies leverage scaffolding via data types from Hi-C groups like Phase Genomics and optical mapping from Bionano Genomics for chromosome-scale assemblies used in projects coordinated by Darwin Tree of Life Project and Vertebrate Genomes Project.
Illumina–PacBio hybrid sequencing is applied to de novo genome assembly efforts at Wellcome Trust Sanger Institute and Broad Institute, clinical sequencing workflows at Mayo Clinic and Massachusetts General Hospital, cancer genomics studies at Memorial Sloan Kettering Cancer Center and Dana-Farber Cancer Institute, and microbial surveillance at Centers for Disease Control and Prevention and European Centre for Disease Prevention and Control. Agricultural research at CIMMYT and USDA uses hybrids for crop genomes, while conservation genomics by Zoological Society of London and Conservation Genomics Consortium benefits from long-read scaffolding. Functional studies at Stanford University, Harvard University, and University of Cambridge integrate transcript isoform resolution and methylation detection for epigenetics research supported by groups such as ENCODE and Roadmap Epigenomics Project.
Hybrid Illumina–PacBio assemblies typically achieve higher continuity than short-read only approaches used by groups employing SOAPdenovo or earlier Illumina HiSeq pipelines, and greater base-level accuracy after polishing compared with raw long-read assemblies from early PacBio RS II runs. Benchmarks published by teams at Broad Institute, Wellcome Sanger Institute, and Genome Institute at Washington University show improvements in contig N50 and reduction of misassemblies compared with single-technology assemblies. Comparative evaluations involving tools from European Bioinformatics Institute and metrics set by Genome in a Bottle demonstrate strengths in resolving tandem repeats, segmental duplications, and structural variants detected alongside callers developed at Broad Institute and UCSD.
Challenges for Illumina–PacBio workflows include sample quantity and quality constraints noted by protocols at Cold Spring Harbor Laboratory and Wellcome Trust Sanger Institute, cost and throughput considerations debated at National Institutes of Health grant panels, and computational resource demands highlighted by bioinformatics groups at MIT and Carnegie Mellon University. Data integration issues arise when combining datasets from vendors such as Illumina and Pacific Biosciences and when harmonizing analyses across consortia like Earth BioGenome Project and Vertebrate Genomes Project. Regulatory and data-sharing complexities intersect with policies from European Commission and U.S. Food and Drug Administration for clinical applications deployed at Johns Hopkins Hospital and Mayo Clinic.
Hybrid strategies emerged in the 2010s as long-read platforms from Pacific Biosciences matured alongside high-throughput short reads from Illumina, driven by collaborative efforts at institutions including Broad Institute, Wellcome Trust Sanger Institute, and J. Craig Venter Institute. Early demonstrations combining technologies were published by consortia such as 1000 Genomes Project and groups at Washington University in St. Louis and University of California, Santa Cruz, with software innovations from teams at Broad Institute, European Bioinformatics Institute, and UC Berkeley enabling practical adoption. Investments by funders like the National Institutes of Health and philanthropic initiatives from Bill & Melinda Gates Foundation accelerated applications in public health and agriculture.
Future directions include tighter integration with ultra-long reads from platforms developed by Oxford Nanopore Technologies teams, increased use of high-fidelity HiFi reads from Pacific Biosciences Sequel IIe combined with ultra-deep short reads from Illumina NovaSeq 6000 for population genomics by Human Genome Project-Write collaborators, and adoption in large-scale initiatives like the Earth BioGenome Project and Darwin Tree of Life Project. Innovations from computational groups at Broad Institute, European Bioinformatics Institute, and Stanford University aim to reduce costs and computational demands while improving haplotype phasing used in studies at Wellcome Sanger Institute and Broad Institute. Regulatory frameworks from U.S. Food and Drug Administration and data standards from Global Alliance for Genomics and Health will influence clinical translations at Mayo Clinic and Massachusetts General Hospital.
Category:Sequencing methods