LLMpediaThe first transparent, open encyclopedia generated by LLMs

WES

Note: This article was automatically generated by a large language model (LLM) from purely parametric knowledge (no retrieval). It may contain inaccuracies or hallucinations. This encyclopedia is part of a research project currently under review.
Article Genealogy
Parent: Diploma Supplement Hop 5 terminal

This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.

WES
NameWES
FieldsGenetics, Genomics, Molecular biology
InstitutionsBroad Institute, Wellcome Sanger Institute, National Institutes of Health
Notable worksThe 100,000 Genomes Project, ClinVar, Exome Aggregation Consortium

WES

WES is the systematic high-throughput sequencing and analysis of the protein-coding regions of the human genome, rapidly adopted across NHGRI, NIH Clinical Center, Mayo Clinic, Johns Hopkins Hospital and in consortia such as The Cancer Genome Atlas and International Cancer Genome Consortium. It complements whole-genome efforts like projects at the Broad Institute and the Wellcome Sanger Institute by focusing on exons to identify variants relevant to BRCA1, CFTR, TP53, APOE and other clinically important genes. Clinical laboratories accredited by College of American Pathologists and regulated under Clinical Laboratory Improvement Amendments implement WES for diagnostics, research hospitals use it alongside databases such as ClinVar and gnomAD for variant interpretation.

Overview

WES targets approximately 1–2% of the genome encompassing coding exons annotated in resources like RefSeq, Ensembl, UCSC Genome Browser, and captures variants including single-nucleotide variants, small insertions and deletions, and splice-site changes affecting genes such as HBB, FBN1, MECP2, DMD, SCN1A. Library preparation pipelines often reference kits from vendors competing with Agilent Technologies, Illumina, Twist Bioscience, and use instruments such as Illumina NovaSeq, HiSeq, Thermo Fisher Ion Proton; bioinformatics workflows leverage tools from GATK, BWA-MEM, SAMtools, Picard and annotation frameworks like VEP, ANNOVAR and databases including dbSNP, ClinVar, OMIM.

History

Early targeted exon sequencing builds on methods from groups at Sanger Institute and laboratories affiliated with Cold Spring Harbor Laboratory and Cambridge University; key advances occurred after publications from Allan Bradley-era projects and milestones like the launch of the Exome Aggregation Consortium and the inclusion of exome assays in The 100,000 Genomes Project. Development of capture hybridization techniques was influenced by work at Agilent Technologies and Roche NimbleGen; adoption accelerated with decreases in cost reported by analysts at NHGRI and investment by clinical centers such as Mayo Clinic and university hospitals including Massachusetts General Hospital.

Applications

WES is used in rare disease diagnosis at centers like Boston Children's Hospital and Great Ormond Street Hospital, cancer research in collaborations with Dana-Farber Cancer Institute and MD Anderson Cancer Center, pharmacogenomics studies referencing alleles in CYP2D6 and CYP2C19 through initiatives at FDA and CPIC, and population genetics analyses drawing on cohorts from UK Biobank, 1000 Genomes Project, and All of Us Research Program. It supports newborn screening pilots in partnership with March of Dimes and carrier screening programs offered by commercial laboratories such as Invitae and Myriad Genetics.

Methodologies

Sample processing follows protocols standardized by bodies like CLSI and often employs exon capture using platforms developed by Agilent, Roche, Illumina, followed by sequencing on instruments from Illumina or Thermo Fisher. Read alignment and variant calling pipelines use BWA, GATK Best Practices, joint genotyping strategies developed in projects like The Cancer Genome Atlas, and quality-control metrics informed by standards from Genome in a Bottle and benchmarking by NIST. Annotation and interpretation integrate resources including ClinVar, OMIM, HGMD and evidence frameworks from ACMG and AMP for variant classification.

Performance and Limitations

WES offers high sensitivity for coding single-nucleotide variants in well-captured exons demonstrated in comparisons with whole-genome sequencing performed at Broad Institute and benchmarking studies from NIST; however, it has limitations in detecting structural variants, repeat expansions implicated in Huntington's disease and Fragile X syndrome, and deep intronic or regulatory variants uncovered by whole-genome approaches used in projects at Wellcome Sanger Institute. Capture bias affects genes with high GC content such as MUC genes and pseudogene interference complicates interpretation for loci like PMS2 and CYP2D6; copy-number variant detection requires specialized algorithms validated against arrays from Affymetrix and Illumina.

Clinical and research deployment raises consent and return-of-results issues debated in forums convened by NIH, HHS, World Health Organization, and advocacy groups including Genetic Alliance and Global Alliance for Genomics and Health. Incidental findings policies often reference the ACMG recommendations; privacy concerns intersect with regulations such as HIPAA and with data-sharing platforms like dbGaP and EGA managed by EBI and NCBI. Equity challenges involve underrepresentation of populations from regions associated with H3Africa, All of Us, and indigenous communities, prompting initiatives at Wellcome Trust and Bill & Melinda Gates Foundation to promote inclusion.

Notable Projects and Initiatives

Major efforts employing exome strategies include The 100,000 Genomes Project, Exome Aggregation Consortium (ExAC), Deciphering Developmental Disorders Study, clinical networks at Genomics England, diagnostic programs at Mayo Clinic, and research consortia such as Consortium on Mendelian Genomics and Undiagnosed Diseases Network. International collaborations include contributions to databases like gnomAD, variant curation collaborations with ClinGen, and large-scale sequencing in cohorts managed by UK Biobank, All of Us, and disease-focused networks at American College of Medical Genetics and Genomics.

Category:Genomics