This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.
| GeneMark | |
|---|---|
| Name | GeneMark |
| Author | Mark Borodovsky |
| Developer | Georgia Institute of Technology; Softberry; various research groups |
| Released | 1993 |
| Programming language | C, C++ |
| Operating system | Unix, Linux, Windows |
| License | academic, commercial |
GeneMark GeneMark is a family of computational tools for prokaryotic and eukaryotic gene prediction developed to identify protein-coding regions in DNA sequences. It integrates probabilistic models, species-specific parameterization, and heuristic strategies to annotate genomic sequences for research groups, biotechnology firms, and public databases. The software has influenced genome projects, comparative analyses, and annotation pipelines across institutions such as National Institutes of Health, European Molecular Biology Laboratory, and Wellcome Trust.
GeneMark implements statistical methods to detect coding potential in nucleotide sequences, designed for organisms ranging from bacteria to vertebrates and viruses. It has been applied in projects at Broad Institute, Sanger Institute, Cold Spring Harbor Laboratory, European Bioinformatics Institute, and National Center for Biotechnology Information for draft and finished genome annotation. The package supports single-genome prediction, metagenomic fragments, and models trained on trusted annotations from resources like GenBank, RefSeq, UniProt, Ensembl, and DDBJ.
GeneMark relies on inhomogeneous Markov chains and hidden Markov models (HMMs) to represent coding and non-coding regions, with model orders chosen to capture nucleotide dependencies found in organisms such as Escherichia coli, Saccharomyces cerevisiae, Arabidopsis thaliana, Drosophila melanogaster, and Homo sapiens. Training employs unsupervised and supervised strategies drawing on annotated data from repositories like Swiss-Prot and curated sets from GenBank submitters, using maximum likelihood and Bayesian approaches influenced by work at Stanford University and Massachusetts Institute of Technology. GeneMark incorporates translation initiation site models, codon usage bias profiles, and models for intron/exon structure inspired by statistical frameworks used at Johns Hopkins University and University of California, Berkeley.
Multiple variants address diverse needs: a prokaryotic version used in projects at Joint Genome Institute, a eukaryotic version applied by Broad Institute teams, a metagenomic fragment predictor adopted by JGI and MG-RAST pipelines, and a viral/taxonomic-aware iteration used in surveillance by Centers for Disease Control and Prevention and World Health Organization laboratories. Implementations exist in command-line C/C++ form, web services hosted by entities such as SoftBerry and academic servers at Georgia Institute of Technology, and integrations into annotation platforms like Apollo (genome annotation tool), Galaxy (platform), and workflow systems developed at European Molecular Biology Laboratory. Ports and wrappers have been created for environments at Argonne National Laboratory and Lawrence Berkeley National Laboratory.
GeneMark has been used to annotate microbial genomes sequenced by consortia including Human Microbiome Project, 1000 Genomes Project, Human Genome Project, and environmental surveys by Tara Oceans. It supports bacterial genome publishing in journals from Nature, Science, and PLoS Biology and is used in industrial pipelines at companies like Illumina and Roche Diagnostics. Researchers at institutions such as MIT, Caltech, and University of Cambridge apply GeneMark output in comparative genomics, gene family analyses with tools from Pfam and InterPro, and pathway reconstructions using databases like KEGG and Reactome.
Benchmarks compare GeneMark to predictors such as Glimmer, Prodigal, AUGUSTUS, SNAP, and GeneID across reference sets from RefSeq, Ensembl Genomes, and curated bacterial collections at NCBI RefSeq Bacteria. Studies published in venues like Genome Research, Nucleic Acids Research, and Bioinformatics evaluate sensitivity, specificity, and precision on genomes including Mycobacterium tuberculosis, Bacillus subtilis, Caenorhabditis elegans, and Zea mays. Performance depends on model training, sequence composition, and presence of atypical genomic islands documented in analyses at European Nucleotide Archive and surveillance reports by Centers for Disease Control and Prevention.
GeneMark faces challenges when annotating sequences with horizontal gene transfer from taxa such as Bacteroides fragilis or Streptococcus pneumoniae, high GC content genomes like Streptomyces coelicolor, and fragmented metagenomic contigs from studies by MG-RAST and EMBL-EBI. Predicting small open reading frames discovered in research at Harvard University and non-canonical translation initiation sites reported by groups at Max Planck Institute can reduce accuracy. Integration with RNA-seq-based evidence from platforms developed at Illumina and single-molecule sequencing from Pacific Biosciences and Oxford Nanopore Technologies poses opportunities and computational challenges highlighted in comparative work from European Molecular Biology Laboratory teams.
Conceived and led by Mark Borodovsky with collaborators at Georgia Institute of Technology and influenced by statistical genetics work at Carnegie Mellon University and University of Pennsylvania, GeneMark originated in the early 1990s and evolved through contributions from groups at SoftBerry, National Institutes of Health, and European partners. Iterative releases responded to genome projects at Sanger Institute, adoption in microbial genomics at Joint Genome Institute, and demands from large-scale initiatives like Human Microbiome Project and 1000 Genomes Project. Academic dissemination occurred via publications in Nucleic Acids Research, Genome Research, and presentations at conferences hosted by International Society for Computational Biology and Cold Spring Harbor Laboratory.
Category:Bioinformatics software