LLMpediaThe first transparent, open encyclopedia generated by LLMs

Broad Institute Single Cell Portal

⚠Note: This article was automatically generated by a large language model (LLM) from purely parametric knowledge (no retrieval). It may contain inaccuracies or hallucinations. This encyclopedia is part of a research project currently under review.
Article Genealogy

This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.

Broad Institute Single Cell Portal
NameBroad Institute Single Cell Portal
Established2018
LocationCambridge, Massachusetts
Operating organizationBroad Institute
Typebiomedical research

Broad Institute Single Cell Portal The Single Cell Portal is a data-sharing and visualization platform developed by the Broad Institute to host, explore, and disseminate high-dimensional single-cell RNA sequencing datasets produced by consortia, academic laboratories, and industry partners. The Portal integrates tools for interactive visualization, metadata management, and controlled access to support collaborative projects spanning institutions such as Massachusetts Institute of Technology, Harvard University, and multi-center initiatives including the Human Cell Atlas, BRAIN Initiative, and disease-focused consortia. It connects researchers, clinicians, and computational biologists with resources from major funders like the National Institutes of Health and the Wellcome Trust.

Overview

The Portal provides an online environment where investigators can upload processed single-cell matrices and associated metadata from experiments conducted across platforms such as 10x Genomics, Fluidigm C1, and Drop-seq. Users interact with integrated viewers for dimensionality reduction methods including t-distributed stochastic neighbor embedding and uniform manifold approximation and projection, alongside gene-expression violin plots, heatmaps, and cluster annotations commonly used in publications from groups at Broad Institute collaborators like Dana-Farber Cancer Institute and Massachusetts General Hospital. The service supports data citation practices aligned with standards from organizations such as Digital Object Identifier agencies and repository norms championed by European Bioinformatics Institute and National Center for Biotechnology Information.

History and Development

The Portal emerged from efforts within the Broad Institute and affiliated labs to standardize sharing for single-cell experiments after rapid adoption of droplet-based technologies first popularized by teams at institutions like Harvard Medical School and companies such as 10x Genomics. Early development aligned with community initiatives including the Human Cell Atlas roadmap and leveraged software patterns from projects at European Genome-phenome Archive and Gene Expression Omnibus. Funding and collaborative governance involved stakeholders including National Human Genome Research Institute, philanthropic partners like the Chan Zuckerberg Initiative, and clinical partners such as Brigham and Women's Hospital. Over successive releases the Portal incorporated features and contributions from open-source projects maintained by organizations like Broad Institute sibling groups and university labs at Stanford University and University of California, San Francisco.

Features and Functionality

Key features include an interactive study viewer for cluster navigation, gene-expression queries, and cell-level metadata filtering used by researchers at Cold Spring Harbor Laboratory and Salk Institute. Visualization modules support embedding viewers for t-SNE and UMAP, expression violin plots, dot plots, and spatial overlays that accommodate data from spatial transcriptomics approaches developed at 10x Genomics and NanoString Technologies. The Portal integrates access controls and embargo options to coordinate preprint workflows with journals like Nature, Cell, and Science, and supports dataset DOIs for citation in repositories such as Zenodo and data portals run by the European Molecular Biology Laboratory. User roles mirror models used by consortia such as GTEx Project and The Cancer Genome Atlas for tiered access.

Data Content and Submission

Datasets span tissue and disease areas represented by contributors from Stanford Medicine, Yale School of Medicine, and international centers like Wellcome Sanger Institute and Max Planck Society laboratories. Submitters upload processed count matrices, cell metadata, and cluster labels consistent with standards advocated by the Human Cell Atlas metadata working group and the Global Alliance for Genomics and Health. The submission pipeline accepts annotations for cell type nomenclature used by resources such as Cell Ontology and integrates donor-level metadata aligned with policies from National Institutes of Health. Data owners can designate public release, controlled access coordinated with institutional review boards at institutions like Harvard Medical School and data use committees modeled on dbGaP workflows.

Software Architecture and Technology

The Portal is built on web frameworks and scalable components comparable to platforms developed by groups at Broad Institute and relies on technologies including Python (programming language), R (programming language), and JavaScript libraries such as React (JavaScript library) and D3.js. Backend services integrate scalable storage solutions used by cloud providers like Amazon Web Services and object stores similar to those operated by Google Cloud Platform, while authentication and access control mirror federated models adopted by InCommon and OAuth 2.0-based systems. The Portal leverages containerization concepts popularized by Docker (software) and orchestration practices from Kubernetes clusters common to computational biology deployments at institutions such as University of California, Berkeley.

Access, Licensing, and Privacy

Public datasets are freely browsable, enabling reuse consistent with licensing schemes used by projects at European Bioinformatics Institute and data-sharing policies of funders like the National Institutes of Health and Wellcome Trust. Controlled-access datasets follow data-use agreements and privacy safeguards analogous to procedures established by dbGaP and institutional review boards at major hospitals. Authentication options support institutional single sign-on patterns used across universities such as Massachusetts Institute of Technology and research consortia identity federations, while licensing choices allow submitters to apply terms like those from Creative Commons or bespoke consortium agreements.

Impact and Use Cases

The Portal has accelerated collaborative research in single-cell biology across studies from cancer centers including Memorial Sloan Kettering Cancer Center, developmental biology groups at Max Planck Institute for Developmental Biology, and neuroscience consortia linked to the BRAIN Initiative. Investigators use Portal-hosted datasets to reproduce analyses published in journals such as Nature Methods, integrate multi-study atlases comparable to efforts at the Human Cell Atlas consortium, and to build machine-learning models in labs at Stanford University and MIT. The platform supports education and training workflows employed in courses at institutions like Harvard University and workshops organized by organizations such as the International Society for Computational Biology.

Category:Biological databases