This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.
| POMACS | |
|---|---|
| Name | POMACS |
POMACS is a computational toolkit and framework for parallelized observational and molecular analysis with cross-disciplinary interfaces. It integrates high-throughput data processing, visualization, and simulation coupling to support research workflows across astronomy, chemistry, climatology, and bioinformatics. The project emphasizes modularity, distributed execution, and reproducible pipelines for large-scale experiments.
POMACS serves as an orchestrator for data ingestion from instruments and archives such as Hubble Space Telescope, Chandra X-ray Observatory, Kepler, James Webb Space Telescope, Large Hadron Collider, Square Kilometre Array, Human Genome Project, Protein Data Bank, NOAA, and European Space Agency missions, while interfacing with compute backends including Apache Hadoop, Apache Spark, Kubernetes, SLURM, and HTCondor. It supports visualization through integrations with ParaView, VisIt, Blender, Matplotlib, and D3.js. The architecture enables coupling to workflow managers like Cromwell, Nextflow, Snakemake, and Airflow. POMACS targets researchers at institutions such as NASA, European Southern Observatory, CERN, Max Planck Society, and Los Alamos National Laboratory.
POMACS originated from collaborations among research groups at Massachusetts Institute of Technology, Stanford University, University of California, Berkeley, and Princeton University to address scaling issues experienced in projects like Event Horizon Telescope and ENIGMA. Early prototypes borrowed ideas from platforms developed at Argonne National Laboratory and Oak Ridge National Laboratory and were influenced by middleware efforts such as Globus, Condor Project, and MPI. Major milestones include adoption by teams working on LIGO Scientific Collaboration data challenges, deployment in multi-institutional campaigns led by European Southern Observatory and NOAA, and partnerships with industrial labs including IBM Research, Microsoft Research, Google Research, and Amazon Web Services. The project’s roadmap intersected with initiatives like Human Cell Atlas, Square Kilometre Array Organization, and Intergovernmental Panel on Climate Change assessments.
POMACS employs a modular microservice architecture inspired by systems such as Docker Swarm and Service Fabric, with messaging patterns influenced by Apache Kafka, RabbitMQ, and ZeroMQ. Storage integrations support Amazon S3, Google Cloud Storage, OpenStack Swift, Ceph, and archival systems used by European Organisation for the Exploitation of Meteorological Satellites and National Oceanic and Atmospheric Administration. For authentication and identity it integrates with OAuth 2.0, OpenID Connect, and identity providers like Okta and Microsoft Azure Active Directory. The framework supports computational kernels from languages and runtimes including Python (programming language), C++, Fortran, R (programming language), Julia (programming language), and CUDA for GPU acceleration. Interconnects leverage networking technologies such as InfiniBand, RDMA, and cloud virtual networking from Amazon Web Services, Google Cloud Platform, and Microsoft Azure.
POMACS offers pipeline composition, provenance tracking, and metadata catalogs interoperable with standards from International Virtual Observatory Alliance, Dublin Core, and DataCite. It provides adapters for simulation engines like GROMACS, NAMD, LAMMPS, FLASH (software), and ENZO (code), and links to analysis tools including scikit-learn, TensorFlow, PyTorch, XGBoost, and LightGBM. Visualization and interactive analysis integrate with Jupyter Notebook, JupyterLab, RStudio, Plotly, and Bokeh. Data quality and calibration modules draw on algorithms used by teams behind Planck (spacecraft), Gaia (spacecraft), and Sloan Digital Sky Survey. Security and compliance features reference practices from NIST, ISO/IEC 27001, and General Data Protection Regulation implementations in institutional deployments.
Benchmark suites for POMACS have been run on supercomputers and cloud platforms including Summit (supercomputer), Sierra (supercomputer), Frontera (supercomputer), Tianhe-2, and public clouds from Amazon Web Services, Google Cloud Platform, and Microsoft Azure. Performance evaluations compared scalability against systems like Hadoop Distributed File System, CephFS, and Lustre, and parallel processing benchmarks referenced standards from SPEC and HPCG. Reported improvements in throughput and latency were demonstrated on cosmology simulations similar to those produced by teams at Max Planck Institute for Astrophysics and molecular dynamics workloads analogous to Broad Institute research, often measured using profiling tools such as gprof, perf (Linux tool), and Intel VTune.
POMACS has been applied in observational campaigns for facilities like Very Large Telescope, Atacama Large Millimeter Array, and ALMA, and in multi-messenger astronomy projects coordinated with IceCube Neutrino Observatory and LIGO Scientific Collaboration. In climate and Earth science it has been used alongside datasets from IPCC assessments, Copernicus Programme, and USGS remote sensing archives. Bioinformatics and structural biology applications include pipelines for projects such as Human Genome Project, Protein Data Bank, and Cryo-EM facilities affiliated with European Molecular Biology Laboratory and National Institutes of Health. Industry deployments have supported pharmaceutical research at Pfizer, Roche, and Novartis, as well as materials discovery collaborations with Toyota, BASF, and Intel Corporation.
Development of POMACS is driven by contributors from universities, national labs, and corporations including MIT Computer Science and Artificial Intelligence Laboratory, Lawrence Berkeley National Laboratory, Sandia National Laboratories, IBM Research, Microsoft Research, Google Research, Amazon Web Services, and NVIDIA. The community organizes workshops and sessions at conferences like Supercomputing Conference, International Conference for High Performance Computing, Networking, Storage and Analysis, NeurIPS, ICML, AAAI Conference on Artificial Intelligence, AGU Fall Meeting, and American Astronomical Society meetings. Documentation and training have been provided through tutorials at institutions such as Coursera, edX, and university extension programs at Harvard University and Stanford University. Governance models reference practices from open-source projects like Linux Kernel, Apache Software Foundation, and OpenStack Foundation.
Category:Scientific software