LLMpediaThe first transparent, open encyclopedia generated by LLMs

ERDIS

Note: This article was automatically generated by a large language model (LLM) from purely parametric knowledge (no retrieval). It may contain inaccuracies or hallucinations. This encyclopedia is part of a research project currently under review.
Article Genealogy
Parent: Universities in Italy Hop 5 terminal

This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.

ERDIS
NameERDIS
TypeDistributed information system
DeveloperConsortium of research labs and corporations
First release2019
Latest release2024
Programming languageC++, Rust, Python
Operating systemCross-platform
LicenseMixed proprietary and open-source components

ERDIS

ERDIS is a distributed information and data-integration system designed for large-scale indexing, querying, and federated analytics. It combines technologies from projects and institutions such as Apache Hadoop, Elasticsearch, PostgreSQL, Cassandra (database), Dremio and Google Bigtable to provide high-throughput storage, flexible query planning, and multi-tenant orchestration. ERDIS has been used by research centers, enterprises, and public institutions including collaborations with teams from Massachusetts Institute of Technology, Stanford University, University of California, Berkeley, IBM, Microsoft Research, and Amazon Web Services.

Introduction

ERDIS targets scenarios requiring integrated handling of structured, semi-structured, and unstructured data across heterogeneous sources. Influenced by systems such as Apache Spark, Presto (SQL query engine), Hadoop Distributed File System, TensorFlow, and Apache Kafka, ERDIS implements a layered architecture to decouple storage, compute, and metadata. The project emphasizes interoperability with standards and tools developed at organizations like World Wide Web Consortium, Open Geospatial Consortium, IEEE, and IETF.

History and Development

Development of ERDIS began in a research consortium that included teams from Lawrence Berkeley National Laboratory, Los Alamos National Laboratory, European Organization for Nuclear Research, and private labs such as Google Research and Facebook AI Research. Early prototypes drew on academic work from researchers associated with Carnegie Mellon University, University of Washington, ETH Zurich, and Max Planck Society. The system’s roadmap incorporated lessons from deployments in projects including Human Genome Project, Large Hadron Collider, Square Kilometer Array, and enterprise migrations led by Goldman Sachs and Walmart. Major milestones coincided with conferences like SIGMOD, VLDB, NeurIPS, and ICML, and code contributions were presented at workshops hosted by USENIX and ACM.

Architecture and Design

ERDIS adopts a modular design inspired by architectures from Google File System and MapReduce, while integrating ideas from Lambda architecture and Kappa architecture alternatives used at companies such as Twitter and LinkedIn. Its core layers include a distributed storage layer compatible with Ceph, MinIO, and object stores like Amazon S3; a query and execution engine influenced by Apache Flink and ClickHouse; and a metadata management service interoperable with Apache ZooKeeper and etcd. The system supports pluggable connectors patterned after interfaces used by ODBC and JDBC, and adopts serialization formats like Apache Avro, Apache Parquet, and ORC.

Features and Functionality

ERDIS provides federated query planning, adaptive ingestion pipelines, columnar analytics, and vectorized execution for workloads similar to those found in deployments at Netflix, Uber, Airbnb, and Spotify. It offers machine-learning model serving integrations that mirror patterns from Kubeflow, MLflow, and TensorFlow Serving. The platform supports role-based access modeled on practices from OAuth 2.0 adopters and integrates auditing features used by auditors from firms such as Deloitte and PwC. Additional capabilities include time-series optimization influenced by InfluxDB and Prometheus, geospatial indexing comparable to PostGIS, and graph query support using paradigms from Neo4j and JanusGraph.

Applications and Use Cases

ERDIS has been applied in genomics pipelines akin to workflows in NCBI projects, astrophysics data analysis similar to workflows at European Southern Observatory, financial risk analytics modeled on systems at JPMorgan Chase and BlackRock, and supply-chain telemetry comparable to implementations at Maersk and FedEx. Public-sector deployments reference standards used by agencies like NASA and European Space Agency for satellite telemetry, and health-data integrations that parallel architectures employed by National Institutes of Health initiatives. Startups and enterprises use ERDIS for real-time personalization akin to systems at Google Ads and Facebook Ads, fraud detection inspired by analysts at Mastercard and Visa, and compliance reporting aligned with requirements encountered at SEC and European Banking Authority.

Security and Privacy

Security design in ERDIS follows threat models and controls advocated by NIST, ISO/IEC 27001, and guidance from Center for Internet Security. Encryption at rest and in transit uses algorithms and key-management patterns employed by AWS KMS, HashiCorp Vault, and Azure Key Vault. Authentication and authorization integrate with identity providers comparable to Okta and Microsoft Azure Active Directory, and privacy-preserving computation features draw on techniques discussed in research from Differential privacy literature and implementations similar to projects at OpenMined and Google Privacy Sandbox. ERDIS also supports data masking and tokenization methods used by PCI DSS-compliant systems and audit trails formatted for intake by tools from Splunk and Elastic (company).

Reception and Criticism

ERDIS received praise in community discussions at KDD and IEEE BigData for its interoperability and performance tuning, and was highlighted in case studies involving partners such as Siemens and Bayer. Critics, including voices from Electronic Frontier Foundation-adjacent forums and independent consultants from Gartner and Forrester Research, raised concerns about complexity of configuration, vendor lock-in risks associated with proprietary connectors, and operational challenges similar to those observed in earlier systems like Hadoop. Academic reviewers compared ERDIS to experimental platforms showcased at CNRS and Riken and suggested further work on ease-of-use, deterministic reproducibility, and transparency in optimization heuristics.

Category:Distributed data systems