LLMpediaThe first transparent, open encyclopedia generated by LLMs

Graph 500

Note: This article was automatically generated by a large language model (LLM) from purely parametric knowledge (no retrieval). It may contain inaccuracies or hallucinations. This encyclopedia is part of a research project currently under review.
Article Genealogy

This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.

Graph 500
NameGraph 500
Established2010
DisciplineHigh-performance computing benchmarks
OrganizerLawrence Berkeley National Laboratory; Oak Ridge National Laboratory; National Energy Research Scientific Computing Center
CountryInternational

Graph 500 is an international benchmark initiative for evaluating high-performance computing systems on data-intensive graph problems, developed by researchers at Lawrence Berkeley National Laboratory, Oak Ridge National Laboratory, National Energy Research Scientific Computing Center, and partners including IBM, Intel Corporation, and Cray Inc.. It complements traditional benchmarks such as the TOP500 list and seeks to measure throughput and scalability for graph algorithms used in applications from cybersecurity to bioinformatics and social network analysis. The project involves academic teams, national laboratories, and commercial vendors who submit results for community-driven comparison at conferences like the International Supercomputing Conference and workshops associated with the SC Conference.

Overview

Graph 500 was announced in 2010 by teams at Lawrence Berkeley National Laboratory and Oak Ridge National Laboratory with the intent to focus attention on irregular memory access and communication patterns not captured by the LINPACK benchmark used for the TOP500. The initiative is governed by an organizing committee with contributors from Sandia National Laboratories, Los Alamos National Laboratory, Argonne National Laboratory, National Institute of Standards and Technology, and industry partners such as NVIDIA and AMD. It is presented at venues including the IEEE International Parallel and Distributed Processing Symposium, ACM Symposium on High-Performance Parallel and Distributed Computing, and associated meetings of the High Performance Computing community.

Benchmark Specifications

The Graph 500 specification defines problem sizes, input graph generators, and kernel definitions referenced to implementations from groups like University of Tennessee, University of California, Berkeley, and Georgia Institute of Technology. The benchmark prescribes a Kronecker graph generator inspired by work from Alexei Krakovski and methodologies related to Stochastic Kronecker Graphs used in network science research associated with Leskovec Lab and Stanford University. Testbeds and submission rules reflect testing practices familiar to participants from Lawrence Livermore National Laboratory and European Centre for Medium-Range Weather Forecasts. The specification is versioned and maintained with contributions from researchers affiliated with University of California, Davis, Princeton University, University of Illinois Urbana-Champaign, and ETH Zurich.

Performance Metrics and Scoring

Scoring in Graph 500 is measured primarily in traversed edges per second (TEPS), a metric comparable in role to GFLOPS in LINPACK but focused on BFS-style workloads inspired by algorithmic research from ULAM Project and graph algorithm studies at Massachusetts Institute of Technology. Results are reported for defined problem sizes (e.g., "scale" parameters) with additional metrics such as energy efficiency measured in TEPS per watt, an approach used in analyses at Oak Ridge National Laboratory and Argonne National Laboratory. Comparative scoring often appears alongside rankings from Green500 and TOP500 in analyses published by InsideHPC and presented at the International Supercomputing Conference. The scoring methodology has been discussed in workshops held at IEEE and summits hosted by U.S. Department of Energy laboratories.

Implementation and Reference Kernels

The benchmark provides reference kernels for breadth-first search (BFS) and other graph primitives; reference implementations have been produced by researchers at Sandia National Laboratories, Los Alamos National Laboratory, University of California, Santa Barbara, and industry teams from IBM Research and Intel Labs. Implementations exploit technologies such as MPI standards from the Message Passing Interface Forum, shared-memory approaches using OpenMP from OpenMP Architecture Review Board, and accelerator programming models from NVIDIA (CUDA) and OpenCL by the Khronos Group. Community-maintained repositories and test harnesses have contributors from GitHub organizations tied to Lawrence Berkeley National Laboratory and International Exascale Software Project participants.

Historical Results and Records

Since its inception, notable records have been set on systems including Fugaku-class prototypes, Summit at Oak Ridge National Laboratory, Sierra at Lawrence Livermore National Laboratory, and national supercomputers at Lawrence Berkeley National Laboratory and Argonne National Laboratory. Landmark results have been reported by vendors like Cray Inc., Hewlett Packard Enterprise, and IBM and discussed in venues such as the SC Conference and International Supercomputing Conference. Historic entries often correlate with shifts in architecture, for example the adoption of GPUs from NVIDIA or many-core CPUs from Intel Corporation and AMD, and with software optimizations from groups at University of California, Berkeley and Massachusetts Institute of Technology.

Impact and Applications

Graph 500 has influenced system procurement, architecture research, and algorithm development across institutions such as Los Alamos National Laboratory, Sandia National Laboratories, Argonne National Laboratory, and industrial research centers like IBM Research and Intel Labs. Its emphasis on irregular workloads has affected designs in exascale systems funded by U.S. Department of Energy programs and informed software stacks used by projects at CERN, Human Genome Project-related centers, and social media analytics teams at Facebook-affiliated research labs. The benchmark has been cited in academic work from Stanford University, Carnegie Mellon University, University of Michigan, and University of Texas at Austin.

Criticisms and Limitations

Critics from academic groups at University of California, Berkeley, Princeton University, and ETH Zurich have noted limitations including the synthetic nature of Kronecker-generated graphs versus real-world datasets used by Google Research or Facebook AI Research, and concerns about optimizations that exploit microbenchmarks rather than application-level behavior—points also raised in analyses by ACM and IEEE committees. Other limitations cited by teams at Lawrence Livermore National Laboratory and Oak Ridge National Laboratory include difficulties in reproducing results across diverse software stacks and the benchmark's narrow focus on specific kernels versus broader suites like those from SPEC or the PARSEC benchmark suite. Some vendors and researchers have proposed alternate graph benchmarks developed at Stanford University and University of California, San Diego to address these concerns.

Category:High-performance computing benchmarks