This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.
| QUDA | |
|---|---|
| Name | QUDA |
| Developer | NVIDIA Corporation; Fermilab; University of Edinburgh; University of Wuppertal; Lawrence Berkeley National Laboratory |
| Released | 2009 |
| Latest release | (varies) |
| Programming language | C++; CUDA; MPI |
| Operating system | Linux; macOS (experimental); Windows (via WSL) |
| License | BSD-style |
QUDA
QUDA is a library for accelerating lattice field theory computations on graphics processing units developed to target lattice quantum chromodynamics workloads on NVIDIA hardware. It provides GPU-optimized implementations of solvers, linear algebra, and data-layout transformations commonly used in projects at institutions such as Fermi National Accelerator Laboratory, Lawrence Berkeley National Laboratory, CERN, University of Edinburgh, and University of Wuppertal. QUDA is widely used by collaborations connected to experiments and theory groups including MILC, RBC-UKQCD, HotQCD, CLS, and ETMC.
QUDA originated to exploit the massively parallel architecture of NVIDIA GPUs for sparse linear systems arising in lattice field theory, particularly for the Dirac operator inversion central to Lattice QCD computations. Early development drew on contributions from researchers associated with Fermi National Accelerator Laboratory and academic groups collaborating with projects like USQCD and SciDAC. The project sits at the intersection of efforts tied to initiatives such as CUDA adoption, HPC facility procurement involving Oak Ridge National Laboratory and Argonne National Laboratory, and algorithmic advances that partnered with teams from Stanford University and MIT.
QUDA's architecture maps lattice degrees of freedom and gauge field representations onto GPU memory hierarchies, optimizing for streaming bandwidth and on-chip shared memory usage. Its design integrates components familiar to developers at NVIDIA and HPC centers such as NERSC and Princeton University's HPC group, including occupancy tuning inspired by research from Los Alamos National Laboratory and memory-coalescing techniques advocated by teams at Lawrence Livermore National Laboratory. The codebase uses a modular structure to separate Dirac operator implementations (Wilson, Clover, Domain Wall, Staggered) from solver frameworks, allowing interoperability with middleware developed at Columbia University and University of Arizona, and with scheduling frameworks used at National Energy Research Scientific Computing Center.
QUDA implements a variety of Krylov-subspace and multigrid algorithms including Conjugate Gradient, BiCGStab, GMRES, and algebraic multigrid methods developed in collaboration with groups at Brookhaven National Laboratory, University of Illinois Urbana-Champaign, and University of Regensburg. It incorporates mixed-precision techniques influenced by numerical analysis work at ETH Zurich and University of Cambridge, and communication-avoidant algorithms that reflect strategies reported from Lawrence Berkeley National Laboratory and University of Edinburgh. Optimizations include gauge-field compression used in studies from University of Glasgow, spin-projection techniques consistent with formulations by Columbia University researchers, and autotuning modules similar to tools from ATLAS (software) and the BLAS ecosystem.
Primarily developed for NVIDIA CUDA-capable GPUs such as those deployed at Oak Ridge Leadership Computing Facility and Argonne Leadership Computing Facility, QUDA supports architectures from Fermi (microarchitecture) through Volta (microarchitecture), Turing (microarchitecture), and Ampere (microarchitecture). It integrates with MPI stacks common at centers like NERSC and Jülich Research Centre, and can be used within container environments advocated by teams at CERN and EMBL. Interfacing layers allow integration with software used at Fermilab and universities that run clusters based on systems from vendors including IBM, HPE, and Dell EMC.
QUDA is commonly integrated into community codes and workflows such as MILC, Chroma (lattice QCD), CPS (Columbia Physics System), and bespoke analysis pipelines used by collaborations like RBC-UKQCD and Hadron Spectrum Collaboration. Users link QUDA into build systems maintained with tools familiar to groups at GitHub and Bitbucket, and use continuous integration practices similar to those at Travis CI and Jenkins. Job submission and resource management typically occur via schedulers used at Cray-based centers and batch systems employed by PRACE partners.
Benchmarking reports from collaborations running on clusters at Fermilab, SLAC National Accelerator Laboratory, and Brookhaven National Laboratory demonstrate order-of-magnitude speedups over CPU-only implementations for key kernels. Performance scales with GPU generation and interconnect; studies referencing technologies from NVLink and InfiniBand networks reported by groups at Cornell University and University of California, Berkeley show improved multi-GPU strong-scaling. Comparative analyses published in proceedings associated with conferences such as SC (conference), Lattice (conference), and ICFP document throughput metrics, solver convergence behavior, and energy efficiency measures in contexts overseen by institutions like DOE laboratories.
Development of QUDA is driven by a distributed community of researchers from national labs and universities including Fermi National Accelerator Laboratory, University of Edinburgh, University of Wuppertal, Brookhaven National Laboratory, and Lawrence Berkeley National Laboratory. The project engages with workshops and working groups associated with USQCD, ILDG, and the Lattice Field Theory community, and contributors frequently present at venues such as Lattice (conference), SC (conference), and domain-specific summer schools organized by Jefferson Lab and TRIUMF. Maintenance follows practices similar to collaborative scientific software projects at CERN and GitHub, with issue tracking, code review, and periodic releases coordinated among academic and laboratory partners.
Category:Scientific software