This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.
| Maui Cluster Scheduler | |
|---|---|
| Name | Maui Cluster Scheduler |
| Operating system | Unix-like |
| Genre | Batch scheduler, job scheduler |
| License | Academic/Moab-related |
Maui Cluster Scheduler is a high-performance batch scheduler designed to manage job queues and resource allocation on high-performance computing clusters. Maui provides advanced scheduling features such as backfilling, fair-share, reservations, and quality-of-service, and is often deployed alongside resource managers to optimize throughput for scientific computing, engineering simulations, and grid environments. Maui has been used in research centers, national laboratories, and academic institutions to coordinate workloads for applications in climate modeling, computational chemistry, and astrophysics.
Maui originated in the late 1990s as part of efforts at research institutions to improve upon basic queuing systems developed at places like Lawrence Livermore National Laboratory and Sandia National Laboratories. Early development was influenced by projects and collaborations involving National Center for Supercomputing Applications, Petascale Computing Research, and the emerging needs of the High Performance Computing community. Over successive releases, Maui incorporated ideas from academic literature on backfilling and fair-share drawn from conferences such as the International Conference for High Performance Computing, Networking, Storage and Analysis and workshops associated with ACM SIGOPS and IEEE. The project intersected with commercial initiatives exemplified by Adaptive Computing and research efforts funded by agencies such as National Science Foundation and Department of Energy.
Maui’s architecture separates policy from execution and typically interfaces with resource managers like TORQUE (software), PBS Professional, and batch systems influenced by Portable Batch System. Core components include the scheduler daemon, policy engine, accounting subsystem, and command-line utilities for administrators and users. The scheduler daemon communicates with host-level daemons on compute nodes and follows designs similar to resource management systems used in Los Alamos National Laboratory and other national labs. Accounting integrates with databases and logging systems inspired by standards from Open Grid Forum and monitoring systems comparable to Ganglia and Nagios. Maui’s modular design allows optional modules for reservations, advance reservations influenced by practices at European Grid Infrastructure centers, and plugins developed in the style of extensions for SLURM or HTCondor ecosystems.
Maui implements a suite of scheduling policies including conservative and aggressive backfilling, priority queuing, fair-share algorithms, and multi-factor priority calculations. These models are informed by theoretical work from researchers at University of California, Berkeley, Massachusetts Institute of Technology, and Princeton University on job packing and resource fragmentation. Maui’s fair-share uses usage decay and entitlement calculations akin to techniques described in publications from Sandia National Laboratories and Argonne National Laboratory. Reservation policies enable both static and dynamic reservations for maintenance windows and time-sensitive workflows, reflecting practices used at facilities like Oak Ridge National Laboratory and Lawrence Berkeley National Laboratory. Maui also supports class-based scheduling and quality-of-service tiers similar to mechanisms found in Amazon Web Services and commercial batch services.
Configuration is file-driven with parameters controlling priorities, limits, reservations, and policies; common files and command utilities mirror administrative patterns used in UNIX system administration and cluster management frameworks from vendors such as Dell EMC and Hewlett-Packard Enterprise. Administrators integrate accounting and user data from identity services including LDAP and site policies shaped by institutional governance at organizations like European Organization for Nuclear Research. Operational procedures for tuning backfill windows and node granularity derive from runbooks developed at supercomputing centers such as National Energy Research Scientific Computing Center and Texas Advanced Computing Center. Maui’s command-line tools and scripting hooks allow automation via configuration management systems like Ansible, Puppet (software), and SaltStack.
Maui has been applied to large-scale scientific workloads in domains represented by collaborations such as Coupled Model Intercomparison Project for climate, Human Genome Project-era genomics workflows, and computational fluid dynamics used in aerospace partnerships with NASA. Integration patterns include coupling with MPI job launches from Open MPI, hybrid workflows orchestrated with HTCondor, and ticketing or incident flows tied to enterprise systems like ServiceNow. Maui has been used in grids and federated environments alongside middleware from Globus and scheduling overlays like those demonstrated in XSEDE resource allocations.
Maui’s performance has been benchmarked in production centers handling thousands to tens of thousands of simultaneous jobs, with tuning guidance informed by studies from NERSC and scalability reports published by Oak Ridge Leadership Computing Facility. Scalability considerations include scheduler decision latency, accounting database throughput, and the overhead of complex priority calculations; sites mitigate these via hierarchical scheduling and partitioning strategies used in deployments at Argonne Leadership Computing Facility and large university clusters. Maui’s backfilling algorithms can dramatically increase utilization metrics compared to FIFO systems, a result echoed in benchmarking literature from ACM and IEEE venues.
Maui’s licensing and distribution history intersects with academic licenses and collaborations with commercial entities such as Adaptive Computing and community projects influenced by foundations like Open Source Initiative. Development activity has waxed and waned as institutions adopted alternative schedulers like SLURM and Kubernetes for containerized workloads; nonetheless, legacy deployments remain in many national laboratories and universities informed by long-term support models at organizations such as Cray and IBM. Ongoing maintenance typically occurs within site-specific teams and through collaborations among research computing centers participating in forums like the Supercomputing Conference.
Category:Job scheduling software