Ai2 GPU Scheduler Delivers 98% of Budgeted Compute at Full Occupancy
Ai2’s budget-based GPU scheduler delivers 98% of owed compute, cuts debug waits to 30 seconds and human-assisted repairs 74% at 98% occupancy for AI research.
Summary
On October 9, 2026, Ai2 detailed replacing priority scheduling across thousands of NVIDIA H100, B200 and B300 GPUs arranged in 88 to 1,024 GPU clusters serving about 150 researchers in LLM and VLM training, robotics reinforcement learning and scientific agent post-training. Demand runs two to three times supply. The old opt-out preemption model produced squatting, universal HIGH priority, starvation of lower levels and maintenance negotiations, while dedicated team monopolies stranded seasonal capacity.
Ai2 now assigns hierarchical GPU time budgets through research management to the CEO; one example project claims 35% of capacity. A fair-share scheduler tracks charged occupancy over a seven-day window; funded jobs declare a protected minimum runtime capped at eight hours, then resumable work can be preempted and requeued. Unallocated jobs are free, immediately preemptible and keep idle cycles occupied. Simulations forecast debug p90 waits dropping from about six hours to five minutes. Cluster-by-cluster rollout began at July’s end, and Chris Clark said burst recovery felt like 30% more compute.
Over a 30-day test, teams received 98% of GPU hours owed; 13 of 15 allocations received at least 95%, and the worst received 90%. Occupancy remained 98% while demand stayed two to three times capacity; unallocated work supplied 18% of delivered time. Actual debug p90 wait fell from two hours to 30 seconds, although the baseline sample was smaller; on the largest H100 cluster, median wait dropped from five minutes to 24 seconds and p90 from 2.8 to 1.8 hours. Automatic draining cut repairs needing people by 74%. Confusing terminology and inconsistent early behavior required live sessions and new visualizations. Eight-hour protection disrupts stateful interactive sessions, prompting a CPU-only cluster beside on-prem storage and restorable sessions. Ai2 is testing possible fragmentation delays for the largest jobs, then targets bootstrapping, checkpointing and training utilization.
Positives
- Teams received 98% of owed GPU hours during the 30-day test, with 13 of 15 allocations receiving at least 95%.
- Cluster occupancy held at 98%, with unallocated, preemptible workloads supplying 18% of delivered GPU time.
- Debug workload p90 queue time fell from two hours to 30 seconds after the new scheduler launched.
- Human-assisted repairs declined 74% because unhealthy hosts automatically drain workloads after protected runtimes expire.
- Chris Clark estimated that reclaiming unused allocations made the system feel like it provided 30% more compute.
Risks & concerns
- Demand remains two to three times available capacity across Ai2’s GPU clusters.
- Eight-hour protected-runtime limits can interrupt interactive sessions and force researchers to rebuild volatile working state.
- Changed priority terminology and inconsistent behavior during the incremental rollout confused researchers until Ai2 added live training and visualizations.
- Capacity fragmentation may increase queue waits for the largest distributed workloads because protected jobs cannot always be interrupted together.
- The worst-performing team allocation received 90% of owed GPU hours, below the 95% reached by 13 of 15 allocations.