UK AISI Publishes Reproducible AI Benchmark Results on EvalEval
UK AISI is publishing verified results for five benchmarks and six frontier models through EvalEval's open Evaluation Cards, improving reproducibility.
Summary
On September 22, 2026, the UK AI Security Institute began sharing publicly reported evaluation methods and findings through EvalEval’s Evaluation Cards where appropriate. The release provides verified results, context and configurations for HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0 across Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. Cyber CTFs and The Last Ones are also included, using a different, partially overlapping model set. The data accompany How Inference Compute Shapes Frontier LLM Evaluation.
The research shows that Humanity’s Last Exam performance changes with evaluation protocol and inference-time compute. Curves track the cumulative share of attempted tasks solved within each token count using the earliest observed success, while oracle correctness feedback after every attempt enables models to solve additional tasks as token use rises. Publishing transcript-level setup information creates reference points for diagnosing score differences when other reports omit reproducibility details or are too expensive to rerun.
The collaboration began at a joint NeurIPS 2025 workshop, with AISI feedback shaping EvalEval’s Every Eval Ever schema. Evaluation Cards combines benchmark, run and model metadata, complementing AISI’s OptStop efficiency work, HiBayES statistical methods, transcript analysis and capability elicitation standards. Developers are invited to submit verified results and EEE-formatted data, while evaluation, governance and policy researchers can compare benchmarks and models. AISI operates within the UK Department for Science, Innovation and Technology, and further standardisation with EvalEval and other evaluation organisations is planned.
Positives
- Five benchmarks now have verified results, context and configuration data available through Evaluation Cards.
- Six frontier models from Anthropic and OpenAI are covered, spanning Claude Opus 4 to 4.6 and GPT-5 to 5.4.
- Two cyber evaluations, Cyber CTFs and The Last Ones, broaden the release beyond the paper’s main experiment.
- Every Eval Ever gives developers a shared schema for reporting benchmark and evaluation-run data.
- Transcript-level setup information helps researchers diagnose why apparently similar model scores differ across evaluation conditions.
- AISI’s OptStop and HiBayES work adds efficiency and statistical rigor to the broader standardisation effort.
Risks & concerns
- Many AI evaluation results remain scattered across incompatible formats and omit information required for reproduction.
- Repeating evaluations can be prohibitively expensive, leaving poorly documented findings difficult to verify.
- Humanity’s Last Exam scores change with protocol and inference compute, undermining comparisons that ignore setup differences.
- Cyber CTFs and The Last Ones use a different, only partially overlapping model set, limiting direct comparisons.
- Evaluation science still lacks consensus on documenting when benchmarks are applicable and useful.