How UK AISI and EvalEval Are Making Benchmark Results Reproducible
AISI and EvalEval are publishing open, reproducible benchmark results for top LLMs, making evaluation data transparent and comparable.
Why Now
The UK AI Security Institute (AISI) has started using EvalEval’s infrastructure to share evaluation results publicly, following earlier collaboration at a NeurIPS 2025 workshop.
What Happened
AISI released verified results for five benchmarks—HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench, and Terminal-Bench 2.0—across six frontier models (Claude Opus 4‑4.6, GPT‑5‑5.4). The release includes evaluation cards with context, configuration, and transcript-level transparency. The data accompany AISI’s paper on how inference compute shapes frontier LLM evaluation, showing performance varies with token budget and oracle feedback.
Why It Matters
Open, standardized reporting lets researchers and practitioners compare models under identical conditions, diagnose gaps in evaluation practices, and build more reliable meta‑research. It also supports policy makers by providing verifiable evidence of model capabilities and limitations.
The Limitation
The released results cover only a subset of benchmarks and models; broader coverage and long‑term reproducibility depend on wider community adoption of the Every Eval Ever schema.
What You Can Do
Explore the Evaluation Cards on Hugging Face to compare model performance under different evaluation setups.
Source
Read original sourceWhy we picked this
UK AISI and EvalEval reproducibility work is meaningful AI news.