transcript
show notes
When a flagship AI model scores 25.4% on a benchmark and gets crushed by a dumb baseline, something's wrong. This episode explores why error bars matter, why almost nobody in AI computes them, and why your "state-of-the-art" result might just be statistical noise. A deep dive into Evan Miller's preprint on bringing rigor to AI evaluations.
00:00 - The Court Prediction Model That Failed
06:30 - When Your Result Meets a Coin Flip
12:15 - Why Error Bars Are Missing from AI Benchmarking
16:45 - The Statisticians Were Right All Along
---
Sources & further reading:
• All fetched and quoted on 2026-09-17.
• [preprint] Evan Miller (Anthropic), *Adding Error Bars to Evals: A Statistical Approach to
• Language Model Evaluations*, arXiv:2411.00640, 2024-11-01: https://arxiv.org/abs/2411.00640
• the five recommendations, the clustered-SE table (DROP 1.34 vs 0.44; MGSM 1.88×), the paired
• difference recommendation, the temperature warning, the K-resampling arithmetic, and the n≈969 /
• 1,000-question power result.
• Bowyer, Ivanova, Aitchison, *Position: Don't use the CLT in LLM evals with fewer than a few
• hundred datapoints*, arXiv:2503.01747, ICML 2025 — the "fairly catastrophic failure" quote and
• the Wilson/Bayesian alternatives.
• Reuel et al., BetterBench, arXiv:2411.12990, 2024-11-20, NeurIPS 2024 D&B — 24 benchmarks
• 46 practices, 14-of-24, and MMLU scoring lowest. Living site: https://betterbench.stanford.edu/
• Biderman, Schoelkopf, Sutawika, Gao et al., arXiv:2405.14782 — lm-eval already reporting standard
• errors, and the call for statistical practice.
• Hochlehnert, Bhatnagar, Udandarao, Albanie, Prabhu, Bethge, *A Sober Look at Progress in
• Language Model Reasoning*, arXiv:2504.07086, 2025-04-09, COLM 2025 — 5–15 point seed SD, the
• one-question sensitivity, K ≥ 30, the gains-inside-variance conclusion, and the cross-cluster
• hardware gap.
• Lu, Bartolo, Moore, Riedel, Stenetorp, Fantastically Ordered Prompts and Where to Find Them
• arXiv:2104.08786, ACL 2022 — the near-SOTA-to-random ordering effect and the <1% fine-tuning
• contrast. Their "30%" is relative gain from prompt selection, not an accuracy spread.
• Mizrahi et al., State of What Art?, arXiv:2401.00595, TACL — 21 of 25 tasks with significant
• prompt effects; the 1st-to-9th rank move.
• Madaan et al., Quantifying Variance in Evaluation Benchmarks, arXiv:2406.10229, 2024-06-14
• seed variance across 280 models; the cloze reformulation raising monotonicity 0.09 → 0.95; and
• benchmarks sitting at chance after 210B tokens.
• Hugging Face, Open LLM Leaderboard v2 post, late June 2024
• normalisation against the random: https://huggingface.co/spaces/open-llm-leaderboard/blog
• baseline, the worked A-vs-B example, and the GPQA/MuSR near-chance notes. Normalisation mechanics
• Zheng, Pang, Du et al., Cheating Automatic LLM Benchmarks, arXiv:2410.07137, ICLR 2025 Oral
• the constant-response 86.5% LC win rate.
• Huang, Shen, Wei, Broderick, *Dropping Just a Handful of Preferences Can Change Top Large
• Language Model Rankings*, arXiv:2508.11847, 2025-08-16 — the two-vote flip, the MT-bench
• contrast, and the 77%-of-random-1%-deletions caveat.
• [internal] GetTheJob/research/gamesenser-technical-profile.md — Wall A at 25.4% (16/63), the
• 28.2% (20/71) constant-zone floor at z = 0.36 / p = 0.72, the 16.7%-not-11.1% chance-floor
• correction, and EXP-39's 9.4% geometry model with 188 of 200 random permutations beating it
• (p = 0.945) — .
This podcast episode was fully generated by AI — research, script, voices, and production. Built with Claude, Piper TTS, and automated pipeline tooling.