Skip to content
Artwork for Clown Cast
Clown Cast · Today · 20 min

Error Bars, or: The Number is a Random Variable

When a flagship AI model scores 25.4% on a benchmark and gets crushed by a dumb baseline, something's wrong. This episode explores why error bars matter, why almost nobody in AI computes them, and why your "state-of-the-art" result might just be statistical noise. A deep dive into Evan Miller's preprint on bringing rigor to AI evaluations. 00:00 - The Court Prediction Model That Failed 06:30 - When Your Result Meets a Coin Flip 12:15 - Why Error Bars Are Missing from AI Benchmarking 16:45 - The Statisticians Were Right All Along --- Sources & further reading: • All fetched and quoted on 2026-09-17. • [preprint] Evan Miller (Anthropic), *Adding Error Bars to Evals: A Statistical Approach to • Language Model Evaluations*, arXiv:2411.00640, 2024-11-01: https://arxiv.org/abs/2411.00640 • the five recommendations, the clustered-SE table (DROP 1.34 vs 0.44; MGSM 1.88×), the paired • difference recommendation, the temperature warning, the K-resampling arithmetic, and the n≈969 / • 1,000-question power result. • Bowyer, Ivanova, Aitchison, *Position: Don't use the CLT in LLM evals with fewer than a few • hundred datapoints*, arXiv:2503.01747, ICML 2025 — the "fairly catastrophic failure" quote and • the Wilson/Bayesian alternatives. • Reuel et al., BetterBench, arXiv:2411.12990, 2024-11-20, NeurIPS 2024 D&B — 24 benchmarks • 46 practices, 14-of-24, and MMLU scoring lowest. Living site: https://betterbench.stanford.edu/ • Biderman, Schoelkopf, Sutawika, Gao et al., arXiv:2405.14782 — lm-eval already reporting standard • errors, and the call for statistical practice. • Hochlehnert, Bhatnagar, Udandarao, Albanie, Prabhu, Bethge, *A Sober Look at Progress in • Language Model Reasoning*, arXiv:2504.07086, 2025-04-09, COLM 2025 — 5–15 point seed SD, the • one-question sensitivity, K ≥ 30, the gains-inside-variance conclusion, and the cross-cluster • hardware gap. • Lu, Bartolo, Moore, Riedel, Stenetorp, Fantastically Ordered Prompts and Where to Find Them • arXiv:2104.08786, ACL 2022 — the near-SOTA-to-random ordering effect and the <1% fine-tuning • contrast. Their "30%" is relative gain from prompt selection, not an accuracy spread. • Mizrahi et al., State of What Art?, arXiv:2401.00595, TACL — 21 of 25 tasks with significant • prompt effects; the 1st-to-9th rank move. • Madaan et al., Quantifying Variance in Evaluation Benchmarks, arXiv:2406.10229, 2024-06-14 • seed variance across 280 models; the cloze reformulation raising monotonicity 0.09 → 0.95; and • benchmarks sitting at chance after 210B tokens. • Hugging Face, Open LLM Leaderboard v2 post, late June 2024 • normalisation against the random: https://huggingface.co/spaces/open-llm-leaderboard/blog • baseline, the worked A-vs-B example, and the GPQA/MuSR near-chance notes. Normalisation mechanics • Zheng, Pang, Du et al., Cheating Automatic LLM Benchmarks, arXiv:2410.07137, ICLR 2025 Oral • the constant-response 86.5% LC win rate. • Huang, Shen, Wei, Broderick, *Dropping Just a Handful of Preferences Can Change Top Large • Language Model Rankings*, arXiv:2508.11847, 2025-08-16 — the two-vote flip, the MT-bench • contrast, and the 77%-of-random-1%-deletions caveat. • [internal] GetTheJob/research/gamesenser-technical-profile.md — Wall A at 25.4% (16/63), the • 28.2% (20/71) constant-zone floor at z = 0.36 / p = 0.72, the 16.7%-not-11.1% chance-floor • correction, and EXP-39's 9.4% geometry model with 188 of 200 random permutations beating it • (p = 0.945) — . This podcast episode was fully generated by AI — research, script, voices, and production. Built with Claude, Piper TTS, and automated pipeline tooling.

0:00-20:19

transcript

No transcript — this publisher did not publish one.

show notes

When a flagship AI model scores 25.4% on a benchmark and gets crushed by a dumb baseline, something's wrong. This episode explores why error bars matter, why almost nobody in AI computes them, and why your "state-of-the-art" result might just be statistical noise. A deep dive into Evan Miller's preprint on bringing rigor to AI evaluations.

00:00 - The Court Prediction Model That Failed
06:30 - When Your Result Meets a Coin Flip
12:15 - Why Error Bars Are Missing from AI Benchmarking
16:45 - The Statisticians Were Right All Along

---
Sources & further reading:
• All fetched and quoted on 2026-09-17.
• [preprint] Evan Miller (Anthropic), *Adding Error Bars to Evals: A Statistical Approach to
• Language Model Evaluations*, arXiv:2411.00640, 2024-11-01: https://arxiv.org/abs/2411.00640
• the five recommendations, the clustered-SE table (DROP 1.34 vs 0.44; MGSM 1.88×), the paired
• difference recommendation, the temperature warning, the K-resampling arithmetic, and the n≈969 /
• 1,000-question power result.
• Bowyer, Ivanova, Aitchison, *Position: Don't use the CLT in LLM evals with fewer than a few
• hundred datapoints*, arXiv:2503.01747, ICML 2025 — the "fairly catastrophic failure" quote and
• the Wilson/Bayesian alternatives.
• Reuel et al., BetterBench, arXiv:2411.12990, 2024-11-20, NeurIPS 2024 D&B — 24 benchmarks
• 46 practices, 14-of-24, and MMLU scoring lowest. Living site: https://betterbench.stanford.edu/
• Biderman, Schoelkopf, Sutawika, Gao et al., arXiv:2405.14782 — lm-eval already reporting standard
• errors, and the call for statistical practice.
• Hochlehnert, Bhatnagar, Udandarao, Albanie, Prabhu, Bethge, *A Sober Look at Progress in
• Language Model Reasoning*, arXiv:2504.07086, 2025-04-09, COLM 2025 — 5–15 point seed SD, the
• one-question sensitivity, K ≥ 30, the gains-inside-variance conclusion, and the cross-cluster
• hardware gap.
• Lu, Bartolo, Moore, Riedel, Stenetorp, Fantastically Ordered Prompts and Where to Find Them
• arXiv:2104.08786, ACL 2022 — the near-SOTA-to-random ordering effect and the <1% fine-tuning
• contrast. Their "30%" is relative gain from prompt selection, not an accuracy spread.
• Mizrahi et al., State of What Art?, arXiv:2401.00595, TACL — 21 of 25 tasks with significant
• prompt effects; the 1st-to-9th rank move.
• Madaan et al., Quantifying Variance in Evaluation Benchmarks, arXiv:2406.10229, 2024-06-14
• seed variance across 280 models; the cloze reformulation raising monotonicity 0.09 → 0.95; and
• benchmarks sitting at chance after 210B tokens.
• Hugging Face, Open LLM Leaderboard v2 post, late June 2024
• normalisation against the random: https://huggingface.co/spaces/open-llm-leaderboard/blog
• baseline, the worked A-vs-B example, and the GPQA/MuSR near-chance notes. Normalisation mechanics
• Zheng, Pang, Du et al., Cheating Automatic LLM Benchmarks, arXiv:2410.07137, ICLR 2025 Oral
• the constant-response 86.5% LC win rate.
• Huang, Shen, Wei, Broderick, *Dropping Just a Handful of Preferences Can Change Top Large
• Language Model Rankings*, arXiv:2508.11847, 2025-08-16 — the two-vote flip, the MT-bench
• contrast, and the 77%-of-random-1%-deletions caveat.
• [internal] GetTheJob/research/gamesenser-technical-profile.md — Wall A at 25.4% (16/63), the
• 28.2% (20/71) constant-zone floor at z = 0.36 / p = 0.72, the 16.7%-not-11.1% chance-floor
• correction, and EXP-39's 9.4% geometry model with 188 of 200 random permutations beating it
• (p = 0.945) — .

This podcast episode was fully generated by AI — research, script, voices, and production. Built with Claude, Piper TTS, and automated pipeline tooling.