Skip to content
Artwork for Clown Cast
Clown Cast · September 24 · 16 min

Grading the Grader: When Tests Fail to See

How good is your test suite really? Host A plants bugs while Host B watches tests pass anyway—some changes harmless, others cutting critical safety systems. They explore the mutation testing framework, kill rates, and the uncomfortable truth: a test suite that passes everything might be hiding catastrophic blind spots. 00:00 - Introduction: The bug-planting experiment 01:30 - First invisible bug: The harmless edit no test catches 05:00 - Second invisible bug: Removed backup interlock 08:00 - The paradox: Two invisible bugs, opposite meanings 10:00 - What is grading the grader? Mutation testing framework 12:00 - Episode 8 callback: The Compilation Illusion 14:00 - Kill rate scoring and the 90% benchmark --- Sources & further reading: • BrewSys repo (C:\Users\musse\Projects\BrewSys), branch master at bf60802. Read-only. • docs/podcast/LEARNING-CONTEXT.md §1: the only source for project claims (who did what, Phase • 0, the don't-say list). • docs/results/2026-09-23-brite-transfer-hand.md: run 4, hand suite, commit 44b06db. • Headline, per-operator table, union, 35 survivors and aggregate triage. • docs/results/2026-09-23-brite-transfer-llm.md: run 4, FDS-only LLM suite, commit 44b06db. • Headline, per-operator table, the drop_cond gap description, 86 survivors. • docs/results/README.md: what a record must carry, STALE rule, survivor grouping rule. • docs/backlog.md: B-004 (lint rule for redundant interlock layers; status ready), B-005 • B-006. • brewsys/mutate.py: operator definitions and the kill rule. • Not used, on purpose: anything in out/; docs/results/*-run6.md (STALE). • Outside • [peer-reviewed] Jia and Harman, *An Analysis and Survey of the Development of Mutation • Testing*, IEEE TSE 37(5):649–678, 2011, doi:10.1109/TSE.2010.62 — authors' copy • read 2026-09-24: http://crest.cs.ucl.ac.uk/fileadmin/crest/sebasepaper/JiaH10.pdf • (killed/survived definition, equivalent mutants, undecidability, mutation-score definition • quoted verbatim above). • [preprint] Evan Miller, *Adding Error Bars to Evals: A Statistical Approach to Language Model • Evaluations*, arXiv:2411.00640 — — checked 2026-09-24: https://arxiv.org/abs/2411.00640 • (quote and submission date). • [internal] data/series/industrial-controls/ep-08-can-an-llm-write-ladder.md (the setup this • pays off); data/series/ai-benchmarking/ep-06-the-number-is-a-random-variable.md (the • callback). This podcast episode was fully generated by AI — research, script, voices, and production. Built with Claude, Piper TTS, and automated pipeline tooling.

0:00-16:27

transcript

No transcript — this publisher did not publish one.

show notes

How good is your test suite really? Host A plants bugs while Host B watches tests pass anyway—some changes harmless, others cutting critical safety systems. They explore the mutation testing framework, kill rates, and the uncomfortable truth: a test suite that passes everything might be hiding catastrophic blind spots.

00:00 - Introduction: The bug-planting experiment
01:30 - First invisible bug: The harmless edit no test catches
05:00 - Second invisible bug: Removed backup interlock
08:00 - The paradox: Two invisible bugs, opposite meanings
10:00 - What is grading the grader? Mutation testing framework
12:00 - Episode 8 callback: The Compilation Illusion
14:00 - Kill rate scoring and the 90% benchmark

---
Sources & further reading:
• BrewSys repo (C:\Users\musse\Projects\BrewSys), branch master at bf60802. Read-only.
• docs/podcast/LEARNING-CONTEXT.md §1: the only source for project claims (who did what, Phase
• 0, the don't-say list).
• docs/results/2026-09-23-brite-transfer-hand.md: run 4, hand suite, commit 44b06db.
• Headline, per-operator table, union, 35 survivors and aggregate triage.
• docs/results/2026-09-23-brite-transfer-llm.md: run 4, FDS-only LLM suite, commit 44b06db.
• Headline, per-operator table, the drop_cond gap description, 86 survivors.
• docs/results/README.md: what a record must carry, STALE rule, survivor grouping rule.
• docs/backlog.md: B-004 (lint rule for redundant interlock layers; status ready), B-005
• B-006.
• brewsys/mutate.py: operator definitions and the kill rule.
• Not used, on purpose: anything in out/; docs/results/*-run6.md (STALE).
• Outside
• [peer-reviewed] Jia and Harman, *An Analysis and Survey of the Development of Mutation
• Testing*, IEEE TSE 37(5):649–678, 2011, doi:10.1109/TSE.2010.62 — authors' copy
• read 2026-09-24: http://crest.cs.ucl.ac.uk/fileadmin/crest/sebasepaper/JiaH10.pdf
• (killed/survived definition, equivalent mutants, undecidability, mutation-score definition
• quoted verbatim above).
• [preprint] Evan Miller, *Adding Error Bars to Evals: A Statistical Approach to Language Model
• Evaluations*, arXiv:2411.00640 — — checked 2026-09-24: https://arxiv.org/abs/2411.00640
• (quote and submission date).
• [internal] data/series/industrial-controls/ep-08-can-an-llm-write-ladder.md (the setup this
• pays off); data/series/ai-benchmarking/ep-06-the-number-is-a-random-variable.md (the
• callback).

This podcast episode was fully generated by AI — research, script, voices, and production. Built with Claude, Piper TTS, and automated pipeline tooling.