transcript
show notes
How good is your test suite really? Host A plants bugs while Host B watches tests pass anyway—some changes harmless, others cutting critical safety systems. They explore the mutation testing framework, kill rates, and the uncomfortable truth: a test suite that passes everything might be hiding catastrophic blind spots.
00:00 - Introduction: The bug-planting experiment
01:30 - First invisible bug: The harmless edit no test catches
05:00 - Second invisible bug: Removed backup interlock
08:00 - The paradox: Two invisible bugs, opposite meanings
10:00 - What is grading the grader? Mutation testing framework
12:00 - Episode 8 callback: The Compilation Illusion
14:00 - Kill rate scoring and the 90% benchmark
---
Sources & further reading:
• BrewSys repo (C:\Users\musse\Projects\BrewSys), branch master at bf60802. Read-only.
• docs/podcast/LEARNING-CONTEXT.md §1: the only source for project claims (who did what, Phase
• 0, the don't-say list).
• docs/results/2026-09-23-brite-transfer-hand.md: run 4, hand suite, commit 44b06db.
• Headline, per-operator table, union, 35 survivors and aggregate triage.
• docs/results/2026-09-23-brite-transfer-llm.md: run 4, FDS-only LLM suite, commit 44b06db.
• Headline, per-operator table, the drop_cond gap description, 86 survivors.
• docs/results/README.md: what a record must carry, STALE rule, survivor grouping rule.
• docs/backlog.md: B-004 (lint rule for redundant interlock layers; status ready), B-005
• B-006.
• brewsys/mutate.py: operator definitions and the kill rule.
• Not used, on purpose: anything in out/; docs/results/*-run6.md (STALE).
• Outside
• [peer-reviewed] Jia and Harman, *An Analysis and Survey of the Development of Mutation
• Testing*, IEEE TSE 37(5):649–678, 2011, doi:10.1109/TSE.2010.62 — authors' copy
• read 2026-09-24: http://crest.cs.ucl.ac.uk/fileadmin/crest/sebasepaper/JiaH10.pdf
• (killed/survived definition, equivalent mutants, undecidability, mutation-score definition
• quoted verbatim above).
• [preprint] Evan Miller, *Adding Error Bars to Evals: A Statistical Approach to Language Model
• Evaluations*, arXiv:2411.00640 — — checked 2026-09-24: https://arxiv.org/abs/2411.00640
• (quote and submission date).
• [internal] data/series/industrial-controls/ep-08-can-an-llm-write-ladder.md (the setup this
• pays off); data/series/ai-benchmarking/ep-06-the-number-is-a-random-variable.md (the
• callback).
This podcast episode was fully generated by AI — research, script, voices, and production. Built with Claude, Piper TTS, and automated pipeline tooling.