
When a Fake Dashboard Makes an AI Agent Just as Confident
When a Fake Dashboard Makes an AI Agent Just as Confident Source: Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable Paper was published on August 27, 2026 This episode was AI-generated on August 29, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Show a language model a market panel where every single number is fabricated, and it commits to a prediction just as often as when the data is real — 37.6% versus 36.8%. The models know these questions are unanswerable; they say so 90% of the time when you ask directly. This episode traces exactly which component breaks, why auditing stated confidence can't see it, and what a 540-example fix does and doesn't fix. Key Takeaways Why commitment climbs from 6.5% with a bare question to 54% with a full technical panel — and collapses to 3.5% when the panel obviously belongs to another company The diagnostic that rules out incapacity (99.9% accuracy reading the same panel), belief change (three points of movement), and missing judgment (90% correctly labeled irreducible) — leaving a disconnected decision gate Why the committed forecasts score AUROC 0.346 — worse than chance, pointing at the wrong outcome two times in three — while the mean stated probability of 49.1% would pass any aggregate calibration audit Where the pooled headline breaks down: three of twelve models carry almost the whole effect, and restricted to the responsive seven the equivalence claim no longer holds How 540 synthetic dice-and-coin examples drive a 3B model's commitment to 0.0% on unseen stock cases — and why the same gate collapses under a rigid structured-output format Why the author's five-percentage-point equivalence margin was chosen after seeing the point estimate, and what that means for how you read the result 00:00 — Every number on the screen was invented The cold open: an agent commits to a ten-day stock call just as readily on a fully fabricated dashboard as on a real one. 01:59 — The comforting 2022 result this overturns Why 'models mostly know what they know' shaped evaluation practice, and why these models knowing the question is unanswerable ~90% of the time makes the failure stranger, not safer. 03:58 — Informed by the panel, or impressed by it? Tyler steelmans the reading that technical indicators carry weak real signal — and Juniper explains why only intervening on the content, not observing outputs, can separate the two. 05:57 — How to build a question with no answer The construction: aleatoric versus epistemic uncertainty, balanced test sets, post-cutoff dates, sealed outcomes, and a three-option menu where declining is explicitly on the table. 07:56 — The commitment ladder, and the costume test Commitment climbs 6.5% to 14.8% to 54% as the panel gets richer — then drops to 3.5% when the panel belongs to the wrong company. 09:55 — Swap the numbers, watch nothing move The scrambled-panel and fully-fabricated arms land at 38.3% and 36.8% against a real-panel 37.6%, plus the equivalence test and its post-hoc margin. 11:54 — A dial, not a switch — and only three models Commitment scales with panel density (0.0% to 50%), but the pooled headline hides that three Claude models carry nearly the whole effect and scale doesn't predict who fails. 13:54 — The sensor works, the wire isn't connected Four explanations ruled out: models read the panel at 99.9% accuracy, barely change stated belief, and label the question irreducible 90% of the time — the judgment simply never reaches the decision. 16:48 — A compass that reliably points south The 257 committed forecasts score Brier 0.281 (worse than a flat 50%) and AUROC 0.346 — systematically inverted — while the aggregate mean of 49.1% would pass a standard calibration audit. 17:52 — 540 dice problems, zero stocks Fine-tuning a 3B model on 540 synthetic dice, coin and jar examples drops commitment to 0.0% on stock cases and transfers to crypto, sports, and weather. 19:51 — The format that switches the gate off The trained gate survives two prompt framings and collapses under rigid wrapper tags — zero of 288 responses contain reasoning, and one variant commits on 48 of 48 unknowable items. 21:50 — Licensing the decision, not informing it The closing argument that belief calibration and action calibration come apart, plus the open question of whether the gate belongs in the model or the scaffolding. Recommended Reading Language Models (Mostly) Know What They Know — The 2022 result the episode explicitly positions itself against — the source of the 'audit the stated confidence' framing that this paper argues sails right past action-level failure. Towards Understanding Sycophancy in Language Models — The closest existing account of models being swayed by the social packaging of input rather than its content, which is the mechanism the fabricated-dashboard experiment isolates. Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting — Documents models producing fluent, confident reasoning driven by prompt features they never acknowledge — the same dissociation seen in the transcript reasoning correctly over numbers that describe nothing. Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models — Direct empirical backing for the episode's section-eight warning that rigid structured-output formats suppress reasoning — the exact condition under which the trained refusal gate collapsed.
- Transcript