Skip to content
Artwork for Clown Cast
Clown Cast · August 24 · 20 min

The Model That Couldn't Tell Time

How do video models understand when things happen? In this episode, we crack open a model that could describe every pixel and every action in a video—but completely failed at measuring duration. The failure becomes our roadmap: by watching how it breaks, we discover exactly how temporal understanding is encoded in the architecture. A lesson in useful failure. 00:00 - Recap: frames as patch tokens and the position problem 05:20 - The model that could see everything but not tell time 11:45 - Inside the mechanism: time markers and temporal encoding 17:30 - Why this failure reveals how duration gets computed --- Sources & further reading: • game-senser-fable repo — docs/retrospectives/2026-08-23-detection-is-solved.md • §2.1 (audio ablation), §2.2 (uniform tilings, both excuses removed, honest timestamps) • §2.3 (93.8% classification, the bad kill, confidence at 0.98), §4 (shared phantom • locations) • game-senser-fable repo — docs/touch-arc/v91/TEST-C-RERUN-RESULT.md (three prompt arms • the control reproducing, "a replay shows volleyball being played", the fraction arm • closing the graded-output idea) • game-senser-fable repo — src/geminirally/promptssetoptics.py (clip-local coordinates and • the 68.6% figure; gap-scan +13.2pp, 70.6 → 83.8) • game-senser-fable repo — JOURNAL.md 2026-08-24 (stride table, stage-2 boundary MAE) and • 2026-08-23 overnight (the 1.7/1.7/1.9 vs 3.0s gap-fill trace) • game-senser-fable repo — docs/touch-arc/v91/RUN-1-MODEL-MIGRATION.md (thinking-budget • hazard, hallucinated out-of-range windows) • game-senser-fable repo — data/qwen-keystone/probeprocessorshape.py (the literal • markers decoded out of the model's own input) This podcast episode was fully generated by AI — research, script, voices, and production. Built with Claude, Piper TTS, and automated pipeline tooling.

0:00-20:19

transcript

No transcript — this publisher did not publish one.

show notes

How do video models understand when things happen? In this episode, we crack open a model that could describe every pixel and every action in a video—but completely failed at measuring duration. The failure becomes our roadmap: by watching how it breaks, we discover exactly how temporal understanding is encoded in the architecture. A lesson in useful failure. 00:00 - Recap: frames as patch tokens and the position problem 05:20 - The model that could see everything but not tell time 11:45 - Inside the mechanism: time markers and temporal encoding 17:30 - Why this failure reveals how duration gets computed --- Sources & further reading: • game-senser-fable repo — docs/retrospectives/2026-08-23-detection-is-solved.md • §2.1 (audio ablation), §2.2 (uniform tilings, both excuses removed, honest timestamps) • §2.3 (93.8% classification, the bad kill, confidence at 0.98), §4 (shared phantom • locations) • game-senser-fable repo — docs/touch-arc/v91/TEST-C-RERUN-RESULT.md (three prompt arms • the control reproducing, "a replay shows volleyball being played", the fraction arm • closing the graded-output idea) • game-senser-fable repo — src/geminirally/promptssetoptics.py (clip-local coordinates and • the 68.6% figure; gap-scan +13.2pp, 70.6 → 83.8) • game-senser-fable repo — JOURNAL.md 2026-08-24 (stride table, stage-2 boundary MAE) and • 2026-08-23 overnight (the 1.7/1.7/1.9 vs 3.0s gap-fill trace) • game-senser-fable repo — docs/touch-arc/v91/RUN-1-MODEL-MIGRATION.md (thinking-budget • hazard, hallucinated out-of-range windows) • game-senser-fable repo — data/qwen-keystone/probeprocessorshape.py (the literal • markers decoded out of the model's own input) This podcast episode was fully generated by AI — research, script, voices, and production. Built with Claude, Piper TTS, and automated pipeline tooling.