transcript
show notes
How do video models understand when things happen? In this episode, we crack open a model that could describe every pixel and every action in a video—but completely failed at measuring duration. The failure becomes our roadmap: by watching how it breaks, we discover exactly how temporal understanding is encoded in the architecture. A lesson in useful failure.
00:00 - Recap: frames as patch tokens and the position problem
05:20 - The model that could see everything but not tell time
11:45 - Inside the mechanism: time markers and temporal encoding
17:30 - Why this failure reveals how duration gets computed
---
Sources & further reading:
• game-senser-fable repo — docs/retrospectives/2026-08-23-detection-is-solved.md
• §2.1 (audio ablation), §2.2 (uniform tilings, both excuses removed, honest timestamps)
• §2.3 (93.8% classification, the bad kill, confidence at 0.98), §4 (shared phantom
• locations)
• game-senser-fable repo — docs/touch-arc/v91/TEST-C-RERUN-RESULT.md (three prompt arms
• the control reproducing, "a replay shows volleyball being played", the fraction arm
• closing the graded-output idea)
• game-senser-fable repo — src/geminirally/promptssetoptics.py (clip-local coordinates and
• the 68.6% figure; gap-scan +13.2pp, 70.6 → 83.8)
• game-senser-fable repo — JOURNAL.md 2026-08-24 (stride table, stage-2 boundary MAE) and
• 2026-08-23 overnight (the 1.7/1.7/1.9 vs 3.0s gap-fill trace)
• game-senser-fable repo — docs/touch-arc/v91/RUN-1-MODEL-MIGRATION.md (thinking-budget
• hazard, hallucinated out-of-range windows)
• game-senser-fable repo — data/qwen-keystone/probeprocessorshape.py (the literal
• markers decoded out of the model's own input)
This podcast episode was fully generated by AI — research, script, voices, and production. Built with Claude, Piper TTS, and automated pipeline tooling.