“Mech Interp is a Verifiable Task” by Logan Riggs
transcript
show notes
If we think parts of MLP0-MLP3 are computing [a sorting algorithm], we can replace those parts with [a sorting algorithm] and check reconstruction loss. However, reconstruction loss is not enough.
Suppose we replace MLP0 with two things:
- Its mean activation - simple, but poor reconstruction
- MLP0 - perfect reconstruction, but no reduction in complexity
We can visualize this as a pareto frontier trading off reconstruction with "simplicity". Ideally we achieve perfect reconstruction with perfect simplicity. For more intuition on the pareto frontier, we could have an MLP that clusters all inputs in two clusters: "early positions" and "late positions", which would be slightly more complex than the mean.
We can make this an RLVR environment, if only we could clearly...
Define "Simplicity"
Defining simplicity has been complex. But what do we want from a perfectly decomposed model? If we've "perfectly decomposed" a model, then I'd expect ideal circuits to fall out, with "ideal" meaning:
- Help predict OOD behavior
- Given an [addition] circuit, we can know which types of inputs it'll succeed & fail on (and why)
- Be extractable & minimal
- The smallest part of the model that does [addition]
- Be removable w/ minimal harm to [...]
---
Outline:
(01:16) Define "Simplicity"
(03:16) Tensor Networks Don't Solve This Issue
(03:54) Red Herrings of Simplicity
(05:42) How to Gain Tractability
(06:09) Tract 1: Death Success by 1000 Circuits
(07:02) Tract 2: QK OV Circuits but for Everything
(08:45) Tract 3: Interpreting Small Models
(09:50) Big if True
The original text contained 6 footnotes which were omitted from this narration.
---
First published:
September 21st, 2026
Source:
https://www.lesswrong.com/posts/QxHuKtfGfzn8uokoR/mech-interp-is-a-verifiable-task
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.