Skip to content
Artwork for LessWrong (30+ Karma)
LessWrong (30+ Karma) · September 21 · 10 min

“Mech Interp is a Verifiable Task” by Logan Riggs

If we think parts of MLP0-MLP3 are computing [a sorting algorithm], we can replace those parts with [a sorting algorithm] and check reconstruction loss. However, reconstruction loss is not enough. Suppose we replace MLP0 with two things: Its mean activation - simple, but poor reconstruction MLP0 - perfect reconstruction, but no reduction in complexity We can visualize this as a pareto frontier trading off reconstruction with "simplicity". Ideally we achieve perfect reconstruction with perfect simplicity. For more intuition on the pareto frontier, we could have an MLP that clusters all inputs in two clusters: "early positions" and "late positions", which would be slightly more complex than the mean. We can make this an RLVR environment, if only we could clearly... Define "Simplicity" Defining simplicity has been complex. But what do we want from a perfectly decomposed model? If we've "perfectly decomposed" a model, then I'd expect ideal circuits to fall out, with "ideal" meaning: Help predict OOD behavior Given an [addition] circuit, we can know which types of inputs it'll succeed & fail on (and why) Be extractable & minimal The smallest part of the model that does [addition] Be removable w/ minimal harm to [...] --- Outline: (01:16) Define "Simplicity" (03:16) Tensor Networks Don't Solve This Issue (03:54) Red Herrings of Simplicity (05:42) How to Gain Tractability (06:09) Tract 1: Death Success by 1000 Circuits (07:02) Tract 2: QK OV Circuits but for Everything (08:45) Tract 3: Interpreting Small Models (09:50) Big if True The original text contained 6 footnotes which were omitted from this narration. --- First published: September 21st, 2026 Source: https://www.lesswrong.com/posts/QxHuKtfGfzn8uokoR/mech-interp-is-a-verifiable-task --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

0:00-10:58

transcript

No transcript — this publisher did not publish one.

show notes

If we think parts of MLP0-MLP3 are computing [a sorting algorithm], we can replace those parts with [a sorting algorithm] and check reconstruction loss. However, reconstruction loss is not enough.

Suppose we replace MLP0 with two things:

  1. Its mean activation - simple, but poor reconstruction
  2. MLP0 - perfect reconstruction, but no reduction in complexity

We can visualize this as a pareto frontier trading off reconstruction with "simplicity". Ideally we achieve perfect reconstruction with perfect simplicity. For more intuition on the pareto frontier, we could have an MLP that clusters all inputs in two clusters: "early positions" and "late positions", which would be slightly more complex than the mean.

We can make this an RLVR environment, if only we could clearly...

Define "Simplicity"

Defining simplicity has been complex. But what do we want from a perfectly decomposed model? If we've "perfectly decomposed" a model, then I'd expect ideal circuits to fall out, with "ideal" meaning:

  1. Help predict OOD behavior
    1. Given an [addition] circuit, we can know which types of inputs it'll succeed & fail on (and why)
  2. Be extractable & minimal
    1. The smallest part of the model that does [addition]
  3. Be removable w/ minimal harm to [...]

---

Outline:

(01:16) Define "Simplicity"

(03:16) Tensor Networks Don't Solve This Issue

(03:54) Red Herrings of Simplicity

(05:42) How to Gain Tractability

(06:09) Tract 1: Death Success by 1000 Circuits

(07:02) Tract 2: QK OV Circuits but for Everything

(08:45) Tract 3: Interpreting Small Models

(09:50) Big if True

The original text contained 6 footnotes which were omitted from this narration.

---

First published:
September 21st, 2026

Source:
https://www.lesswrong.com/posts/QxHuKtfGfzn8uokoR/mech-interp-is-a-verifiable-task

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Curved path between two dots, star at top right.

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

links4