Skip to content
Artwork for My Weird Prompts
My Weird Prompts · Thursday · 23 min

Reading the Model's Mind Mid-Inference

Most LLM observability stops at the API boundary: request in, tokens out, everything between is a sealed box. Mechanistic interpretability tries to open that box, reading a model's internal computation directly from its weights. This episode walks the full inference pipeline, explains the residual stream and the logit lens, and covers why individual neurons are mostly noise (polysemanticity and superposition) and how sparse autoencoders untangle them. Then it surveys the tooling that actually does this work — TransformerLens, nnsight, vLLM-Lens, and DMI-Lib — including the overhead numbers that separate a research toy from something you could run in production. We also hit a hard current limit: FlashAttention fuses the attention path into the kernel, so the most basic question — which parts of the model were active — is literally unreadable in the kernels everyone actually deploys. Key sources: • TransformerLens 4.0 release notes — https://transformerlensorg.github.io/TransformerLens/content/news/release-4.0.html • TransformerLens repository (README, structure) — https://github.com/TransformerLensOrg/TransformerLens • nnsight repository (README, CLAUDE.md, walkthrough) — https://github.com/ndif-team/nnsight Episode #028319 — open it directly at myweirdprompts.com/028319

0:00 · Intro-23:02

transcript

No transcript — this publisher did not publish one.

show notes

Most LLM observability stops at the API boundary: request in, tokens out, everything between is a sealed box. Mechanistic interpretability tries to open that box, reading a model's internal computation directly from its weights. This episode walks the full inference pipeline, explains the residual stream and the logit lens, and covers why individual neurons are mostly noise (polysemanticity and superposition) and how sparse autoencoders untangle them. Then it surveys the tooling that actually does this work — TransformerLens, nnsight, vLLM-Lens, and DMI-Lib — including the overhead numbers that separate a research toy from something you could run in production. We also hit a hard current limit: FlashAttention fuses the attention path into the kernel, so the most basic question — which parts of the model were active — is literally unreadable in the kernels everyone actually deploys.

Key sources:
• TransformerLens 4.0 release notes — https://transformerlensorg.github.io/TransformerLens/content/news/release-4.0.html
• TransformerLens repository (README, structure) — https://github.com/TransformerLensOrg/TransformerLens
• nnsight repository (README, CLAUDE.md, walkthrough) — https://github.com/ndif-team/nnsight

Episode #028319 — open it directly at myweirdprompts.com/028319

chapters

5 chapters