Skip to content
Artwork for AI Article Readings
AI Article Readings · Tuesday · 36 min

God Help Us, Let’s Try To Learn About Mechanistic Interpretability Techniques - By Scott Alexander

Scott Alexander works through the tools researchers use to investigate what happens inside language models, from linear probes and sparse autoencoders to activation verbalizers and the Jacobian lens. A technical tour of how these methods work, what they reveal, and the questions they leave open. * 00:00 Introduction * 04:03 Linear Probes * 11:46 Sparse Autoencoders * 14:56 Activation Verbalizers * 18:55 Natural Language Autoencoders * 19:28 Emotion Vectors * 26:01 The Jacobian Lens * 32:19 You Go To War With The Weapons You Have, Not The Weapons You Want * 35:27 Outro https://open.substack.com/pub/astralcodexten/p/god-help-us-lets-try-to-learn-about?r=67y1h&utm_campaign=post-expanded-share&utm_medium=web Get full access to Askwho Casts AI at askwhocastsai.substack.com/subscribe

0:00-36:04

transcript

No transcript — this publisher did not publish one.

show notes

Scott Alexander works through the tools researchers use to investigate what happens inside language models, from linear probes and sparse autoencoders to activation verbalizers and the Jacobian lens. A technical tour of how these methods work, what they reveal, and the questions they leave open.

* 00:00 Introduction

* 04:03 Linear Probes

* 11:46 Sparse Autoencoders

* 14:56 Activation Verbalizers

* 18:55 Natural Language Autoencoders

* 19:28 Emotion Vectors

* 26:01 The Jacobian Lens

* 32:19 You Go To War With The Weapons You Have, Not The Weapons You Want

* 35:27 Outro

https://open.substack.com/pub/astralcodexten/p/god-help-us-lets-try-to-learn-about?r=67y1h&utm_campaign=post-expanded-share&utm_medium=web



Get full access to Askwho Casts AI at askwhocastsai.substack.com/subscribe
links2