
AI Article Readings · Tuesday · 36 min
God Help Us, Let’s Try To Learn About Mechanistic Interpretability Techniques - By Scott Alexander
0:00-36:04
transcript
show notes
Scott Alexander works through the tools researchers use to investigate what happens inside language models, from linear probes and sparse autoencoders to activation verbalizers and the Jacobian lens. A technical tour of how these methods work, what they reveal, and the questions they leave open.
* 00:00 Introduction
* 04:03 Linear Probes
* 11:46 Sparse Autoencoders
* 14:56 Activation Verbalizers
* 18:55 Natural Language Autoencoders
* 19:28 Emotion Vectors
* 26:01 The Jacobian Lens
* 32:19 You Go To War With The Weapons You Have, Not The Weapons You Want
* 35:27 Outro
Get full access to Askwho Casts AI at askwhocastsai.substack.com/subscribe
links2





