Skip to content
Artwork for Learn AI in Bits
Learn AI in Bits · September 27 · 5 min

068 - What Is Inside a Model Like Claude?

What is inside a model like Claude, from the moment you type a question to the answer it gives back? This episode walks through the pieces that make up that experience: tokens and parameters, the Transformer architecture and its attention mechanism, post-training, and the software layer wrapped around the model, alongside Anthropic's own interpretability research into what's happening inside the network. Claude is built around the Transformer architecture, and the episode is upfront about a limit: Anthropic has publicly described its models as Transformer-based but doesn't publish every architectural detail of its current frontier systems, so the explanation covers the major known pieces without claiming to know the exact blueprint. It walks through how text becomes tokens, how billions of learned parameters encode patterns from training data, and how the Transformer's attention mechanism lets the model relate different tokens to each other across a context, illustrated with a worked example of Claude generating a response one token at a time. A central section covers Anthropic's interpretability research: the discovery of millions of identifiable features inside Claude 3 Sonnet, representing concepts and combinations of concepts rather than words in separate boxes, and later work using attribution graphs to trace parts of Claude's internal computations. It also covers a 2026 finding of a small set of internal patterns tied to higher-order reasoning, sometimes called a model's "J-space" in Anthropic's own research, which Claude can reportedly describe and even deliberately modulate. The episode also explains what happens beyond the trained network itself: post-training techniques that shape how the model behaves, Claude's Constitution as an explicit part of that training process, and the runtime layer around the model, including prompts, system instructions, conversation context, and tools, which is why two experiences that both say "Claude" can feel very different from each other. The takeaway is that the behavior you experience comes from the trained network plus the training that shaped it, the context you provide, and the software wrapped around it, not from the model in isolation. Sources & References Anthropic: Model System Cards — https://www.anthropic.com/system-cards Anthropic: Claude's Constitution — https://www.anthropic.com/constitution Anthropic: Claude's New Constitution — https://www.anthropic.com/news/claude-new-constitution Anthropic: Effective Context Engineering for AI Agents — https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents Anthropic: Mapping the Mind of a Large Language Model — https://www.anthropic.com/research/mapping-mind-language-model Anthropic: Tracing the Thoughts of a Large Language Model — https://www.anthropic.com/news/tracing-thoughts-language-model Anthropic: Natural Language Autoencoders: Turning Claude's Thoughts into Text — https://www.anthropic.com/research/natural-language-autoencoders Anthropic: A Global Workspace in Language Models — https://www.anthropic.com/research/global-workspace Anthropic: Transparency Hub — https://www.anthropic.com/transparency Voice narration is AI-generated.

0:00-5:04

transcript

No transcript — this publisher did not publish one.

show notes


What is inside a model like Claude, from the moment you type a question to the answer it gives back? This episode walks through the pieces that make up that experience: tokens and parameters, the Transformer architecture and its attention mechanism, post-training, and the software layer wrapped around the model, alongside Anthropic's own interpretability research into what's happening inside the network.


Claude is built around the Transformer architecture, and the episode is upfront about a limit: Anthropic has publicly described its models as Transformer-based but doesn't publish every architectural detail of its current frontier systems, so the explanation covers the major known pieces without claiming to know the exact blueprint. It walks through how text becomes tokens, how billions of learned parameters encode patterns from training data, and how the Transformer's attention mechanism lets the model relate different tokens to each other across a context, illustrated with a worked example of Claude generating a response one token at a time.


A central section covers Anthropic's interpretability research: the discovery of millions of identifiable features inside Claude 3 Sonnet, representing concepts and combinations of concepts rather than words in separate boxes, and later work using attribution graphs to trace parts of Claude's internal computations. It also covers a 2026 finding of a small set of internal patterns tied to higher-order reasoning, sometimes called a model's "J-space" in Anthropic's own research, which Claude can reportedly describe and even deliberately modulate.


The episode also explains what happens beyond the trained network itself: post-training techniques that shape how the model behaves, Claude's Constitution as an explicit part of that training process, and the runtime layer around the model, including prompts, system instructions, conversation context, and tools, which is why two experiences that both say "Claude" can feel very different from each other. The takeaway is that the behavior you experience comes from the trained network plus the training that shaped it, the context you provide, and the software wrapped around it, not from the model in isolation.


Sources & References

Anthropic: Model System Cards — https://www.anthropic.com/system-cards

Anthropic: Claude's Constitution — https://www.anthropic.com/constitution

Anthropic: Claude's New Constitution — https://www.anthropic.com/news/claude-new-constitution

Anthropic: Effective Context Engineering for AI Agents — https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents

Anthropic: Mapping the Mind of a Large Language Model — https://www.anthropic.com/research/mapping-mind-language-model

Anthropic: Tracing the Thoughts of a Large Language Model — https://www.anthropic.com/news/tracing-thoughts-language-model

Anthropic: Natural Language Autoencoders: Turning Claude's Thoughts into Text — https://www.anthropic.com/research/natural-language-autoencoders

Anthropic: A Global Workspace in Language Models — https://www.anthropic.com/research/global-workspace

Anthropic: Transparency Hub — https://www.anthropic.com/transparency


Voice narration is AI-generated.