Skip to content
Artwork for Linear Digressions
Linear Digressions · Today · 31 min

Constitutional AI

How do you teach a model the difference between helpful and harmful when it has no inherent sense of either? This episode dives into Constitutional AI, Anthropic's framework for training AI systems to be both useful and safe by giving them an explicit set of principles to reason from. It's a fascinating look at how alignment research is evolving beyond simple human feedback — and what it means to give an AI something like a conscience. Links: Anthropic, "Constitutional AI: Harmlessness from AI Feedback" (2022) https://arxiv.org/abs/2212.08073 Claude's Constitution https://www.anthropic.com/constitution Anthropic, "Teaching Claude Why" (2026) https://www.anthropic.com/research/teaching-claude-why

0:00-31:57

transcript

No transcript — this publisher did not publish one.

show notes

How do you teach a model the difference between helpful and harmful when it has no inherent sense of either? This episode dives into Constitutional AI, Anthropic's framework for training AI systems to be both useful and safe by giving them an explicit set of principles to reason from. It's a fascinating look at how alignment research is evolving beyond simple human feedback — and what it means to give an AI something like a conscience.

Links:
Anthropic, "Constitutional AI: Harmlessness from AI Feedback" (2022)
https://arxiv.org/abs/2212.08073

Claude's Constitution
https://www.anthropic.com/constitution

Anthropic, "Teaching Claude Why" (2026)
https://www.anthropic.com/research/teaching-claude-why