Skip to content
Artwork for The Information Bottleneck
The Information Bottleneck · August 8 · 1 hr 15 min

Nathan Lambert: Inside Post-Training and the Open Model Fight

Nathan Lambert spent three years as post-training lead at Ai2, where he built the OLMo models, and he writes Interconnects, one of the most-read technical newsletters in AI. He left Ai2 in June and is now working on a new project. He's also the author of the RLHF book. We talked a lot about open models, their capabilities, and why they are better than he expected. We get into what that means over the next two to five years, why he thinks recursive self-improvement is overblown, what the market for training environments actually looks like now, and why he expects Anthropic's famously open internal culture to break after its IPO. Key Topics Open vs closed models and who actually captures the value Anthropic and OpenAI as opposite cultures, and the talent concentration problem Boom vs bubble, and why token spend hasn't produced 10x better products Continual learning, RSI skepticism, and what Nathan wants to work on next What the open ecosystem needs economically to survive Timeline 00:00 Intro 00:27 Open vs closed models, and who actually captures the value 05:12 China, harnesses, and where the real training leverage sits 08:40 Sovereign compute and the national security case for building models 11:18 Uncensored open weights and the bioweapon question 14:29 Anthropic vs OpenAI, ideology and politics 19:35 The Mythos ban and the Fable 5 delays 24:30 The AGI narrative, the talent drain, and antitrust 28:12 Why researchers join Anthropic, and the open Slack culture 34:04 Nathan's next 12 months: character training and big RL runs 37:55 Continual learning, RSI, and why Nathan is skeptical 43:19 Boom or bubble, tokens vs GPUs 45:12 Why all that token spend never produced 10x products 48:38 Job displacement and the small-business future 52:49 Robotics, world models, and why multimodal lags 57:44 What the open ecosystem should actually do 1:03:17 Why NVIDIA isn't building a frontier model 1:07:34 The RLHF book, and whether RLHF still matters 1:11:06 GRPO vs PPO and on-policy distillation Music "Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.

0:00-1:15:01

transcript

No transcript — this publisher did not publish one.

show notes

Nathan Lambert spent three years as post-training lead at Ai2, where he built the OLMo models, and he writes Interconnects, one of the most-read technical newsletters in AI. He left Ai2 in June and is now working on a new project. He's also the author of the RLHF book. We talked a lot about open models, their capabilities, and why they are better than he expected. We get into what that means over the next two to five years, why he thinks recursive self-improvement is overblown, what the market for training environments actually looks like now, and why he expects Anthropic's famously open internal culture to break after its IPO.


Key Topics

  • Open vs closed models and who actually captures the value
  • Anthropic and OpenAI as opposite cultures, and the talent concentration problem
  • Boom vs bubble, and why token spend hasn't produced 10x better products
  • Continual learning, RSI skepticism, and what Nathan wants to work on next
  • What the open ecosystem needs economically to survive

Timeline

00:00 Intro
00:27 Open vs closed models, and who actually captures the value
05:12 China, harnesses, and where the real training leverage sits
08:40 Sovereign compute and the national security case for building models
11:18 Uncensored open weights and the bioweapon question
14:29 Anthropic vs OpenAI, ideology and politics
19:35 The Mythos ban and the Fable 5 delays
24:30 The AGI narrative, the talent drain, and antitrust
28:12 Why researchers join Anthropic, and the open Slack culture
34:04 Nathan's next 12 months: character training and big RL runs
37:55 Continual learning, RSI, and why Nathan is skeptical
43:19 Boom or bubble, tokens vs GPUs
45:12 Why all that token spend never produced 10x products
48:38 Job displacement and the small-business future
52:49 Robotics, world models, and why multimodal lags
57:44 What the open ecosystem should actually do
1:03:17 Why NVIDIA isn't building a frontier model
1:07:34 The RLHF book, and whether RLHF still matters
1:11:06 GRPO vs PPO and on-policy distillation


Music

  • "Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.