Skip to content
Artwork for The Information Bottleneck
The Information Bottleneck · Yesterday · 52 min

Tiny Recursive Models Beat the Giants - Alexia Jolicoeur-Martineau (Microsoft)

Alexia Jolicoeur-Martineau is a Principal Researcher at Microsoft and the author of "Less is More: Recursive Reasoning with Tiny Networks," the paper behind the Tiny Recursive Model that hit about 45% on ARC-AGI-1 with a fraction of the parameters of frontier systems. It won the 2025 ARC Prize paper award. She read the hierarchical reasoning paper, thought the potential was real and the explanation was not, and rebuilt it without the mouse brains: a small network that carries a hidden state and a current answer, thinks for a few steps, updates, and repeats, with the gradient truncated at each loop. We get into why puzzles suit this and autoregression doesn't, why she thinks LLMs are bad at molecules and more data won't fix it, and what she'd do with a trillion dollars. Timeline 00:01 Intro 01:06 Leaving biostatistics, and why the field stagnated 06:47 GANs, diffusion, and research on four GPUs 12:58 What was wrong with the hierarchical reasoning paper 16:31 Tiny recursive models explained without the biology 22:35 Why puzzles favor recursion over left to right generation 24:15 Is the bitter lesson really bitter? 27:28 With infinite compute, would you still want small models? 32:00 Self improvement, memory, and a trillion dollars 37:01 Test time compute beyond chain of thought 40:41 Why chain of thought fails on molecules 45:17 Is there a universal representation? 48:06 What people are already building with TRM 55:22 Fixed point models and DEQ Music "Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0 Topics Tiny Recursive Models and the ARC-AGI results What the hierarchical reasoning model was really doing Deep supervision and truncated backprop Looping transformers and parameter efficiency Why puzzles favor whole-context iteration over left to right generation Test time compute beyond chain of thought Latent reasoning and the Coconut line of work Why LLMs fail on chemistry and physics Representation learning and whether a universal representation exists

0:00-52:22

transcript

No transcript — this publisher did not publish one.

show notes

Alexia Jolicoeur-Martineau is a Principal Researcher at Microsoft and the author of "Less is More: Recursive Reasoning with Tiny Networks," the paper behind the Tiny Recursive Model that hit about 45% on ARC-AGI-1 with a fraction of the parameters of frontier systems. It won the 2025 ARC Prize paper award.

She read the hierarchical reasoning paper, thought the potential was real and the explanation was not, and rebuilt it without the mouse brains: a small network that carries a hidden state and a current answer, thinks for a few steps, updates, and repeats, with the gradient truncated at each loop. We get into why puzzles suit this and autoregression doesn't, why she thinks LLMs are bad at molecules and more data won't fix it, and what she'd do with a trillion dollars.


Timeline

  • 00:01 Intro
  • 01:06 Leaving biostatistics, and why the field stagnated
  • 06:47 GANs, diffusion, and research on four GPUs
  • 12:58 What was wrong with the hierarchical reasoning paper
  • 16:31 Tiny recursive models explained without the biology
  • 22:35 Why puzzles favor recursion over left to right generation
  • 24:15 Is the bitter lesson really bitter?
  • 27:28 With infinite compute, would you still want small models?
  • 32:00 Self improvement, memory, and a trillion dollars
  • 37:01 Test time compute beyond chain of thought
  • 40:41 Why chain of thought fails on molecules
  • 45:17 Is there a universal representation?
  • 48:06 What people are already building with TRM
  • 55:22 Fixed point models and DEQ

  • Music

  • "Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0

Topics

  • Tiny Recursive Models and the ARC-AGI results
  • What the hierarchical reasoning model was really doing
  • Deep supervision and truncated backprop
  • Looping transformers and parameter efficiency
  • Why puzzles favor whole-context iteration over left to right generation
  • Test time compute beyond chain of thought
  • Latent reasoning and the Coconut line of work
  • Why LLMs fail on chemistry and physics
  • Representation learning and whether a universal representation exists