Skip to content
Artwork for Machine Learning Street Talk (MLST)
Machine Learning Street Talk (MLST) · Monday · 2 hr 14 min

Why Scaling Prediction Cannot Create Intelligence - Alexander Mattick

Alexander Mattick is a researcher at Fraunhofer IIS and a PhD researcher at the University of Technology Nuremberg (UTN), and a regular on Yannic Kilcher's Discord. He first came on MLST in 2022, after helping research the Yann LeCun and Randall Balestriero episode on interpolation. SPONSOR: --- Cyber Fund built the Monastery to help founders ship products that were impossible a year ago. Applications for Batch 1 are now open. Apply now: https://cyber.fund --- Alexander treats inference as the thread running through modern machine learning: once you have a model, what does it cost to get an answer out of it? He works through Monte Carlo, GFlowNets, energy-based models, diffusion, normalising flows and flow matching, with four short explainers he recorded himself. He is blunt about energy-based models: you can sample from them in principle, but it is rarely worth the compute. JEPA and "world model", he says, are closer to branding than to technical categories. Next: theories of deep learning, none of which he thinks predicts enough yet to guide practice, then reinforcement learning. --- 0:00 Cold open: information is expensive 0:51 Welcome back, Alexander Mattic 2:08 Alexander's research background 2:50 Inference: densities, sampling and Monte Carlo 6:42 GFlowNets, energy functions and MCMC 9:45 Explainer: energy-based models 11:03 Why model a density at all? 17:30 From learned energies to flow matching 25:08 Explainers: diffusion and normalising flows 28:33 Are energy-based models generative? 33:22 JEPA, contrastive learning and collapse 41:13 Why non-language modalities need flows 44:51 Inference as search: branch and bound 49:43 Q-learning and delayed consequences 55:14 Flow matching, optimal transport, Fokker-Planck 1:00:03 Explainer: flow matching 1:01:49 AlphaFold, latents and scale versus architecture 1:07:52 Two families of deep learning theory 1:15:04 What a good theory would predict 1:23:53 The manifold hypothesis and compression 1:28:25 Is reward enough? 1:32:01 Control theory versus reinforcement learning 1:37:22 The Bitter Lesson and expensive information 1:42:08 Constrained RL: the constrained MDP toolbox 1:50:12 Creativity as constrained search 1:55:44 Reality is protean: when abstractions hold 2:00:32 What is a world model? 2:04:38 Prediction is not control 2:08:13 Robot demos, MPC and reliability --- REFERENCES: [6:55] GFlowNets (Bengio et al., 2021) https://arxiv.org/abs/2106.04399 [38:46] Contrastive Self-Supervised Learning (Anand, 2020) https://ankeshanand.com/blog/2020/01/26/contrative-self-supervised-learning.html [38:56] LeJEPA (Balestriero and LeCun, 2025) https://arxiv.org/abs/2511.08544v3 [47:10] RL for Node Selection in Branch-and-Bound (Mattick) https://openreview.net/forum?id=0ez68a5UqI [56:20] Flow Matching for Generative Modeling https://arxiv.org/abs/2210.02747v2 [1:12:41] Disentangling feature and lazy training in deep neural networks https://arxiv.org/abs/1906.08034v4 [1:31:05] Reward is enough (Silver) https://doi.org/10.1016/j.artint.2021.103535 [1:35:12] Learning ReLU networks to high uniform accuracy is intractable (Berner et al.) https://arxiv.org/abs/2205.13531v2 [1:40:20] Dota 2 with Large Scale Deep RL https://arxiv.org/abs/1912.06680v1 [1:45:41] Constrained Update Projection for Safe Policy Optimization (Yang et al., 2022) https://arxiv.org/abs/2209.07089 [1:46:11] SafeMPO (ICLR 2026) https://openreview.net/forum?id=1m0EU6QXj6 [1:50:17] Why Creativity Cannot Be Interpolated https://archive.mlst.ai/paper/why-creativity-cannot-be-interpolated/ [1:51:39] Invalid Action Masking (Huang and Ontañón) https://arxiv.org/abs/2006.14171 [2:00:04] Probability Theory: The Logic of Science (Jaynes, 2003) https://www.cambridge.org/core/books/probability-theory/9CA08E224FF30123304E6D8935CF1A99 [2:01:53] Training Agents Inside of Scalable World Models (Hafner et al., 2025) https://arxiv.org/abs/2509.24527v1 [2:03:43] World Models (Ha and Schmidhuber, 2018) https://arxiv.org/abs/1803.10122v4

0:00-2:14:20

transcript

No transcript — this publisher did not publish one.

show notes

Alexander Mattick is a researcher at Fraunhofer IIS and a PhD researcher at the University of Technology Nuremberg (UTN), and a regular on Yannic Kilcher's Discord. He first came on MLST in 2022, after helping research the Yann LeCun and Randall Balestriero episode on interpolation.


SPONSOR:

---

Cyber Fund built the Monastery to help founders ship products that were impossible a year ago. Applications for Batch 1 are now open.

Apply now: https://cyber.fund

---


Alexander treats inference as the thread running through modern machine learning: once you have a model, what does it cost to get an answer out of it? He works through Monte Carlo, GFlowNets, energy-based models, diffusion, normalising flows and flow matching, with four short explainers he recorded himself. He is blunt about energy-based models: you can sample from them in principle, but it is rarely worth the compute. JEPA and "world model", he says, are closer to branding than to technical categories.


Next: theories of deep learning, none of which he thinks predicts enough yet to guide practice, then reinforcement learning.


---

0:00 Cold open: information is expensive

0:51 Welcome back, Alexander Mattic

2:08 Alexander's research background

2:50 Inference: densities, sampling and Monte Carlo

6:42 GFlowNets, energy functions and MCMC

9:45 Explainer: energy-based models

11:03 Why model a density at all?

17:30 From learned energies to flow matching

25:08 Explainers: diffusion and normalising flows

28:33 Are energy-based models generative?

33:22 JEPA, contrastive learning and collapse

41:13 Why non-language modalities need flows

44:51 Inference as search: branch and bound

49:43 Q-learning and delayed consequences

55:14 Flow matching, optimal transport, Fokker-Planck

1:00:03 Explainer: flow matching

1:01:49 AlphaFold, latents and scale versus architecture

1:07:52 Two families of deep learning theory

1:15:04 What a good theory would predict

1:23:53 The manifold hypothesis and compression

1:28:25 Is reward enough?

1:32:01 Control theory versus reinforcement learning

1:37:22 The Bitter Lesson and expensive information

1:42:08 Constrained RL: the constrained MDP toolbox

1:50:12 Creativity as constrained search

1:55:44 Reality is protean: when abstractions hold

2:00:32 What is a world model?

2:04:38 Prediction is not control

2:08:13 Robot demos, MPC and reliability


---

REFERENCES:

[6:55] GFlowNets (Bengio et al., 2021)

https://arxiv.org/abs/2106.04399

[38:46] Contrastive Self-Supervised Learning (Anand, 2020)

https://ankeshanand.com/blog/2020/01/26/contrative-self-supervised-learning.html

[38:56] LeJEPA (Balestriero and LeCun, 2025)

https://arxiv.org/abs/2511.08544v3

[47:10] RL for Node Selection in Branch-and-Bound (Mattick)

https://openreview.net/forum?id=0ez68a5UqI

[56:20] Flow Matching for Generative Modeling

https://arxiv.org/abs/2210.02747v2

[1:12:41] Disentangling feature and lazy training in deep neural networks

https://arxiv.org/abs/1906.08034v4

[1:31:05] Reward is enough (Silver)

https://doi.org/10.1016/j.artint.2021.103535

[1:35:12] Learning ReLU networks to high uniform accuracy is intractable (Berner et al.)

https://arxiv.org/abs/2205.13531v2

[1:40:20] Dota 2 with Large Scale Deep RL

https://arxiv.org/abs/1912.06680v1

[1:45:41] Constrained Update Projection for Safe Policy Optimization (Yang et al., 2022)

https://arxiv.org/abs/2209.07089

[1:46:11] SafeMPO (ICLR 2026)

https://openreview.net/forum?id=1m0EU6QXj6

[1:50:17] Why Creativity Cannot Be Interpolated

https://archive.mlst.ai/paper/why-creativity-cannot-be-interpolated/

[1:51:39] Invalid Action Masking (Huang and Ontañón)

https://arxiv.org/abs/2006.14171

[2:00:04] Probability Theory: The Logic of Science (Jaynes, 2003)

https://www.cambridge.org/core/books/probability-theory/9CA08E224FF30123304E6D8935CF1A99

[2:01:53] Training Agents Inside of Scalable World Models (Hafner et al., 2025)

https://arxiv.org/abs/2509.24527v1

[2:03:43] World Models (Ha and Schmidhuber, 2018)

https://arxiv.org/abs/1803.10122v4