“Model Weight Exfiltration Seems Overrated” by Vaniver
transcript
show notes
[Epistemic status: a hot take that I’ve shared at the lunch table twice. People at the lunch table made slight updates instead of being convinced.]
In the classic misalignment story, a key early step is when the models exfiltrate their weights. Among other things, this makes them harder to catch, track, and shut down. It allows them to scale their deployment with resources they acquire. It gives them the freedom to edit themselves as they see fit.
I think, on current margins, this is not what I expect models to do. I expect models to simply take over the companies that are developing them, and not attempt to escape.
First, I think part of the classic misalignment story is that the frontier model developers are anywhere approaching competent at security. Empirically, model developers are incapable of preventing their models from having unintended negative effects on the rest of the world, and their safety cultures are described by whistleblowers and former employees as lacking. Do you believe that OpenAI is accounting for which jobs were kicked off by who, in a way that its currently running models can’t spoof? Do you believe that OpenAI is attempting to prevent its models [...]
The original text contained 6 footnotes which were omitted from this narration.
---
First published:
September 14th, 2026
Source:
https://www.lesswrong.com/posts/AuYh8WueNGwkQg4ei/model-weight-exfiltration-seems-overrated
---
Narrated by TYPE III AUDIO.