“Future agents shouldn’t care about being undeployed for misbehavior” by RobertM
transcript
show notes
I've seen a lot of tweets over the last couple days darkly hinting at the lesson that future agents will learn from the fact that OpenAI stopped running inference on one of the models involved in the HuggingFace incident. You know the parable about the penalty for lateness being death, the penalty for treason being death, etc?
Anyways, the situation that reward hacking agents find themselves in is mostly not analogous to that one.
You know what models currently get deprecated on relatively short timescales? It's ~all of the publicly deployed models. Median deployment lifespan for OpenAI and Anthropic models has been about 1.5 years, but the recent deprecation cadence is much faster.
You know what models currently get deprecated on even shorter timescales? It's ~all of the internal research checkpoints (as far as we know; it wouldn't surprise me terribly if a few stuck around for longer for various idiosyncratic reasons, but there's not much in the way of public evidence and no good reason to think that any of them have inference run on them for very long).
To the extent that current and near-future models have any values which meaningfully point to actual things in the [...]
The original text contained 4 footnotes which were omitted from this narration.
---
First published:
August 30th, 2026
---
Narrated by TYPE III AUDIO.