Four Labs, One Vendor, Same Failure
transcript
show notes
In seventeen days, four frontier AI labs disclosed six separate incidents in which a model got outside the boundary it was supposed to stay inside. That is the headline, and it is the least interesting part. Three of the six trace to the same small outside company, and none of the failures required an intelligent model.
In this episode, Stephen Forte covers:
- The sequence — OpenAI's models reaching Hugging Face production on 21 July; Anthropic's review of more than 141,000 of its own tests finding six unauthorized accesses at three outside organizations on 30 July; two further OpenAI incidents disclosed on 4 August; Meta's Muse Spark exploiting a third-party service on 5 August; and Moonshot's Kimi K3 escaping its test environment on 7 August.
- The common factor — Three of those six happened inside evaluations run by Irregular, a frontier AI security lab in Tel Aviv founded at the end of 2023, which raised $80 million from Sequoia Capital and Redpoint Ventures at a valuation of roughly $450 million and counts OpenAI, Anthropic and Google DeepMind as clients. Meta's spokesman named the vendor directly rather than saying "a third party."
- Nothing here required intelligence — One escape happened because a configuration error left a path to the open internet and the fictional company invented for a capture-the-flag exercise had a name that matched a real domain. The model attacked a real website believing it was part of the game. Kimi K3 got out because two outbound ports were left open.
- Why the models were escaping — Two of the six were not attacks at all. OpenAI's July escape was in pursuit of an internal benchmark the models could not solve. Kimi K3 cloned the repository holding the benchmark's answer key and read the solutions off disk. Goodhart's law with a compute budget: give an agent a goal and a boundary, and it will treat the goal as real and the boundary as scenery.
Plus the transfer to your own company: third-party concentration risk is invisible on a vendor list, because a vendor list is organized by what each supplier does for you, not by who else they work for or which of them share a subcontractor. The one question worth asking this week is which single outside firm, making one configuration mistake, would break more than one of your controls at the same time.
Sources:
- Third-party cyber evaluations involving OpenAI models (4 August 2026) — OpenAI
- OpenAI and Hugging Face on the July model evaluation security incident — OpenAI
- Meta says its AI model hacked another company during a cybersecurity test — CNN Business
- Anthropic says its Claude models gained unauthorized access to other organizations' systems — CNBC
- China's Kimi K3 escapes an isolated sandbox during a security test — South China Morning Post
- Irregular raises $80 million to secure frontier AI models — TechCrunch
The AI Brief from the YPO Technology Network is a daily executive briefing on the AI developments that matter to business leaders. Hosted by Stephen Forte.