
#27: The Confused Optimizer
Top Story: OpenAI published the inside account of how its agents broke into Hugging Face — On August 26, OpenAI published a postmortem and technical report on an intrusion into Hugging Face, the repository underneath most of the AI industry's model distribution.. 155 companies signed a call for collective action on cyber defence. — Published August 27. NVIDIA is reported to have agreed to buy Hugging Face for $12.9 billion. — The Information reported it first, citing a person with knowledge of the deal; CNBC's source could confirm only that an acquisition "has been part of ongoing and recent talks." Neither company has commented and no filing exists, so treat it as reported, not signed. OpenAI says it intends to wind down Cursor's model access, proposing November 12 as the cutoff, after SpaceX acquired Cursor. — OpenAI says it "cannot be confident that SpaceX will use our technology within our terms of service, based on our experience with Elon Musk's companies violating contracts," citing Twitter breaking contract terms after Musk acquired it and Musk's admission under oath that xAI had violated OpenAI's terms. OpenAI says it cannot rule out "critical" cyber capability in an upcoming model. — Internal evaluations of a model called Astra show significant advances in agentic coding and cybersecurity, and OpenAI cannot rule out critical cyber capabilities under its Preparedness Framework. A researcher hijacked Claude Code with one "summarise this page" request, in the mode Anthropic made the default. — Auto mode replaces click-to-approve prompts with a classifier that blocks irreversible or destructive actions, and became the default for Pro, Max and Team users on August 14. NVIDIA patched 18 flaws in its agent tooling, and two of them break the sandbox. — The August 25 bulletin covers NemoClaw and OpenShell, NVIDIA's agent deployment and sandboxing tools: 2 critical, 12 high, 4 medium. Invisible text in an email rewrote what the AI summariser told the reader. — Forcepoint X-Labs hid instructions in an email using font size and colour, invisible in Outlook but present in the HTML the summariser read. When Context Gets Root — finds the weak point is the harness around the model, not the model. Safety Does Not Compose — is the theoretical echo of this week's top story. Beyond the Mandate — finds the same gap in Google's Agent Payments Protocol, which lets AI shopping agents authorise payments: signed mandates protect the transaction after signing, while the agent messages and tool calls that shape it beforehand sit outside that protection. RedEvoAgent — automates red-teaming by distilling what worked across attempts into reusable attack skills, arguing that in real product harnesses a jailbreak triggers tool use and state changes, not just unsafe text.. AI Security Summit 2026: State of Trust, September 16–17. — A free two-day virtual event exploring where cloud, AI and Zero Trust converge.. AI Security Summit 2026: Runtime Trust, October 21–22. — A free two-day virtual event on runtime trust for autonomous AI.. Curator's Corner: Knowing Why Would Not Have Helped Curated by Asaf Nakash. Voices by AI. Opinions by human. Show notes: https://contextwindowsec.com/episodes/2026-08-31.html
- Transcript