
The Agent Can Rewrite Itself. So Who Controls It?
transcript
show notes
In this episode of Attention Span, I look inside Muse Secure VM and the security architecture surrounding the agent.
We get into Sentinel, the separate system that decides what Muse is allowed to do; why Muse can use your accounts without ever seeing the real credentials; how “tainted egress” changes permissions after a process touches private data; and why reading an email can give an agent considerably more power than “read access” suggests.
And somehow, all of this takes us back to computer-security ideas from 1975.
The larger question is one we are going to encounter everywhere as agents become more capable:
How much freedom can we give an agent to discover new ways of doing things without allowing it to expand its own authority?
*Watch it*, and tell me where you would draw that boundary.
*IN THIS EPISODE*
→ What “rogue agent” actually means technically
→ How Muse Secure VM isolates the agent
→ Why Sentinel sits outside the environment Muse can modify
→ How an agent can write a new tool without granting that tool permission
→ How surrogate tokens keep real credentials away from the model
→ What “tainted egress” means and why permission may need memory
→ Why “read my email” can unlock much more than email
→ Least privilege and complete mediation, 51 years later
→ Where Muse's security architecture still has unresolved problems
→ Secure from whom? The agent, other users, or Meta itself
→ DeepSeek Harness + Muse: capability can become fluid; authority cannot
*META MUSE*
→ Introducing Muse Meta's product announcement and overview of Muse Secure VM. https://about.fb.com/news/2026/09/introducing-muse-personal-ai-agent/
→ How We Built Safety Into Muse The technical deep dive. This is where Meta explains Sentinel, Linux isolation, surrogate tokens, credential insertion, eBPF-based taint tracking and network egress controls. https://research.meta.ai/blog/security-and-safety-for-ai-agents-our-approach-with-muse
*THE SECURITY IDEAS BEHIND IT*
→ Melanie Mitchell: Misleading Metaphors and Real Risks A useful grounding discussion of what we actually mean when we say an agent “escaped” or “went rogue.” https://aiguide.substack.com/p/misleading-metaphors-and-real-risks
→ Saltzer & Schroeder: The Protection of Information in Computer Systems The 1975 paper behind principles such as least privilege and complete mediation that suddenly look very current again. https://web.mit.edu/Saltzer/www/publications/protection/Basic.html
*RELATED ATTENTION SPAN*
→ Everything Is a Plugin: DeepSeek Harness Our previous episode on agents that can create and modify their own tools. https://www.youtube.com/watch?v=jtyV7O4Pt0s
→ AI Escaped? The OpenAI–Hugging Face Incident What actually happened when cybersecurity agents found routes outside their intended environment. https://www.youtube.com/watch?v=RGeZ2moLkIc
*MORE FROM TURING POST*
→ We Don't Know What They Know Our deeper look at the problem of understanding what increasingly capable models know and how they will use it. https://www.turingpost.com/p/we-don-t-know-what-they-know
→ Turing Post https://www.turingpost.com
→ Instagram @turingpost_tv https://www.instagram.com/turingpost_tv
→ TikTok @turingpost_tv https://www.tiktok.com/@turingpost_tv
→ Interviews: @realturingpost