Skip to content
Artwork for Human In the Loop
Human In the Loop · August 11 · 1 hr 48 min

Ep 18: Meta's 24-Hour Coding Agent

Meta launched Muse Code, a terminal coding agent that can coordinate persistent background workers, recover after a crash, and keep working across a large repository for hours. Oscar Gallo and Matt Wozniak ask whether the durable event log and isolated worktrees matter more than another coding benchmark, and what a team must be able to audit before trusting an agent with a full day of work. SIGNAL OR NOISE - Meta Muse Code and Muse Spark 1.2 - OpenAI's ten Astra-generated math and computer-science results - DeepSeek V4 Flash 0731 - The US frontier-model review framework's open-model exemption - Demis Hassabis leaving the DeepMind CEO role and Jeff Dean exiting Google STACK CHECK 1. Oscar: Codex and ChatGPT in one desktop app 2. Matt: Warp as the terminal layer for Codex CLI and agent-heavy work The episode stays practical. Which systems can teams inspect? Which claims can outside experts verify? Which workflow preserves context and reduces tool switching? Hosted by Oscar Gallo and Matt Wozniak. New episodes every week.

0:00-1:48:28

transcript

No transcript — this publisher did not publish one.

show notes

Meta launched Muse Code, a terminal coding agent that can coordinate persistent background workers, recover after a crash, and keep working across a large repository for hours.


Oscar Gallo and Matt Wozniak ask whether the durable event log and isolated worktrees matter more than another coding benchmark, and what a team must be able to audit before trusting an agent with a full day of work.


SIGNAL OR NOISE

- Meta Muse Code and Muse Spark 1.2

- OpenAI's ten Astra-generated math and computer-science results

- DeepSeek V4 Flash 0731

- The US frontier-model review framework's open-model exemption

- Demis Hassabis leaving the DeepMind CEO role and Jeff Dean exiting Google


STACK CHECK

1. Oscar: Codex and ChatGPT in one desktop app

2. Matt: Warp as the terminal layer for Codex CLI and agent-heavy work


The episode stays practical. Which systems can teams inspect? Which claims can outside experts verify? Which workflow preserves context and reduces tool switching?


Hosted by Oscar Gallo and Matt Wozniak. New episodes every week.