Claude Code Cast

The Agent Benchmark That Should Scare Managers

May 29 · 19 min · 37.5 MB
0:00-19:24

Streams straight from the publisher. podnod never proxies or re-hosts episode audio.

Agentic coding tools are moving into enterprise workflows, but the week's most useful signal is a benchmark where frontier models still struggle below 50% on real IT tasks. Alex and Sam unpack Microsoft Learn grounding, agent deception, Copilot data leaks, and the practical harness every team should build before handing agents production authority.