Skip to content
Artwork for Claude Code Cast
Claude Code Cast · July 3 · 18 min

Your AI Coding Benchmarks Are Lying To You

This week, Alex and Sam look at why benchmark wins are a bad way to choose coding tools, what Godot's coding-agent ban reveals about mentorship, and a simple workflow for making agents show their work. If your team is still asking "which model scored highest?", this episode gives you a better test.

0:00-18:34

transcript

No transcript — this publisher did not publish one.

show notes

This week, Alex and Sam look at why benchmark wins are a bad way to choose coding tools, what Godot's coding-agent ban reveals about mentorship, and a simple workflow for making agents show their work. If your team is still asking "which model scored highest?", this episode gives you a better test.