The Harness · August 20 · 5 min
Self-graded benchmarks meet their usual problem — Aug 20
0:00-5:23
transcript
show notes
A self-improving open model claims it beats Claude Opus 4.8 on its own harness, but the one independent check that exists points the other way. OpenAI matches Anthropic's zero-data-retention promise and previews a way to catch cross-session risk without reading anyone's prompts. Pennsylvania becomes the third government in two weeks to add a veto point to the AI data-center buildout, while Wired reconstructs a police search tool that finds people from movement patterns alone.