

AGENTSCOPE: Why Bigger Models Do Not Fix Agent Debugging
When an AI agent fails after dozens of steps, the final error rarely reveals where the problem began. AGENTSCOPE turns long execution traces into structured reasoning-action graphs, then checks them against ten neural invariants covering reasoning, control flow, and tool use. On the new AgentErrata benchmark, it raised exact failure-step localization from 1.32% to 31.35% with GPT-5.1 and more than doubled failure-type accuracy over a direct LLM judge. Yet the best exact localization score remains only 34.98%, and AgentErrata relies on injected, manually verified failures rather than organic production incidents. Inspired by the work of Jiayi Bi, Yanjie Gao, Yuanmin Xie, Liqun Li, Tianyin Xu, Fan Yang, and Mao Yang, this episode was created using Google's NotebookLM. Read the original paper here: https://arxiv.org/abs/2609.02371


















