
#22 — From Agent Traces to Better Evals with Nikita Kabardin
An agent can pass its evals and still disappoint customers. Production traces help only when someone turns real successes and corrected failures into tests worth trusting. Andrey Devyatkin, Vladimir Samoylov and Fernando Piviani join Langfuse engineer Nikita Kabardin to discuss golden datasets, human review of evals, reliable pull request workflows, context for autonomous agents, and permission boundaries between customer data and public fixes. What you will learn: Build a golden dataset from successful traces and corrected failures Recognize when an agent improves its scores by changing the grading Move pull request state transitions from prompts into deterministic code Understand why unattended agents need context beyond local Markdown files Keep private customer traces out of public development workflows Maybe try B.O.R.I.S, our context layer for AI agents: https://www.getboris.ai Episode page, show notes and links