
transcript
show notes
Anthropic 推出 build-eval 與 hillclimb,讓 Claude 替程式出模型評測考卷再爬分數。爬分時留一組沒看過的題抓過擬合,裁判不該是被測模型,分數卡住先回頭查爛題。
⭐ 文章深度讀:別再只看模型評測分數!Anthropic 讓 Claude 先查考卷
→ https://heymaibao.com/check-the-eval-before-trusting-scores-c093dc/
⚡ 章節重點
分數背後的那張考卷 00:00
出考卷與爬分數的新工具 02:09
看過與沒看過的題 03:01
寫進流程的查考卷 04:59
拿自己的任務比一比 06:43
📝 懶人包
∙ 修掉 bug、換框架後,Opus 4.5 分數跳到 95%
∙ 爬分要分兩組題,看過的漲、沒看過的持平就是過擬合警訊
∙ 客服案例在沒看過的工單上 90.5% 對 78.6%,成本約五分之一
∙ 我的觀點:最值錢的是把別信自己的考卷寫成流程
📚 參考資料
Automating eval design and hillclimbing with Claude
→ https://claude.dev/blog/automating-eval-design-and-hillclimbing/
Demystifying evals for AI agents
→ https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents
skills/skills/claude-api at main · anthropics/skills
→ https://github.com/anthropics/skills/tree/main/skills/claude-api
- https://heymaibao.com/check-the-eval-before-trusting-scores-c093dc/heymaibao.com
- https://claude.dev/blog/automating-eval-design-and-hillclimbing/claude.dev
- https://www.anthropic.com/engineering/demystifying-evals-for-ai-agentsanthropic.com
- https://github.com/anthropics/skills/tree/main/skills/claude-apigithub.com