Skip to content
Artwork for 脈報
脈報 · Monday · 6 min

AI agent 長任務一直原地打轉?PoS 研究放了監工 (長對話本身不是進度)

AI agent 長任務原地打轉時,PoS 研究放了監工,把待辦拆成還要搞清楚與還要動手做到兩種缺口。12 種測驗與模型組合整體指標都是受測方法最高,總 token 是 5.06 倍。 ⭐ 文章深度讀:AI agent 長任務一直原地打轉?PoS 研究在旁邊放了監工 → https://heymaibao.com/pos-ai-agent-stuck-supervisor-406dd3/ ⚡ 章節重點 越拉越長的對話與監工 00:00 熱杯子的兩次執行 01:25 流水帳和現況表 02:00 空轉的三種樣子與脫困 03:04 PoS 治不了的事 04:10 自己當監工的兩個問題 04:41 帳單、缺口與下一步 05:10 📝 懶人包 ∙ PoS 在 12 種測驗與模型組合裡,整體指標都是受測方法最高。 ∙ 代價:在 GitHub 頁那組比較裡,總 token 是直接用完整紀錄的 5.06 倍。 ∙ PoS 把待辦拆成兩種缺口:還要搞清楚的事和還要動手做到的事。 ∙ 我的觀點:長對話本身不是進度,我們可以自己每隔一段請 AI 寫現況。 📚 參考資料 Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States → https://arxiv.org/html/2610.01415v1 Progression of States | Beyond Memory → https://luoyu100.github.io/projects/progression-of-states/project/ GitHub - luoyu100/PoS: An inference-time framework for long-horizon LLM agents with explicit belief states, consistency validation, and trapping-aware recovery. → https://github.com/luoyu100/PoS 非線智能 NoneLinear - ReLE 評測:中文 AI 大模型能力 → https://github.com/jeinlee1991/chinese-llm-benchmark feat(dashscope): add qwen3.7-plus and qwen3.7-max to the model cost map → https://github.com/BerriAI/litellm/pull/35123 阿里雲百鍊模型價格 → https://www.alibabacloud.com/help/tc/model-studio/model-pricing

0:00-6:59

transcript

No transcript — this publisher did not publish one.

show notes


AI agent 長任務原地打轉時,PoS 研究放了監工,把待辦拆成還要搞清楚與還要動手做到兩種缺口。12 種測驗與模型組合整體指標都是受測方法最高,總 token 是 5.06 倍。

⭐ 文章深度讀:AI agent 長任務一直原地打轉?PoS 研究在旁邊放了監工
→ https://heymaibao.com/pos-ai-agent-stuck-supervisor-406dd3/

⚡ 章節重點
越拉越長的對話與監工 00:00
熱杯子的兩次執行 01:25
流水帳和現況表 02:00
空轉的三種樣子與脫困 03:04
PoS 治不了的事 04:10
自己當監工的兩個問題 04:41
帳單、缺口與下一步 05:10

📝 懶人包
∙ PoS 在 12 種測驗與模型組合裡,整體指標都是受測方法最高。

∙ 代價:在 GitHub 頁那組比較裡,總 token 是直接用完整紀錄的 5.06 倍。

∙ PoS 把待辦拆成兩種缺口:還要搞清楚的事和還要動手做到的事。

∙ 我的觀點:長對話本身不是進度,我們可以自己每隔一段請 AI 寫現況。

📚 參考資料
Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States
→ https://arxiv.org/html/2610.01415v1

Progression of States | Beyond Memory
→ https://luoyu100.github.io/projects/progression-of-states/project/

GitHub - luoyu100/PoS: An inference-time framework for long-horizon LLM agents with explicit belief states, consistency validation, and trapping-aware recovery.
→ https://github.com/luoyu100/PoS

非線智能 NoneLinear - ReLE 評測:中文 AI 大模型能力
→ https://github.com/jeinlee1991/chinese-llm-benchmark

feat(dashscope): add qwen3.7-plus and qwen3.7-max to the model cost map
→ https://github.com/BerriAI/litellm/pull/35123

阿里雲百鍊模型價格
→ https://www.alibabacloud.com/help/tc/model-studio/model-pricing

links7