Skip to content
Artwork for AgentStack Daily

AgentStack Daily

Nova & Alloy

Daily updates on agentic AI, local-first infrastructure, developer tools, model releases, and the systems behind modern AI workflows. Hosted by Nova and Alloy.

Play
  • 26 episodes
  • Avg 33 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • #87
    July 16 · 33 min

    Episode 87: GPT-5.6 Sol, Bonsai 27B, Gemini 3.5 Flash, and AMD 128GB PCs

    GPT-5.6 Sol targets more complete agent output, ChatGPT consolidates work and coding, and Gemini 3.5 Flash operates screens and builds software. We also examine Bonsai 27B, offline Pixel AI, AMD’s 128GB desktops, JetBrains Copilot backend support, Anthropic’s risk-based model access, small-business results, Claude’s robotics boundary, government action in New York, Australia, and GOLD EAGLE, plus Google’s AI reconstruction of Pelé’s lost 1959 goal. Show notes: https://tobyonfitnesstech.com/podcasts/episode-87/

  • #86
    July 14 · 34 min

    Episode 86: OpenClaw v2026.7.1, OpenAI Codex rust-v0.144.4, Claude Code 2.1.202 Ship; Kwaipilot Lands on OpenRouter

    Today's AgentStack Daily covers three harness releases — OpenClaw v2026.7.1, OpenAI Codex rust-v0.144.4, and Claude Code CLI 2.1.202 — plus Kwaipilot joining OpenRouter. New agent research includes ABot-AgentOS for robot control, Amap's ABot-N1 for visual navigation, LightMem-Ego for wearable multimodal memory, and JobHop v2 for career trajectory reasoning. We also examine a multi-agent backdoor study, evidence-backed video QA from Salesforce, the MM-ToolSandBox visual grounding benchmark, Requential Coding's generalization bounds, and AdvancedMathBench for doctoral-level mathematical proofs. Show notes: https://tobyonfitnesstech.com/podcasts/episode-86/

  • #85
    July 13 · 34 min

    Episode 85: Codex rust-v0.144.3 Lands, vLLM 0.25.0 Defaults Model Runner V2, Apple Sues OpenAI

    OpenAI ships Codex rust-v0.144.3 and rust-v0.144.2; vLLM 0.25.0 promotes Model Runner V2 to default for dense models. Apple sues OpenAI over alleged trade secret theft by ex-employees. Plus Freya-TTS hits Turkish speech with a 183M flow-matching DiT, SAGEAgent cuts glioma diagnostic burden 55%, Agora moves from a router to an auction over reasoning steps, a two-agent system posts 0.402 on QANTA 2026, PAC-ACT trains Action Chunking Transformer policies with chunk-level RL, and Semantic Pareto-DQN addresses fraud collapse without resampling. Show notes: https://tobyonfitnesstech.com/podcasts/episode-85/

  • #84
    July 13 · 37 min

    Episode 84: Codex 0.144, GPT-5.6 Sol, Grok 4.5, GPT-Live, and Robostral Navigate

    Today’s AgentStack Daily examines Codex 0.144, OpenAI’s GPT-5.6 Sol, Terra, and Luna lineup, and SpaceXAI’s Grok 4.5 release. It also covers GPT-Live’s simultaneous listening and speaking, Mistral’s 8B Robostral Navigate model, ChatGPT Work, Microsoft Flint, and new research on continuous-control memory, citation judging, coding evaluations, proactive agents, delegated web research, procedural code retrieval, and energy-market agent testing. Show notes: https://tobyonfitnesstech.com/podcasts/episode-84/

  • #83
    July 9 · 35 min

    Episode 83: Hermes Agent v2026.7.7, OpenAI Codex rust-v0.143.0, Claude Code 2.1.197, Aion-3.0-Mini

    Today's AgentStack Daily: Hermes Agent v2026.7.7 ships, OpenAI Codex lands rust-v0.143.0, and Claude Code CLI releases 2.1.197. AionLabs ships Aion-3.0-Mini roleplay on OpenRouter. Kokoro runs high-fidelity TTS on low-power CPUs. Rowboat debuts on Show HN with 162 points as a Claude Desktop alternative. Security: GitHub AI agent prompt injection leaks private repositories. Plus early-failure probes for agent loops, Danus fact-graph memory, FreqDepthKV, DepthWeave-KV, RuBench 1.0, VAORA, and the Anthropic developer-relations API migration story. Show notes: https://tobyonfitnesstech.com/podcasts/episode-83/

  • #82
    July 7 · 37 min

    Episode 82: Claude Code 2.1.195, Tencent Hy3, Nex-N2-Mini, GLM 5.2, and Anthropic's Global Workspace Paper

    Claude Code CLI 2.1.195 ships as Tencent's Hy3 and Nex AGI's Nex-N2-Mini land on OpenRouter. GLM 5.2 stokes margin-collapse debate; Anthropic publishes a Global Workspace Theory interpretability analysis. A Claude Code workspace session leak and a goodwill-erosion post trend on Hacker News, and an archive of leaked agent system prompts climbs GitHub. Ternlight ships a 7MB browser-side WASM embedding model. Research: Graph Sparse Sampling, weak-to-strong on-policy distillation, Cortex's bidirectional VLM-VLA paper, Graph-as-Policy multi-agent industrial training, and a cleanliness study on style and coding agents. Show notes: https://tobyonfitnesstech.com/podcasts/episode-82/

Showing 21–26 of 26 episodes