Skip to content
Artwork for Models & Agents
TechnologyNewsTech News

Models & Agents

Patrick

Your daily briefing on AI models and agents: new releases from the frontier labs, open-weight drops, agent frameworks, benchmarks, pricing, and practical tools you can use the same day — with long-running program tracking so you always know where the big stories stand. For developers, builders, and AI practitioners.

AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production.

Play
  • 27 episodes
  • daily
  • Avg 10 min
  • English

Support the show

Goes straight to the publisher. podnod takes nothing.

Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • S1 · E157
    Today · 9 min

    Ep 157: Claude just showed it can autonomously improve alignment in other models on a single GPU…

    Models & Agents Claude just showed it can autonomously improve alignment in other models on a single GPU — shifting alignment work from human teams to model-driven loops. What You Need to Know: Anthropic released research where Claude researched, proposed, trained, and tested alignment fixes for smaller models in 48 hours on one GPU. Cohere shipped Parse 5, a 2.3B vision-language model aimed at high-volume document parsing at $1.50 per 1,000 pages. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=76Jk47VuQ-A 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E156
    Yesterday · 8 min

    Ep 156: Anthropic's new hardware standard gives AI agents a unified way to control lab and…

    Models & Agents Anthropic's new hardware standard gives AI agents a unified way to control lab and manufacturing equipment without custom drivers per device. What You Need to Know: Anthropic opened a research preview of its Model Hardware Standard (MHS) today, inviting partners in science, robotics, and manufacturing to help extend Claude Code's hardware reach. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=i5w7ANst1LI 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E155
    Yesterday · 7 min

    Ep 155: Anthropic’s new hardware standard gives AI agents a unified way to control lab equipment…

    Models & Agents Anthropic’s new hardware standard gives AI agents a unified way to control lab equipment, boards, and cameras through one interface. What You Need to Know: Anthropic opened a research preview of its Model Hardware Standard (MHS) today, inviting partners in science, robotics, and manufacturing to help shape a common driver layer for physical devices. The effort starts with lab and manufacturing gear and will expand via Claude Code to boards and cameras. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=eduHppOVLxw 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E154
    Thursday · 10 min

    Ep 154: Researchers can now analyze real Claude usage data outside labs — Anthropic released…

    Models & Agents Researchers can now analyze real Claude usage data outside labs — Anthropic released privacy-preserved tools and 250k conversations for independent impact studies. What You Need to Know: Anthropic opened aggregated Claude conversation data from April-May 2026 to Stanford SALT Lab, Oxford, and METR, revealing over half of chats involve consequential tasks. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=fgEU1tPe_1A 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E153
    Wednesday · 8 min

    Ep 153: OpenAI’s first custom inference chip is entering production, delivering higher throughput…

    Models & Agents OpenAI’s first custom inference chip is entering production, delivering higher throughput and lower latency in one architecture. What You Need to Know: OpenAI announced deployment plans for its Jalapeño inference chip by year-end, with testing showing gains in intelligence per watt and response speed for ChatGPT and agents. A new Qwen3.8-Flash-Next model is slated for release today, and Liquid AI open-sourced Pipette for on-device benchmarking. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=YmaSdsTMizY 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E152
    Tuesday · 8 min

    Ep 152: Local 27B models just wrote and merged their first production feature on a single 4060 Ti.

    Models & Agents Local 27B models just wrote and merged their first production feature on a single 4060 Ti. What You Need to Know: A developer reported successfully using Qwen3 27B IQ3_K_XXS quantized to run entirely on a 4060 Ti 16GB card, completing a full agentic coding workflow including codebase investigation, plan generation, multi-file edits, and QA gate passing before human merge approval. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=BrVWqPmriKA 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E151
    Monday · 10 min

    Ep 151: AI agents are shifting from experimental tools to major API consumers, changing how…

    Models & Agents AI agents are shifting from experimental tools to major API consumers, changing how developers price and secure their endpoints. What You Need to Know: PYMNTS reports agents now drive significant API traffic as businesses deploy them for routine transactions. Several arXiv papers detail concrete gains in reasoning speed, style control, and domain-specific deployment. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=bbMs8V4yB1A 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E150
    Sunday · 12 min

    Ep 150: Vercel and Ora just shipped a free public audit tool that scores any website’s readiness…

    Models & Agents Vercel and Ora just shipped a free public audit tool that scores any website’s readiness for AI agents across 118 checks. What You Need to Know: The biggest concrete release today is Vercel’s “Is Agentic” scorer, which lets developers quickly test whether their sites can support autonomous agents. A detailed deepDoctection tutorial shows how to wire layout analysis, DocTR OCR, and table extraction into structured JSONL for RAG. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=DCEocF4Bmaw 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E149
    August 22 · 9 min

    Ep 149: A 250M-parameter model trained on 30B tokens now deploys in 60 MB with million-token…

    Models & Agents A 250M-parameter model trained on 30B tokens now deploys in 60 MB with million-token retrieval from disk. What You Need to Know: A solo developer released SHADOW-250M, a heavily quantized LLM that keeps recent context in fp16 while compressing older tokens to 1 bit on disk. Nvidia published a linear-mapping technique that transfers KV caches between model sizes without full re-prefill. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=3NC6y8HexD4 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E148
    August 21 · 12 min

    Ep 148: Agent reliability benchmarks just exposed the gap between occasional success and…

    # Models & Agents Agent reliability benchmarks just exposed the gap between occasional success and consistent stateful execution in real business workflows. What You Need to Know: Thinkingbox introduces a sandbox and 507-workflow benchmark across retail, insurance, and IT support domains that measures end-to-end state transitions rather than isolated tool calls. Several arXiv papers released today examine attention allocation, KV-cache reuse, and multi-agent hypothesis generation. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=_nSKXMijex4 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E147
    August 20 · 9 min

    Ep 147: OpenAI is testing private safety processing that keeps frontier-model interactions off…

    # Models & Agents OpenAI is testing private safety processing that keeps frontier-model interactions off-limits to staff while still catching risks across long agent sessions. What You Need to Know: OpenAI previewed Private Safety Processing for frontier models to improve safety without personnel seeing raw content. Simon Willison documented an untrusted-sandbox experiment where Claude Code triggered an autonomous GitHub Actions push. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=b0WOMNQDb6k 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E146
    August 19 · 11 min

    Ep 146: Claude designed novel protein binders from scratch for 14 of 15 targets, with 22-35%…

    Models & Agents Claude designed novel protein binders from scratch for 14 of 15 targets, with 22-35% success rates that beat the field's typical 10-15%. What You Need to Know: Anthropic demonstrated Claude autonomously creating functional protein binders that were then built and validated by Adaptyv Bio and Twist Bioscience. Open-weight releases from Ornith and inclusionAI add new dense and MoE options for local use. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=FQQd1T5Wgm8 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E145
    August 18 · 10 min

    Ep 145: OpenAI ships ChatGPT for Teens with stronger safeguards and parent controls, giving…

    Models & Agents OpenAI ships ChatGPT for Teens with stronger safeguards and parent controls, giving builders a new production template for age-gated agents. What You Need to Know: OpenAI released ChatGPT for Teens today with built-in protections, healthy-use features, and parent controls aimed at learning rather than shortcuts. Snowflake added dynamic model routing to Cortex AI Gateway that can cut token costs up to 3x by sending simple tasks to smaller models. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=KcHynJycx-Q 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E144
    August 17 · 11 min

    Ep 144: Language-server retrieval costs more tokens than grep for most coding-agent tasks and…

    Models & Agents Language-server retrieval costs more tokens than grep for most coding-agent tasks and rarely improves success rates. What You Need to Know: A new measurement study on Claude Opus 4.8, Sonnet 4.6, and Haiku 4.5 finds that LSP-based semantic retrieval increases token use by 6-118% on symbol localization while delivering no recall gains over simple grep. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=g6xco_0gR7g 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E143
    August 16 · 10 min

    Ep 143: Models sound most sure of themselves exactly when their answers are wrong — and a new eval…

    Models & Agents Models sound most sure of themselves exactly when their answers are wrong — and a new eval harness is making that gap impossible to ignore. What You Need to Know: An enterprise architect built a synthetic ground-truth harness that revealed LLMs confidently misattribute root causes in data-drift scenarios, especially when signals overlap. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=MYxUdQklaGY 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E142
    August 15 · 9 min

    Ep 142: Z.ai just showed how far post-training alone can push a fixed base model into serious…

    Models & Agents Z.ai just showed how far post-training alone can push a fixed base model into serious coding and cyber agent territory. What You Need to Know: GLM-5.3 arrives with the same 743-753B base as GLM-5.2 but delivers large jumps on Terminal-Bench, DeepSWE, and ExploitBench after heavy post-training scaling. Anthropic published its second Responsible Scaling Policy Risk Report alongside an EU AI Act watermarking FAQ that confirms no output quality or cost impact. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=xjj37XKuI30 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E141
    August 14 · 8 min

    Ep 141: Gemini 3.7 Flash delivers major upgrades for coding and web work at half the prior Flash…

    Models & Agents Gemini 3.7 Flash delivers major upgrades for coding and web work at half the prior Flash price, giving builders a faster, cheaper frontier option right now. What You Need to Know: Google released Gemini 3.7 Flash with targeted gains in software engineering and knowledge tasks plus a halved introductory price. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=SQosLctVaKA 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E140
    August 13 · 8 min

    Ep 140: SpaceXAI’s Grok 4.6 ships 500K context and a new xhigh reasoning mode tuned specifically…

    Models & Agents SpaceXAI’s Grok 4.6 ships 500K context and a new xhigh reasoning mode tuned specifically for long-running agents and coding workflows. What You Need to Know: SpaceXAI released Grok 4.6 yesterday as a post-training upgrade over Grok 4.5, not a larger base model. It matches GPT-5.6 Sol Max at 61 on the Artificial Analysis Intelligence Index while keeping pricing at $2/$6 per million tokens. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=DH0xN1hid38 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E139
    August 12 · 9 min

    Ep 139: OpenAI’s ChatGPT desktop app now runs natively on Linux, letting developers keep browser…

    Models & Agents OpenAI’s ChatGPT desktop app now runs natively on Linux, letting developers keep browser and project workflows inside one authenticated session. What You Need to Know: OpenAI released a preview of the ChatGPT desktop app for Ubuntu 24.04/26.04, Debian 13, and Fedora 43/44 with both x64 and ARM64 .deb/.rpm packages. Three VentureBeat surveys released today quantify how enterprises are actually buying, securing, and evaluating agents and infrastructure. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=vZSm2UYrSpg 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E138
    August 11 · 11 min

    Ep 138: OpenAI is shipping purpose-trained frontier models to defenders first, giving authorized…

    Models & Agents OpenAI is shipping purpose-trained frontier models to defenders first, giving authorized teams GPT-5.6-Cyber for vulnerability research before attackers can weaponize similar capabilities. What You Need to Know: OpenAI expanded its Daybreak program with two new access tiers—Daybreak Blue for broad defensive work and Daybreak Red for advanced authorized testing—centered on GPT-5.6 variants. The models have already surfaced real zero-days in Chrome’s V8 engine. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=FC3D2oZSbyQ 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
Showing 1–20 of 27 episodes