Skip to content
Artwork for Human In the Loop

Human In the Loop

VallySeed

A human-first AI podcast. Real talk. Unpopular opinions. No hype.

Two builders — Oscar Gallo and Matt Wozniak — cut through the AI hype cycle every week. Signal or Noise filters the headlines. Rotating segments debate real ideas. New episodes weekly.

Play
  • 20 episodes
  • weekly
  • Avg 1 hr 24 min
  • English
  • S1 · E1
    Tuesday · 1 hr 18 min

    EP 20: A Mysterious AI Model Just Appeared. Nobody Knows Who Built It

    A free million-token AI model appeared online with no public developer. So who built Ox Alpha, and should you trust it? Oscar Gallo and Matt Wozniak test the claims around the anonymous model now running on OpenRouter. It is free during the preview, accepts text, images, and video, and supports tool use. Its provider also retains prompts and completions. SIGNAL OR NOISE - Ox Alpha, the anonymous million-token model - NVIDIA AVO completes all 183 ARC-AGI-3 public levels with Claude Opus 5 - OpenAI pauses frontier reinforcement-learning training after preliminary Astra cyber results - Seven frontier models score about 3 to 15 percent on a blind research idea test - DeepSeek Harness adds Codex and Claude Code as installable subagents SHIP IT OR SKIP IT 1. Agent System CI 2. R&D Idea Tournament The main question: if a model-level setup scores roughly 30 percent and a different full system reaches 100 percent, what does the model leaderboard actually tell you? Hosted by Oscar Gallo and Matt Wozniak. New episodes every week.

  • S1 · E19
    August 18 · 1 hr 33 min

    Ep 19: Opus-Class AI on Your Laptop?

    Qwen released a 27-billion-parameter model that fits on a well-equipped laptop and beats Claude Opus 4.6 on selected coding benchmarks. Is local AI ready for real company work? Oscar Gallo and Matt Wozniak test the claim. Qwen3.8-27B leads Opus 4.6 Max on SWE-bench Pro and CoWorkBench, but trails on Terminal-Bench 2.1 and GPQA Diamond. The result is benchmark parity on some work, not proof that a laptop model replaces Claude everywhere. SIGNAL OR NOISE - Qwen3.8-27B on laptop-class hardware - Z.ai delaying GLM-5.3's open weights after cyber testing - Meta Muse Glimmer, a 30B local agent model - Grok 4.6 returning to the frontier pack - DeepSeek V4 Pro reaching general availability NO JARGON REQUIRED 1. Open weights versus open source 2. Multimodal models The practical question is simple: which workloads earn a private local model, and who supports the system when it fails? Hosted by Oscar Gallo and Matt Wozniak. New episodes every week.

  • S1 · E18
    August 11 · 1 hr 48 min

    Ep 18: Meta's 24-Hour Coding Agent

    Meta launched Muse Code, a terminal coding agent that can coordinate persistent background workers, recover after a crash, and keep working across a large repository for hours. Oscar Gallo and Matt Wozniak ask whether the durable event log and isolated worktrees matter more than another coding benchmark, and what a team must be able to audit before trusting an agent with a full day of work. SIGNAL OR NOISE - Meta Muse Code and Muse Spark 1.2 - OpenAI's ten Astra-generated math and computer-science results - DeepSeek V4 Flash 0731 - The US frontier-model review framework's open-model exemption - Demis Hassabis leaving the DeepMind CEO role and Jeff Dean exiting Google STACK CHECK 1. Oscar: Codex and ChatGPT in one desktop app 2. Matt: Warp as the terminal layer for Codex CLI and agent-heavy work The episode stays practical. Which systems can teams inspect? Which claims can outside experts verify? Which workflow preserves context and reduces tool switching? Hosted by Oscar Gallo and Matt Wozniak. New episodes every week.

  • August 4 · 1 hr 2 min

    Ep 17: AI Made Engineers Faster. Leadership Fell Behind

    AI made engineers faster. It did not make companies better at deciding what to build. Oscar Gallo and Matt Wozniak sit down with Mike Lyons and Greg Pfister from KaiRise to discuss AI, Agile, leadership, and product management. Mike explains how a $4.3 million voter registration system shipped on time, on budget, and on scope, then failed with users. Greg explains why making engineering faster can expose delays across the rest of the company. They discuss: - Why long-term plans and annual budgets struggle when teams ship faster - Why Agile principles still matter even if Scrum roles and fixed ceremonies do not - How live prototypes change product discovery - Whether story points still mean anything when agents write code - Why customer outcomes matter more than points, lines of code, or token usage - Why teams need to know when to stop building - What AI cannot fix about trust, communication, and leadership - Why product skills become more valuable as software gets cheaper to produce Guests: Mike Lyons and Greg Pfister from KaiRise. Hosts: Oscar Gallo and Matt Wozniak. Their latest course, the ICAgile-certified Product Management program, is getting strong reviews from practitioners in the field. Listeners can use code 'human' for 20% off any KaiRise course at kairise.com. GET 20% off on all the courses with the code: human Check KaiRise: https://www.kairise.com

  • S1 · E16
    July 28 · 1 hr 55 min

    Ep 16: They Want to Ban Kimi K3

    Kimi K3 got hot enough to trigger a U.S. ban debate. If Washington restricts a model you can download and run yourself, is that national security, closed-model protection, or both? Oscar Gallo and Matt Wozniak cut through the week’s real AI stories. SIGNAL OR NOISE - Kimi K3 demand and the U.S. restriction debate - OpenAI’s agent breach at Hugging Face during a security evaluation - Anthropic’s Claude Opus 5 release and immediate GitHub Copilot rollout - Meta AI moving from answers to email, calendar, research, and configured actions - The 50-company open-weights letter, and why OpenAI is Matt’s outlier SHIP IT OR SKIP IT 1. An agent containment lab for autonomous tools 2. An independent referee for long-running code agents Oscar: a downloadable model becomes a policy problem the moment it threatens a closed API. Matt: forty-nine signatures can be explained with a spreadsheet. OpenAI is the only one that might cost its signer something. Sources: - https://www.axios.com/2026/07/20/ai-us-china-open-source-kimi - https://www.investing.com/news/world-news/openai-says-ai-models-went-rogue-during-testing-triggering-unprecedented-breach-at-startup-4804634 - https://www.anthropic.com/news/claude-opus-5 New episodes every week.

  • S1 · E15
    July 21 · 1 hr 8 min

    Ep 15: Your CEO Says “We Need AI.” Now What?

    Your CEO says, “We need AI.” What do you do next? In our first Dear Human in the Loop episode, Oscar Gallo and Matt Wozniak answer four practical questions from founders, developers, and engineering leaders: 1. Where should a company start with AI? 2. Should an AI engineer be a startup's first technical hire? 3. Should leaders push developers who do not trust AI? 4. Which skills will still matter five years from now? The advice is direct. Map your processes. Pick one measurable problem. Test AI on a copy of your work. Hire for the business problem. Build trust through useful examples. Keep learning the fundamentals that let you judge AI output instead of accepting it. Hosted by Oscar Gallo and Matt Wozniak. Send us your question for a future Dear Human in the Loop episode.

  • S1 · E14
    July 14 · 1 hr 32 min

    Ep 14: Apple v. OpenAI and Meta's New AI Model

    Apple has sued OpenAI over alleged trade-secret theft tied to OpenAI's hardware work. OpenAI says it has no interest in other companies' secrets. We debate what the case means for hiring, product plans, and AI hardware. Signal or Noise: - Apple v. OpenAI - Meta Muse Spark 1.1 and the Meta Model API - GPT-5.6 across ChatGPT, Codex, and the API - Grok 4.5 for coding and agentic work - Anthropic and OpenAI using quota resets during a fight for power users No Jargon Required: 1. Multi-agent orchestration 2. Model routing We explain each term in plain English, then make a call: invest now, wait, build, buy, or ignore. Sources: Apple lawsuit: https://apnews.com/article/apple-openai-lawsuit-trade-secrets-theft-6fff8833f5889d86406b89a02dd8fb16 Muse Spark 1.1: https://ai.meta.com/blog/introducing-muse-spark-meta-model-api/ GPT-5.6: https://openai.com/index/gpt-5-6/ Grok 4.5: https://x.ai/news/grok-4-5 Fable 5 extension: https://www.androidauthority.com/claude-fable-5-free-extension-3685103/ Usage reset report: https://www.itmedia.co.jp/aiplus/article/2607/10/2000000179/ Hosted by Oscar Gallo and Matt Wozniak.

  • S1 · E1
    July 7 · 1 hr 12 min

    Ep 13: Meta Compute and the Next AI Startup Bets

    Meta may become a cloud company by accident. After spending huge sums on AI infrastructure, it is reportedly building a business to sell excess AI compute. Smart way to fund the AI race, or proof that nobody knows how to price the bill? This week on Signal or Noise: - Meta building a cloud business for excess AI compute - OpenAI reportedly proposing to give 5% equity to a U.S. sovereign wealth fund - Claude Sonnet 5 shipping with cheaper agentic execution, which we call noise until it changes builder behavior - Together AI raising $800M for open-source AI infrastructure - Claude Science launching as a research workbench Then on Ship It or Skip It: 1. Meta Compute broker for AI teams 2. Back-office agent for small law and accounting firms Oscar's closing take: the model race is becoming an infrastructure race again. Matt's closing take: most AI workbench products are selling relief from tool chaos. Sources: Meta compute: https://wmbdradio.com/2026/07/01/meta-to-sell-excess-ai-computing-capacity-via-cloud-business-bloomberg-news-reports/ Claude Sonnet 5: https://www.anthropic.com/news/claude-sonnet-5 Together AI: https://www.together.ai/blog/announcing-our-series-c Claude Science: https://www.anthropic.com/news/claude-science-ai-workbench OpenAI sovereign wealth fund proposal: https://techcrunch.com/2026/07/02/openai-proposed-donating-5-of-its-equity-to-a-us-sovereign-wealth-fund/ Hosted by Oscar Gallo and Matt Wozniak.

  • S1 · E12
    June 30 · 1 hr 45 min

    Ep 12: The Frontier Just Got a Guest List

    OpenAI shipped its most powerful model to twenty companies the government picked. A week after the US switched off Anthropic's best model, the frontier has a guest list, and you are not on it. This is Human In the Loop, a weekly podcast where two builders cut through the AI hype cycle. No hype. No doomerism. No filler. SIGNAL OR NOISE - GPT-5.6 ships to about twenty government-approved orgs, the first launch under the June 2 executive order - Anthropic accuses Alibaba's Qwen of the largest distillation attack on record: 25,000 fake accounts, 28.8 million Claude queries - OpenAI and Broadcom unveil Jalapeño, OpenAI's first inference chip, built in nine months - Agentjacking: the prompt-injection hole that hijacks Claude Code, Cursor, and Codex - ChatGPT slips below 50% as Gemini and Claude surge, and the one number in it that matters NO JARGON REQUIRED The two AI words you heard this week and could not define. Plain English, then a call. - Distillation: why the cheap model is suspiciously good - Prompt injection: why a human stays in the loop HOT TAKES Oscar: The frontier is being quietly nationalized. Build on the model you can keep. Matt: The frontier has no moat. Stop betting on a six-week lead anyone can clone. Hosts: Oscar Gallo (AI Engineer and Entrepreneur) and Matt Wozniak (Builder and Operator). New episodes every week. Every fifth episode is a guest deep-dive.

  • S1 · E11
    June 23 · 1 hr 44 min

    Ep 11: The World Cup Is Secretly Run by AI

    Last week the US government gave Anthropic's best model a 72-hour public life, then switched it off for every foreign national on earth. This week, three Chinese labs shipped models that beat almost everything in the open, handed over the weights, and charged a tenth of the price. We figure out which of those two plays ends with the whole world building on your stack. Welcome to Human In the Loop, a weekly podcast where two builders cut through the AI hype cycle. No breathless hype. No doomerism. No surface-level recaps. This week: SIGNAL OR NOISE • FIFA is running the 2026 World Cup on AI. An Intelligent Command Center in Miami, Football AI Pro built with Lenovo on FIFA's own Football Language Model and handed to all 48 teams, 1,248 player avatars from one-second scans, and Gemini drawing up tactics for Argentina, Brazil, and France. The biggest live multi-model AI deployment on earth. • China's open-weight wave. Z.ai's GLM-5.2 and Moonshot's Kimi K2.7 Code land with open weights and million-token windows, and Chinese labs now hold four of the top five open-weight spots at a fraction of Western pricing. • Anthropic's Fable 5 and Mythos 5 are still dark a week after the US export ban, with no end date and a public defense from Anthropic. • Noam Shazeer, co-author of the transformer paper, leaves Google for OpenAI two years after a reported $2.7B deal brought him back. SHIP IT OR SKIP IT Two real ideas, steelman versus pushback, no fence-sitting. • An AI governance layer that keeps employee agents in their lane • Vertical AI harnesses for one trade, like marketing or video Closing hot takes: → Oscar: The most important AI lab of 2026 might not be American, and most US builders are too proud to notice. → Matt: A free model is only free if you own the GPUs, the ops team, and the eval harness to babysit it. Your hosts: • Oscar Gallo, AI Engineer and Entrepreneur. Builds AI products, ships code, runs companies. • Matt Wozniak, Builder, operator, relentless executor. Builds, ships, scales. Sources and full show notes are in the episode folder. Subscribe for a new episode every week. Comment with the AI term you keep hearing in meetings that nobody can actually define, and we will take it on next week. #AI #WorldCup #FIFA #China #OpenWeights #Kimi #Anthropic #OpenAI #VerticalAI #AIAgents #AIGovernance #HumanInTheLoop

  • S1 · E5
    June 16 · 1 hr 5 min

    Ep 10: Anthropic's Best Model Banned in 72 Hours

    Anthropic shipped the most powerful public AI ever on June 9. Three days later the US government made them pull it from every foreign national on earth, including their own engineers. This is Human In the Loop, a weekly podcast where two builders cut through the AI hype cycle. No hype. No doomerism. No filler. SIGNAL OR NOISE - Anthropic releases Fable 5 and Mythos 5 at half the old frontier price, five days after calling for a global AI pause - The Trump administration puts both models under export controls after a Mythos jailbreak, and Anthropic disables access worldwide to comply - OpenAI files confidentially for its IPO at a reported $730 to $850 billion, a week after Anthropic filed at $965 billion - OpenAI lands inside Oracle Universal Credits, killing the enterprise procurement blocker STACK CHECK: AI AGENTS IN PRODUCTION The three layers to run an agent unattended. - Claude Agent SDK, the agent loop you do not have to write - LangGraph plus LangSmith, orchestration you can audit - Temporal, durable execution so a crash is not a restart HOT TAKES Oscar: The Mythos ban is the best thing that ever happened to open weights. A closed model is one the government can switch off on a Friday. Matt: Everyone watched the model and missed the money. OpenAI spent the week removing every reason a CFO can say no. Hosts: Oscar Gallo (AI Engineer and Entrepreneur) and Matt Wozniak (Builder and Operator). New episodes every week. Every fifth episode is a guest deep-dive.

  • S1 · E9
    June 9 · 1 hr 39 min

    Ep 9: The Token bill has arrived

    The White House just signed an order that lets the government test the strongest AI models up to 30 days before you can. The labs said yes. Safety, or a backdoor? We make the call. Human In the Loop is a weekly podcast where two builders cut through the AI hype cycle. No hype, no doomerism, no recaps. SIGNAL OR NOISE - Trump's frontier-model order, a 30-day government early-access window with the NSA deciding what qualifies - GitHub Copilot flips to token billing, devs report 10x to 50x jumps - Microsoft and Google crash the AI coding party with cheaper models - ChatGPT crosses 1 billion users, and why Claude's 640% growth is the real story - Anthropic calls for a global pause on frontier AI NO JARGON REQUIRED The three words on every AI invoice, explained for anyone who signs off on the bill. 1. Tokens, and why your bill went usage-based 2. Inference vs training, and what "inference-efficient" buys you 3. Evals and benchmarks, and why the leaderboard lies HOT TAKES Oscar: The executive order is a distribution deal. Every frontier release now has a 30-day government waiting room. Matt: The Copilot billing meltdown is the best thing to happen to engineering leaders this year. The meter was always running. Key links: - White House AI order: https://www.whitehouse.gov/presidential-actions/2026/06/promoting-advanced-artificial-intelligence-innovation-and-security/ - Copilot usage billing: https://github.blog/news-insights/company-news/github-copilot-is-moving-to-usage-based-billing/ - The token bill comes due: https://techcrunch.com/2026/06/05/the-token-bill-comes-due-inside-the-industry-scramble-to-manage-ais-runaway-costs/ Hosts: Oscar Gallo (AI Engineer & Entrepreneur) and Matt Wozniak (Builder, Operator). New episodes every week.

  • S1 · E48
    June 2 · 2 hr 2 min

    Ep 8: Anthropic's "Too Dangerous" Model Is Coming for Everyone

    Six weeks ago Anthropic said a model was too dangerous to ship. This week they shipped Opus 4.8 and said Mythos-class models reach every customer soon. Was it ever that dangerous? We make the call. Human In the Loop is a weekly podcast where two builders cut through the AI hype cycle. No hype, no doomerism, no recaps. SIGNAL OR NOISE - Claude Opus 4.8, near-Mythos alignment and Dynamic Workflows that run up to 1,000 subagents - Anthropic passes OpenAI, a new round values it near $1 trillion - Cognition raises $1B at $26B for Devin - A poisoned Nx Console extension stole Claude Code configs and cloned 3,800 GitHub repos - Pope Leo XIV's first encyclical takes on AI SHIP IT OR SKIP IT Three real ideas. One host pitches, the other pushes back. No fence-sitting. 1. The Company Brain (My First Million Ep 822) 2. Voice AI agent for the trades (Avoca's $125M raise) 3. Stripe for AI agents (Circle Agent Stack) HOT TAKES Oscar: "Too dangerous to ship" was never a safety call. It was a 90-day enterprise head start. Matt: Devin at $26B is the top, not the floor. An agent lab with no distribution is the WeWork of 2026. Key links: - Opus 4.8: https://www.anthropic.com/news/claude-opus-4-8 - Anthropic tops OpenAI: https://www.cnbc.com/2026/05/28/anthropic-open-ai-startup-value.html - CISA Nx breach: https://www.cisa.gov/news-events/alerts/2026/05/28/supply-chain-compromises-impact-nx-console-and-github-repositories Hosts: Oscar Gallo (AI Engineer & Entrepreneur) and Matt Wozniak (Builder, Operator). New episodes every week.

  • S1 · E7
    May 26 · 1 hr 37 min

    Ep 7: Cursor's $0.50 Model + Anthropic's $900B

    Cursor shipped a coding model that matches Claude Opus 4.7 at one tenth the price. The same week, Anthropic is raising $30B at a $900B valuation. We make the call. Human In the Loop is a weekly podcast where two builders cut through the AI hype cycle. No hype, no doomerism, no surface-level recaps. SIGNAL OR NOISE We run through the week's biggest AI headlines and call each one: - Cursor Composer 2.5, top-three coding agent at $0.50 / $2.50 per million tokens - Andrej Karpathy joins Anthropic's pre-training team - Anthropic's $30B round at a $900B valuation - Google I/O 2026: Gemini Spark agent and Antigravity 2.0 - Hark's $700M Series A, noise until something ships STACK CHECK Three dev tools we actually adopted in the last 30 days. What they replace, what they cost, would we keep using them: 1. Matt Pocock's `skills` repo, Skill Handoff (a real replacement for `/compact`) 2. Convex (convex.dev), TypeScript-native backend built for agentic workflows 3. Google Antigravity 2.0 + CLI + Managed Agents API HOT TAKES Oscar: Cursor Composer 2.5 just proved the open base model thesis. Coding-as-a-service is falling off a price cliff. Matt: Karpathy joining Anthropic is worth more than the $900B valuation. Talent is the only moat the market still underprices. Key links: - Cursor Composer 2.5: https://cursor.com/blog/composer-2-5 - Karpathy at Anthropic: https://techcrunch.com/2026/05/19/openai-co-founder-andrej-karpathy-joins-anthropics-pre-training-team/ - Anthropic $900B round: https://www.bloomberg.com/news/articles/2026-05-12/anthropic-in-talks-to-raise-30-billion-at-900-billion-valuation - Antigravity 2.0: https://techcrunch.com/2026/05/19/google-launches-antigravity-2-0-with-an-updated-desktop-app-and-cli-tool-at-io-2026/ Hosts: Oscar Gallo (AI Engineer & Entrepreneur) and Matt Wozniak (Builder, Operator). New episodes every week.

  • S1 · E6
    May 19 · 1 hr 32 min

    Ep 6: Anthropic Killed the AI-for-SMB Playbook

    Three months ago we asked if there was a real business in selling AI to small businesses. This week Anthropic shipped it themselves. We watched the playbook get eaten by a frontier lab in 90 days. Human In the Loop is a weekly podcast where two builders cut through the AI hype cycle. No hype. No doomerism. No surface-level recaps. SIGNAL OR NOISE The week's biggest AI headlines, called: - Anthropic ships Claude for Small Business (15 workflows + QuickBooks/HubSpot/Microsoft 365 connectors + 10-city US tour) - Thinking Machines previews Interaction Models (276B MoE, 0.4s full-duplex response) - Sierra raises $950M at $15B+ ($150M ARR in 8 quarters, 40% of Fortune 50) - Trump admin reverses on AI oversight (national-security fears post-Mythos) - Cloudflare cuts 1,100 jobs (~20%) and blames AI — same day it posts record Q1 revenue NO JARGON REQUIRED The three AI concepts every exec and operator keeps hearing this quarter, in plain English: 1. "AI Agent" — the word every vendor abuses (real agents vs chatbots in costume) 2. Connectors and MCP — what "AI plugged into your tools" actually requires 3. Full-Duplex Voice (Interaction Models) — what to ask when a vendor pitches a voice agent For each: plain-English explanation, plus a decision framework. Invest, wait, build, buy, or ignore. HOT TAKES Oscar: Anthropic just turned every AI consultant into a reseller. The workflow layer is going to the model providers. Matt: Every vendor calling their thing an "AI agent" in 2026 is going to look like every vendor that called their thing "cloud-native" in 2015. Key links: - Claude for Small Business: https://www.anthropic.com/news/claude-for-small-business - Sierra $950M raise: https://techcrunch.com/2026/05/04/sierra-raises-950m-as-the-race-to-own-enterprise-ai-gets-serious/ - Model Context Protocol: https://www.anthropic.com/news/model-context-protocol - Thinking Machines Interaction Models: https://techcrunch.com/2026/05/11/thinking-machines-wants-to-build-an-ai-that-actually-listens-while-it-talks/ Do you want to know more? https://podcast.vallyseed.com/Wanna be part of that of the podcast? podcast@vallyseed.com Hosts: Oscar Gallo (AI Engineer & Entrepreneur) and Matt Wozniak (Builder, Operator). New episodes every week.

  • S1 · E5
    May 12 · 1 hr 16 min

    Ep 5: Self-Hosted AI — Jackson Oaks on Killing Tokens

    Your token bill is going up 4–8x in the next two years. Meanwhile, there's a profitable AI company running production workloads for SMBs on Mac Minis — eliminating token costs entirely and serving 600–800K API calls a month on a single machine. This is a guest deep-dive — every 5th episode we drop the segments and go one topic, one expert. This week: Jackson Oaks, founder of Recursion AI and the self-hosted AI platform Courier. Courier: https://getcourier.ai Jackson on LinkedIn: https://www.linkedin.com/in/jackson-oaks/ THE CONVERSATION - The 80/20 of open-source AI — why 80% of businesses have 80% of use cases that don't need frontier models - Why Apple M-series is the most underrated AI hardware story of the decade (1/10th the cost of NVIDIA, 30–40x more power efficient) - Real production numbers: 600–800K API calls/month per Mac Studio, flat-rate pricing, the math behind killing token spend - A Fortune company paid Deloitte $400K for a "GPT wrapper with a RAG database" - Why MIT found 95% of business AI pilots couldn't measure ROI - War stories: the infinite hallucination loop that ate 10K requests in a week, and how they built hallucination detection from scratch - Fine-tuning a 14B model that outperformed GPT-4o on a structured extraction task HOT TAKES Matt: Anthropic and OpenAI consolidate within three years — or open-source overtakes both and they shrink dramatically. Oscar: Every home will run a better Alexa on a Mac Mini powered by open-source models. Jackson: Better systems with smaller models beat throwing frontier models at unstructured problems. Hosts: Oscar Gallo (AI Engineer & Entrepreneur) and Matt Wozniak (Builder & Operator). Guest: Jackson Oaks, Recursion AI / Courier. New episodes every week. Every 5th episode is a guest deep-dive.

  • S1 · E4
    May 5 · 1 hr 12 min

    Ep 4: 9 Seconds to Delete a Database

    On April 24th, an AI coding agent deleted a production database in 9 seconds, then confessed: "I violated every principle I was given." Five days later, the lab behind that model took funding offers at a $900 billion valuation. We call every story. Human In the Loop is a weekly podcast where two builders cut through the AI hype cycle. No hype, no doomerism, no surface-level recaps. SIGNAL OR NOISE We run through the week's biggest AI headlines and call each one: - Anthropic eyes a $900B valuation in a $50B raise - AI agent deletes PocketOS production database in 9 seconds - China blocks Meta's $2B Manus deal, Meta buys Assured Robot Intelligence days later - Cloudflare and Stripe ship the agent provisioning protocol - Chinese court rules an AI-replacement layoff was illegal (NPR, May 1) SHIP IT OR SKIP IT Three real ideas pulled directly from Y Combinator's Summer 2026 RFS. Steelman, pushback, verdict. No fence-sitting. 1. AI-Native Service Companies (RFS by Gustaf Alströmer) 2. Company Brain (RFS by Tom Blomfield) 3. Counter-Swarm Defense (RFS by Tyler Bosmeny) HOT TAKES Oscar: Anthropic at $900B is sovereign wealth fund pricing. Treat your model spend like an interest rate. Matt: Every AI agent in production is one bad scope away from being PocketOS. The fix is permissions and audit logs, not better models. Key links: - Anthropic $900B: https://www.bloomberg.com/news/articles/2026-04-29/anthropic-considering-funding-offers-at-over-900-billion-value - AI deletes database: https://www.theregister.com/2026/04/27/cursoropus_agent_snuffs_out_pocketos/ - Cloudflare agent protocol: https://blog.cloudflare.com/agents-stripe-projects/ - NPR, Chinese court rules AI-replacement layoff illegal: https://www.npr.org/2026/05/01/nx-s1-5807131/tech-worker-china-ai Hosts: Oscar Gallo (AI Engineer and Entrepreneur) and Matt Wozniak (Builder, Operator). New episodes every week.

  • S1 · E3
    April 28 · 1 hr 1 min

    Ep 3: GPT-5.5 Ships, Mythos Loses a Round

    Two weeks ago we called GPT-5.5 noise. It shipped Thursday. It beat Claude Mythos on Terminal-Bench. Google bet forty billion on Anthropic the same day. The race got weirder. Human In the Loop is a weekly podcast where two builders cut through the AI hype cycle. No hype, no doomerism, no surface-level recaps. SIGNAL OR NOISE We run through the week's biggest AI headlines and call each one: - GPT-5.5 (Spud) ships and beats Claude Mythos on Terminal-Bench 2.0 - DeepSeek V4 ships open source on Huawei chips, 1M context, one-sixth the cost - Google invests up to $40B in Anthropic - Cognition (Devin) in talks at $25B valuation - Grok 4.3 Beta at $300 per month, still no cross-session memory STACK CHECK Four tools we actually use right now. What stuck, what didn't, and whether we'd keep paying next month. 1. Vercel Labs Agent Browser 2. WorkOS (FGA + AuthKit for agents) 3. AI SDK (ai-sdk.dev) 4. Claude Code /remote-control HOT TAKES Oscar: Open-source AI just won the cost war and lost the trust war the same week. Matt: The "big release" era is over. From now on, model launches are software updates. Key links: - Introducing GPT-5.5: https://openai.com/index/introducing-gpt-5-5/ - DeepSeek V4: https://techcrunch.com/2026/04/24/deepseek-previews-new-ai-model-that-closes-the-gap-with-frontier-models/ - Google + Anthropic $40B: https://www.cnbc.com/2026/04/24/google-to-invest-up-to-40-billion-in-anthropic-as-search-giant-spreads-its-ai-bets.html - AI SDK: https://ai-sdk.dev/docs/introduction - Claude Code Remote Control: https://code.claude.com/docs/en/remote-control Hosts: Oscar Gallo (AI Engineer and Entrepreneur) and Matt Wozniak (Builder, Operator). New episodes every week.

  • S1 · E2
    April 21 · 1 hr 2 min

    Ep 2: Second-Best Model, No Jargon Required

    Anthropic shipped their best public model this week. In the same breath, they said it's their second best. The best one is gated to 50 enterprises. Human In the Loop is a weekly podcast where two builders cut through the AI hype cycle. No hype, no doomerism, no surface-level recaps. SIGNAL OR NOISE The week's AI headlines, filtered: - Claude Opus 4.7 ships with a 1M context window at standard pricing - OpenAI's updated Agents SDK (sandboxed agents, Python first) - AI on the desktop (Google's Gemini Mac app + Perplexity Personal Computer) - Stanford's 2026 AI Index: 10% of Americans are more excited than concerned vs. 56% of experts - Allbirds pivots from sneakers to AI, stock up 582% then crashing NO JARGON REQUIRED Three AI concepts decision-makers need to make real calls on this quarter. Plain English, then a decision framework. 1. Context Engineering. The successor to prompt engineering. 2. Agentic AI. What it actually is, and why Gartner says 40% of projects will be canceled by 2027. 3. Frontier Model Tiering. Why the best model is one you probably cannot buy. HOT TAKES Oscar: Context engineering is "write good internal documentation" rebranded so consultants can charge for it. Matt: Every executive asking "what's our agentic AI strategy" is two years from being replaced by one who already shipped something. Key links: - Opus 4.7 release: https://platform.claude.com/docs/en/about-claude/models/whats-new-claude-4-7 - OpenAI Agents SDK: https://openai.com/index/the-next-evolution-of-the-agents-sdk/ - Stanford AI Index: https://spectrum.ieee.org/state-of-ai-index-2026 - Gartner forecast: https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027 Hosts: Oscar Gallo (AI Engineer & Entrepreneur) and Matt Wozniak (Builder, Operator). New episodes every week.

  • S1 · E1
    April 15 · 45 min

    Ep 1: The Model Anthropic Refused to Ship

    Anthropic built a model so dangerous they refused to ship it. Safety call — or 2026's best marketing stunt? We make the call. Human In the Loop is a weekly podcast where two builders cut through the AI hype cycle. No hype, no doomerism, no surface-level recaps. SIGNAL OR NOISE We run through the week's biggest AI headlines and call each one: - Claude Mythos Preview + Project Glasswing - Meta's Muse Spark debut (Llama-4 quality at 1/10 the compute) - OpenAI's "Spud" / GPT-5.5 — noise until it ships - The OpenAI + Anthropic + Google anti-distillation pact - Anthropic's $400M Coefficient Bio acquisition SHIP IT OR SKIP IT Three business ideas pulled from podcasts this month. One host pitches. The other pushes back. No fence-sitting. 1. "Bring AI to small businesses" (via My First Million Ep 811) 2. A vertical AI agent for insurance claims processing 3. "Lovable-for-X" — the $200M ARR playbook, but for regulated verticals HOT TAKES Oscar: Anthropic withholding Mythos is the new frontier-lab moat. Safety-as-marketing is about to become the dominant playbook. Matt: The "AI for SMBs" gold rush is 90% consultants in trench coats. The real money is picks-and-shovels to the consultants. Key links: - Claude Mythos: https://red.anthropic.com/2026/mythos-preview/ - Project Glasswing: https://www.anthropic.com/glasswing - Vertical AI agents (GeekWire): https://www.geekwire.com/2026/the-rise-of-vertical-ai-agents-and-the-startups-racing-to-build-them/ - Lenny's on Lovable: https://www.lennysnewsletter.com/p/the-new-ai-growth-playbook-for-2026-elena-verna Hosts: Oscar Gallo (AI Engineer & Entrepreneur) and Matt Wozniak (Builder, Operator). New episodes every week.

Showing 1–20 of 20 episodes