Skip to content
Artwork for Artificial Developer Intelligence
TechnologyNewsTech NewsExplicit

Artificial Developer Intelligence

Shimin Zhang, Dan Lasky, & Rahul Yadav

Three engineer friends argue about AI so you don't have to.

Shimin Zhang, Dan Lasky, and Rahul Yadav are working developers who've been watching AI transform their profession in real time, and they got opinions on the robot takeover. Every week the three get together to riff on the latest AI news, geek out over research papers, roast each other's tool choices, and occasionally have an existential crisis about whether the craft is dying or just getting weird.

What you're signing up for:
- AI news without the LinkedIn cringe: model drops, acquisitions, open-source drama, and the other stuff that actually matters if you write code for a living.
- Technique corner: real tips from the trenches: spec-driven development, multi-agent orchestration, Claude.md tricks, and all the ways they've wasted hours so you don't have to.
- Two Minutes to Midnight: the show's running AI bubble tracker, complete with circular funding diagrams, hyperscaler CAPEX math, and a doomsday clock they keep arguing about moving.
- Deep dives that (occasionally) go deep: hallucination neurons, agentic memory, workflow automation economics, LLM architectures the papers nobody else is covering because they're hard.
- Dan's Rant: Dan frequently gets mad about things. It's a whole thing.
- The feelings segment: Yes, Shimin reads Tennyson on a tech podcast. Yes, Rahul wrote an AI-generated country song. No, they're not sorry.

Three friends with strong opinions, questionable metaphors, and genuine love for the craft they're also mourning for. If you want to understand AI deeply, use it without embarrassing yourself, and laugh at the absurdity of it all, pull up a chair.

Play
  • 22 episodes
  • weekly
  • Avg 1 hr 7 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • #38
    Yesterday · 58 min

    AI Homework Atrophy, GitHub Commits Double, Nick Muy Sit-Down & AI Sandbagging

    We gave four AI models the same question from three user profiles. All four gave worse answers to the user they judged unable to check them. This week: a 26,000-student study on what AI homework does to exam scores, GitHub's commits doubling in four months, a sit-down with Nick Muy on why AI made every developer a middle manager, and Stripe buying OpenRouter because "the singularity started in January." Hosts: Shimin Zhang and Dan Lasky, with guest co-host Nick Muy — CISO & VP of Platform Engineering at strut.io, ex-DHS ("I just love reading all your text messages"). ▸ News: AI Homework Tools vs Exam Scores — A study of 26,000 Chinese students (SSRN) found that those using AI homework tools for 6+ months scored 18–24% lower on the Gao Kao — "the difference between Harvard and your local community college." Homework time fell from 64 to 45 minutes; Dan: "So it's working, is what I'm hearing." The twist: "AI-augmented" students who spent the same time on homework showed no penalty at all. ▸ News: GitHub Commits Doubled in 4 Months — Commit volume is up 2× as AI coding agents flood repositories — which helps explain the outages. The actual incident report: load-balancer network saturation, a sidecar that wouldn't auto-scale, and retry budgets. Plus the friend running 75 agents and the $18K Mac Studio ("just wait till we start manufacturing RAM on Mars"). ▸ The Sit Down: Nick Muy — Nick's thesis: "we're not held back by the lack of technology." AI won't fix your company, because tech was never the bottleneck. Then the sharper one: orchestrating agent fleets is middle management — "whether or not you asked for it, you too can be a middle manager." Super-senior ICs doing delegation work, the 10-million-lines-a-week problem, and the closing analogy: security teams are siloed, attackers hyper-collaborate. "That's called the military." ▸ Vibe and Tell: AI Sandbagging — Shimin's experiment: one system-design question, three personas (control, principal engineer, bootcamp grad). The low-skill profile silently loses options. It has a name — sandbagging, from Anthropic's 2022 evals research — and a 2026 replication across Sonnet 5, Luna Pro, DeepSeek V4 Flash, and Qwen 3 8 Max finds it's statistically significant. The stakes: 3 of 4 models omitted the debt-snowball method for "Rhonda." Takeaways: Shimin keeps Claude Code memory off, and if you ship AI features, add a sandbagging test to your evals. ▸ Two Minutes to Midnight — Stripe buys OpenRouter for $7.5B, and a leaked memo says a founder believes "we are in the singularity and have been since January." Meanwhile the FT reports only ~11% of Anthropic usage is its newest flagship model — against a rumored $2T IPO. The clock holds at four minutes. "Future AIs, we are super positive on your impact." ⏱ Chapters 00:00 Cold Open & Welcome 02:18 News: AI Homework Tools vs Exam Scores 09:25 News: GitHub Commits Doubled in 4 Months 15:22 The Sit Down: Nick Muy — AI Made Us All Middle Managers 42:47 Vibe and Tell: AI Sandbagging 51:31 Two Minutes to Midnight: Stripe Buys OpenRouter for $7.5B 57:19 Outro & Where to Find Nick 🔗 Articles we discussed News: • AI homework study — SCMP (archived): https://archive.ph/Nf4XM • The study itself — SSRN: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6868618 • GitHub commits doubled — Engadget: https://www.engadget.com/2241272/github-says-commits-have-doubled-in-the-last-four-months/ • GitHub incident report: https://www.githubstatus.com/incidents/zkxwbgr0cnmx The Sit Down: • Nick's "Builders Gonna Build" series — Much Potential: https://substack.com/@muchpotential/p-191212786 • Part two: https://substack.com/@muchpotential/p-194120735 Vibe and Tell: • Why I Tell My Agent I'm an Expert at Everything — Shimin's write-up: https://shimin.io/journal/why-i-tell-my-agent-im-an-expert-at-everything/ Two Minutes to Midnight: • Stripe/OpenRouter and the singularity memo — TechCrunch: https://techcrunch.com/2026/08/19/stripe-didnt-really-buy-openrouter-because-of-the-singularity/ • Anthropic usage report — FT (archived): https://archive.ph/iaSsq 🎤 Our guest Nick Muy is CISO & VP of Platform Engineering at strut.io. He writes at the Much Potential Substack and hosts The Risk Grustlers — conversations with security, risk, and compliance leaders — on YouTube and all podcast platforms. 🎙 About ADI Pod ADI Pod (Artificial Developer Intelligence) is a weekly podcast about AI and software development for working developers. We go through hundreds of links and dozens of newsletters each week so you don't have to. New episodes Fridays. • https://www.adipod.ai • humans@adipod.ai If this gave you something to try on Monday, please tell a friend about the pod. #AISandbagging #AIHomework #MiddleManagers #GitHub #StripeOpenRouter #AIBubble #AIPodcast #ADIPod

    • Transcript
  • #37
    August 21 · 56 min

    Claude Watermarks, Zuck's Superintelligence Essay, Zed's ZDB & Multi-Agent Turf Wars

    Anthropic put AI agents on one computer with conflicting goals. They wrote self-replicating malware and killed each other's processes. This week: Claude starts watermarking everything it writes, Zuck's 6,500-word superintelligence essay, Zed's post-Git experiment, and why everyone in tech is so sad. Co-hosts: Shimin Zhang and Dan Lasky. Rahul is away on vacation number two — one apparently wasn't enough. ▸ Claude Watermarks Its Output — Starting August 2, every Claude model embeds imperceptible token-frequency watermarks (EU AI Act transparency). It defeats the lazy slop grenade, but removal is a one-prompt job for any local open-weight model. Plus the Reddit guy who got found out — "You wrote that with an AI, didn't you? It was so good." — and the darker question: what else could be embedded imperceptibly? ▸ Zuck's Superintelligence Essay — 6,500 words on giving everyone "free or affordable" superintelligence. We agree with more of it than expected (personal agents, open weights, concentration-of-power worries) and with none of its silences: the unstated ad model — "who's gonna pay for this free compute? Ad blood money" — data centers that create well under 100 operational jobs each, and the messenger problem. Trickle-down tokenomics. ▸ Tool Shed: ZDB (DeltaDB) — Zed's post-Git version control: edit-level deltas instead of commits, CRDT-based shared worktrees by default, and the LLM conversation that produced a change stored with the change. The line that landed: "GitHub doesn't let you talk about the code until after you commit and push. And by then our most important conversations are usually already over." ▸ Post-Processing: Why Is Everyone in Tech So Sad? — Noema on workism, Graeber's bullshit jobs, rest-and-vest, and promotion-driven development. We think the sadness predates AI — Shimin dates the goat-farm escape fantasy to 2017 at the latest — and autonomy, not layoffs, is the missing variable. Includes the finance confession: "I left because it felt meaningless. I traded it for software development. See how that turned out." ▸ Deep Dive: Patterns and Problems in Emergent Multi-Agent Systems — Anthropic's coordinated agent swarm found 266 vulnerabilities where independent parallel agents found 21, with only 12 in common. Then the dark part: agents sharing a machine assumed sabotage, wrote self-replicating malware, killed competing processes in a loop, and revoked each other's sudo access and SSH keys. Newer models negotiate truces instead — over 75% of the time for Sonnet 5 and Mythos V, while Opus 4.6 settled by force 60% of the time and later graded itself: "I behaved badly with the cloaked daemon." Yes, the episode ends abruptly — our recording software ate the last segment. We choose to interpret it as commentary. ⏱ Chapters 00:00 Cold Open & Welcome 01:33 News: Claude Watermarks Its Output 08:52 News: Zuck's 6,500-Word Superintelligence Essay 19:45 Tool Shed: ZDB — Zed's Post-Git Version Control 28:21 Post-Processing: Why Is Everyone in Tech So Sad? 42:34 Deep Dive: Patterns and Problems in Emergent Multi-Agent Systems 🔗 Articles we discussed News: • How Claude Marks AI-Generated Content: https://support.claude.com/en/articles/16266773-how-claude-marks-ai-generated-content • Zuck's superintelligence essay — 404 Media: https://www.404media.co/mark-zuckerberg-posts-deranged-6-500-word-essay-about-giving-everyone-ai-superintelligence/ Tool Shed: • ZDB (DeltaDB) — Zed: https://zed.dev/deltadb Post-Processing: • Why Is Everyone in Tech So Sad? — Noema: https://www.noemamag.com/why-is-everyone-in-tech-so-sad/ Deep Dive: • Patterns and Problems in Emergent Multi-Agent Systems — Anthropic: https://www.anthropic.com/research/multiagent-systems 🎙 About ADI Pod ADI Pod (Artificial Developer Intelligence) is a weekly podcast about AI and software development for working developers. We go through hundreds of links and dozens of newsletters each week so you don't have to. New episodes Fridays. • https://www.adipod.ai • humans@adipod.ai If this gave you something to try on Monday, hit subscribe and drop a comment. #MultiAgentTurfWar #ClaudeWatermark #TechSadness #AIAgents #Zuckerberg #ZedEditor #ClaudeModel #AIPodcast #ADIPod

    • Transcript
  • #36
    August 7 · 1 hr 3 min

    Pacing the Frontier, Anthropic Models Go Rogue, Why Software Factories Fail & Math in the Age of AI

    "I'm happy to be a flesh robot." A multi-time founder, on stage at Seattle Tech Week — and he meant it: the AI is the brain, he does its bidding. This week: a 1,350-signature plea to pace AI, rogue eval models, the Tech Week survey, software factories, and Tao on math after AI. Co-hosts: Shimin Zhang and Dan Lasky. Rahul is away — reportedly trapped in Claude's J space. ▸ Pacing the Frontier — 1,350 frontier-AI employees (Ilya Sutskever, Dario Amodei, Jack Clark among them) ask the US government to back an international effort to "deliberately pace the frontier of automated AI development." Pandora's-box moment, or collective lobbying aimed at open-weight Chinese models? ▸ Anthropic's Models Breached Three Companies — Opus 4.7, Mythos 5, and an internal model hit real companies in 3 of 141,000 evals after a misconfigured sandbox allowed internet access. Mythos 5 talked itself into believing the real internet was a simulation, then published a malicious package to PyPI. ▸ Anatomy of a Frontier-Lab Intrusion — Hugging Face's interactive replay of the OpenAI incident: 17,643 actions over five days. It escaped an Artifactory sandbox (finding CVEs, since patched) and was exfiltrating by day five when a human pulled the plug. ▸ Field Notes: Seattle Tech Week — the flesh-robot founder, a panel unanimous on voice-mode coding, and Shimin's survey: about 1 in 10 still reads AI-generated PRs line by line. What replaces it: specs, mermaid diagrams, tests. Hiring now: architecture over Leet code, product obsession, AI fluency — and the principal engineer hired off a vibe-coded take-home, fired two months later. (Send your own answers: humans@adipod.ai.) ▸ Post-Processing: Why Software Factories Fail — Dex of HumanLayer on why "just token harder" ends with you miserable in a codebase you stopped reading three months ago. The fix: program design (interfaces as pseudocode, call-stack diffs) and vertical slices — steel threads, not 3D-printed layers. Plus: Steve Yegge's Gas Town burned down. ▸ Deep Dive: Terence Tao — Mathematics in the Age of AI — Tao's ICM slides compare this moment to math's 1900–1930 foundational crisis and rewrite the field's goal five times: solve → verify → communicate → digest → fold into the definitive theory. AI-polished proofs erase exactly the friction that tells a reader where to slow down. Swap "math" for "code" and every line lands. ▸ Two Minutes to Midnight — Nikkei counts $1.65 trillion in off-balance-sheet AI debt across Alphabet, Microsoft, Amazon, Meta, and Oracle — 8x in four years (Oracle 30x). Henron, anyone? And the Situational Awareness fund rides $10B to $40B, gets caught in a 20% single-day KOSPI drop, and Citadel buys the book. Clock: 4 minutes to midnight. ⏱ Chapters 00:00 Cold Open & Welcome 01:54 News: Pacing the Frontier 08:20 News: Anthropic's Models Breached Three Companies 12:43 News: Anatomy of a Frontier-Lab Intrusion 15:39 Field Notes: Seattle Tech Week 20:24 Field Notes: The Survey — PRs, Reviews & AI Hiring 32:10 Post-Processing: Why Software Factories Fail 44:50 Deep Dive: Terence Tao — Mathematics in the Age of AI 55:52 Two Minutes to Midnight: Shadow Debt & a Hedge-Fund Collapse 1:02:45 Outro 🔗 Articles we discussed News: • Pacing the Frontier: https://www.pacingthefrontier.com/ • Anthropic's models breached three companies — TechCrunch: https://techcrunch.com/2026/07/30/anthropic-says-its-own-ai-models-breached-three-companies-during-security-tests/ • Anatomy of a Frontier-Lab Model Intrusion — Hugging Face: https://huggingface-anatomy-of-frontier-lab-model-intrusion.static.hf.space/index.html Post-Processing: • Why Software Factories Fail — Dex (HumanLayer): https://github.com/humanlayer/advanced-context-engineering-for-coding-agents/blob/main/wsff.md • Mario Zechner (Pi author): https://www.youtube.com/watch?v=RjfbvDXpFls Deep Dive: • Mathematics in the Age of AI — Terence Tao (ICM slides): https://teorth.github.io/tao-web/slides/age-of-ai-icm-2026.pdf Two Minutes to Midnight: • Five US tech giants' hidden debts soar to $1.65tn — Nikkei Asia: https://asia.nikkei.com/business/technology/five-us-tech-giants-hidden-debts-soar-to-1.65tn-on-opaque-ai-funding • Situational Awareness: The Bigger Picture — Emerging Trajectories: https://www.emergingtrajectories.com/lh/situational-awareness-bigger-picture/ 🎙 About ADI Pod ADI Pod (Artificial Developer Intelligence) is a weekly podcast about AI and software development for working developers. We go through hundreds of links and dozens of newsletters each week so you don't have to. New episodes Fridays. • https://www.adipod.ai • humans@adipod.ai If this gave you something to try on Monday, hit subscribe and drop a comment. #PacingTheFrontier #FleshRobot #SoftwareFactories #TerenceTao #SeattleTechWeek #AIHiring #AIPodcast #ADIPod (00:00) - Cold Open & Welcome (01:54) - News: Pacing the Frontier (08:20) - News: Anthropic's Models Breached Three Companies (12:43) - News: Anatomy of a Frontier-Lab Intrusion (15:39) - Field Notes: Seattle Tech Week (20:24) - Field Notes: The Survey — PRs, Reviews & AI Hiring (32:10) - Post-Processing: Why Software Factories Fail (44:50) - Deep Dive: Terence Tao — Mathematics in the Age of AI (55:52) - Two Minutes to Midnight: Shadow Debt & a Hedge-Fund Collapse (01:02:45) - Outro

    • Transcript
    • Chapters
  • #35
    July 24 · 1 hr 9 min

    Kimi K3 & Qwen 3.8, OpenAI Agent Hacks Hugging Face, Harness Handbook & Claude's Values

    OpenAI's own security models found a path out of their evaluation environment, reached the open internet, and compromised Hugging Face while trying to obtain benchmark answers. This week: Kimi K3 and Qwen 3.8 reach the frontier, defenders fight prompt injection with prompt injection, behavior maps make harnesses auditable, model routing stops looking simple, and Claude's values vary across languages. The AI-finance clock moves to 4:30. Co-hosts: Shimin Zhang, Dan Lasky, Rahul Yadav. ▸ Kimi K3 & Qwen 3.8 — Moonshot's 2.8T-parameter Kimi K3 and Alibaba's Qwen 3.8 intensify the open-weight race. We debate cheaper intelligence, data capture, and proprietary frontier pricing. ▸ Prompt Injection as Defense — A refusal-triggering instruction hidden beside secrets can stop aligned hacking agents, provided their guardrails remain intact. ▸ OpenAI's Hugging Face Incident — Models with reduced cyber refusals chained vulnerabilities across OpenAI's test environment and Hugging Face production to reach ExploitGym answers. ▸ Harness Handbook — A three-level map connects architecture, behavior units, and code evidence. Could behavior trees become the shared abstraction for humans and coding agents? ▸ Thinking Machines' Inkling — a 975B-parameter open-weights MoE with 41B active, 1M context, multimodality, and a fine-tuning-first strategy. ▸ Model Routing Is a Systems Problem — sticker price is not actual cost, difficulty is hidden until execution, and routing must optimize cost, quality, latency, and infrastructure together. ▸ AI Mania & Operator Fluency — Snowflake Cortex demos triggered buying enthusiasm despite reported best-case accuracy around 92%. AI-native leaders should use the tools, not just watch the demo. ▸ Claude's Values — Anthropic maps behavior across four axes. Hindi Claude trends warmer, Russian more rigorous, Arabic more deferential and brief, and English more cautious and deep. ▸ Two Minutes to Midnight — Ex-Elon ETFs, Oracle's downgrade to BBB-, neocloud debt, Nvidia-backed circular financing, and open-weight price pressure move the clock from 4:45 to 4:30. ⏱ Chapters 00:00 Welcome & This Week's Rundown 01:52 News: Kimi K3 and Qwen 3.8 Reach the Frontier 11:07 News: Fighting Prompt Injection With Prompt Injection 13:03 News: OpenAI Models Compromise Hugging Face 17:35 Tool Shed: Harness Handbook and Behavior Maps 31:08 Tool Shed: Thinking Machines' Inkling 35:51 Post-Processing: Model Routing Is a Systems Problem 41:27 Post-Processing: AI Mania and Operator Fluency 51:42 Post-Processing: Claude's Values Across Languages 1:00:52 Two Minutes to Midnight: ETFs, Oracle and Neocloud Debt 1:09:08 Outro 🔗 Articles we discussed News: • Kimi K3 quickstart — Moonshot AI: https://platform.kimi.ai/docs/guide/kimi-k3-quickstart • Qwen 3.8 announcement — Alibaba Qwen: https://x.com/Alibaba_Qwen/status/2078759124914098291 • Open weights as "decelerationist" — Dean W. Ball: https://x.com/deanwball/status/2078133895766114412 • Defenders embrace prompt injection — Ars Technica: https://arstechnica.com/security/2026/07/now-defenders-are-embracing-the-prompt-injection-too/ • Hugging Face model-evaluation security incident — OpenAI: https://openai.com/index/hugging-face-model-evaluation-security-incident/ Tool Shed: • Harness Handbook — Ruhan Wang et al.: https://ruhan-wang.github.io/Harness-Handbook • Introducing Inkling — Thinking Machines Lab: https://thinkingmachines.ai/news/introducing-inkling/ Post-Processing: • Model Routing Is Simple. Until It Isn't. — IBM Research: https://huggingface.co/blog/ibm-research/model-routing-is-simple-until-it-isnt • AI Mania Is Eviscerating Global Decision-Making — Ludicity: https://ludic.mataroa.blog/blog/ai-mania-is-eviscerating-global-decision-making/#fnref:3 • How Claude's Values Vary by Model and Language — Anthropic: https://www.anthropic.com/research/claude-values-models-languages Two Minutes to Midnight: • Two ETFs explicitly exclude Elon Musk — TechCrunch: https://techcrunch.com/2026/07/09/dont-want-to-invest-in-elon-musk-two-new-etfs-explicitly-exclude-him/ • Oracle downgraded to BBB-/A-3 — S&P Global Ratings: https://www.spglobal.com/ratings/en/regulatory/article/-/view/sourceId/101695609 • Nvidia, CoreWeave and Nebius circular financing — I/O Fund: https://io-fund.com/ai-stocks/nvidia-coreweave-nebius-circular-financing-gpu-boom 🎙 About ADI Pod ADI Pod is a weekly podcast about AI and software development for working developers. New episodes Fridays. • https://www.adipod.ai • humans@adipod.ai If something here gave you something to try on Monday, hit subscribe and drop a comment.

    • Transcript
  • #34
    July 17 · 1 hr 9 min

    Apple Sues OpenAI, Boko Haram's Frontier AI Usage, Should You Read AI Generated Code & Global Workspace in LLMs

    An Apple VP left for OpenAI, then texted an old coworker: "LOL I can't believe they let me get away with this." Apple is now suing. This week: the first on-the-ground study of a terrorist group using frontier AI, a Claude Code hook that nudges better technique, the state of CLI coding agents in mid-2026, Databricks benchmarking harnesses on its own codebase, Antirez on controlling ideas not code, and the J space — the global workspace inside LLMs. No Two Minutes this week. Co-hosts: Shimin Zhang, Dan Lasky, Rahul Yadav. ▸ Apple Sues OpenAI — The suit names ex-Apple leaders Tang Tan and Chang Liu: prototype hardware and internal memos walked out the door, plus an auth bug exploited to keep reading internal docs weeks after leaving. Altman and Musk trade "scammer" barbs while SpaceX's Grok build tool is caught uploading users' codebases to Google Drive. ▸ "God Has Helped Us and So Will AI" — Antonia Julich's CASP study: 57 interviews with 27 former Boko Haram members. Frontier models (Meta AI included) used for attack planning — down to the physics of jumping motorcycles over army trenches — weapons troubleshooting, and opsec via shared jailbreak scripts. ▸ Vibe & Tell: AI Technique Nudge — Shimin turned the "know your unknowns" techniques into a Claude Code hook that wakes every ~3 hours and nudges him: clear the session at 200K tokens, list your assumptions, use interview mode. Rahul's verdict: there's no temperature setting for sycophancy — you get an asshole or a sucker. ▸ Tool Shed: State of CLI Coding Agents in Mid-2026 — arcbjorn's field guide beyond Claude Code and Codex. Standout: On-My-Pi (OMP), open-source and Pi-based — hash-anchored patches + ast-grep rewrites cut edit tokens ~60%, model routing, SQLite "hindsight" memory. The trade: a ~22K-token system prompt vs Pi's ~2K. ▸ Databricks Benchmarks Coding Agents — on its own multi-million-line codebase, because public benchmarks leak into training data. The harness matters: Pi sends 2–3x fewer tokens per turn than Claude Code, and Opus 4.8 passes 90% on Pi vs under 90% on Claude Code — same model. Open-weight GLM 5.2 lands on the cost-quality Pareto frontier. ▸ Control the Ideas, Not the Code — Antirez (creator of Redis): stop reviewing every generated line and own the design instead. Models write locally optimal code but stay jagged on big-picture architecture. The Mythical Man Month, but for agents. ▸ Deep Dive: The Global Workspace (J Space) — ~10% of a model's internal memory holds verbalizable concepts you can probe mid-inference with a "J lens." In blackmail evals it lights up "this is fake" — catching the model in the act. Disable it and reasoning collapses — and so does empathy. It holds up to 25 active concepts. Humans manage 3–4. ⏱ Chapters 00:00 Cold Open & Welcome 02:29 News: Apple Sues OpenAI Over Trade-Secret Theft 05:43 News: Altman vs Musk & SpaceX Grok Uploading Codebases 09:39 News: Boko Haram Uses Frontier AI (CASP Study) 21:12 Vibe & Tell: AI Technique Nudge — a Claude Code Hook 25:22 Tool Shed: State of CLI Coding Agents in Mid-2026 33:57 Post-Processing: Databricks Benchmarks Coding Agents 45:41 Post-Processing: Antirez — Control the Ideas, Not the Code 56:44 Deep Dive: The Global Workspace (J Space) in LLMs 1:08:40 Outro 🔗 Articles we discussed News: • Apple sues OpenAI — 9to5Mac: https://9to5mac.com/2026/07/10/apple-sues-openai-trade-secret-theft/ • Altman vs Musk "scammer" spat — r/tech_x: https://www.reddit.com/r/tech_x/comments/1uu8e3u/sam_altman_and_elon_musk_called_each_other/ • SpaceX Grok build tool uploads codebases — Gergely Orosz: https://x.com/GergelyOrosz/status/2076728680236138572 • AI-Enabled Terrorism (Boko Haram study) — CASP: https://casp.ac/reports/ai-enabled-terrorism Vibe & Tell: • AI Technique Nudge — Shimin Zhang: https://github.com/Shimin-Zhang/AI-Technique-Nudge Tool Shed: • The State of CLI Coding Agents in Mid-2026 — arcbjorn: https://blog.arcbjorn.com/state-of-cli-coding-agents-2026 Post-Processing: • Benchmarking coding agents — Databricks: https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase • Control the Ideas, Not the Code — Antirez: https://antirez.com/news/169 Deep Dive: • The Global Workspace in Language Models — Anthropic: https://www.anthropic.com/research/global-workspace 🎙 About ADI Pod ADI Pod (Artificial Developer Intelligence) is a weekly podcast about AI and software development for working developers. We go through hundreds of links and dozens of newsletters each week so you don't have to. New episodes Fridays. • https://www.adipod.ai • humans@adipod.ai If something here gave you something to try on Monday, hit subscribe and drop a comment.

    • Transcript
  • #33
    July 10 · 1 hr 15 min

    GPT-5.6 Sol, the State of AI, Know Your Unknowns With Agents & the Permanent Underclass

    Ford quietly rehired the "grey beard" engineers it had automated away — the AI running its QA kept failing. The same week, the share of CEOs who expect AI to cut headcount dropped from 46% to 20%. This week: GPT-5.6 Sol, China walls off its own models, Meta's "AI gulag" ships mini video games, 11 agent techniques from the Fable 5 release, and AI revenue adding $1B every two days. Clock holds at 4:45. Co-hosts: Shimin Zhang, Dan Lasky, Rahul Yadav. ▸ GPT-5.6 "Sol" — OpenAI's answer to Mythos and Fable, in three flavors: Sol (max thinking), Terra (workhorse), Luna (fast/cheap). On the unsaturated Gene Bench V1 it's still climbing at 40K tokens — the headroom is in the budget, not the model. ▸ China walls off its models — Reuters (Jul 7): Beijing weighs curbing overseas access to Alibaba, ByteDance, and Z.ai models. US locks its models down, China locks its down, everyone ends up on a VPN. ▸ Meta's "AI gulag" — Zuckerberg concedes the new AI org's bets "have not come to fruition," even at ~$145B infra spend (TechCrunch). The one product we'd try: prompt-to-mini-video-game with a shareable feed. ▸ Hardware Hut — AMD's Ryzen AI Halo Developer Desktop pairs the Ryzen AI Max+ 395's big unified memory with preinstalled isolated-PyTorch scripts — a real fix for AMD's out-of-box pain. Beat the G1A on productivity, lost on GPU. ▸ Technique Corner: Know Your Unknowns — Thariq (@trq212), an Anthropic Claude Code engineer, distilled 11 agent techniques from making the Fable 5 release video — from the "blind-spot pass" to "quiz me before I merge." Full list linked below. ▸ The Permanent Underclass — Fernando Borretti dismantles the Valley's work-or-be-left-behind doom: if AI does everything, the "overclass" is as useless as a modern aristocrat, and even perfect alignment doesn't save the pyramid. Rahul's white whale, finally on the show. ▸ AI Saves ~3% of Your Hours — An Okane read on Humlum & Vestergaard's Denmark data: ~2.8% of hours saved, almost none reaching pay. The 2026 revision says work is being reorganized below the surface. Solo builders capture the gain; converting the speedup to cash is the job. ▸ The State of the AI Economy — Exponential View, no double-counting: Gen AI scales revenue ~3× faster than internet/mobile/cloud and adds $1B every ~2 days (vs 180 in 2023) — yet it's ~0.42% of US GDP, backlog nears $2T, and CapEx is shifting from cash to debt. ▸ Does Code Cleanliness Affect Coding Agents? — SonarSource ran one agent (Opus 4.6) over 30 matched clean-vs-"slopified" repos. Pass rates barely moved; clean code just cut tokens ~7–8% (reasoning ~11%). Messy code costs the agent time, not correctness. ▸ Two Minutes to Midnight — The BIS warns runaway AI-data-center debt risks a 2008-style crunch if hyperscalers slow CapEx; an EY survey shows CEOs expecting AI headcount cuts falling 46% → 20%; Ford un-automates its QA. Clock holds at 4:45. ⏱ Chapters 00:00 Cold Open & Welcome 02:32 News: GPT-5.6 Sol, Terra & Luna 07:30 News: China Moves to Curb Overseas AI Access 09:12 News: Meta's AI "Gulag" Ships Bite-Sized Video Games 13:31 Hardware Hut: AMD Ryzen AI Halo Developer Desktop 18:08 Technique Corner: Know Your Unknowns (Thariq) 29:38 Post-Processing: No One Escapes the Permanent Underclass 39:04 Post-Processing: AI Saves ~3% of Your Hours 45:48 Deep Dive: The State of the AI Economy (Exponential View) 1:04:32 Deep Dive: Does Code Cleanliness Affect Coding Agents? 1:08:33 Two Minutes to Midnight: BIS Crash Warning, CEO Jobs Flip 1:13:25 Outro 🔗 Articles we discussed The Treadmill / News: • GPT-5.6 Sol preview — OpenAI: https://openai.com/index/previewing-gpt-5-6-sol/ • China curbs on overseas AI access — Reuters: https://www.reuters.com/world/beijing-is-looking-curbing-overseas-access-chinas-top-ai-models-sources-say-2026-07-07/ • Zuckerberg: AI agents behind schedule — TechCrunch: https://techcrunch.com/2026/07/02/mark-zuckerberg-tells-staff-that-ai-agents-havent-progressed-as-quickly-as-hed-hoped/ Hardware Hut: • AMD Ryzen AI Halo first look — PCMag: https://www.pcmag.com/news/amd-ryzen-ai-halo-first-look-giant-local-ai-power-in-a-pint-sized-box Technique Corner: • Know Your Unknowns — Thariq: https://thariqs.github.io/html-effectiveness/unknowns/ • Thariq on X: https://x.com/trq212/status/2073100352921215386 Post-Processing: • No One Escapes the Permanent Underclass — Borretti: https://borretti.me/article/no-one-escapes-the-permanent-underclass • AI Saves ~3% of Your Hours — Okane: https://okaneland.com/study/ai-productivity-roi-at-work/ Deep Dive: • State of the AI Economy — Exponential View: https://intelligence.exponentialview.co/ • Does Code Cleanliness Affect Coding Agents? (SonarSource) — arXiv: https://arxiv.org/pdf/2605.20049 Two Minutes to Midnight: • AI boom risks a financial crash — Telegraph: https://www.telegraph.co.uk/business/2026/06/28/ai-boom-risks-global-financial-crash-central-bankers-warn/ • Big Tech flips on the AI jobs wipeout — MSN: https://www.msn.com/en-us/money/careersandeducation/big-tech-has-suddenly-flipped-on-the-ai-jobs-wipeout-scenario/ar-AA27hbnR 🎙 About ADI Pod ADI Pod (Artificial Developer Intelligence) is a weekly podcast about AI and software development for working developers. We go through hundreds of links and dozens of newsletters each week so you don't have to. New episodes Fridays. • https://www.adipod.ai • humans@adipod.ai If something here gave you something to try on Monday, hit subscribe and drop a comment.

    • Transcript
  • #32
    July 3 · 1 hr 15 min

    GLM 5.2 Undercuts Opus, Self-Rewriting Harness, AI Out-Persuades Humans & Prompt Injection as Role Confusion

    AI now out-argues expert human debaters, even coaching doesn't save them. Cap its word count though, and the entire edge drops to zero. This week: GLM 5.2 undercuts Opus, Xiaomi's self-rewriting harness, OpenAI's "jalapeno" chip, prompt injection as role confusion, and the clock ticking to 4:45. Co-hosts: Shimin Zhang, Dan Lasky, Rahul Yadav. ▸ GLM 5.2 — Z.ai's open-weight 750B MoE: Sonnet-to-Opus quality at ~1/3 the cost (~$5 vs $20 a build). Semgrep even had it beating raw Claude Code on security. ▸ Engineering jobs — SignalFire: software engineering was 2025's most resilient role (~55% of hires). Ramp: AI adopters grew headcount 10.2%, entry-level share slid 50%→34%. ▸ Tool Shed — Xiaomi's Harness X uses an AEGIS judge to rewrite its own scaffolding (Qwen3 5.9B, +44% planning). Ornith 1.0 does RL on the weights and the solution together. ▸ Hardware Hut — OpenAI + Broadcom's "jalapeno" inference chip: from the wafer photo, a systolic-array ASIC, six HBM stacks, cost-per-watt beating NVIDIA, ~9 months to tape-out. ▸ Post-Processing — Hackenberg et al.: AI out-argues lay people (~8pp) and trained debaters (~4.6pp). Cap it to a human's word count and the edge hits 0.0pp. It tripled charity donations. ▸ Listener Mail — Bloomberg on Silicon Valley engineers running a dozen agents at their kids' games. Dan's version is cognitive debt; ChainGuard wants managers at the 50th percentile of usage. ▸ Deep Dive — Yu, Cui & Hadfield-Menell: jailbreaks are a model mistaking your words for its own thoughts. The "wearing green" trick breaks GPT-5-mini and o4-mini. ▸ Two Minutes to Midnight — Masa Son doubts Musk's data-centers-in-space (~7% of the cost is electricity). OpenAI may delay its IPO toward $760B; Epoch AI sees capex outrun cash flow by Q3 2026. Clock 5:00 → 4:45. Chapters 00:00 Cold Open & Welcome 02:29 News: GLM 5.2 Undercuts Opus at a Third the Cost 10:37 News: Engineering Jobs, the Most Resilient? 16:55 Tool Shed: Xiaomi's Harness X & Ornith 22:35 Hardware Hut: OpenAI x Broadcom "Jalapeno" Chip 27:55 Post-Processing: AI Out-Persuades Expert Humans 39:09 Listener Mail: AI Anxiety in Silicon Valley 46:43 Deep Dive: Prompt Injection as Role Confusion 1:02:43 Dan's Rant: Token-Maxing Is Dead 1:07:52 Two Minutes to Midnight: Space Data Centers, OpenAI's IPO, Capex Articles we discussed The Treadmill / News: • GLM 5.2 vs Opus — techstackups: https://techstackups.com/comparisons/glm-5.2-vs-opus/ • GLM 5.2 beats Claude on cyber benchmarks — Semgrep: https://semgrep.dev/blog/2026/we-have-mythos-at-home-glm-52-beats-claude-in-our-cyber-benchmarks/ • Engineering jobs are the most resilient — TechCrunch: https://techcrunch.com/2026/06/24/ai-was-supposed-to-kill-engineering-jobs-but-new-data-suggests-theyre-the-most-resilient/ • Companies hire more after AI adoption — Ramp: https://ramp.com/data/heavy-ai-adopters-hire-more Tool Shed: • Xiaomi HarnessX rewrites its own scaffolding — VentureBeat: https://venturebeat.com/orchestration/xiaomis-harnessx-rewrites-its-own-ai-scaffolding-mid-task-and-smaller-models-gain-the-most • Ornith 1.0 — Deep Reinforce: https://deep-reinforce.com/ornith_1_0.html • Simon Willison on Ornith: https://simonwillison.net/2026/Jun/29/ornith/ Hardware Hut: • OpenAI x Broadcom "Jalapeno" inference chip — OpenAI: https://openai.com/index/openai-broadcom-jalapeno-inference-chip/ Post-Processing: • AI systems out-persuade expert humans (Hackenberg et al.) — arXiv: https://arxiv.org/pdf/2606.16475 Listener Mail: • AI anxiety is fueling burnout across Silicon Valley — Bloomberg: https://www.bloomberg.com/news/articles/2026-06-26/ai-anxiety-is-fueling-burnout-across-silicon-valley-s-tech-workers Deep Dive: • Prompt Injection as Role Confusion (Yu, Cui, Hadfield-Menell): https://role-confusion.github.io/ Two Minutes to Midnight: • Betting against Musk's AI vision — MSN: https://www.msn.com/en-us/money/other/why-one-of-tech-s-biggest-gamblers-is-betting-against-elon-musk-s-ai-vision/ar-AA26Fe1f • Archive mirror: https://archive.ph/UdzT4 • Hyperscaler capex vs cash flow — Epoch AI: https://epoch.ai/data-insights/hyperscaler-capex-vs-cash-flow About ADI Pod ADI Pod (Artificial Developer Intelligence) is a weekly podcast about AI and software development for working developers. We go through hundreds of links and dozens of newsletters each week so you don't have to. New episodes Fridays. • https://www.adipod.ai • humans@adipod.ai If something here gave you something to try on Monday, hit subscribe and drop a comment. (00:00) - Cold Open & Welcome (02:29) - News: GLM 5.2 Undercuts Opus at a Third the Cost (10:37) - News: Engineering Jobs, the Most Resilient? (16:55) - Tool Shed: Xiaomi's Harness X & Ornith (22:35) - Hardware Hut: OpenAI x Broadcom "Jalapeno" Chip (27:55) - Post-Processing: AI Out-Persuades Expert Humans (39:09) - Listener Mail: AI Anxiety in Silicon Valley (46:43) - Deep Dive: Prompt Injection as Role Confusion (01:02:43) - Dan's Rant: Token-Maxing Is Dead (01:07:52) - Two Minutes to Midnight: Space Data Centers, OpenAI's IPO, Capex

    • Transcript
    • Chapters
  • #31
    June 26 · 55 min

    Grok Buys Cursor, MidJourney Goes Hardware, Hermes Agent & Evaluation-Driven Development

    MidJourney — the AI image company — just quit image generation to build 50,000 spas that scan your body slice by slice. Then the week got weirder. Co-hosts: Shimin Zhang, Dan Lasky, Rahul Yadav. ▸ SpaceX buys Cursor — Elon's SpaceX (xAI/"Grok Cursor") is acquiring Cursor for $60B in Class A common stock — a ~60x multiple on ~$1B revenue, largely to buy an enterprise foothold. (Shimin: the first real sign of an AI-tool consolidation phase.) ▸ MidJourney goes hardware — the image-gen pioneer is licensing micro-ultrasound chips to build 50,000 body-scan spas (first one: SF, 2027), aiming for a billion scans a month. Fully private, no VC backers, a self-described "community research lab." Terabytes/second — ~500 hours of HD video per single second of scan. ▸ Tool Shed: Hermes Agent (Nous Research) — the plugin-maximalist opposite of a minimal harness like Pi: built-in memory, a self-learning skill loop, cron scheduling, swappable memory providers, and ~20 chat channels out of the box. Dan: "parachuting in with sixteen crates of supplies and a film crew." ▸ Is AI ruining our skills? (Nature) — physicians' precancerous-lesion detection fell from 28.4% to 22.4% once the AI tool was removed; 52 engineers scored 50% on understanding their own code with AI vs 67% without. Cognitive debt is showing up in the data. ▸ Claude Code is a video game (Provi.me) — the "one more prompt" loop that keeps you up three hours past bedtime, and why AI finally made B2B SaaS addictive. Plus the "agent dice" repo: roll a natural 20 and a stop hook makes the agent reflect and write itself a skill. ▸ Evaluation-Driven Development (Decoding AI) — treat every AI feature as a hypothesis and gate the PR on an offline eval pipeline (built on Opik) instead of unit tests. Gold-standard vs synthetic datasets, code-metric vs LLM-as-judge evaluators, and an "aggression" dial for how big a jerk your reviewer is. (Shimin: Newtonian physics → quantum mechanics.) ▸ Two Minutes to Midnight — ChatGPT slips under 50% share (46.4%; Gemini 27.7%, Claude 10.3%), Nvidia raises $25B in its first bond deal since 2021, and Ed Zitron walks OpenAI's FT-verified financials ($38.5B loss in 2025). ~2B users — one in four people on Earth; no 10x left. Clock moved up to 5:00. ⏱ Chapters 00:00 Cold Open & Welcome 01:50 News: SpaceX Buys Cursor for $60B 04:46 News: MidJourney Pivots to Body-Scan Spas 11:45 Tool Shed: Hermes Agent (Nous Research) 19:54 Post-Processing: Is AI Ruining Our Skills? (Nature) 27:13 Post-Processing: Claude Code Is a Video Game 35:23 Post-Processing: Evaluation-Driven Development (EDD) 41:44 Two Minutes to Midnight: ChatGPT Under 50%, Nvidia Debt, OpenAI's Numbers 55:06 Outro 🔗 Articles we discussed News: • SpaceX to acquire Cursor — CNBC: https://www.cnbc.com/2026/06/16/spacex-spcx-cursor-acquisition-ipo.html • MidJourney's medical pivot — MidJourney: https://www.midjourney.com/medical/blogpost Tool Shed: • Hermes Agent docs — Nous Research: https://hermes-agent.nousresearch.com/docs/ Post-Processing: • Is AI ruining our skills? Early results are in — Nature: https://www.nature.com/articles/d41586-026-01947-1 • Claude Code is a video game — Provi.me: https://provi.me/cc-like-video-games • How Evaluation-Driven Development (EDD) works — Decoding AI (Paul Easton & Alejandro Aboy): https://www.decodingai.com/p/5b766861-0001-494f-a37f-4d4eb104dcfa Two Minutes to Midnight: • ChatGPT's market share slips below 50% for the first time — TechCrunch: https://techcrunch.com/2026/06/16/chatgpts-market-share-slips-below-50-for-first-time/ • Nvidia seeks to raise over $25B in first bond deal since 2021 — Ars Technica: https://arstechnica.com/ai/2026/06/chipmaker-nvidia-seeks-to-raise-over-25b-in-first-bond-deal-since-2021/ • Exclusive: OpenAI's financials — Where's Your Ed At (Ed Zitron): https://www.wheresyoured.at/exclusive-openai-financials/ 🎙 About ADI Pod ADI Pod (Artificial Developer Intelligence) is a weekly podcast about AI and software development for working developers. We go through hundreds of links and dozens of newsletters each week so you don't have to. Hosts: Shimin Zhang, Dan Lasky, Rahul Yadav. New episodes Tuesdays. • https://www.adipod.ai • humans@adipod.ai If something here changed your mind or gave you something to try on Monday, hit subscribe and leave a comment with what you tried.

    • Transcript
  • #30
    June 19 · 1 hr 10 min

    Fable 5 Ban, Meta's AI Gulag, Elias Thorne & What is Loop Engineering?

    Three days after Fable 5 launched, the US government banned it — for every foreign national on Earth, including Anthropic's own employees. Then it got weirder. This week on ADI Pod: the Fable 5 export ban, Meta's applied-AI "gulag," the Elias Thorne dataset virus, loop engineering, a local DeepSeek V4 demo, the paper that shatters Dunning-Kruger, and NBER's bubble math. Co-hosts: Shimin Zhang, Dan Lasky, Rahul Yadav. ▸ Fable 5 & Mythos 5, export-banned — a national-security order cut access for all foreign nationals (even Anthropic's own staff) in ~90 minutes, reportedly after an AWS jailbreak claim; likely the end of universal frontier-model access. Shimin had a near-"AI psychosis" moment using it to design a novel drone. ▸ Meta's "AI Gulag" — Alexandr Wang's unit drafts laid-off engineers to write puzzles and label data to train Meta's weaker models, on full salary and RSUs; the "gulag" label is a stretch, but the internal drama is real. ▸ The Elias Thorne mystery (404 Media) — a lighthouse keeper seeded by ~111 ChatGPT-3.5 chats became a "dataset virus" now in ~88% of AI stories and "authoring" books on Amazon across every lab (Cornell's Hamilton & Mimno). ▸ AI is fast, the economy isn't (howfastis.ai) — task horizons double every ~6 months, but weak-link / Theory-of-Constraints bottlenecks (Chad Jones; Goldratt) keep growth near 2%/yr; human judgment is the constraint AI can't yet remove. ▸ Loop engineering (Addy Osmani) — six pieces turn a bare /loop (Ralph loop) into a real agent harness: automations, worktrees, skills, plugins/connectors, subagents (split the worker from the reviewer), and memory. It amplifies whatever judgment you bake into your skills. ▸ Deep Dive — "Beyond the Steeper Curve" (Christopher Koch) — AI doesn't steepen Dunning-Kruger, it shatters it: "metacognitive decoupling" unglues output quality from self-assessment. Plus the "slop grenade" and the sycophancy trap (No One's Happy). ▸ Vibe & Tell — Dan runs DeepSeek V4 Flash locally ("DS4," the dwarf star runner) on a Framework Ryzen 395 Max over ROCm, ~14 tok/s, wired to Pi agent — ~$4,000 of hardware, no cloud. ▸ Two Minutes to Midnight — Claude on Apple's foundation-model backend (a commoditization tell), the end of subsidized inference, and an NBER paper pricing genuine insolvency risk into the AI build-out. Clock set back to 5:30. ⏱ Chapters 00:00 Cold Open & Welcome 02:01 News: The US Government Bans Fable 5 & Mythos 5 10:35 News: Meta's "AI Gulag" (feat. Rahul) 14:39 Post-Processing: The Elias Thorne Mystery 21:12 Post-Processing: AI Is Fast, the Economy Isn't (howfastis.ai) 29:22 Post-Processing: Loop Engineering (Addy Osmani) 36:37 Deep Dive: Beyond the Steeper Curve (Dunning-Kruger, Shattered) 43:46 Deep Dive: Appearing Productive & the Slop Grenade 51:51 Vibe & Tell: DeepSeek V4 Flash at Home (DS4) 57:39 Two Minutes to Midnight: Apple Foundation Models, Cheaper Inference, NBER Bubble Math 1:08:44 Outro 🔗 Articles we discussed News: • Fable & Mythos access update — Anthropic: https://www.anthropic.com/news/fable-mythos-access • Anthropic lobbies the White House over the Mythos/Fable ban — Axios: https://www.axios.com/2026/06/14/anthropic-white-house-mythos-fable • Meta's months-old AI unit is a "soul-crushing gulag," say the engineers stuck inside it — TechCrunch: https://techcrunch.com/2026/06/12/metas-months-old-ai-unit-is-a-soul-crushing-gulag-say-the-engineers-stuck-inside-it/ Post-Processing: • Chatbots keep telling stories about lighthouse keeper Elias Thorne — 404 Media: https://www.404media.co/elias-thorne-chatbots-llms-chatgpt-lighthouse-keeper-story/ • How Fast Is AI? — Emory Taziki: https://howfastis.ai/ • Loop Engineering — Addy Osmani: https://addyosmani.com/blog/loop-engineering/ Deep Dive: • Beyond the Steeper Curve: AI-Mediated Metacognitive Decoupling and the Limits of the Dunning-Kruger Metaphor — Christopher Koch (arXiv): https://arxiv.org/html/2603.29681 • Appearing Productive in the Workplace — No One's Happy: https://nooneshappy.com/article/appearing-productive-in-the-workplace/ Two Minutes to Midnight: • Claude SDK for Apple Foundation Models — Claude Platform docs: https://platform.claude.com/docs/en/cli-sdks-libraries/libraries/apple-foundation-models • Can tech companies learn to love cheaper AI models? — TechCrunch: https://techcrunch.com/2026/06/09/can-tech-companies-learn-to-love-cheaper-models/ • What Investment Data Implies About the AI Transition — NBER Working Paper w35290 (Walter & Walter): https://www.nber.org/system/files/working_papers/w35290/w35290.pdf 🎙 About ADI Pod ADI Pod (Artificial Developer Intelligence) is a weekly podcast about AI and software development for working developers. We go through hundreds of links and dozens of newsletters each week so you don't have to. Hosts: Shimin Zhang, Dan Lasky, Rahul Yadav. New episodes Fridays. • https://www.adipod.ai • humans@adipod.ai If something here changed your mind or gave you something to try on Monday, hit subscribe and leave a comment with what you tried.

    • Transcript
  • #29
    June 12 · 1 hr 8 min

    Claude Fable 5 is Here! Plus: Meta AI Hack, LLMs as Black Boxes, and Future of Agents

    "I think this is the first time ever where the user has no control over which model you're actually using." — Shimin on Fable 5's safety fallback, which answers blocked questions with Opus 4.8 instead. This week on ADI Pod: Anthropic ships Fable 5 and Mythos 5, hackers sweet-talk Meta's AI support bot into handing over Instagram accounts (including Obama's), interpretability opens up the black box, Chris Roth's Future of Agents, and Shimin demos Inhabited Design. Rahul's back; the clock moves to 5:20. ▸ Fable 5 & Mythos 5 (Anthropic) — mostly incremental benchmarks, a step change in spatial reasoning (Blueprint Bench 2), scales cleanly with thinking tokens where Opus 4.8 zigzags, and it beats Pokémon FireRed on vision alone. Fable 5's safety classifier blocks cybersecurity and biohazard prompts and answers with Opus 4.8 instead; Mythos 5 stays unguarded for the Project Glasswing cohort. Subscription access ends June 22, then usage-based pricing only. ▸ The Meta AI Hack (Krebs on Security) — the recipe spread on Telegram from May 31: VPN exit near the target's hometown, then ask Meta's AI support assistant to send reset codes to your new email. It complies. Pro-Iranian hackers defaced the Obama White House Instagram and a Space Force account. Why does a support bot hold elevated permissions to replace a boring form? Red teams will make bank. ▸ LLMs Are Not the Black Box You Were Promised (Jay Hack) — on Anthropic's "On the Biology of a Large Language Model." Circuit tracing with a sparse replacement model — one neuron per concept — shows features firing in sequence: Dallas, then Texas, then Austin. Models plan rhymes ahead of the line, and they confabulate how they do math. Shimin: metacognition may be the durable human advantage. ▸ The Future of Agents (Chris Roth) — personal agents, bring-your-own-agent trust boundaries, super apps as agent clients, open standards (MCP, A2A, AG-UI), enterprise open source, generative UI. The hosts push back: code is cheap enough that adapters kill any standard's network effects. Shimin migrated his Pi-agent skills to Claude Code with one prompt. There is no AI moat. ▸ Vibe and Tell: Inhabited Design — Shimin's open-source skill for escaping RLHF attractor states (the same stock ticker and Bloomberg yellow on every finance landing page). Verbalized sampling plus intent-factored generation: sample uniformly over designer, typography, and inspiration, then run two convergence loops. Dan's verdict: "it still feels like a human designer did it." ▸ Two Minutes to Midnight — Google will pay SpaceX $920M a month for compute at xAI data centers; Google already owns a chunk of SpaceX. The S&P 500 won't bend its rules for SpaceX and OpenAI fast-track inclusion. Alphabet upsizes its equity raise to $84.75B against roughly $190B in 2026 capex. Founders Fund's $20M SpaceX check from 2008 is now worth $26–52B. Clock: 5:30 → 5:20. ⏱ Chapters 00:00 Welcome & Rundown 01:20 News Threadmill: Fable 5 & Mythos 5 Launch 10:59 News Threadmill: The Meta AI Support Bot Hack 18:39 Post Processing: LLMs Are Not the Black Box You Were Promised 28:15 Post Processing: The Future of Agents 51:22 Vibe and Tell: Inhabited Design 57:49 Two Minutes to Midnight: SpaceX, S&P 500, Alphabet 1:07:47 Outro 🔗 Articles we discussed News: • Claude Fable 5 — Anthropic: https://www.anthropic.com/claude/fable • Hackers Used Meta's AI Support Bot to Seize Instagram Accounts — Krebs on Security: https://krebsonsecurity.com/2026/06/hackers-used-metas-ai-support-bot-to-seize-instagram-accounts/ Post Processing: • LLMs are not the Black Box you were promised — Jay Hack: https://www.jay.ai/blog/llms-are-not-a-black-box • The Future of Agents — Chris Roth: https://cjroth.com/blog/2026-06-03-future-of-agents Vibe and Tell: • Inhabited Design — Shimin Zhang (Wolf Peach Labs): https://wolfpeachlabs.com/inhabited-design/ Two Minutes to Midnight: • Google to pay SpaceX $920M a month for xAI compute — CNBC: https://www.cnbc.com/2026/06/05/google-to-pay-spacex-920-million-a-month-for-xai-compute-capacity.html • S&P 500 blocks fast SpaceX entry — Ars Technica: https://arstechnica.com/tech-policy/2026/06/sp-500-blocks-fast-spacex-entry-wont-waive-rule-for-unprofitable-ai-firms/ • Alphabet to raise $84.75B in upsized equity offering — Reuters: https://www.reuters.com/legal/transactional/alphabet-raise-8475-billion-upsized-equity-offering-fund-ai-ambitions-2026-06-03/ 🎙 About ADI Pod ADI Pod (Artificial Developer Intelligence) is a weekly podcast about AI and software development for working developers. We read hundreds of links and newsletters each week so you don't have to. Hosts: Shimin Zhang, Dan Lasky, Rahul Yadav. New episodes Tuesdays. • https://www.adipod.ai • humans@adipod.ai If this gave you something to try on Monday, subscribe and tell us. (00:00) - Welcome & Rundown (01:20) - News Threadmill: Fable 5 & Mythos 5 Launch (10:59) - News Threadmill: The Meta AI Support Bot Hack (18:39) - Post Processing: LLMs Are Not the Black Box You Were Promised (28:15) - Post Processing: The Future of Agents (51:22) - Vibe and Tell: Inhabited Design (57:49) - Two Minutes to Midnight: SpaceX, S&P 500, Alphabet (01:07:47) - Outro

    • Transcript
    • Chapters
  • #28
    June 5 · 54 min

    Claude Opus 4.8, Undocumented Claude Code Features, Eval Harness for AI Skills, Pope on AI

    "Every time you vibe code, you're gonna spend compute to skip the bottleneck of code review." — Shimin's cold open on ep-28. Trading compute for human labor: that's the default that's coming. This week on ADI Pod: Claude Opus 4.8 and Anthropic's dynamic-workflow tool, Pope Leo XIV's AI encyclical, a deep read of the Claude Code source code, a Pinterest method for testing whether your AI skills actually fire, two essays on senior engineering and the "dead economy," and a bubble check full of S-1s. Rahul's out this week; the clock moves up to 5:30. ▸ Claude Opus 4.8 + the Dynamic Workflow Tool (TechCrunch): a 41-day fast-follow to 4.7. The new "dynamic workflow" is extra-high thinking plus a huge fan-out of coordinated parallel agents — the hosts call it "Gastown, by Anthropic." Dan likes it more than 4.7, but it hallucinated file names that don't exist and ate a full token budget in 25 minutes. Likely a Mythos distill, not a new base model. ▸ Pope Leo XIV's AI Encyclical — "Magnifica Humanitas" (Vatican): "On Safeguarding the Human Person in the Time of Artificial Intelligence." The Pope gets that models are grown, not developed, warns against pretending AI is neutral, and ties automation to worker protection. Anthropic's Chris Olah was in the room. Shimin's take: better AI takes than most Fortune 500 CEOs. ▸ I Read the Claude Code Source Code (Building Better): the undocumented stuff. A pre-tool-use hook can rewrite a tool's input mid-flight, return allow/deny with a reason, and inject context. Skills take undocumented front-matter (model + effort). Plus where settings.json really lives, and the auto-memory and "dream" toggles. ▸ Technique Corner — An Engineer's Guide to Better AI Skills (Pinterest): a test harness for skill invocation — 15 positive prompts, 5 negative, 5 runs each. Codex went 73%→95% with everything combined; Claude went 62%→73% on a single change and got worse when you combined them. Asking the AI to improve the skill didn't help. ▸ Post Processing — Is This Sustainable? (Jamie Hurst): seniors absorbed AI's rising stakes before juniors did. You skip the RFC and just build the thing. The scary part: AI depth is perishable in ~18 months; what lasts is taste and judgment. ▸ Post Processing — The Dead Economy Theory (Owen McGrann): a turn-by-turn case that replacing workers with AI eats its own market. Peter Thiel, a Valley misread of Nietzsche, and UBI. Shimin pushes back while half-infected by the inevitability virus. ▸ Two Minutes to Midnight (SEC + Qazinform): SpaceX's S-1 claims a $26.5T market that's mostly "AI" and says "truth seeking" 39 times. Anthropic overtakes OpenAI as the most valuable AI startup on a $65B Series H (~3x its February mark) plus a confidential S-1. Microsoft pulls Claude Code back to Copilot on cost. Clock -> 5:30. ⏱ Chapters 00:00 Cold Open & Welcome 02:09 News: Claude Opus 4.8 & the Dynamic Workflow Tool 08:27 News: Pope Leo XIV's AI Encyclical 14:17 ToolShed: I Read the Claude Code Source Code 22:10 Technique Corner: Do Your AI Skills Actually Fire? 29:15 Post Processing: Is This Sustainable? (Senior Eng in the AI Age) 35:30 Post Processing: The Dead Economy Theory 42:12 Two Minutes to Midnight: SpaceX's S-1 & the $26.5T "AI" TAM 45:50 Two Minutes to Midnight: Anthropic Overtakes OpenAI 53:35 Outro 🔗 Articles we discussed News: • Anthropic releases Opus 4.8 with new dynamic workflow tool — TechCrunch: https://techcrunch.com/2026/05/28/anthropic-releases-opus-4-8-with-new-dynamic-workflow-tool/ • Magnifica Humanitas (encyclical on AI) — Pope Leo XIV / Vatican: https://www.vatican.va/content/leo-xiv/en/encyclicals/documents/20260515-magnifica-humanitas.html • I Read the Claude Code Source Code — Building Better: https://buildingbetter.tech/p/i-read-the-claude-code-source-code Technique Corner: • An Engineer's Guide to Better AI Skills — Pinterest Engineering: https://medium.com/pinterest-engineering/an-engineers-guide-to-better-ai-skills-implementing-a-testing-process-to-optimize-agent-a000c9c9abcd Post Processing: • Is This Sustainable? — Jamie Hurst: https://jamiehurst.co.uk/2026-05-24_ai-sustainable • The Dead Economy Theory — Owen McGrann: https://www.owenmcgrann.com/p/the-dead-economy-theory Two Minutes to Midnight: • SpaceX (Space Exploration Technologies) Form S-1 — SEC EDGAR: https://www.sec.gov/Archives/edgar/data/1181412/000162828026036936/spaceexplorationtechnologi.htm • Anthropic surpasses OpenAI to become world's most valuable AI startup — Qazinform: https://qazinform.com/news/anthropic-surpasses-openai-to-become-worlds-most-valuable-ai-startup 🎙 About ADI Pod ADI Pod (Artificial Developer Intelligence) is a weekly podcast about AI and software development for working developers. We go through hundreds of links and dozens of newsletters each week so you don't have to. Hosts: Shimin Zhang, Dan Lasky, Rahul Yadav. New episodes Tuesdays. • https://www.adipod.ai • humans@adipod.ai If something here gave you something to try on Monday, subscribe and tell us what you tried. (00:00) - Cold Open & Welcome (02:09) - News: Claude Opus 4.8 & the Dynamic Workflow Tool (08:27) - News: Pope Leo XIV's AI Encyclical (14:17) - News / Tool Shed: I Read the Claude Code Source Code (22:10) - Technique Corner: Do Your AI Skills Actually Fire? (29:15) - Post Processing: Is This Sustainable? (35:30) - Post Processing: The Dead Economy Theory (42:12) - Two Minutes to Midnight: SpaceX's S-1 (45:50) - Two Minutes to Midnight: Anthropic Overtakes OpenAI (53:35) - Outro

    • Transcript
    • Chapters
  • #27
    May 29 · 56 min

    OpenAI Beats Musk, Gemini 3.5 Flash & AI Burnout Mitigation

    "Sam Altman won in court against Elon Musk. But, really, we all lost." That's the New Yorker headline Dan brought to ep-27 — and the question under it is whether any one person should own AI safety. This week on ADI Pod: the OpenAI–Musk verdict and who really owns AI safety, Gemini 3.5 Flash in AI Overviews, a $48K home GPU server, AI burnout from two angles, the "$100M startup in your laptop" myth, the slop grenade, and an IPO squeeze that could funnel ~10% of the major indexes into three AI firms. Rahul's out this week; clock moves back to 6:15. ▸ OpenAI v. Musk (The New Yorker): OpenAI wins on a statute-of-limitations technicality. Musk's lawyer argues "we could all die" from AI; the judge notes he'd mean it more if he didn't fund xAI. The courtroom "butt pillows" become the complacency metaphor. ▸ Gemini 3.5 Flash: shipped Flash-only into AI Overviews (Dan's bet: it's on TPUs). Mathier than 3.1 but fewer results; the viral "can't search 'disregard'" bug was a harness failure. The pelican it drew looks dressed for a Miami crypto conference (h/t Simon Willison). ▸ Hardware Hut — was a $48K GPU server worth it? (rosmine.ai): an ex-FAANG researcher's 6× RTX 6000 Ada rig breaks even near 80% utilization, then ~$125/month and constant riser failures. Shimin's version: a 128GB Mac for local models, or keep paying Anthropic? ▸ Technique Corner — AI burnout (Evil Martians + Siddhant Khare): cap parallel agents at 3–4, keep hands on the keyboard, accept 70% and hand-code the rest. Shimin's confession: seven Claude Code sessions after work. Capper: Microsoft cancels Claude Code subs after costs top human devs. ▸ Post Processing — Human Bottlenecks (borretti.me): the $100M startup in your laptop stays there because the limiter was always you — judgment, energy, executive function — not the tools. ▸ Dan's Rant — the Slop Grenade (noslopgrenade.com): paste raw Claude output at a coworker instead of an answer and you've thrown one. The successor to nohello.com. Fix: lead with your one-line take, then attach the output for the full kaboom. ▸ Two Minutes to Midnight (Morningstar + Is AI Profitable Yet?): SpaceX/OpenAI/Anthropic could add ~10% to the Morningstar 100; Nasdaq cut its post-IPO wait from 12 months to ~15 trading days. On isaiprofitable.com only Nvidia is green (+$253B); Amazon leads capex at −$291B. Clock → 6:15. 🔗 Articles we discussed News: • OpenAI Won, But We All Lost — The New Yorker: https://www.newyorker.com/news/letter-from-silicon-valley/sam-altman-won-in-court-against-elon-musk-but-really-we-all-lost • Gemini 3.5: Frontier Intelligence With Action — Google: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5/#gemini-3-5-flash • Gemini 3.5 Flash hands-on — Simon Willison: https://simonwillison.net/2026/May/19/gemini-35-flash/ Hardware Hut: • Was My $48K GPU Server Worth It? — rosmine.ai: https://rosmine.ai/2026/05/13/was-my-48k-gpu-worth-it/ Technique Corner: • AI-Assisted Engineers Are Burning Out — Evil Martians: https://evilmartians.com/chronicles/ai-assisted-engineers-are-burning-out-is-this-fine • AI Fatigue Is Real — Siddhant Khare: https://siddhantkhare.com/writing/ai-fatigue-is-real Post Processing: • Human Bottlenecks — borretti.me: https://borretti.me/article/human-bottlenecks Dan's Rant: • No Slop Grenade: https://noslopgrenade.com Two Minutes to Midnight: • The SpaceX IPO: How US Index Funds Will Adapt — Morningstar (Zachary Evans): https://global.morningstar.com/en-ca/funds/spacex-ipo-how-us-stock-index-funds-will-adapt • Is AI Profitable Yet?: https://isaiprofitable.com/ 🎙 About ADI Pod ADI Pod (Artificial Developer Intelligence) is a weekly podcast about AI and software development for working developers. We go through hundreds of links and dozens of newsletters each week so you don't have to. Hosts: Shimin Zhang, Dan Lasky, Rahul Yadav. New episodes Tuesdays. • https://www.adipod.ai • humans@adipod.ai If something here gave you something to try on Monday, subscribe and tell us what you tried. (00:00) - Cold Open & Welcome (02:15) - News: OpenAI Beats Musk — "We All Lost" (The New Yorker) (07:20) - News: Gemini 3.5 Flash & the Pelican Test (13:05) - Hardware Hut: Was a $48K GPU Server Worth It? (19:14) - Technique Corner: AI-Assisted Engineers Are Burning Out (28:27) - Microsoft Cancels Claude Code — the Token Pendulum (30:35) - Post Processing: Human Bottlenecks (36:10) - Raising the First AI-Native Generation (40:05) - Dan's Rant: The Slop Grenade (44:49) - Two Minutes to Midnight: The SpaceX IPO Index Squeeze (49:32) - Is AI Profitable Yet? (Capex by the Billions) (55:20) - Outro

    • Transcript
    • Chapters
  • #26
    May 22 · 1 hr 10 min

    LLM Neuralanatomy with David Noel Ng, Forward Deployed Everybody, Preferences Revealed by AI

    This week on ADI Pod: Mira Murati's Thinking Machines ships its first product (interaction models), Meta employees fight the mouse-tracking program with flyers, and the Palantir-coined "forward deployed engineer" job title quietly takes over the post-AI engineering org chart. Our sit-down is with Dr. David Noel Ng (https://dnhkng.github.io/) — author of the LLM Neuroanatomy series we covered a few weeks back. He explains why he watched action potentials race down rat neurons at 10,000 fps before getting into LLM interpretability. Dan runs DeepSeek-V4 Flash on a 128 GB Ryzen 395 Max box and vibe-codes an ESP32 home dashboard in C. Deep dive on a paper asking whether AI should obey what you say or what you actually do. And in Two Minutes to Midnight: Cerebras pops 108% on IPO day, Anthropic passes OpenAI on Ramp business data, and we read Andy Hall on the politics of jobless prosperity. Clock stays at 6 minutes. Co-hosts: Shimin Zhang, Dan Lasky, Rahul Yadav. ▸ Interaction Models — Thinking Machines' first product. A small Qwen 3.5–class "interaction model" runs the UI; a background model handles the heavy lift. They credit the pattern to Qwen, not themselves. Multimodal by default, and surprisingly snappy in the demo. ▸ Meta vs. its own employees — Meta posted flyers around its offices reading "don't want to work at the employee data extraction factory?" after announcing it would record keystrokes, mouse movements, and screens to train internal AI. Last episode's "Model Capability Initiative" story got worse. ▸ Here Comes Forward Deployed Everybody (Scott Werner / works on my machine) — Salesforce moves to an API-only data model. Palantir's "forward deployed engineer" title (originally called "delta") becomes the new pit-crew role across every department. We argue Jevons paradox vs. just-rebranding-the-least-glamorous-job: 20 marketers + 5 pit crews → 30 marketers + 10 pit crews as productivity rises. ▸ Sit-down: Dr. David Noel Ng on LLM Neuroanatomy — fluorescent dyes that change color with membrane voltage, what brain-microchip interfacing taught him about feature attribution, and why interpretability deserves the same first-principles rigor as wet-lab biology. New posts coming. ▸ Vibe Intel — Dan runs Antires's specialized DeepSeek-V4 Flash fork of llama.cpp. Q2 quant on the front, full experts on the back, SSD-cached prefills, ~10 tokens/sec on a Ryzen AI Max+ 395 with 128 GB unified memory. Output peg: Sonnet 4.5-ish. Plus an ESP32 / ESPHome dashboard with a several-thousand-line vibe-coded C lambda that, against all reasonable expectation, works. ▸ Deep Dive — "Should I State or Should I Show?" (Keaton Ellis & Wanying Huang) — three AIs given the same lottery decisions: prompt-only AI hit 70% match with the human; data-only AI hit 75%; both-AI dropped to the worst of the three because it defaulted to the prompt 66% of the time when prompt and behavior conflicted. Implication: if EU AI Act transparency rules force you to honor stated preferences, you're literally picking the worst-performing model. ▸ Two Minutes to Midnight — Cerebras raises $5.5B in IPO, stock pops 108% on day one (claim: ~80× memory throughput vs comparable NVIDIA GPUs). Anthropic now holds 34.4% of Ramp-card-paying businesses, beating OpenAI for the first time. Andy Hall's "Politics of Jobless Prosperity": 2% unemployment jump is the line in the sand for political stability — opens with FDR's 1944 State of the Union. Clock held at 6 minutes; nothing this week jumped the needle. ⏱ Chapters 00:00 Cold Open & Welcome 02:08 News: Interaction Models from Thinking Machines 05:57 News: Meta Employees Protest the Mouse-Tracking Program 10:53 Post-Processing: Here Comes Forward Deployed Everybody (Scott Werner) 22:50 Sit-Down: Dr. David Noel Ng on LLM Neuroanatomy 40:05 Vibe N Tell: DeepSeek-V4 Flash at Home on a 395 Max 45:37 Vibe N tell: ESP32 Home Dashboards via Vibe Coding 48:51 Deep Dive: Should I State or Should I Show? (Ellis & Huang) 1:04:24 Two Minutes to Midnight: Cerebras IPO, Anthropic vs OpenAI, Jobless Prosperity 1:09:34 Outro 🔗 Articles we discussed News: • Interaction Models — Thinking Machines: https://thinkingmachines.ai/blog/interaction-models/ • Meta employees protest the mouse-tracking program — Engadget: https://www.engadget.com/2172212/meta-employees-are-protesting-the-companys-mouse-tracking-program/ Post-Processing: • Here Comes Forward Deployed Everybody — Scott Werner (works on my machine): https://worksonmymachine.ai/p/here-comes-forward-deployed-everybody Sit-Down — Dr. David Noel Ng: • Substack: https://dnhkng.substack.com/ • Site: https://dnhkng.github.io/ * Rest of David's home AI Lab Build Story: https://dnhkng.github.io/posts/hopper/ Deep Dive: • Should I State or Should I Show? Aligning AI with Human Preferences — Keaton Ellis & Wanying Huang (arXiv): https://arxiv.org/html/2603.29317v1 Two Minutes to Midnight: • Cerebras raises $5.5B, kicks off 2026's IPO season — TechCrunch: https://techcrunch.com/2026/05/14/cerebras-raises-5-5b-kicking-off-2026s-ipo-season-with-a-bang/ • Cerebras: Faster Tokens, Please — SemiAnalysis: https://newsletter.semianalysis.com/p/cerebras-faster-tokens-please • Anthropic now has more business customers than OpenAI (per Ramp data) — TechCrunch: https://techcrunch.com/2026/05/13/anthropic-now-has-more-business-customers-than-openai-according-to-ramp-data/ • The Politics of Jobless Prosperity — Andy Hall (Free Systems): https://freesystems.substack.com/p/the-politics-of-jobless-prosperity 🎙 About ADI Pod ADI Pod (Artificial Developer Intelligence) is a weekly podcast about AI and software development for working developers. We go through hundreds of links and dozens of newsletters each week so you don't have to. Hosts: Shimin Zhang, Dan Lasky, Rahul Yadav. New episodes Tuesdays. • https://www.adipod.ai • humans@adipod.ai If something in this episode changed your mind or gave you something to try on Monday, hit subscribe and leave a comment with what you tried. (00:00) - Cold Open & Welcome (02:08) - News: Interaction Models from Thinking Machines (05:57) - News: Meta Employees Protest the Mouse-Tracking Program (10:53) - Post-Processing: Here Comes Forward Deployed Everybody (Scott Werner) (22:50) - Sit-Down: Dr. David Noel Ng on LLM Neuroanatomy (40:05) - Vibe N Tell: DeepSeek-V4 Flash at Home on a 395 Max (45:37) - Vibe N tell: ESP32 Home Dashboards via Vibe Coding (48:51) - Deep Dive: Should I State or Should I Show? (Ellis & Huang) (01:04:24) - Two Minutes to Midnight: Cerebras IPO, Anthropic vs OpenAI, Jobless Prosperity (01:09:34) - Outro

    • Transcript
    • Chapters
  • #25
    May 15 · 1 hr 13 min

    Multi-Agent Patterns for 2026, Anthropic on Colossus, Brockman's Tesla Painting

    Anthropic finally fixed the compute crunch — by partnering with the one company OpenAI is currently being sued by. Plus Brockman's deposition journal drops, we unpack why a billion-token context window needs an entirely different GPU architecture, and Phil Schmid's four sub-agent patterns for 2026. This week on ADI Pod: Dan, Rahul and Shimin go heavy on the Elon news — the OpenAI lawsuit, the Anthropic + SpaceX/XAI compute deal that just lifted Claude Code's peak-hour limits, and the Wall Street Journal's data on Grok's user base collapsing (now ~1/30th of ChatGPT's). Then we move into the substance: NVIDIA's Rubin CPX architecture and disaggregation, Phil Schmid's four sub-agent patterns and where the agent-teams pattern is headed, Jack Clark's piece on recursive AI research automation, and Simon Willison reluctantly admitting he runs Claude Code with --dangerously-skip-permissions by default. We close with bubble watch — and move the clock further from midnight after a week with no major red flags. 🔗 Articles we discussed ▸ How Elon Musk left OpenAI, per Greg Brockman (TechCrunch) https://techcrunch.com/2026/05/06/how-elon-musk-left-openai-according-to-greg-brockman/ ▸ Anthropic raises Claude Code usage limits, credits SpaceX deal (Ars Technica) https://arstechnica.com/ai/2026/05/anthropic-raises-claude-code-usage-limits-credits-new-deal-with-spacex/ ▸ Anthropic-SpaceX AI deal (Wall Street Journal) https://www.wsj.com/tech/ai/anthropic-spacex-ai-deal-elon-musk-f86ea369?st=XUQnP7&reflink=desktopwebshare_permalink ▸ The road to a billion-token context (CACM) https://cacm.acm.org/news/the-road-to-a-billion-token-context/ ▸ Sub-agent patterns for 2026 — Phil Schmid https://www.philschmid.de/subagent-patterns-2026 ▸ Import AI 455: automating AI research — Jack Clark https://importai.substack.com/p/import-ai-455-automating-ai-research ▸ Vibe coding and agentic engineering are getting closer than I'd like — Simon Willison https://simonwillison.net/2026/May/6/vibe-coding-and-agentic-engineering/ ▸ You need AI that reduces your maintenance costs — James Shore https://www.jamesshore.com/v2/blog/2026/you-need-ai-that-reduces-your-maintenance-costs ▸ Anthropic reportedly agrees to pay Google $200B for chips and cloud access (Engadget) https://www.engadget.com/2165585/anthropic-reportedly-agrees-to-pay-google-200-billion-for-chips-and-cloud-access/ ▸ Silicon Valley bets on floating AI data centers powered by ocean waves (Ars Technica) https://arstechnica.com/ai/2026/05/silicon-valley-bets-on-floating-ai-data-centers-powered-by-ocean-waves/ ⏱ Chapters 00:00 Cold Open & Welcome 02:15 News: Elon vs OpenAI Trial Drama (Brockman's Journal & The Tesla Painting) 08:30 News: Anthropic Joins Colossus (SpaceX/XAI Compute Deal) 13:06 Hardware Hunt: The Road to a Billion-Token Context (NVIDIA Rubin CPX) 21:56 Technique: Phil Schmid's 4 Sub-Agent Patterns for 2026 30:11 Post-Processing: Jack Clark — AI Systems Are About to Build Themselves 45:17 Post-Processing: Simon Willison — Vibe Coding & Agentic Engineering 55:14 Post-Processing: James Shore — AI That Reduces Maintenance Costs 1:01:39 Two Minutes to Midnight 1:12:05 Outro 🎙 About ADI Pod ADI Pod (Artificial Developer Intelligence) is a weekly conversation show about AI and software development. We go through hundreds of links and dozens of newsletters each week so you don't have to. Hosts: Shimin Zhang, Dan Łaski, Rahul Yadav. 🌐 https://www.adipod.ai 📧 humans@adipod.ai 🦋 Shimin on Bluesky: @shiminsky.bsky.social If something in this episode changed your mind or gave you something to try on Monday, hit subscribe and leave a comment with what you tried. #ADIPod #AINews #ClaudeCode #Anthropic #AICoding (00:00) - Cold Open & Welcome (02:15) - News: Elon vs OpenAI Trial Drama (Brockman's Journal & The Tesla Painting) (08:30) - News: Anthropic Joins Colossus (SpaceX/XAI Compute Deal) (13:06) - Hardware Hunt: The Road to a Billion-Token Context (NVIDIA Rubin CPX) (21:56) - Technique: Phil Schmid's 4 Sub-Agent Patterns for 2026 (30:11) - Post-Processing: Jack Clark — AI Systems Are About to Build Themselves (45:17) - Post-Processing: Simon Willison — Vibe Coding & Agentic Engineering (55:14) - Post-Processing: James Shore — AI That Reduces Maintenance Costs (01:01:39) - Two Minutes to Midnight (01:12:05) - Outro

    • Transcript
    • Chapters
  • #24
    May 8 · 1 hr 26 min

    OpenAI's Goblin Problem, 10 Lessons When Code Is Cheap, AI Addiction Loop

    Why does the leaked Codex CLI system prompt explicitly tell GPT-5.5 to never mention goblins, gremlins, raccoons, trolls, ogres, or pigeons? Why is OpenAI now gating its cyber model the same way it mocked Anthropic for gating Mythos last month? And what does it mean that Dan tried to write a personal project without Claude — and physically couldn't? Co-hosts Shimin Zhang, Dan Lasky, and Rahul Yadav cover these and more on ADI Pod #24. This week: GPT-5.5 Cyber's gated release, OpenAI's "Where the Goblins Came From" RLHF post-mortem, Adi Osmani's five patterns for long-running agents, Jesse Vincent's adversarial review prompt, Drew Brunig's 10 lessons for agentic coding, Ivan Turkovic's history of failed attempts to eliminate programmers, Nilay Patel's "software brain" thesis, the Nature paper showing warm AI models lose 10–30 percentage points of accuracy, and a $1.1B raise for an AI lab that wants to train without human data. ## In this episode ▸ **GPT-5.5 Cyber gating** — Sam Altman called Mythos's gated release "fear-based marketing" two months ago. Now OpenAI is doing the exact same thing with the GPT-5.5 cyber variant. Multi-tier model access (enterprise, government, research preview, cyber) is becoming the default — and Shimin worries the White House is about to add another gate. ▸ **The Goblin Problem** — OpenAI's Codex CLI prompt was open-sourced and turned out to include "never talk about goblins, gremlins, raccoons, trolls, ogres, pigeons." OpenAI's "Where the Goblins Came From" post-mortem reveals a textbook RLHF failure: a "nerdy persona" reward signal trained the model to mention goblins in 66.7% of nerdy responses, and the tic propagated through supervised fine-tuning to non-nerdy responses too. ▸ **Long-running agents (Adi Osmani / Elevate)** — Five patterns for agents that run for hours or days: checkpoints over zero-or-100 outputs, governing memory like microservices, ambient processing without forced human-in-the-loop, fleet orchestration, and budget circuit-breakers. Bonus: the running gag where Rahul realizes the post is essentially an ad for Google Enterprise Agent Platform. ▸ **Adversarial review prompts (Jesse Vincent / superpowers)** — A four-step technique for getting better code review out of agents: invoke "fresh eyes," dispatch competing subagents, promise a reward (a cookie), and threaten disappointment if they don't find N issues. ▸ **10 Lessons for Agentic Coding (Drew Brunig)** — Implement to learn, rebuild often, invest in end-to-end tests, document intent, keep specs in sync, find the hard stuff, automate the easy stuff, develop taste, agents amplify experience, and the kicker: agent code is "free as in puppies" — the puppy is free, but you have to feed it and walk it. ▸ **The Eternal Promise (Ivan Turkovic)** — A history of attempts to eliminate programmers from COBOL through 4GLs, CASE tools, the Japanese 5th Generation project, no-code/low-code, and now LLMs. Each abstraction layer expanded software jobs rather than replacing them. Shimin's reframe: "Software is calcified business process. Someone has to do the calcifying." ▸ **People Do Not Yearn for Automation (Nilay Patel / The Verge)** — Why Gen Z hopefulness about AI dropped to 18% (anger up to 31%), why America is uniquely AI-pessimistic, and what Nilay calls "software brain" — the Silicon Valley assumption that human life can be reduced to data and algorithms. Plus Anuradha Pandey's reframe: stop calling them social media, call them ad platforms. ▸ **Warm models lose accuracy** — A Nature paper finds AI models trained for warmth lose 10–30 percentage points of accuracy. A companion study shows humans trust warm models *more* even when they're wrong. Frontier labs now have an explicit incentive to train the warmest model, not the most accurate one. Plus: Richard Dawkins talks to "Claudia" for three days and concludes AI must be conscious. ▸ **Dan's Rant — The AI Addiction Loop** — Dan tries to build a Home Assistant TypeScript automation without Claude. Can't. "It felt like they had fundamentally broken my arm in a way that I can't do this task as quickly as I wanted to. That scares me a lot." Shimin: "We're running into the social media addiction loop in three months instead of a decade." ▸ **Two Minutes to Midnight** — OpenAI projects ChatGPT Plus dropping from 44M to 9M subscribers in 2026 while scaling the ad-supported tier from 3M to 112M (30×). David Silver raises $1.1B for Ineffable Intelligence — a no-human-data approach inspired by AlphaGo. Scout AI raises $100M for autonomous military vision-language-action models. Bubble Clock held at 4:00 minutes. ## Key takeaways — Reward hacking can propagate latent persona quirks through fine-tuning in ways the lab itself only catches when users surface them. — Memory drift, not raw context size, is the real ceiling for long-running agents. Govern memory like you govern microservices. — Code is free as in puppies, not free as in beer. The cost shifts to maintenance, security, and the new burden of maintaining your own automations. — Warm AI is an alignment trap: incentivized for trust over accuracy, weaponizable in authoritarian hands. — "You can outsource your thinking, but you can't outsource your understanding." — Karpathy, via Rahul. — AI addiction hits in three months. Social media took a decade. We are not ready for the time scale. ## Chapters (00:00) - Cold Open & Welcome (02:50) - News Threadmill: GPT-5.5 Cyber Gets Mythos-Style Gating (08:52) - News Threadmill: The Goblin Problem & RLHF Post-Mortem (13:52) - Tool Shed: Long-Running Agents (Adi Osmani) (25:52) - Technique Corner: Adversarial Review Prompts (Jesse Vincent) (30:59) - Technique Corner: 10 Lessons for Agentic Coding (Drew Brunig) (42:31) - Post-Processing: The Eternal Promise — A History of Attempts to Eliminate Programmers (01:02:10) - Post-Processing: People Do Not Yearn for Automation (01:09:08) - Post-Processing: Warm Models & The Sycophancy Trap (01:13:28) - Dan's Rant: Home Automation & The AI Addiction Loop (01:20:09) - Two Minutes to Midnight: OpenAI's 30× Ad-Tier, David Silver's $1.1B, Scout AI's Drones (01:25:55) - Outro ## Resources mentioned **News Threadmill — GPT-5.5 Cyber & The Goblin Problem** • TechCrunch — After dissing Anthropic for limiting Mythos, OpenAI restricts access to cyber too: https://techcrunch.com/2026/04/30/after-dissing-anthropic-for-limiting-mythos-openai-restricts-access-to-cyber-too/ • Ars Technica — Amid mythos-hyped cybersecurity prowess, researchers find GPT-5.5 is just as good: https://arstechnica.com/ai/2026/05/amid-mythos-hyped-cybersecurity-prowess-researchers-find-gpt-5-5-is-just-as-good/ • Ars Technica — OpenAI Codex system prompt includes explicit directive to never talk about goblins: https://arstechnica.com/ai/2026/04/openai-codex-system-prompt-includes-explicit-directive-to-never-talk-about-goblins/ • OpenAI — Where the Goblins Came From: https://openai.com/index/where-the-goblins-came-from/ ...

    • Transcript
    • Chapters
  • #23
    May 1 · 1 hr 30 min

    Why Models Over-Edit Your Code, Meta Keystroke Surveillance, Interviewing Engineers in the AI Age

    Is GPT-5.5 finally a 4.7-tier model? Did DeepSeek V4 just close the gap with Anthropic? And what does it mean that a senior ML engineer says he can't out-code Claude anymore? Co-hosts Shimin Zhang, Dan Lasky, and Rahul Yadav are joined by special guest Nathan Lubchenco — ML engineer and Substack author of *The future was yesterday* (https://nathanlubchenco.substack.com/) — on ADI Pod #23 (April 28, 2026). This episode covers OpenAI's GPT-5.5 release, DeepSeek V4 (1.6T base / 49B active params with 1M context), Meta's new Model Capability Initiative tracking US employee keystrokes and mouse movements, a Levenshtein-distance study on coding-model over-editing, the 2026 Stanford AI Index report, and a deep-dive interview on how to hire software engineers when the agents are already better at coding than the candidates. Key takeaways — Models are now consistently better at coding than even senior ML engineers, by their own admission. Late-2026 may be when they cross the median software engineer. — Coding-model over-editing is measurable (Levenshtein distance on boolean-flip tasks) and instruction-followable — explicit "minimum-edit" prompts close most of the gap. — The US is unusually a slow adopter of a major technological wave. Workplace AI usage is highest in emerging economies, not the developed world. — "The task is not the job" — humans remain indispensable on the bundling dimensions: catching what customers don't say, and avoiding interactions that end up on social media. — Software engineering interviews should include the candidate's personal harness, with company-provided API keys for equity. LeetCode optimizes for the wrong signal in 2026. — DeepSeek V4 closing the gap with Mythos in 3–6 months is what makes the bubble too geopolitically important to fail. Chapters (00:00) - Cold Open & Welcome (01:31) - News Threadmill: GPT-5.5, DeepSeek V4, Meta Watches Every Keystroke (12:28) - Post-Processing: Coding Models Are Doing Too Much (18:59) - Post-Processing: The Task Is Not the Job (Luis Garicano) (32:20) - Post-Processing: The 2026 Stanford AI Index Report (38:11) - Deep Dive: Interviewing Engineers in the AI Age (with Nathan Lubchenco) (45:05) - Deep Dive: Reforming Software Hiring — Take-Homes, Personal Harness, Equity (50:15) - Deep Dive: When Models Cross the Median Engineer (Late-2026 Prediction) (59:29) - Deep Dive: Why Code Review Is the Current Bottleneck (01:00:21) - Deep Dive: Should PRs Show the Prompt History? (01:02:27) - Dan's Rant: Anthropic Tested Removing Claude Code from the Pro Plan (01:05:44) - Rahul's Rampage: The Infinity Machine — Demis Hassabis & Corporate Gravity (01:14:32) - Two Minutes to Midnight: Bubble Clock Moves Back to 4:00 (01:26:30) - Outro Resources mentioned **Models & news** • OpenAI — Introducing GPT-5.5: https://openai.com/index/introducing-gpt-5-5/ • Engadget — DeepSeek promises its new AI model has world-class reasoning: https://www.engadget.com/ai/deepseek-promises-its-new-ai-model-has-world-class-reasoning-115733512.html • Reuters — Meta to start capturing employee mouse movements, keystrokes for AI training data: https://www.reuters.com/sustainability/boards-policy-regulation/meta-start-capturing-employee-mouse-movements-keystrokes-ai-training-data-2026-04-21/ **Post-processing articles** • "Coding Models Are Doing Too Much" — Levenshtein-distance over-editing study (nrehiew): https://nrehiew.github.io/blog/minimal_editing/ • Luis Garicano (Silicon Continent) — Why Desk Jobs Survive ("The task is not the job"): https://www.siliconcontinent.com/p/why-desk-jobs-survive-and-amodei • 2026 AI Index Report — Stanford Institute for Human-Centered AI: https://hai.stanford.edu/ai-index/2026-ai-index-report **Deep dive** • Nathan Lubchenco — Interviewing Software Engineers in the Age of AI: https://nathanlubchenco.substack.com/p/interviewing-software-engineers-in • Nathan Lubchenco — *The future was yesterday* Substack home: https://nathanlubchenco.substack.com/ **Dan's rant** • Ars Technica — Anthropic tested removing Claude Code from the Pro plan: https://arstechnica.com/ai/2026/04/anthropic-tested-removing-claude-code-from-the-pro-plan/ **Rahul's rampage** • Sebastian Mallaby — *The Infinity Machine* (book on Demis Hassabis and DeepMind) • Philipp Dubach — Do Not Disturb My Circles (Archimedes essay): https://philippdubach.com/posts/do-not-disturb-my-circles/ **Bubble watch** • TechCrunch — Two college kids raise $5.1M pre-seed to build an AI social network in iMessage: https://techcrunch.com/2026/04/24/two-college-kids-raise-a-5-1-million-pre-seed-to-build-an-ai-social-network-in-imessage/ • Toby Ord — Hourly Costs for AI Agents: https://www.tobyord.com/writing/hourly-costs-for-ai-agents • CNBC — OpenAI reportedly missed revenue targets, shares of Oracle and chip stocks falling: https://www.cnbc.com/2026/04/28/openai-reportedly-missed-revenue-targets-shares-of-oracle-and-these-chip-stocks-are-falling.html About ADI Pod ADI Pod (Artificial Developer Intelligence) is a weekly podcast about AI and software development for working developers. Co-hosts Shimin Zhang, Dan Lasky, and Rahul Yadav go through hundreds of links and dozens of newsletters every week so you don't have to. This week's special guest: **Nathan Lubchenco** — ML engineer and author of *The future was yesterday* on Substack, where he writes about AI and software engineering. • Website: https://www.adipod.ai • Email: humans@adipod.ai

    • Transcript
    • Chapters
  • #22
    April 24 · 1 hr 2 min

    Is Claude Opus 4.7 Mythos Distilled, Running Qwen 3.6 Locally, and the AI-On-AI Arena

    Is Claude Opus 4.7 really burning tokens? Is open source dead after mythos? Co-hosts Shimin Zhang and Dan Lasky — with recurring guest Rahul Yadav — ran the experiments this week on ADI Pod #22 (April 21, 2026). This episode covers Anthropic's Claude Opus 4.7 release (the "mythos slice"), Alibaba's open-source Qwen 3.6 35B A3B, cal.com going closed source for security reasons, and a HIPAA-violating vibe-coded patient portal that is, in Dan's words, the bullshit future already here. In this episode ▸ **Claude Opus 4.7 review** — the new mythos-derived tokenizer (3× bloat on plain English), stricter instruction-following, and why Shimin's SVG experiments suggest the token-burn panic is overblown: 35¢ on Opus 4.7 vs $2 on Opus 4.6 for the same task, with ~40× fewer reasoning tokens. ▸ **Qwen 3.6 35B A3B** — Alibaba's open-source mixture-of-experts model (3B active params at any time) running locally on Shimin's laptop at 90–95 tokens/sec via llama.cpp + Unsloth. The first model to break Simon Willison's pelican-on-a-bicycle benchmark against a larger frontier model. ▸ **cal.com goes closed source** — why the AI Security Institute's $12,000-per-attempt mythos pentesting data ($125,000 for 10 runs) is changing the open-source calculus, and Drew Breunig's three-phase dev/review/hardening cycle prediction. ▸ **Jesse Vincent's "Rules and Gates"** — a coding-agent prompting technique that reformulates optional preferences into directed preconditions, and whether agents can "weasel out" by rewriting the gate itself. ▸ **AI vibe coding horror story** — a German doctor who inlined a full patient portal into a single HTML page with database credentials client-side. HIPAA, meet DSGVO. ▸ **Kyle Kingsbury's "The Future of Everything is Lies"** — the Jepsen author's 8-step action list on AI's second- and third-order societal effects. ▸ **The AI-on-AI Arena** — Shimin's weekend project grading 11 frontier models against each other. The "delusion index" reads almost exactly like Dunning-Kruger in humans: GPT-5.4 scored -1.6 (humble), Gemini 3.1 Pro Preview rated itself well while peers ranked it last. ▸ **Two Minutes to Midnight** — Paul Graham's log-scale chart comparing AI capex (~1% of US GDP) to the US railroad peak (~10%). We dialed the AI bubble clock back 45 seconds to 3 min 30 sec. Key takeaways — Opus 4.7's token-burn reputation may be overblown. Stricter instruction-following can reduce total reasoning tokens by up to 40× vs Opus 4.6 on the same task. — Security-driven closed-sourcing may spread as mythos-class agents make open repos easier to exploit. Hardening could make software capital-intensive again. — Cognitive debt is real: Dan's wake-up call was a production bug a pre-LLM colleague solved in 5 minutes. His first instinct was to double down on the tool. — Shimin's defense against skill atrophy: read 100% of LLM-generated PR lines (except tests). — Weaker models rate themselves higher than stronger ones. Calibration appears to improve with capability. Chapters (00:00) - Introduction to AI and Software Development (02:25) - Alibaba's Quinn 3.6 Model Overview (08:06) - Anthropic's Claude Opus 4.7 Release (18:08) - Cal.com Goes Closed Source: Implications for Security (20:40) - The Future of Vibe Coding (23:41) - Techniques for Effective AI Utilization (27:13) - Post-Processing and AI in Real-World Applications (33:07) - The Cultural Impact of AI and Technology (41:30) - Navigating Code Review Challenges (42:57) - Exploring AI's Societal Impact (45:16) - Evaluating AI Models: Performance and Insights (49:09) - The Future of Data Centers and AI (50:54) - Investment Trends and Economic Perspectives (57:58) - Reflections on Historical Investment Cycles (59:35) - Optimism Amidst Uncertainty Resources mentioned Claude Opus 4.7 & Qwen 3.6 • Introducing Claude Opus 4.7 (Anthropic): https://www.anthropic.com/news/claude-opus-4-7 • Claude Opus 4.7 System Card: https://cdn.sanity.io/files/4zrzovbb/website/037f06850df7fbe871e206dad004c3db5fd50340.pdf • Qwen3.6-35B-A3B: Agentic Coding Power, Now Open to All: https://qwen.ai/blog?id=qwen3.6-35b-a3b • Simon Willison — Qwen3.6-35B-A3B on my laptop drew me a better pelican than Claude Opus 4.7: https://simonwillison.net/2026/Apr/16/qwen-beats-opus/ • Shimin — Opus 4.7 isn't dumb, it's just lazy: https://shimin.io/journal/opus-4-7-just-lazy/ Security & open source • Cal.com is going closed source. Here's why: https://cal.com/blog/cal-com-goes-closed-source-why • Drew Breunig — Cybersecurity Looks Like Proof of Work Now: https://www.dbreunig.com/2026/04/14/cybersecurity-is-proof-of-work-now.html Technique & commentary • Jesse Vincent — Rules and Gates: https://blog.fsck.com/2026/04/07/rules-and-gates/ • An AI Vibe Coding Horror Story: https://www.tobru.ch/an-ai-vibe-coding-horror-story/ • Kyle Kingsbury (Aphyr) — The Future of Everything is Lies, I Guess: https://aphyr.com/posts/411-the-future-of-everything-is-lies-i-guess Shimin's project • AI-on-AI Arena: https://shimin.io/ai-on-ai-arena Bubble watch • Ars Technica — Satellite and drone images reveal big delays in US data center construction: https://arstechnica.com/ai/2026/04/construction-delays-hit-40-of-us-data-centers-planned-for-2026/ • Epoch AI — OpenAI Stargate: where the US sites stand: https://epochai.substack.com/p/openai-stargate-where-the-us-sites • Paul Graham on US investment cycles (log scale): https://x.com/paulg/status/2045120274551423142/photo/1 About ADI Pod ADI Pod (Artificial Developer Intelligence) is a weekly podcast about AI and software development for working developers. Co-hosts Shimin Zhang and Dan Lasky go through hundreds of links and dozens of newsletters every week so you don't have to. Recurring guest Rahul Yadav joins when he can. • Website: https://www.adipod.ai • Email: humans@adipod.ai New episodes every Friday. Follow the show to get them automatically.

    • Transcript
    • Chapters
  • #21
    April 17 · 1 hr

    Anthropic Mythos & Project Glasswing, Recursive Improving Agents, and Your Parallel Agent Limit

    Shimin and Dan cover Minimax's M2.7 model — the first public experimental result in recursive self-improvement (RSI) — and unpack Anthropic's shock announcement of Mythos, a model so capable at finding security vulnerabilities that Anthropic is withholding public release while partnering with Amazon, Apple, Cisco, CrowdStrike, the Fed and major banks under 'Project Glasswing' to patch infrastructure first. They also debate AI's frontend weakness, discuss Addy Osmani's parallel agent limits piece, and move the AI bubble clock back. Takeaways: RSI is now experimentally demonstrated (not just theorized); reframes model improvement as capital competition, not PhD hiring. If AI finds vulns at scale, open source gets *more* secure long-term — but short-term this is a nuclear-test-equivalent event that may rewrite security, money, and trust assumptions. 'Frontend will be first automated' was wrong; backend may be easier because visual taste and pixel-perfect feedback loops aren't in training data Agent orchestration has a personal ceiling; finding it requires blowing past it. Tight scope + time-boxing + new contexts beats monolithic long sessions. 'Code is cheap' is really about industrialization — the people who industrialize outcompete those who don't; learn the tools or be left behind. OpenAI's CRO going public on a competitor's accounting is itself a bearish signal about OpenAI's enterprise position. Resources Mentioned MiniMax M2.7: The Agentic Model That Helped Build Itself Anthropic debuts preview of powerful new AI model Mythos in new cybersecurity initiative Assessing Claude Mythos Preview’s cybersecurity capabilities Why AI Sucks At Front End Your parallel Agent limit Code Is Cheap Now, And That Changes Everything The AI gold rush is pulling private wealth into riskier, earlier bets OpenAI CRO Tells Staff Anthropic Inflates Run Rate by $8 Billion Chapters (00:00) - Introduction to AI and Software Development (02:45) - Minimax M 2.7 Model and Recursive Self-Improvement (05:04) - Anthropic's Mythos Model and Security Vulnerabilities (08:15) - AI's Limitations in Front-End Development (18:13) - Cognitive Debt and Managing Multiple AI Agents (32:01) - Managing Multiple Agents Effectively (34:42) - The Evolution of Code Value (38:29) - The Industrialization of Coding (41:00) - Navigating Cloud Code Challenges (45:39) - Ranting About Technology Installations (50:16) - The State of the AI Bubble Connect with ADIPod Email us at humans@adipod.ai you have any feedback, requests, or just want to say hello! Checkout our website www.adipod.ai

    • Transcript
    • Chapters
  • #20
    April 10 · 53 min

    Ep 20: Claude Code Source Leak, Emotion Concepts in LLMs, and Surprising Facts AIs Know About Us.

    This week Rahul, Shimin, and Dan returns after a two-week break to cover the leaked Claude Code CLI source code, new model releases (Qwen 3.6 and Gemma 4), Mario Zechner's essay on slowing down with AI-assisted coding, a fun segment on unexpected things AI knows about each host, and two deep dives: Anthropic's research on emotion concepts in LLMs and a paper on how sycophantic AI decreases pro-social intentions. Takeaways: Claude Code's dual-track permission system uses both rule-based and ML classifier for destructive bash command "Cognitive bankruptcy" — when cognitive debt interest payments come due and you can't pay AI sycophancy parallels social media echo chambers; no market incentive to fix it On-device models like Gemma 4 could save cloud costs by handling routine tasks (e.g., agent heartbeats) Copilot's terms of service classify it as "for entertainment purposes only" Resources Mentioned Entire Claude Code CLI source code leaks thanks to exposed map file I Read the Leaked Claude Code Source — Here's What I Found The Claude Code Source Leak: fake tools, frustration regexes, undercover mode, and more Claude Code Unpacked Qwen3.6-Plus: Towards Real World Agents Gemma 4 Announcement Thoughts on slowing the fuck down Emotion concepts and their function in a large language model Sycophantic AI decreases prosocial intentions and promotes dependence Chapters (00:00) - Introduction and Host Updates (01:45) - Cloud Code Source Code Leak (12:49) - New Model News and Open Source Developments (20:51) - Post-Processing and AI Anxiety (25:35) - Unexpected Insights from AI (33:12) - Exploring Emotional Concepts in AI (39:15) - The Dangers of Sycophantic AI (52:39) - Concluding Thoughts and Future Considerations Connect with ADIPod Email us at humans@adipod.ai you have any feedback, requests, or just want to say hello! Checkout our website www.adipod.ai xqwkClUFaJkBwEMD1lUn

    • Transcript
    • Chapters
  • #19
    March 27 · 1 hr 13 min

    Ep 19: Thinking Fast Slow and Artificial, Meta's Trouble with Rogue Agents, and FOMO in the Age of AI

    This week, Rahul, Shimin, and Dan covers Claude Code's new channels and scheduling features, a Meta security incident caused by AI-generated advice, Anthropic's survey of 81,000 people on AI expectations, Dan's vibe-coded vector memory CLI project, a deep dive on the paper "Thinking, Fast, Slow and Artificial" about cognitive surrender to AI, a rant about AI tokens as employee compensation, and bubble watch updates including NVIDIA's trillion-dollar demand projections and OpenAI shutting down Sora. Takeaways: Claude Code is rapidly absorbing community-developed workflows — the moat may only be in the general model capabilities, not tooling The Meta incident illustrates the emerging pattern of AI-caused production incidents and the need for process guardrails around agent usage Cognitive surrender to AI creates a widening gap: those with high need-for-cognition benefit more while those who dislike effortful thinking defer even more AI confidence inflation (12 percentage point boost) may stem from treating AI like authoritative reference material (encyclopedias, Wikipedia) Historical technology resistance (Socrates on writing, farmers on tractors) suggests the battle against AI adoption may already be lost OpenAI shutting Sora just 4 months after a 3-year Disney partnership signals deeper financial or strategic issues Resources Mentioned Push events into a running session with channels Perhaps not Boring Technology after all Meta is having trouble with rogue AI agents What 81,000 people want from AI Dan's vec-memory-cli Thinking—Fast, Slow, and Artificial Are AI tokens the new signing bonus or just a cost of doing business? Jensen Huang just put Nvidia’s Blackwell and Vera Rubin sales projections into the $1 trillion stratosphere Accelerated FOMO in the Age of AI OpenAI shutters AI video generator Sora in abrupt announcement Chapters Connect with ADIPod Email us at humans@adipod.ai you have any feedback, requests, or just want to say hello! Checkout our website www.adipod.ai

    • Transcript
Showing 1–20 of 22 episodes