Skip to content
Artwork for Nerd Snipe with Theo and Ben

Nerd Snipe with Theo and Ben

Theo and Ben

"The only dev podcast hosted by real devs" - Theo (not a real dev)

Play
  • 20 episodes
  • weekly
  • Avg 1 hr 31 min
  • English
  • S1 · E20
    Friday · 2 hr 42 min

    Ox Alpha Revealed, OpenAI's Latest Pricing Updates, and Our Coding Model Tier List

    Theo & Ben break down OpenAI's GPT-5.6 Sol price cut, breakdown the "Ox Alpha" stealth model we now know is GLM 5.3 Flash, then take a drink every time they say "Grok" while ranking every current AI model on a tier list! What could go wrong? Thanks to this episode's sponsor, General Translation: General Translation: https://nerdsnipe.link/gt Listen wherever you get your podcasts: Spotify: https://nerdsnipe.link/spotify Apple: https://nerdsnipe.link/apple Elsewhere: https://nerdsnipe.link/listen Sources available on our Substack: https://nerdsnipe.substack.com/ Timestamps 0:00 Intro 3:30 OpenAI vs. Anthropic 14:44 Anthropic’s delayed models 20:53 Kimi K3 and open weights 30:08 The Alpha stealth model 48:26 Model tier list begins 1:09:53 GPT-5.6 Luna 1:13:09 Opus and Sonnet 1:20:25 Gemini models 1:31:15 DeepSeek and local models 1:59:58 Fable vs. Sol 2:28:46 Final rankings

  • S1 · E19
    August 20 · 2 hr 18 min

    Anthropic Doesn't Think You Can Be Trusted, China is closing the gap, SpaceXAI Leaps Ahead, and DEF CON

    Anthropic started watermarking their model's outputs, Dario posted on X, SpaceXAi is getting really good really fast, Meta is shipping again, and we won DEF CON 2026. Thank you PostHog, the all in one suite of product tools for sponsoring today's Episode! Check them out at: nerdsnipe.link/posthog Listen wherever you get your podcasts: - Spotify: nerdsnipe.link/spotify - Apple: nerdsnipe.link/apple - Elsewhere: nerdsnipe.link/listen Sources available on our Substack: nerdsnipe.substack.com Timestamps: 0:00 Intro 2:56 Claude Watermarks 15:04 Meta + Muse 48:48 GLM-5.3 55:15 Qwen + DeepSeek 1:08:48 Grok 4.6 + Bot 1:20:15 Gavin vs Dario 1:43:44 DEFCON 2:15:35 Viewer Q&A

  • S1 · E18
    July 28 · 1 hr 23 min

    Opus 5 Releases, China Catches Up, and the OpenAI Model Sandbox Escape

    Theo and Ben break down why Opus 5 feels like GPT-5.5 crossed with Fable rather than GPT-5.6 crossed with Fable, what Fable found when it audited an Opus thread line by line, where Opus still clearly wins (3D, animation, and Claude Code limits at half Fable's price with no 50% weekly cap), plus Kimi K3, GLM 5.2, the Hugging Face hack, Grok 4.5, Codex, and T3 Code. Thanks to this episode's sponsor, General Translation: General Translation: https://nerdsnipe.link/gt Sources available on our Substack: https://nerdsnipe.substack.com/

  • S1 · E17
    July 14 · 2 hr 6 min

    5 different models dropped last week & the GPT-5.6 usage limits are brutal

    This week Grok 4.5, Muse Spark 1.1, GPT-Live, and GPT 5.6 all dropped. ⁠Theo⁠⁠⁠ and ⁠⁠Ben⁠⁠ break down which models you should care about, how OpenAI fumbled the Codex to ChatGPT app transition, the abysmal usage numbers for 5.6, and the latest drama surrounding OpenAI and Sam Altman. Plus, how many satellites would it take to trap humanity on Earth, and how many data centers to heat up the ocean? Thank you to this episode's sponsors: Composio & WorkOs. https://nerdsnipe.link/composio https://nerdsnipe.link/workos Sources available on Substack: https://nerdsnipe.substack.com/

  • S1 · E16
    July 9 · 1 hr 11 min

    We Tested GPT 5.6 Sol Early

    We've spent six figures in tokens testing OpenAI's 5.6 Sol model to see whether if its better than Fable and what OpenAI have done to make it even better than 5.5. Also, we breakdown why we both moved our agents to Linux boxes, how to actually burn $65k on a single loop, and the Codex vs Claude Code subagent gap that's now bigger than the model gap itself. Thanks to this episode's sponsors Clerk and General Translation: Clerk, the auth platform with the best DX: https://nerdsnipe.link/clerk General Translation: https://nerdsnipe.link/gt Sources Available on our Substack: https://nerdsnipe.substack.com/

  • S1 · E15
    July 8 · 1 hr 16 min

    Fable Is Back...kinda?

    Fable was supposed to give us 14 days and instead we got 3 before the export ban. Now it's back at half the rate limits. Also, we break down why AI-generated Claude Code skills fall apart, what actually triggers Fable-to-Opus rerouting, the Mythos export-control timeline, and whether GLM 5.2 can really touch Sonnet 5. Thank you to PostHog for sponsoring today's episode! PostHog, all in one suite of product tools: ⁠https://nerdsnipe.link/posthog Sources available on our Substack: https://nerdsnipe.substack.com/

  • S1 · E14
    June 30 · 1 hr 11 min

    GPT-5.6 is here! And none of us can use it.

    A new "government-approved rollout" for GPT-5.6 is coming, and it's got us wondering: are we entering the next dark age of AI model access? Additionally, we break down the launch of T3 Code x Grok CLI, Apple's price hikes, the memory supply crunch driving RAM and SSD costs up, a repo-poisoning and AI PR spam wave hitting open source, the GPT-5.6 non-launch, and what frontier model access looks like in a Fable/Mythos-tier world. Thank you to Composio for sponsoring this episode! Composio, connect your agents to everything: https://nerdsnipe.link/composio Sources Available on our Substack: https://nerdsnipe.substack.com/

  • S1 · E13
    June 24 · 1 hr 33 min

    The US Government Banned Claude Fable 5...

    The US government banned Anthropic's Mythos and Fable models just after launch so we break down exactly how Project Glass Wing, a panicked AWS engineer, and Dario's failure to communicate with Washington triggered the chaos. Plus: SpaceX's $60B Cursor acquisition, GLM 5.2 beating Google, the Codex trick for running 200 parallel agents, and why your next GPU will ship with a GPS tracker Thank you to Composio and WorkOS for sponsoring this episode! Composio, connect your agents to everything: https://nerdsnipe.link/composio WorkOS, your enterprise ready solution: https://nerdsnipe.link/workos Sources Available on our Substack: https://nerdsnipe.substack.com/

  • S1 · E12
    June 15 · 1 hr 17 min

    Our impressions of Claude Fable/Mythos (we filmed this before the ban)

    RIP Fable 5. We recorded this before it got taken offline, but it's still worth talking about. The model is incredible. We really miss it. Thank you, Firecrawl, Depot, and Clerk for sponsoring! Firecrawl, the best api for searching and crawling the web: nerdsnipe.link/firecrawl Depot, better CI in every way: nerdsnipe.link/depot Clerk, the best dx in auth: nerdsnipe.link/clerk Sources https://x.com/thsottiaux/status/2043177597434306699 https://cognition.ai/blog/frontier-code https://x.com/paradite_/status/2064585901351792887 https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdf (page 13 has the “prompt modification” quote) https://www.anthropic.com/institute/recursive-self-improvement Timestamps 0:00 Intro 3:29 Fable First Impressions 10:24 Benchmark Drama 26:42 Claude Code & Workflows 36:07 Pricing & June 22 Cutoff 39:54 Data Retention 45:58 Hidden Safeguards 1:07:01 The Claude Constitution

  • S1 · E11
    June 10 · 1 hr 35 min

    Now even Google's buying GPUs from SpaceX?

    Cloudflare buys Void0, Google buying up compute from xAI, and Claude seems to be getting more anxious so we're here to break down everything this week on another episode of Nerd Snipe! Thank you to Composio for sponsoring today's episode! Composio, connect your agents to everything: https://nerdsnipe.link/composio Sources: https://x.com/vite_js/status/2062525206158078047 https://x.com/EdLudlow/status/2062970770612199542 https://x.com/nrehiew_/status/2063099050719846719 https://www.reddit.com/r/EconomyCharts/comments/1lp34n4/china_vs_us_energy/ https://x.com/elonmusk/status/1963443919150330139 https://x.com/AnthropicAI/status/2062568862479208923 https://x.com/HSVSphere/status/2060396271756595666 https://x.com/theo/status/2061018426152530232 https://x.com/Teknium/status/2062522290504613944 01:58 Cloudflare vs Vercel 13:03 Convex angle 19:59 SpaceX compute 35:49 AI self-improvement 44:45 Article reactions 50:28 Claude anxiety 57:42 Guardrails 01:08:21 Cursed image gen 01:19:18 Hermes agents

  • S1 · E10
    June 3 · 1 hr 15 min

    We (mostly) like Claude Opus 4.8

    Opus 4.8 + a ton of new stuff in Claude Code dropped this week, and we actually kinda like it. There's also a new benchmark that's actually good, and we have a lot of thoughts about the future of these AI labs... Thank you PostHog and Clerk for sponsoring! PostHog, the all in one suite of product tools: nerdsnipe.link/posthog Clerk, the best dx in auth: nerdsnipe.link/clerk SOURCES https://www.anthropic.com/news/claude-opus-4-8 https://x.com/theo/status/2060120708815139241 https://x.com/Baconbrix/status/2060065875911422343 ⁠https://x.com/zoink/status/2060769829133721974 https://x.com/maria_rcks/status/2060937270824153251 https://x.com/_catwu/status/2060054180379689074 https://x.com/datacurve/status/2060834005998793199 https://deepswe.datacurve.ai/ https://deepswe.datacurve.ai/blog#results https://github.com/scaleapi/SWE-agent/blob/402a7b8fdac8193f3f255bb53859ba274234f596/config/benchmarks/anthropic_filemap_multilingual.yaml https://deepswe.datacurve.ai/ https://x.com/AnthropicAI/status/2060061347522433422 https://x.com/theo/status/2060199299632472494 https://x.com/theo/status/2060901326058561795 https://x.com/38twelveDaily/status/2060408696945975631 00:00 - Shower Thoughts 02:44 - Deep SWE Benchmark 10:45 - Opus vs GPT-5.5 19:57 - Anthropic’s Huge Raise 25:39 - Token Maxing 40:02 - AI Slot Machine 43:49 - Claude Code Friction 50:01 - Opus, Mythos, and Safety

  • S1 · E9
    May 28 · 1 hr 22 min

    Google is Not a Serious Company

    Not only did Google accidentally ban Railway's account, but their new flagship model Gemini 3.5 Flash is absurdly bad. Oh and apparently Theo's building his own cloud... Thank you Macroscope and GT for sponsoring today's episode! Macroscope: ⁠nerdsnipe.link/macroscope General Translation: nerdsnipe.link/gt SOURCES ⁠https://x.com/theo/status/2057359424378097823⁠ ⁠https://x.com/JustJake/status/2056881510939283776⁠ ⁠https://x.com/KorduGG/status/2059141337895604626⁠ ⁠https://x.com/unboringtech/status/2059144145273491610⁠ TIMESTAMPS 00:00 - Gemini fallout 05:21 - Video gen 10:10 - Google Cloud 14:04 - Windsurf/Cursor 25:10 - Manus/Meta 30:00 - China lock-in 35:00 - Cursor/xAI 45:00 - Cloud workflows 50:32 - Lakebed 65:22 - Hermes/security

  • S1 · E8
    May 20 · 1 hr 9 min

    How the OpenClaw creator uses $1.3 million of tokens

    Peter, the creator of openclaw is apparently going through $1,300,000 worth of tokens every month. Seems like we're not using nearly enough tokens. Oh and the security psychosis is getting worse. (Anthropic is being bad again too) Thank you to today's sponsors! - AgentMail, email inboxes for your agents: nerdsnipe.link/agentmail - Clerk, the best experience for auth, orgs, billing, and more: nerdsnipe.link/clerk TIMESTAMPS 00:00 - Intro 02:44 - Token Spend 06:47 - Token Future 12:12 - Secure Agents 22:50 - Anthropic Rules 31:35 - Security 39:36 - Token Tax 49:49 - AI Psychosis 59:47 - macOS 01:01:43 - AI Video

  • S1 · E7
    May 14 · 1 hr 36 min

    Anthropic solved their compute problem by buying it from Elon?

    Anthropic seems to have finally solved their compute problems (kinda) by buying it from Elon, the security problem is getting so much worse, and apparently Bun's getting re-written in rust? Thank you to PostHog and Composio for sponsoring today's episode! - PostHog, the all in one suite of product tools: nerdsnipe.link/posthog - Composio, connect your agents to everything: nerdsnipe.link/composio Sources/references: https://x.com/claudeai/status/2052060691893227611 https://openai.com/index/elon-musk-wanted-an-openai-for-profit/#december-2018-elon-told-us-to-raise-billions-per-year-immediately-or-forget-it https://x.com/jarredsumner/status/2053391824702898475 https://x.com/jarredsumner/status/2051595933704761618 https://x.com/jarredsumner/status/2053047748191232310 https://ze3tar.github.io/post-zcrx.html https://www.jefftk.com/p/ai-is-breaking-two-vulnerability-cultures https://x.com/thdxr/status/2053570581807722968 https://x.com/badlogicgames/status/2052691176373805534 https://x.com/garrytan/status/2052996691586932783 TIMESTAMPS 00:00:00 - Anthropic/xAI00:08:59 - OpenAI Lawsuit00:25:40 - Bun in Rust00:34:34 - Security00:59:28 - Pottery Coding01:11:11 - Local Models

  • S1 · E6
    May 6 · 1 hr 47 min

    Theo Almost Lost $1 Million

    This week Theo nearly lost a million dollars and Ben got AI psychosis (from gstack)... Thank you to Coderabbit and Clerk for sponsoring today's episode! - Coderabbit, the ultimate AI code reviewer: ⁠nerdsnipe.link/coderabbit - Clerk, the auth platform with the best DX: nerdsnipe.link/clerk Sources/references: - https://x.com/theo/status/2014863266888233193 - https://x.com/theo/status/2050305813894648289 - https://x.com/theo/status/2050314995561611357 - https://x.com/sama/status/2050671161915371998 - https://x.com/davis7/status/2050718508372431026 - https://x.com/thdxr/status/2050719575033983323 - https://x.com/davis7/status/2050762009592148375 - https://x.com/naval/status/2050560057675522500 - https://x.com/theo/status/1952229335416623592 00:00 - Intro / studio return 00:52 - Azure $1M / Microsoft 11:22 - Cloud platform talk 18:08 - Coding agents / SDKs 29:06 - GPT-5.5 pricing debate 45:14 - OpenClaw workflows 01:12:37 - G Stack / G Brain01:29:00 - Dynamic UI / wrap-up

  • S1 · E5
    May 1 · 1 hr 35 min

    We need to talk about OpenAI

    OpenAI and Microsoft are breaking up, Sam's drunk posting, Anthropic is being stupid again, and we still disagree about GPT-5.5Thanks to this episode's sponsors: - Clerk, the auth platform with the best DX: https://nerdsnipe.link/clerk- Coderabbit, the ultimate AI code reviewer: ⁠https://nerdsnipe.link/coderabbit- PlanetScale, the fastest and most scalable cloud databases: https://nerdsnipe.link/planetscale 00:00 Intro 04:05 Sam drunk tweets1 3:13 Anthropic billing woes 28:27 OpenAI divorce 42:34 GitHub can't stop dying 59:40 GPT-5.5 retrospective Sources/references: https://x.com/sama/status/2046808114561974567 https://x.com/sama/status/2046808217133670800 https://x.com/sama/status/2048160404376105179 https://x.com/sama/status/2047403771416940715 https://x.com/om_patel5/status/2048204411986469232 https://x.com/MSFTnews/status/2048749108127506936 https://x.com/GergelyOrosz/status/2048834949667537369 https://x.com/theo/status/2047721472521621991 https://x.com/mitchellh/status/2049213597419774026 https://x.com/kdaigle/status/2047803291988590609 https://x.com/ryanflorence/status/2048538797638599109 https://x.com/davis7/status/2048239401059434710 https://x.com/davis7/status/2048077518725366173 https://x.com/badlogicgames/status/2048444292562026713 https://x.com/0xSero/status/2048744545853030690

  • S1 · E4
    April 23 · 1 hr 44 min

    We've been testing GPT-5.5 for a few weeks now...

    Anthropic month is finally over. We've got a ton to talk about: GPT-5.5 pre-release impressions, the Vercel hack, Cursor + xAI, Qwen models, Kimi k2.6, and so much more...Thank you to today's Sponsors! Depot, truly modern CI: nerdsnipe.link/depot Coderabbit, the ultimate AI code reviewer: nerdsnipe.link/coderabbit Clerk, the auth platform with the best DX: nerdsnipe.link/clerk TIMESTAMPS00:00:00 - Intro, vacation chaos, and episode setup 00:01:34 - Vercel hack explained and security fallout 00:05:38 - Kimi K2.6 and the rise of strong open-weight models 00:14:29 - Cursor + xAI/SpaceX partnership and acquisition option 00:37:24 - GPT Image 2 impressions, strengths, and flaws 00:43:53 - Anthropic Cloud Code Pro pricing controversy 00:49:30 - GPT-5.5 first impressions: split opinions 01:02:25 - Critiques of GPT-5.5 for coding, context, and prompting 01:22:02 - GPT-5.5 Pro

  • S1 · E3
    April 22 · 1 hr 4 min

    Theo gets conspiratorial about Anthropic

    Really hope that Anthropic month is almost over. But at least for now we gotta talk about Opus 4.7, Claude Design, and much much more...Thank you to this episode's sponsor Clerk! The best DX in auth by a mile: https://nerdsnipe.link/clerkOur RSS feed: https://nerdsnipe.link/rssFollow on spotify and everywhere else you get your podcasts! https://nerdsnipe.link/podTIMESTAMPS00:00:00 Anthropic Again...00:07:30 T3 Code Banned?00:16:08 How Claude Caching Works00:24:00 Why Claude Seems Dumber00:38:03 The Conspiracies Begin...

  • S1 · E2
    April 18 · 1 hr 1 min

    We need to talk about gstack

    Diving into Mythos, Claude reasoning, post-AI Uncle Bob, GStack, Thank you to today's sponsors! Coderabbit The AI code reviewer that deeply understands your codebase: ⁠https://nerdsnipe.link/coderabbit⁠ Clerk The auth platform we always either use, or regret not using: ⁠https://nerdsnipe.link/clerk 00:00:00 - Welcome to the TBPNN Show 00:02:50 - Anthropic's Mythos model 00:11:14 - Anthropic nerfed Claude's reasoning effort 00:14:20 - We should be scared of mythos 00:24:20 - Post-AI Uncle Bob 00:29:34 - Static Typing vs Dynamic Typing and Linting 00:36:23 - Claude Code Skills 00:40:49 - GStack is actually good 00:46:53 - Are we boiling in the ocean? 00:51:07 - Everything should be a md file 00:55:23 - Peace Out Nerds

  • S1 · E1
    April 9 · 1 hr 21 min

    Anthropic Can't Stop Making Mistakes

    Talking about everything that went wrong with Anthropic last week (source code leak, rate limits, banning openclaw, and so much more) + our new favorite coding agent Pi. Thank you to this episode's sponsors! Coderabbit The AI code reviewer that deeply understands your codebase: https://nerdsnipe.link/coderabbit Clerk The auth platform we always either use, or regret not using: https://nerdsnipe.link/clerk

Showing 1–20 of 20 episodes