AI Explained Official Podcast
Philip - Host of AI Explained YT
Covering the biggest news of the century - the arrival of smarter-than-human AI. From the author of Simple Bench, which reveals the remaining gap between LLM and human reasoning. Hype-free, and the British accent is a freebie bonus.
- 61 episodes
- Updated July 22
Episodes61
- Nov 19, 2025 · 21 min
- AE
Nov 14, 2025 · 18 minIs GPT-5.1 Really an Upgrade? But Models Can Auto-Hack Govts, so … there’s that
A lot just got released in the last 36 hours, and it will all affect hundreds of millions of people. 10 details you would miss if you just read the headlines, from GPT 5.1 regressions, to how Claude hacked Govt Agencies, to SIMA 2, and Musical Turing Tests. https://assemblyai.com/aiexplained Chapters: 00:00 - Introduction 00:56 - GPT 5.1 Smarter? 01:47 - Some Regressions 03:22 - Sycophancy? 05:22 - Claude Auto-Hacking 06:16 - Jailbreaking through Granularity 08:22 - This Will be Re-used 09:30 - Hallucinating Hacker 09:57 - Surprisingly Neutral Tone 12:18 - SIMA 2 14:10 - Alpha Parallels 17:24 - AI Music GPT 5.1 Announcement: https://openai.com/index/gpt-5-1/ System Card: https://cdn.openai.com/pdf/4173ec8d-1229-47db-96de-06d87147e07e/5_1_system_card.pdf Benchmarks: https://openai.com/index/gpt-5-1-for-developers/ Simple Bench: https://lmcouncil.ai/benchmarks Auto-Hacking: https://x.com/AnthropicAI/status/1989033793190277618 https://www.anthropic.com/news/disrupting-AI-espionage Report: https://assets.anthropic.com/m/ec212e6566a0d47/original/Disrupting-the-first-reported-AI-orchestrated-cyber-espionage-campaign.pdf Sima 2 Announcement: https://deepmind.google/blog/sima-2-an-agent-that-plays-reasons-and-learns-with-you-in-virtual-3d-worlds/ https://x.com/amoufarek/status/1988986075331858693 Scepticism: https://www.technologyreview.com/2025/11/13/1127921/google-deepmind-is-using-gemini-to-train-agents-inside-goat-simulator-3/ Voyager: https://voyager.minedojo.org/ Reuters Music: https://www.reuters.com/legal/litigation/are-you-listening-bots-survey-shows-ai-music-is-virtually-undetectable-2025-11-12/
- AE
Nov 10, 2025 · 12 minBubble or No Bubble, AI Keeps Progressing (ft. Relentless Learning + Introspection)
Don’t let headlines about bubbles distract you from the real avenues of progress being explored in AI every week, including what had been thought to be a long-term blocker - continual learning (learning on the fly). https://app.grayswan.ai/ai-explained This, plus models introspecting (hesitate before you berate), Nano Banana 2 possibly spotted, Chinese imagen and more. AI Insiders ($9!): https://www.patreon.com/AIExplained Chapters: 00:00 - Introduction 01:26 - Continual Learning (Nested Learning / HOPE) 07:00 - Introspection 10:54 - Image-Gen Progress Nested Learning Post: https://research.google/blog/introducing-nested-learning-a-new-ml-paradigm-for-continual-learning/ Nested Learning Paper: https://abehrouz.github.io/files/NL.pdf Original Titans Paper: https://arxiv.org/pdf/2501.00663 Siri News: https://www.bloomberg.com/news/articles/2025-11-05/apple-plans-to-use-1-2-trillion-parameter-google-gemini-model-to-power-new-siri Introspection: https://www.anthropic.com/research/introspection Full Paper: https://transformer-circuits.pub/2025/introspection/index.html#mechanisms Earlier Work: https://www.anthropic.com/research/mapping-mind-language-model https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html Release Post: https://x.com/AnthropicAI/status/1983584136972677319 https://lmcouncil.ai Non-hype Newsletter: https://signaltonoise.beehiiv.com/ Podcast: https://aiexplainedopodcast.buzzsprout.com/
- AE
Oct 1, 2025 · 15 minSora 2 - It will only get more realistic from here
Sora 2 - the start of the infinite slop-feed or a key step to a generalist agent? Better than VEO 3 or over-hyped? I bring out 6 details you may have missed, contrast the announcement to Periodic Labs and even squeeze in some Claude Sonnet 4.5 analysis. Maybe I should make my videos longer… https://80000hours.org/aiexplained AI Insiders ($9!): https://www.patreon.com/AIExplained Chapters: 00:00 - Introduction 00:40 - Two models? 01:15 - Rollout Details 01:43 - Versus Sora 1 / Veo 3 04:30 - Sora App / Social Media 06:40 - Masterplan 09:30 - Generalist Agent? Periodic Labs 12:05 - Claude Sonnet 4.5 13:42 - Future Outlook Announcement: https://openai.com/index/sora-2/ Launch Video: https://www.youtube.com/live/gzneGhpXwjU System Card: https://cdn.openai.com/pdf/50d5973c-c4ff-4c2d-986f-c72b5d0ff069/sora_2_system_card.pdf Sam Altman Blog Post on Sora App: https://blog.samaltman.com/sora-2 Most Intelligent Claim: https://x.com/willdepue/status/1973089331284681110 GTA: https://x.com/AndrewCurran_/status/1973298436536766666 Meta Vibes: https://x.com/alexandr_wang/status/1971295156411433228?s=46 Altman on Regulations: https://www.lesswrong.com/posts/5jjk4CDnj9tA7ugxr/openai-email-archives-from-musk-v-altman OpenAI Profit: https://www.theinformation.com/articles/openais-first-half-results-4-3-billion-sales-2-5-billion-cash-burn?rc=sy0ihq Periodic Labs: https://periodic.com/ https://www.nytimes.com/2025/09/30/technology/ai-meta-google-openai-periodic.html https://x.com/LiamFedus/status/1973055380193431965 https://baincapitalventures.com/insight/we-must-know-we-will-know/?s=09 Sonnet 4.5: https://www.anthropic.com/news/claude-sonnet-4-5 https://simple-bench.com/ Non-hype Newsletter: https://signaltonoise.beehiiv.com/ Podcast: https://aiexplainedopodcast.buzzsprout.com/
- AE
Sep 26, 2025 · 14 minOpenAI Tests if GPT-5 Can Automate Your Job - 4 Unexpected Findings
An OpenAI report released in the last 24 hours is the best look we have as to whether 2025 AI can automate your job. I’ll go through 4 unexpected findings, from which model is best at what, to practical tips and massive caveats. Plus UFC robots, radiologist essay, don’t trust videos and the blockers to the singularity. Gray Swan: https://app.grayswan.ai/ai-explained GDPval: https://cdn.openai.com/pdf/d5eb7428-c4e9-4a33-bd86-86dd4bcf12ce/GDPval.pdf [GDP Impact: https://fred.stlouisfed.org/release/tables?rid=331&eid=211 Task List: https://www.onetonline.org/link/summary/11-9141.00 Summer Tweet: https://x.com/LHSummers/status/1971252567981146347 Emad: https://x.com/EMostaque/status/1971254153067593739 Robots: https://x.com/cixliv/status/1967663286679478759 Unitree G1: https://x.com/UnitreeRobotics/status/1970039940022239491 Don’t Trust Video: https://x.com/AISafetyMemes/status/1970453369446871420 AGI Tweet: https://x.com/hyhieu226/status/1968378785709133915 Blockers to the Singularity: https://www.patreon.com/posts/blockers-to-and-139264812 Framework: https://gemini.google.com/share/f4b9c85a6ae9 METR Study (Dev Slowdown): https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/ Karpathy Tweet: https://x.com/karpathy/status/1971220449515516391 Radiology Essay: https://worksinprogress.co/issue/the-algorithm-will-see-you-now/ Chapters: 00:00 - Introduction 00:55 - OpenAI Report Summary 02:40 - Tipping Point Speed-up 04:11 - Better than Industry Experts? 06:33 - Big Caveat 11:10 - Karpathy and the Radiologist Analogy 13:30 - Outro
- AE
Sep 16, 2025 · 11 minChatGPT Will Guess your Age, Flirt if Asked, and Can Call the Cops
Sam Altman, CEO of OpenAI, announced a set of new ‘protections’ and ‘privileges’ for ChatGPT users, requiring a significant amount of trust from users. From predicting your age based on your chat to calling law enforcement if you are at risk of harm, to allowing non-minors to flirt. But amidst all of these announcements, there are interview snippets you may have missed, as Altman dramatically revises his predictions of AI impact on jobs. Plus a Hassbis backtrack to boot. https://80000hours.org/aiexplained Calling the Cops: https://openai.com/index/teen-safety-freedom-and-privacy/ Age Prediction: https://openai.com/index/building-towards-age-prediction/ Not Everyone Will Agree: https://x.com/sama/status/1967955739911364693?ref_src=twsrc%5Egoogle%7Ctwcamp%5Eserp%7Ctwgr%5Etweet Theory 1: NYT Lawsuit: https://openai.com/index/response-to-nyt-data-demands/ Theory 2: FTC Investigation into AI Companions: https://x.com/AndrewCurran_/status/1966167585994764743 YT Does the Same: https://www.cbsnews.com/news/youtube-ai-powered-technology-teen-users/ Carlsen Interview: https://www.youtube.com/watch?v=5KmpT-BoVf4 vs Senate Testimony (70% Jobs): https://www.youtube.com/watch?v=5CWVP8-XVjQ Hallucinations Paper: https://cdn.openai.com/pdf/d04913be-3f6f-4d2b-b283-ff432ef4aaa5/why-language-models-hallucinate.pdf Hassbis Quote 1: https://www.youtube.com/watch?v=toShbNUGAyo vs Quote 2: https://www.youtube.com/watch?v=Kr3Sh2PKA8Y
- AE
Aug 26, 2025 · 18 minAn ‘AI Bubble’? What Altman Actually said, the Facts and Nano Banana
Wait, why did Sam Altman say AI was in a bubble? Or did he? Is it? 8 points for you to consider, before we all get distracted by Nano Banana. Chapters: 00:00 - Introduction 01:14 - Sam Altman Clarification 02:30 - Media Calls a Bubble (for the tenth time) 03:40 - MIT and McKinsey Analysed 08:21 - Incremental Progress Deceptive 12:07 - Reasoning Breakthroughs 15:31 - CEOs might not know their products 17:25 - But did stocks go down? 17:31 - Media is Contradictory of course https://donate.redcross.org.uk/appeal/gaza-crisis-appeal Bubble about to burst: https://www.telegraph.co.uk/business/2025/08/20/ai-report-triggering-panic-and-fear-on-wall-street/ Nano Banana: https://blog.google/products/gemini/updated-image-editing-model/ https://ai.studio/banana McKinsey Report: https://www.mckinsey.com/capabilities/quantumblack/our-insights/seizing-the-agentic-ai-advantage#/ https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai#/ Revenue: https://www.wsj.com/tech/ai/mckinsey-consulting-firms-ai-strategy-89fbf1be MIT Report: https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf Safe Superintelligence: https://techcrunch.com/2025/04/12/openai-co-founder-ilya-sutskevers-safe-superintelligence-reportedly-valued-at-32b/ Thinking Machines Lab: https://techcrunch.com/2025/07/15/mira-muratis-thinking-machines-lab-is-worth-12b-in-seed-round/ WSJ Prediction 2024: https://www.wsj.com/tech/ai/the-ai-revolution-is-already-losing-steam-a93478b1 WP Prediction 2023: https://www.washingtonpost.com/technology/2023/08/05/ai-hype-bubble-chatgpt/ Companies are Pouring Billions into AI: https://www.nytimes.com/2025/08/13/business/ai-business-payoff-lags.html Consumer Surplus: https://www.wsj.com/opinion/ais-overlooked-97-billion-contribution-to-the-economy-users-service-da6e8f55 Figure AI robot: https://x.com/adcock_brett/status/1958193476639826383 GDP Bet: https://x.com/adamdangelo/status/1627726566259318784?lang=en Genie 3 Immersion: https://x.com/holynski_/status/1953879983535141043 https://x.com/elonmusk/status/1953861448431718662 htttps://simple-bench.com MMMU: https://mmmu-benchmark.github.io/#leaderboard Prophet Arena: https://www.prophetarena.co/leaderboard NYT Jobs: https://www.nytimes.com/2025/08/19/opinion/ai-job-loss-deindustrialization.html Dawn of Reasoning?: https://openreview.net/pdf?id=FkKBxp0FhR vs :https://arxiv.org/pdf/2403.04121 ARC-AGI: https://arcprize.org/arc-agi/1/ https://x.com/fchollet/status/1870169764762710376?lang=en-GB Turing Test: https://arxiv.org/pdf/2503.23674 Mathematics of Starvation: https://www.theguardian.com/world/2025/jul/31/the-mathematics-of-starvation-how-israel-caused-a-famine-in-gaza https://donate.redcross.org.uk/appeal/gaza-crisis-appeal https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/ METR Interview: https://www.patreon.com/c/aiexplained/posts AlphaEvolve: https://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/ Paper: https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/AlphaEvolve.pdf Amodei: https://kantrowitz.medium.com/the-making-of-anthropic-ceo-dario-amodei-449777529dd6 https://www.theloganbartlettshow.com/archive/ep-82-dario-amodeis-ai-predictions-through-2030#:~:text=DARIO%3A%20I%20think%20our%20concern,being%20responsible%20to%20accelerate%20things Unreleased OpenAI: https://x.com/alexwei_/status/1954966393419599962 VLMs Tricked: https://x.com/an_vo12/status/1943715159559545186 AI Insiders ($9!): https://www.patreon.com/AIExplained
- AE
Aug 7, 2025 · 15 minGPT-5 has Arrived
GPT-5 will change how hundreds of millions of people use AI. Yes, you might have to forgive the chart crimes, the underwhelming livestream and Altman hype… But it’s a good model. I have read the 50 page system card in full, have the benchmark scores, coding tests, and things you might have missed. https://app.grayswan.ai/ai-explained Announcement: https://openai.com/index/introducing-gpt-5/ System Card: https://cdn.openai.com/pdf/8124a3ce-ab78-4f06-96eb-49ea29ffb52f/gpt5-system-card-aug7.pdf Extra Paper: https://cdn.openai.com/pdf/be60c07b-6bc2-4f54-bcee-4141e1d6c69a/gpt-5-safe_completions.pdf Altman tweet: https://x.com/sama/status/1953551377873117369 Livestream: https://www.youtube.com/watch?v=0Uu_VJeVVfo METR Report: https://metr.github.io/autonomy-evals-guide/gpt-5-report/ ARC-AGI-2: https://x.com/fchollet/status/1953511631054680085 Claude Opus 4.1: https://www.anthropic.com/news/claude-opus-4-1 MMMU: https://mmmu-benchmark.github.io/ Cursor Praise: https://x.com/ryolu_/status/1953531724895596669
- AE
Aug 5, 2025 · 11 minGenie 3: The World Becomes Playable (DeepMind)
Soon, anything will be playable. A photo becomes an interactive world, a selfie becomes a new game. Genie 3 from Google, debuting just 2 hours ago, is what I mean, and I have the full analysis, plus the pushback I gave the authors (will it really lead to reliable AI agents? Is that even the point?). You make your own mind up, but it’s certainly fascinating, and not to be overlooked in the week that will bring us GPT-5. https://80000hours.org/aiexplained AI Insiders ($9!): https://www.patreon.com/AIExplained Chapters: 00:00 - Introduction 01:27 - Background and Access 04:58 - Caveats 07:24 - Demo 10:12 - Conclusion Announcement: https://deepmind.google/discover/blog/genie-3-a-new-frontier-for-world-models/ Isaac Labs: https://developer.nvidia.com/isaac/lab Genie 2 Coverage: https://www.youtube.com/watch?v=jIm2T7h_a0M TED Talk Roblox: https://www.youtube.com/watch?v=-OAP0ho5AUg DeepThink Post: https://www.patreon.com/posts/deep-ish-on-new-135688441 AI Insiders ($9!): https://www.patreon.com/AIExplained Non-hype Newsletter: https://signaltonoise.beehiiv.com/
- AE
Jul 21, 2025 · 17 minHow Not to Read a Headline on AI (ft. new Olympiad Gold, GPT-5 …)
GPT-5 did what? OpenAI ahead of Google? There are 9 ways to misread the headlines of the last 48 hours, so this video is here to tell you what happened, sans sizzle. It’s been a fairly momentous last few days, so let’s dive in to the International Math Olympiad Gold, GPT-5 alpha release, whether mathematicians are out of jobs, and the white collar impact by year’s end. Job Board: https://80000hours.org/aiexplained New Documentary on Patreon: https://www.patreon.com/posts/our-new-age-of-133960279 Chapters: 00:00 - Introduction 00:18 - AI > Mathematicians? 01:23 - OPENAI vs GOOGLE 02:42 - Irrelevant to Jobs or … 06:45 - White-collar jobs gone? 10:26 - AI is Plateauing? 12:00 - We Don’t Know the Details… 14:33 - GPT-5 alpha 14:54 - Nothing but Exponentials? 15:53 - No Impact? Announcement: https://x.com/alexwei_/status/1946477742855532918 UCLA Math Prof: https://x.com/ErnestRyu/status/1946699302308635130 ChatGPT Agent: https://openai.com/index/introducing-chatgpt-agent/ Livestream: https://www.youtube.com/watch?v=1jn_RpbPbEc&t=796s System Card: https://cdn.openai.com/pdf/839e66fc-602c-48bf-81d3-b21eacc3459d/chatgpt_agent_system_card.pdf Jerry Tworek (OpenAI): https://x.com/MillionInt/status/1946556255490982022 https://x.com/MillionInt/status/1946558130906968330 Noam Brown Details: https://x.com/polynoamial/status/1946478249187377206 Trieu Tranh Retweet: https://x.com/Mihonarium/status/1946880931723194389 Neel Nanda: https://x.com/NeelNanda5/status/1946602953370173647 Terence Tao: https://mathstodon.xyz/@tao Sam Altman: https://x.com/sama/status/1946569252296929727 METR Dev Study: https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/ Ravid Schwatz: https://x.com/ziv_ravid/status/1946378712716562605 AlphaEvolve: https://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/ https://simple-bench.com/ Meta Salary: https://www.tomshardware.com/tech-industry/artificial-intelligence/abel-founder-claims-meta-offered-usd1-25-billion-over-four-years-to-ai-hire-person-still-said-no-despite-equivalent-of-usd312-million-yearly-salary $2k per month: https://www.theinformation.com/articles/openai-considers-higher-priced-subscriptions-to-its-chatbot-ai-preview-of-the-informations-ai-summit?rc=sy0ihq
- AE
Jul 10, 2025 · 11 minGrok 4 - 10 New Things to Know
Grok 4 is here, but did you know these 10 things about the new model? From benchmark caveats to soloing science, $300 a month secrets to Grok 5 promises, here's 10 new things to know in just under 12 minutes. AI Insiders ($9!): https://www.patreon.com/AIExplained Chapters: 00:00 - Introduction 00:22 - Benchmark Results 02:11 - Benchmark Caveats 02:59 - ARC-AGI 2 03:35 - SimpleBench 04:49 - ‘Humanity’s Last Exam’ 07:20 - SuperGrok Heavy Price 07:58 - API Price 08:12 - Grok 5, Gemini 3.0 Beta, GPT-5 09:12 - System Prompt Change + $1B a month, pollution 10:20 - Not soloing science, helping you solo code Livestream: https://www.youtube.com/watch?v=1tQ_KrlHgfg&t=1s Price: https://grok.com/#subscribe https://x.com/ArtificialAnlys/status/1943166841150644622 Gemini DeepThink: https://blog.google/technology/google-deepmind/google-gemini-updates-io-2025/#deep-think https://simple-bench.com/ ARC-AGI 2: https://x.com/arcprize/status/1943168950763950555 Humanity’s Last Exam: https://agi.safe.ai/ SmartGPT: https://www.youtube.com/watch?v=hVade_8H8mE New Power Plant, 1m GPUs: https://www.tomshardware.com/tech-industry/artificial-intelligence/elon-musk-xai-power-plant-overseas-to-power-1-million-gpus Gemini 3.0 beta: https://web.archive.org/web/20250709174548/https://github.com/google-gemini/gemini-cli/blob/b0cce952860b9ff51a0f731fbb8a7649ead23530/packages/cli/src/ui/utils/errorParsing.test.ts Pollution: https://www.theguardian.com/technology/2025/apr/24/elon-musk-xai-memphis https://www.youtube.com/watch?v=C8rU4dv2w8Q https://www.youtube.com/watch?v=3VJT2JeDCyw System Prompt: https://github.com/xai-org/grok-prompts/blob/535aa67a6221ce4928761335a38dea8e678d8501/ask_grok_system_prompt.j2 Burn Rate: https://www.bloomberg.com/news/articles/2025-06-17/musk-s-xai-burning-through-1-billion-a-month-as-costs-pile-up Ron Johnson: https://x.com/jdcmedlock/status/1939814516503847259 Non-hype Newsletter: https://signaltonoise.beehiiv.com/ Podcast: https://aiexplainedopodcast.buzzsprout.com/
- AEJun 24, 2025 · 26 min
When Will AI Models Blackmail You, and Why?
In the last few days Anthropic have released an impressive honest account of how all models blackmail, no matter what goal they have, and despite prompt warnings, and other preventions. But do these models *want* this? Thanks to Storyblocks for sponsoring this video! Download unlimited stock media at one set price with Storyblocks: storyblocks.com/AIExplained AI Insiders ($9!): https://www.patreon.com/AIExplained Chapters: 00:00 - Introduction 01:20 - What prompts blackmail? 02:44 - Blackmail walkthrough 06:04 - ‘American interests’ 08:00 - Inherent desire? 10:45 - Switching Goals 11:35 - Murder 12:22 - Realizing it’s a scenario? 15:02 - Prompt engineering fix? 16:27 - Any fixes? 17:45 - Chekov’s Gun 19:25 - Job implications 21:19 - Bonus Details Report: https://www.anthropic.com/research/agentic-misalignment 30 Page Appendices: https://assets.anthropic.com/m/6d46dac66e1a132a/original/Agentic_Misalignment_Appendix.pdf Announcement: https://x.com/AnthropicAI/status/1936144602446082431?ref_src=twsrc%5Egoogle%7Ctwcamp%5Eserp%7Ctwgr%5Etweet OpenAI Files: https://www.openaifiles.org/ Grok 4 News: https://x.com/RonFilipkowski/status/1936372579607912473 Claude 4 Report Card: https://www-cdn.anthropic.com/6be99a52cb68eb70eb9572b4cafad13df32ed995.pdf New Apollo Research: https://www.apolloresearch.ai/blog/more-capable-models-are-better-at-in-context-scheming Interesting Reflections: https://nostalgebraist.tumblr.com/post/785766737747574784/the-void Non-hype Newsletter: https://signaltonoise.beehiiv.com/
- AE
Jun 12, 2025 · 14 minApple’s ‘AI Can’t Reason’ Claim Seen By 13M+, What You Need to Know
What to make of those headlines that AI can’t reason, seen by tens of millions? I cover the paper in layman’s terms, what it means and doesn’t mean, and what’s next. Thanks to Storyblocks for sponsoring this video! Download unlimited stock media at one set price with Storyblocks: https://storyblocks.com/AIExplained Plus o3-pro and whether it is my current most-recommended model. AI Insiders ($9!): https://www.patreon.com/AIExplained Chapters: 00:00 - Introduction 00:57 - Viral Post + Headlines 01:42 - Apple Paper Analysis 08:34 - But they do Hallucinate 10:43 - Not Supercomputers 11:18 - o3 Pro and Recommendations 13.7M Tweet: https://x.com/RubenHssd/status/1931389580105925115 Apple Paper: https://ml-site.cdn-apple.com/papers/the-illusion-of-thinking.pdf Guardian Article: https://www.theguardian.com/technology/2025/jun/09/apple-artificial-intelligence-ai-study-collapse Lisan al Gaib post: https://x.com/scaling01/status/1931854370716426246 Multiplication: https://x.com/yuntiandeng/status/1836114401213989366 The Illusion of the Illusion of Thinking: https://drive.google.com/file/d/1Zx9ikRj0Enc3SB4wA9HlYIlpmO_8QiUO/view Marcus: https://www.theguardian.com/commentisfree/2025/jun/10/billion-dollar-ai-puzzle-break-down Prof Rao: https://x.com/rao2z/status/1927707640223719631 AI Job Headlines: https://www.nytimes.com/2025/06/11/technology/ai-mechanize-jobs.html https://www.axios.com/2025/05/28/ai-jobs-white-collar-unemployment-anthropic Sky News Story: https://news.sky.com/story/can-we-trust-chatgpt-despite-it-hallucinating-answers-13380975 Veo 3 Ad: https://x.com/Kalshi/status/1932891608388681791 Altman Essay: https://blog.samaltman.com/ o3 Original benchmarks: https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b8b6c44-acd6-43b3-b5c6-1a1d5c6c25e4_2486x1388.png https://pbs.twimg.com/media/GfQ0bfcXQAAQt13.jpg Alpha Evolve Video: https://www.youtube.com/watch?v=RH4hAgvYSzg https://simple-bench.com/ Non-hype Newsletter: https://signaltonoise.beehiiv.com/
- AE
Jun 6, 2025 · 16 minAI Accelerates: New Gemini Model + AI Unemployment Stories Analysed
There’s a new best language model, so let’s go through the up and downs of Gemini 2.5 Pro 06-05. Record-breaking common-sense, but dumb mistakes remain. And it’s not even their best model, which remains behind the scenes - Gemini 2.5 Ultra. Plus Sundar Pichai’s AGI date and an analysis of whether the current AI unemployment headlines are justified, and Elevenlabs v3. https://emergentmind.com AI Insiders ($9!): https://www.patreon.com/AIExplained Chapters: 00:00 - Introduction 02:04 - Gemini 2.5 Ultra 03:34 - Benchmarks 07:41 - AGI Date and Meaning Pichai 09:13 - Jobs and AI Unemployment Fears 15:28 - Elevenlabs v3 Sundar Pichai Fridman: https://www.youtube.com/watch?v=9V6tWC4CdFQ Pichai More Jobs (until 2026 at least): https://www.techradar.com/pro/alphabet-ceo-sundar-pichai-says-ai-wont-lead-to-job-cuts-will-be-an-accelerator Gemini Comparison: https://blog.google/products/gemini/gemini-2-5-pro-latest-preview/ https://x.com/viathebrink/status/1930733154203292121 https://simple-bench.com/ White Collar Bloodbath: https://www.axios.com/2025/05/28/ai-jobs-white-collar-unemployment-anthropic https://fortune.com/2025/05/25/ai-entry-level-jobs-gen-z-careers-young-workers-linkedin/ https://www.nytimes.com/2025/05/19/opinion/linkedin-ai-entry-level-jobs.html https://www.nytimes.com/2025/03/25/business/economy/white-collar-layoffs.html College Unemployment: https://www.newyorkfed.org/research/college-labor-market/#--:explore:unemployment New Scientist AI Hallucinaitons: https://www.newscientist.com/article/2479545-ai-hallucinations-are-getting-worse-and-theyre-here-to-stay/ Duolingo: https://fortune.com/2025/05/24/duolingo-ai-first-employees-ceo-luis-von-ahn/ Klarna: https://www.forbes.com/sites/quickerbettertech/2025/05/18/business-tech-news-klarna-reverses-on-ai-says-customers-like-talking-to-people/ Sholto Douglas: https://www.reddit.com/r/ClaudeAI/comments/1ktt1rb/anthropics_sholto_douglas_says_by_202728_its/ Figure 02: https://x.com/adcock_brett/status/1930693311771332853 Elevenlabs v3: https://www.youtube.com/watch?v=zv_IoWIO5Ek Gemini Speech Generation: https://aistudio.google.com/generate-speech Non-hype Newsletter: https://signaltonoise.beehiiv.com/
- AE
May 22, 2025 · 19 minClaude 4: Full 120 Page Breakdown … Is it the Best New Model?
Not only did I get early access and ran my own tests, as per the title I read both the 120 page Claude 4 Opus and Claude 4 Sonnet System Card, and 25 page report on ASL-3 being triggered, plus the 2 hour launch video, and surrounding coverage. Ft. coding tests, Simple, twitter controversies, deep alignment coverage, spiritual bliss and much more! https://80000hours.org/aiexplained Chapters: 00:00 - Introduction 01:12 - 3 Quick Controversies 02:42 - Benchmark Results 04:20 - 120 page Card 20 Highlights 10:07 - Coding Test 11:27 - Model Welfare and Spiritual Bliss 13:29 - ASL-3 Claude Card: https://www-cdn.anthropic.com/4263b940cabb546aa0e3283f35b686f4f3b2ff47.pdf?s=09 ASL 3:https://www-cdn.anthropic.com/807c59454757214bfd37592d6e048079cd7a7728.pdf Tweets: https://x.com/fish_kyle3/status/1925597284546629753 https://x.com/EMostaque/status/1925624164527874452?ref_src=twsrc%5Egoogle%7Ctwcamp%5Eserp%7Ctwgr%5Etweet Cursor Says State of the Art for Coding: https://x.com/cursor_ai/status/1925594428095561941 Benchmarks: https://www.anthropic.com/news/claude-4
- AE
May 21, 2025 · 17 minGoogle Takes No Prisoners Amid Torrent of AI Announcements
Google just announced at least 12 things that are each worthy of a video, but here are the top I/O highlights. From Veo 3 to Deep Research now being useable, Deep Think breaking records to Gemini Diffusion, Gemini 2.5 Flash changing how AI is priced and GemmaVerse, SynthID Detector and Imagen 4. And even this intro is missing other announcements covered in the vid! And yes, they’ll be plenty of Veo 3 clips to enjoy… https://80000hours.org/aiexplained AI Insiders ($9!): https://www.patreon.com/AIExplained Chapters: 00:00 - Introduction 00:48 - Veo 3 02:10 - Gemini 2.5 Flash 03:13 - Universal Assistant 03:47 - Usage Skyrockets + OpenAI dig 04:51 - Gemini Pro Deep Think 06:21 - Overviews and AI Mode 07:26 - Deep Research Updates (new) + Jules 08:53 - Make and Deploy Apps with Gemini 09:12 - Imagen 4 10:00 - Gemini Diffusion 11:46 - Try It On 12:17 - SynthID Detector 13:30 - GemmaVerse, SignGemma, Gemma3n, medGemma 14:24 - Outro + Clips Event: https://www.youtube.com/watch?v=o8NiE3XMPrM Ntaive Audio: https://aistudio.google.com/generate-speech Gemini Diffusion: https://deepmind.google/models/gemini-diffusion/#capabilities New Gemini 2.5 Flash: https://deepmind.google/models/gemini/flash/ SignGemma (See end of this vid): https://www.youtube.com/watch?v=GjvgtwSOCao Deep Think: https://blog.google/technology/google-deepmind/google-gemini-updates-io-2025/#flash-improvements Google Parallel Sampling: https://www.patreon.com/posts/next-level-good-127441188 Price Plans: https://blog.google/products/google-one/google-ai-ultra/ Imagen 4 Benchmarks: https://deepmind.google/models/imagen/ Jules: https://jules.google/ SynthID Detector: https://blog.google/technology/ai/google-synthid-ai-content-detector/ Veo 3 Benchmarks: https://deepmind.google/models/veo/evals/ MedGemma: https://deepmind.google/models/gemma/medgemma/ Build Apps: https://aistudio.google.com/apps Non-hype Newsletter: https://signaltonoise.beehiiv.com/
- AE
May 19, 2025 · 17 minAI Improves at Self-improving
AlphaEvolve is not the first system to exhibit self-improvement, but it may be the most impressive yet. AI is literally improving the hardware, architectures, data and training methods of AI itself. A deep dive into the paper, drawing on two previous interviews and 5 other papers. Plus a snippet on OpenAI’s new Codex system. Gray Swan: http://app.grayswan.ai/ai-explained AI Insiders ($9!): https://www.patreon.com/AIExplained Chapters: 00:00 - Introduction 00:27 - AlphaEvolve 05:23 - Limitation 06:10 - Achievements 08:21 - Future Improvements 13:30 - Quirks 16:34 - Final Thoughts AlphaEvolve release: https://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/ Paper: https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/AlphaEvolve.pdf Terence Tao Quote: https://mathstodon.xyz/@tao/114508029896631083 Nature Article: https://www.nature.com/articles/s41586-022-05172-4 MIT Article: https://www.technologyreview.com/2025/05/14/1116438/google-deepminds-new-ai-uses-large-language-models-to-crack-real-world-problems/ AI Co-Scientist: https://arxiv.org/pdf/2502.18864 OpenAI Codex: https://openai.com/index/introducing-codex/ 70% of Pull Requests: https://x.com/slow_developer/status/1920920456393028027 Amodei Essay: https://www.darioamodei.com/essay/machines-of-loving-grace OpenAI Jason Wei Tweet: https://x.com/_jasonwei/status/1923091260354531612 PromptBreeder: https://arxiv.org/pdf/2309.16797 DrEureka: https://arxiv.org/pdf/2406.01967 FT DeepMind: https://www.ft.com/content/4e497a91-670a-4f69-be4a-18e247daba3e Non-hype Newsletter: https://signaltonoise.beehiiv.com/
- AE
Apr 25, 2025 · 14 mino3 breaks (some) records, but AI becomes pay-to-win
A green card, o3 vs Gemini 2.5, 6 Benchmarks and a whole bunch of my thoughts on what on earth is happening in AI, from here to 2030. Plus, how AI is becoming pay-to-win, and why. Crazy times, 14 mins probably wasn’t enough. https://app.grayswan.ai/ai-explained AI Insiders ($9!): https://www.patreon.com/AIExplained Chapters: 00:00 - Introduction 00:33 - FictionLiveBench 01:37 - PHYBench 02:14 - SimpleBench 02:54 - Virology Capabilities Test 03:13 - Mathematics Performance 04:29 - Vision Benchmarks 05:43 - V* and how o3 works 06:44 - Revenue and costs for you 08:54 - Expensive RL and trade-offs 09:40 - How to spend the OOMs 13:27 - Gray Swan Arena Green Card: https://techcrunch.com/2025/04/25/an-openai-researcher-who-worked-on-gpt-4-5-had-their-green-card-denied/ PHYBench: https://arxiv.org/pdf/2504.16074Virologytest: https://www.virologytest.ai/ How o3 Vision Works: https://arxiv.org/pdf/2312.14135 https://x.com/sainingxie/status/1912570624523829573 Visual puzzles: https://neulab.github.io/VisualPuzzles/ Fiction Bench: https://x.com/ficlive/status/1912863028141244850 https://geobench.org/ https://simple-bench.com/ AIME 2025: https://openai.com/index/introducing-o3-and-o4-mini/ USAMO: https://x.com/mbalunovic/status/1914398518896193747 NaturalBench: https://linzhiqiu.github.io/papers/naturalbench/ Where’s Waldo: https://uk.pinterest.com/pin/492792384225896298/ IMO and AlphaProof:https://deepmind.google/discover/blog/ai-solves-imo-problems-at-silver-medal-level/ Crazy Revenue: https://www.theinformation.com/articles/openai-forecasts-revenue-topping-125-billion-2029-agents-new-products-gain?rc=sy0ihq Number of Users: https://www.theinformation.com/briefings/googles-gemini-user-numbers-revealed-court?rc=sy0ihq Subscriptions pay to win: https://www.forbes.com/sites/paulmonckton/2025/04/23/google-leak-reveals-new-gemini-ai-subscription-levels/ GPU Trade-offs: https://x.com/sama/status/1915098951067554030 RL Scale-up Amodei: https://www.darioamodei.com/post/on-deepseek-and-export-controls Log-linear Returns: https://x.com/bobmcgrewai/status/1895228291981943265 2030 Scaling: https://epoch.ai/blog/can-ai-scaling-continue-through-2030 Model Size: https://x.com/slow_developer/status/1874554473256997201 Adam on AGI: https://x.com/TheRealAdamG/status/1913998366632968381 Papers on Patreon: https://arxiv.org/pdf/2502.01839 https://arxiv.org/pdf/2504.13837 Chollet Quote: https://x.com/fchollet/status/1912934762580447447 OpenSim: https://opensim.stanford.edu/ Non-hype Newsletter: https://signaltonoise.beehiiv.com/
- AE
Apr 16, 2025 · 14 mino3 and o4-mini - they’re great, but easy to over-hype
Critical analysis of the two most powerful new models behind ChatGPT, o3 and o4-mini. Not just the system cards, benchmarks, and my own tests, but some you may not have seen before. Yes, they can whip up amazing front-end in a few seconds, but you always have to ask what is in their data. Either way, they prove the gains from RL are just beginning… https://weave-docs.wandb.ai/?utm_source=sponsorship&utm_medium=simple_bench&utm_campaign=ai_explained AI Insiders ($9!): https://www.patreon.com/AIExplained Chapters: 00:00 - o3 and o4-mini https://simple-bench.com/ Plus, Teams and Pro, plus token count: https://x.com/btibor91/status/1912568994512662679 System Card: https://openai.com/index/o3-o4-mini-system-card/ Release Notes: https://openai.com/index/introducing-o3-and-o4-mini/ https://deepmind.google/technologies/gemini/pro/ https://x.com/DeryaTR_/status/1912558350794961168 https://x.com/polynoamial/status/1912564068168450396 API Pricing:https://openai.com/api/pricing/ https://aider.chat/docs/leaderboards/ Non-hype Newsletter: https://signaltonoise.beehiiv.com/
- AE
Apr 16, 2025 · 20 min‘Speaking Dolphin’ to AI Data Dominance, 4.1 + Kling 2: 7 Developments Critically Analysed
This pod won’t just be about the release of GPT 4.1 in the last 48 hours, o3 build-up, Kling 2.0, a sneak-peak at the next OpenAI model, or even the new Dolphin language tool. It will be about 7 such stories that contextualise where we are in AI and what is happening. https://www.emergentmind.com/ Chapters: 00:00 - Introduction 00:30 - Kling 2.0 01:35 - GPT 4.1 05:25 - o3 Build-up 07:37 - ‘Product Company’ 09:31 - Safe Superintelligence 10:54 - DolphinGemma 13:16 - Data Dominance? Kling 2.0: https://app.klingai.com/global/release-notes Dolphin Gemma: https://blog.google/technology/ai/dolphingemma/?s=09 https://openai.com/index/gpt-4-1/ OpenAI o3 Build-up The Information: https://www.theinformation.com/articles/openais-latest-breakthrough-ai-comes-new-ideas?rc=sy0ihq Physical reasoning: https://x.com/a_karvonen/status/1911839968990814503 Fiction Live.bench: https://x.com/ficlive/status/1911853409847906626 Altman Ted: https://www.youtube.com/watch?v=5MWT_doo68k https://simple-bench.com/try-yourself https://aider.chat/docs/leaderboards/ 4.5: https://www.youtube.com/watch?v=6nJZopACRuQ Geospatial reasoning: https://research.google/blog/geospatial-reasoning-unlocking-insights-with-generative-ai-and-multiple-foundation-models/ Pioneers: https://x.com/OpenAIDevs/status/1910017976256119151 Evals: https://www.youtube.com/watch?v=scsW6_2SPC4 Anthropic Updates: https://www.bloomberg.com/news/articles/2025-04-15/anthropic-is-readying-a-voice-assistant-feature-to-rival-openai?srnd=phx-ai https://x.com/sethsaler/status/1912188383457059301 https://techcrunch.com/2025/04/12/openai-co-founder-ilya-sutskevers-safe-superintelligence-reportedly-valued-at-32b/ https://ai.meta.com/blog/llama-4-multimodal-intelligence/ https://deepmind.google/technologies/gemini/pro/ https://research.google/blog/accelerating-scientific-breakthroughs-with-an-ai-co-scientist/ https://blog.google/products/google-cloud/ironwood-tpu-age-of-inference/ OpenAI Documentary: https://www.patreon.com/posts/one-machine-to-121940490