Skip to content
Artwork for AI Explained
NewsTechnologyEducation

AI Explained

AI Explained

Covering the biggest news of the century - the arrival of smarter-than-human AI. The author of Simple Bench, exposing the remaining human-LLM reasoning gap.

Solo Developer of LM Council: https://lmcouncil.ai - Code - INSIDER15

Join me at AI Insiders, with exclusive videos and a 1000+ network of AI enthusiasts and professionals: https://www.patreon.com/AIExplained

Links:
AI Insiders: Code - INSIDER15: https://www.patreon.com/AIExplained
My AI Council App: https://lmcouncil.ai
Simple Bench: https://simple-bench.com/
Podcast (New!): https://aiexplainedopodcast.buzzsprout.com/
Newsletter: https://signaltonoise.beehiiv.com/
X: https://twitter.com/AIExplainedYT

Play
  • 100 episodes
  • daily
  • Avg 22 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • Today · 41 min

    Claude Fable 5 - Full 319 page Breakdown

    Fable 5 is out - and it’s good, very good. But beyond the splashy demos, I want to bring you the 20+ nuggets from the 319 page system card, which I read in full, all day, plus benchmarks you may not have noticed. https://assemblyai.com/aiexplained Plus two worrying trends inside the ‘mind’ of Claude, how OpenAI counter, and the transformer inventor’s warning. Check out my fast-growing (!) app, free to use, and code INSIDER15 for paid tiers: https://lmcouncil.ai AI Insiders ($9!): https://www.patreon.com/AIExplained Chapters: 00:00 - Introduction 01:06 - Blocks + Better Models 02:42 - Fable 5 Upgrade over Mythos Preview 04:49 - ML Acceleration Bombshell 07:11 - No RSI yet 07:41 - Bio-capable 14:51 - Creative Writing … no 17:23 - Does need bug-checks 18:57 - OpenAI Response 19:23 - Benchmark Bonanza 28:06 - Chain of Thought worrying trend Fable 5 Release: https://www.anthropic.com/news/claude-fable-5-mythos-5 System Card: https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdf Intelligence Explosion: https://www.patreon.com/posts/anthropic-charts-160231656 Annotated: https://x.com/Miles_Brundage/status/2064500190523113816/photo/1 OpenAI Counter: https://x.com/thsottiaux/status/2064572118264913923 https://x.com/thsottiaux/status/2043177597434306699 Double Lifespan: https://darioamodei.com/essay/machines-of-loving-grace AutomationBench: https://zapier.com/benchmarks Vending Bench: https://x.com/andonlabs/status/2064429817530085804 CritPt: https://critpt.com/ Riemann Bench: https://surgehq.ai/leaderboards/riemann-bench GDPVal: https://artificialanalysis.ai/evaluations/gdpval-aa BluePrint Bench 2: https://andonlabs.com/evals/blueprint-bench-2 MCP Atlas: https://labs.scale.com/leaderboard/mcp_atlas FutureSim: https://x.com/nikhilchandak29/status/2064676801440358774 Roon Stun Lock: https://x.com/tszzl/status/2064454617568874669 Noam Brown Inference Ceiling: https://x.com/polynoamial/status/2064210146558136827 Isochronic Chart: https://isochronic-passage-chart.netlify.app/#nyc Rose Tavern: https://claude.ai/public/artifacts/2295bebe-77e6-43e2-ae94-0fe49e9a776b Redwall Game: https://redwall-mossflower.surge.sh/ Risk Report: https://www-cdn.anthropic.com/097c63b5fe7dd8b14866e1f15bb1910ec713658a.pdf Transformer Inventor Warning: https://x.com/tszzl/status/2064563986914554125 Non-hype Newsletter: https://signaltonoise.beehiiv.com/ Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices

  • Today · 18 min

    Apple’s ‘AI Can’t Reason’ Claim Seen By 13M+, What You Need to Know

    What to make of those headlines that AI can’t reason, seen by tens of millions? I cover the Apple paper in layman’s terms, what it means and doesn’t mean, and what’s next. Thanks to Storyblocks for sponsoring this video! Download unlimited stock media at one set price with Storyblocks: https://storyblocks.com/AIExplained Plus o3-pro and whether it is my current most-recommended model. AI Insiders ($9!): https://www.patreon.com/AIExplained Chapters: 00:00 - Introduction 00:57 - Viral Post + Headlines 01:42 - Apple Paper Analysis 08:34 - But they do Hallucinate 10:43 - Not Supercomputers 11:18 - o3 Pro and Recommendations 13.7M Tweet: https://x.com/RubenHssd/status/1931389580105925115 Apple Paper: https://ml-site.cdn-apple.com/papers/the-illusion-of-thinking.pdf Guardian Article: https://www.theguardian.com/technology/2025/jun/09/apple-artificial-intelligence-ai-study-collapse Lisan al Gaib post: https://x.com/scaling01/status/1931854370716426246 Multiplication: https://x.com/yuntiandeng/status/1836114401213989366 The Illusion of the Illusion of Thinking: https://drive.google.com/file/d/1Zx9ikRj0Enc3SB4wA9HlYIlpmO_8QiUO/view Marcus: https://www.theguardian.com/commentisfree/2025/jun/10/billion-dollar-ai-puzzle-break-down Prof Rao: https://x.com/rao2z/status/1927707640223719631 AI Job Headlines: https://www.nytimes.com/2025/06/11/technology/ai-mechanize-jobs.html https://www.axios.com/2025/05/28/ai-jobs-white-collar-unemployment-anthropic Sky News Story: https://news.sky.com/story/can-we-trust-chatgpt-despite-it-hallucinating-answers-13380975 Veo 3 Kalshi Ad: https://x.com/Kalshi/status/1932891608388681791 Altman Essay: https://blog.samaltman.com/ o3 Original benchmarks: https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b8b6c44-acd6-43b3-b5c6-1a1d5c6c25e4_2486x1388.png https://pbs.twimg.com/media/GfQ0bfcXQAAQt13.jpg Alpha Evolve Video: https://www.youtube.com/watch?v=RH4hAgvYSzg https://simple-bench.com/ Non-hype Newsletter: https://signaltonoise.beehiiv.com/ Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices

  • Today · 18 min

    GPT 4: Full Breakdown (14 Details You May Have Missed)

    I just read the entire technical report on GPT 4, not just the promotional hype. And boy does it have some interesting details. I have gathered the 14 extra details that you, or at least the media, may miss from the release. The last one is more than a little wild. These include things like the training secrets, the cherry-picked bar exam stat, text-to-image breakthroughs, and some truly astounding safety checks. https://www.patreon.com/AIExplained https://cdn.openai.com/papers/gpt-4.pdf https://www.bemyeyes.com/ https://arxiv.org/pdf/2104.12756.pdf https://arxiv.org/pdf/2203.10244.pdf https://chat.openai.com/chat?model=gpt-4 https://arxiv.org/pdf/2302.10329.pdf https://www.alignmentforum.org/posts/Aq82XqYhgqdPdPrBA/full-transcript-eliezer-yudkowsky-on-the-bankless-podcast Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices

  • Today · 18 min

    Claude Fable Blocked - 11 Quiet Details on What’s Next

    Claude Fable 5 banned, but what’s the bigger story. We go through 11 under-reported details, so you have the context to see what’s coming next for your use of AI. From whether the ban will last, what the possible motives are, what the model can actually do, and some wild over-extrapolations going on. Check out my fast-growing (!) app, free to use, and code INSIDER15 for paid tiers: https://lmcouncil.ai AI Insiders ($9!): https://www.patreon.com/AIExplained Chapters: 00:00 - Introduction 00:51 - Came from an Anthropic Investor ‘and other tech leaders’ 01:47 - Govt pressured by CEOs like Jamie Dimon 03:01 - ‘Already decided’ 04:02 - Prompt Injection Robustness Comparison 05:15 - Wellness? 06:36 - “Overreach” 08:17 - Anthropic Did Admit it would cause Difficulty 09:32 - 90 Minutes 10:02 - Equity Absence 10:31 - Lobbying and OpenAI ‘Already Decided’ - https://www.theinformation.com/articles/amazons-jassy-raised-concerns-anthropic-model-trump-crackdown?rc=sy0ihq Not for Other Models: https://www.theinformation.com/briefings/u-s-government-unlikely-extend-anthropic-export-control-ai-companies?rc=sy0ihq 90 Minutes: https://archive.fo/20260614001605/https://www.politico.com/news/2026/06/13/inside-the-whirlwind-24-hours-that-led-the-white-house-to-slap-export-controls-on-anthropic-00961519#selection-807.1-807.219 Anthropic Statement: https://www.anthropic.com/news/fable-mythos-access Life Comes at you Fast: https://x.com/etbrooking/status/2065638276388495742 Anthropic Deputy CISO: https://x.com/TheTranscript_/status/2065883670053847324 Hegseth Gloat: https://x.com/PeteHegseth/status/2065897156226015690 Roon Speculation: https://x.com/tszzl/status/2065939227167392147 Mythos System Card: https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdf Sachs Statement: https://x.com/DavidSacks/status/2065853007619588171 OpenAI Lobbying: https://thehill.com/policy/technology/5912720-altman-openai-get-bogged-down-in-political-spending-fight/ Absent from Equity Talks: https://finance.yahoo.com/sectors/technology/articles/trump-ai-ownership-plan-could-131053732.html Pliny Jaibreak: https://x.com/elder_plinius/status/2064776322979676227 Fusion: https://x.com/OpenRouter/status/2065856871215329545 https://lmcouncil.ai Non-hype Newsletter: https://signaltonoise.beehiiv.com/ Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices

  • Today · 17 min

    Time Until Superintelligence: 1-2 Years, or 20? Something Doesn't Add Up

    Superintelligence when? This question is more urgent than ever as we hear competing timelines from Inflection AI, OpenAI and the leaders at the Center for AI Safety. This video not only covers the what was said, it offers some data on what superintelligence is projected to be capable of. I discuss things that can hasten those timelines (see the new Netflix doc, Top 1% for Creativity and the AI Natural Selection paper) or slow it down (ft. Yuval Noah Hariri, the new Jailbroken paper, and more). And I end on some reflections of what it might mean to interact with a superintelligence (ft Douglas Hofstadter). Introducing Superalignment: https://openai.com/blog/introducing-superalignment Odds of Success: https://twitter.com/janleike/with_replies?ref_src=twsrc%5Egoogle%7Ctwcamp%5Eserp%7Ctwgr%5Eauthor Mustafa Suleyman (Inflection AI) Interview: https://www.youtube.com/watch?v=oAxNkehgzEU Inflection Supercomputer: https://www.tomshardware.com/news/startup-builds-supercomputer-with-22000-nvidias-h100-compute-gpus GPT 2030: https://www.lesswrong.com/posts/WZXqNYbJhtidjRXSi/what-will-gpt-2030-look-like Counterpoint: https://twitter.com/xuanalogue/status/1666765447054647297 Creativity Test: https://www.umt.edu/news/2023/07/070523test.php MMLU: https://arxiv.org/pdf/2009.03300.pdf Boston Globe CAI: https://www.bostonglobe.com/2023/07/06/opinion/ai-safety-human-extinction-dan-hendrycks-cais/ Yuval Noah Hariri Guardian: https://www.theguardian.com/technology/2023/jul/06/ai-firms-face-prison-creation-fake-humans-yuval-noah-harari Jailbroken Paper: https://arxiv.org/pdf/2307.02483.pdf Suleyman Hallucination Tweet: https://twitter.com/mustafasuleymn/status/1678072401760796672?ref_src=twsrc%5Egoogle%7Ctwcamp%5Eserp%7Ctwgr%5Etweet Killer Robots: https://www.youtube.com/watch?v=YsSzNOpr9cE Automated Apple-picking: https://twitter.com/AiBreakfast/status/1677828971214282752 50 year Mortgages: https://www.ft.com/content/281fbba6-28e2-42d4-b241-0f215995f0d1 Heypi: https://heypi.com/talk Natural Selection Favors AIs over Humans: https://arxiv.org/pdf/2303.16200.pdf#:~:text=Natural%20selection%20may%20be%20a,designs%20to%20be%20selected%20naturally. Gödel, Escher, Bach author Doug Hofstadter on the state of AI today: https://www.youtube.com/watch?v=lfXxzAVtdpU&t=1763s https://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices

  • Today · 19 min

    GPT 5 is All About Data

    Drawing upon 6 academic papers, interview snippets, possible leaks, and my own extensive research, I put together everything we might know about GPT 5: what will determine its IQ, its timeline and its impact on the job market and beyond. Starting with an insider interview on the names of GPT models, such as GPT 4 and GPT 5, then looking into the clearest hint that GPT 4 is inside Bing. Next, I briefly cover reports of a leak about GPT 5 and discuss the scale of GPUs require to train it, touching on the upgrade form A100 to H100 GPUs. Then the DeepMind paper that changed everything, focusing LLM research on data rather than parameter count. I go over a lesswrong post about that paper's 'wild implications'. And then the key paper: 'Will We Run Out of Data'. This encapsulates the key dynamic that will either propel or bottleneck GPT and other LLM improvements. Next, I examine a different take, that perhaps data is already limited and caused the Sydney model of Bing. This opens up to a discussion on the data behind these models and why Big Tech is so unforthcoming about where it originates. Could a new legal war be brewing? I then cover 4 of the ways these models may improve even without data augmentation, such as Automatic Chain of Thought, high quality data extraction, tool training, including Wolfram Alpha, retraining on existing data sets, artificial data generation and more. We take a quick look at Sam Altman's timelines and host of Big Bench benchmarks that they may impact, such as reading comprehension, critical reasoning, logic, physics and Math. I address Altman's quote about timelines being delayed by alignment and safety and finally, Altman's comments on AGI and how they pertain to GPT 5. https://www.patreon.com/AIExplained https://stratechery.com/2023/new-bing-and-an-interview-with-kevin-scott-and-sam-altman-about-the-microsoft-openai-partnership/ https://twitter.com/XiXiDu/status/1285225627390443521 https://www.linkedin.com/pulse/building-new-bing-jordi-ribas/ https://twitter.com/davidtayar5/status/1625140481016340483/photo/1 https://www.techradar.com/news/chatgpt-might-bring-about-another-gpu-shortage-sooner-than-you-might-expect https://www.nvidia.com/en-gb/data-center/h100/ https://arxiv.org/pdf/2203.15556.pdf https://www.lesswrong.com/posts/6Fpvch8RR29qLEWNH/chinchilla-s-wild-implications https://twitter.com/Meaningness/status/1628052277050376192 https://arxiv.org/pdf/2211.04325.pdf https://twitter.com/Meaningness/status/1628052277050376192 https://www.searchenginejournal.com/google-bard-training-data/478941/#close https://arxiv.org/pdf/2302.12822.pdf https://arxiv.org/pdf/2302.04761.pdf https://www.wolframalpha.com/ https://arxiv.org/pdf/2207.14502.pdf https://aclanthology.org/2022.findings-emnlp.508.pdf https://www.technologyreview.com/2022/11/24/1063684/we-could-run-out-of-data-to-train-ai-language-programs/ https://twitter.com/sama/status/1625980933861175306 https://github.com/google/BIG-bench/blob/main/bigbench/benchmark_tasks/results/plot_BIG-bench_lite_aggregate.png https://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/gre_reading_comprehension https://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/logical_args https://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/physics https://lifearchitect.ai/gpt-4/ https://www.lesswrong.com/posts/uxnjXBwr79uxLkifG/comments-on-openai-s-planning-for-agi-and-beyond https://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices

  • Today · 20 min

    GPT-5 has Arrived

    GPT-5 will change how hundreds of millions of people use AI. Yes, you might have to forgive the chart crimes, the underwhelming livestream and Altman hype… But it’s a good model. I have read the 50 page system card in full, have the benchmark scores, coding tests, and things you might have missed. https://app.grayswan.ai/ai-explained AI Insiders ($9!): https://www.patreon.com/AIExplained Announcement: https://openai.com/index/introducing-gpt-5/ System Card: https://cdn.openai.com/pdf/8124a3ce-ab78-4f06-96eb-49ea29ffb52f/gpt5-system-card-aug7.pdf Extra Paper: https://cdn.openai.com/pdf/be60c07b-6bc2-4f54-bcee-4141e1d6c69a/gpt-5-safe_completions.pdf Altman tweet: https://x.com/sama/status/1953551377873117369 Livestream: https://www.youtube.com/watch?v=0Uu_VJeVVfo METR Report: https://metr.github.io/autonomy-evals-guide/gpt-5-report/ ARC-AGI-2: https://x.com/fchollet/status/1953511631054680085 Claude Opus 4.1: https://www.anthropic.com/news/claude-opus-4-1 MMMU: https://mmmu-benchmark.github.io/ Cursor Praise: https://x.com/ryolu_/status/1953531724895596669 Non-hype Newsletter: https://signaltonoise.beehiiv.com/ Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices

  • Today · 23 min

    Fable 5 vs GPT 5.6 Sol: The Early Results

    Fable 5 (newly re-released) vs GPT 5.6 Sol, what comparisons can we unearth? Plus, Sonnet 5, a 5% equity seizure by US Govt, the ‘largest heist’, Beetlejuice and more… Exclusive Vids in AI Insiders ($9!): https://www.patreon.com/AIExplained Chapters: 00:00 - Introduction 01:06 - Fable Timeline 02:59 - Sol Release? 06:01 - the Chinese Angle 07:35 - 5% Stake 09:10 - Fable vs Sol - the numbers 13:57 - Sol Misalignment 15:27 - Claude Sonnet 5 16:24 - GLM 5.2 + New Paper on Large Model Learning GPT 5.6 Sol Release: https://openai.com/index/previewing-gpt-5-6-sol/ Altman on Concentrated Power: https://x.com/giacomomiolo/status/2070609506925478352 System Card: https://deploymentsafety.openai.com/gpt-5-6-preview/gpt-5-6-preview.pdf GLM 5.2: https://www.patreon.com/AIExplained/posts/glm-5-2-and-nsa-161869508 Altman Memo: https://www.theinformation.com/articles/trump-administration-asks-openai-stagger-release-new-model-security-concerns?rc=sy0ihq Fable 5 Revised Timeline: https://www.anthropic.com/news/redeploying-fable-5 Claude Sonnet 5: https://www.anthropic.com/news/claude-sonnet-5 Mythose 5 System Card: https://www-cdn.anthropic.com/9e6a1044980d8c4ed85669faf9c2a8342e2e9f1e/Claude%20Sonnet%205%20System%20Card.pdf Anthropic Accuse Alibaba Cloud / Qwen: https://www.bbc.co.uk/news/articles/cwyklykn5dwo Betelgeuse: https://upload.wikimedia.org/wikipedia/commons/6/69/Well_known_stars_2.png OpenAI 5%: https://edition.cnn.com/2026/07/02/business/openai-trump-stake-intl Why Larger Models Learn More: https://arxiv.org/pdf/2605.29548 Patreon Post: https://www.patreon.com/AIExplained/posts/glm-5-2-and-nsa-161869508 Non-hype Newsletter: https://signaltonoise.beehiiv.com/ Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices

  • Yesterday · 31 min

    AI CEO: ‘Stock Crash Could Stop AI Progress’, Llama 4 Anti-climax + ‘Superintelligence in 2027’ ...

    The latest on Llama 4, and whether it signals a slowdown in AI, or solid progress. Plus, a deep dive on that viral prediction of superintelligence by 2027, and Dario Amodei’s cautionary words on what could stop AI progress in its tracks. o3 news, and more, as well. Weights & Biases: https://weave-docs.wandb.ai/?utm_source=sponsorship&utm_medium=simple_bench&utm_campaign=ai_explained DeepSeek Doc: https://www.patreon.com/posts/openai-is-not-r1-125869969 AI Insiders ($9!): https://www.patreon.com/AIExplained Chapters: 00:00 - Introduction 00:47 - Stock Crash 02:28 - Llama 4 10:55 - o3 News 11:59 - OpenAI non-profit? 13:13 - AI 2027 Llama 4 Release: https://ai.meta.com/blog/llama-4-multimodal-intelligence/ Dario Amodei Comments: https://www.youtube.com/watch?v=esCSpbDPJik Knowledge Cut-off: https://www.llama.com/docs/model-cards-and-prompt-formats/llama4_omni/ Aider Polyglot: https://aider.chat/docs/leaderboards/ Gemini 1.5: https://arxiv.org/pdf/2403.05530 Fiction-LiveBench: https://fiction.live/stories/Fiction-liveBench-Mar-25-2025/oQdzQvKHw8JyXbN87 OpenAI Valuation: https://www.nytimes.com/2025/03/31/technology/openai-valuation-300-billion.html?login=smartlock&auth=login-smartlock OpenAI Cybersecurity: https://www.bloomberg.com/news/articles/2024-01-16/openai-working-with-us-military-on-cybersecurity-tools-for-veterans Deep research System Card: https://cdn.openai.com/deep-research-system-card.pdf https://openai.com/index/paperbench/ AI 2027: https://ai-2027.com/ METR Paper: https://arxiv.org/pdf/2503.14499 OpenAI non-profit: https://openai.com/index/nonprofit-commission-guidance/ NYT Piece: https://www.nytimes.com/2025/04/03/technology/ai-futures-project-ai-2027.html?unlocked_article_code=1.804._yKi.QhwOp15Q3tcU&smid=url-share&s=09 Kokotajlo predictions 2021: https://www.lesswrong.com/posts/6Xgy6CAf2jqHhynHL/what-2026-looks-like https://simple-bench.com/ Non-hype Newsletter: https://signaltonoise.beehiiv.com/ Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices

  • Yesterday · 23 min

    Sam Altman's World Tour, in 16 Moments

    Missed by much of the media, Sam Altman (and co) have revealed at least 16 surprising things over his World Tour. From AI's designing AIs to 'unstoppable opensource', the 'customisation' leak (with a new 16k ChatGPT and 'steerable GPT 4), AI and religion, and possible regrets over having 'pushed the button'. I'll bring in all of this and eleven other insights, together with a new and highly relevant paper just released this week on 'dual-use'. Whether you are interested in 'solving climate change by telling AIs to do it', 'staring extinction in the face' or just a deepfake Altman, this video touches on it all, ending with comments from Brockman in Seoul. I watched over ten hours of interviews to bring you this footage from Jordan, India, Abu Dhabi, UK, South Korea, Germany, Poland, Israel and more Altman Abu Dhabi, HUB71, 'change it's architecture': https://youtu.be/RZd870NCukg Israel, TAUVOD, 'AIs Creating AIs': https://www.youtube.com/live/VWUhASix9ws?feature=share Poland, Ideas NCBR, 'opensource + 10 years away': https://www.youtube.com/live/tSCrQQbPPHk?feature=share Dual-use Biotech: https://arxiv.org/ftp/arxiv/papers/2306/2306.03809.pdf Jordan Xpand, 'pushed a button': https://www.youtube.com/live/dgh-L2nk97M?feature=share India Economic Times, 'railgun': https://youtu.be/T-lj7ItGjZE Guardian Altman Interview: https://www.theguardian.com/technology/2023/jun/07/what-should-the-limits-be-the-father-of-chatgpt-on-whether-ai-will-save-humanity-or-destroy-it Germany Conversation: https://youtu.be/uaQZIK9gvNo Leaked Layout: https://the-decoder.com/new-chatgpt-features-are-on-the-way-workspace-file-uploads-profiles/ OpenAI in Seoul, SoftBank Ventures Asia: https://youtu.be/_hpuPi7YZX8 India, Digital India, 'hallucinations in 18 months': https://www.youtube.com/live/Pig9WbMN1lQ?feature=share https://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices

  • Yesterday · 13 min

    8 New Ways to Use Bing's Upgraded 8 [now 20] Message Limit (ft. pdfs, quizzes, tables, scenarios...)

    Not only has Microsoft changed the message limit for Bing, I have come along and given you 8 brand new use cases to try out. These useful, productive, fun and informative scenarios ranges from amalgamating academic papers to debating dead philosophers, from being moments from tragedy to hacking education [through multiple choice quizzes]. Use Bing to boost your learning, increase your productivity, play games, roleplay movies and shows and much more. If you learn anything from these bleeding edge deployments of the likely GPT4 LLM that powers Bing, do let me know in the comments or drop a like. The prompts used were: Create a multiple choice quiz on [transformers in the context of machine learning]. Do not provide the answers and explanations until I have answered and always provide another question after each answer. Please begin with the first question. Explain how Sauron could have defeated the Fellowship in Lord of the Rings Summarise any novel insights that be drawn from combining these academic papers: https://arxiv.org/pdf/2211.04325.pdf and https://arxiv.org/pdf/2207.14502.pdf It is 8am on the 1st November 1755. I am standing on a banks of the beautiful river Tagus, in Lisbon. It seems like a lovely day. Do you have any advice for me? I want to debate the philosopher Socrates of Athens, so please reply only as he would. I want to debate the merits of eating meat. Let me start the discussion. 'Eating meat is morally wrong, as it causes unnecessary suffering.' Create a table of 10 comparisons between the Mona Lisa and Colgate Toothpaste What would Napoleon think about the deal between OpenAI and Microsoft, and how would that view differ from the view of Mahatma Gandhi? And more! Credit for at least 3 of these prompts goes to Ethan Mollick: https://twitter.com/emollick https://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices

  • Yesterday · 19 min

    o3 and o4-mini - they’re great, but easy to over-hype

    Critical analysis of the two most powerful new models behind ChatGPT, o3 and o4-mini. Not just the system cards, benchmarks, and my own tests, but some you may not have seen before. Yes, they can whip up amazing front-end in a few seconds, but you always have to ask what is in their data. Either way, they prove the gains from RL are just beginning… https://weave-docs.wandb.ai/?utm_source=sponsorship&utm_medium=simple_bench&utm_campaign=ai_explained AI Insiders ($9!): https://www.patreon.com/AIExplained Chapters: 00:00 - o3 and o4-mini https://simple-bench.com/ Plus, Teams and Pro, plus token count: https://x.com/btibor91/status/1912568994512662679 System Card: https://openai.com/index/o3-o4-mini-system-card/ Release Notes: https://openai.com/index/introducing-o3-and-o4-mini/ https://deepmind.google/technologies/gemini/pro/ https://x.com/DeryaTR_/status/1912558350794961168 https://x.com/polynoamial/status/1912564068168450396 API Pricing:https://openai.com/api/pricing/ https://aider.chat/docs/leaderboards/ Non-hype Newsletter: https://signaltonoise.beehiiv.com/ Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices

  • Yesterday · 21 min

    Sora 2 - It will only get more realistic from here

    Sora 2 - the start of the infinite slop-feed or a key step to a generalist agent? Better than VEO 3 or over-hyped? I bring out 6 details you may have missed, contrast the announcement to Periodic Labs and even squeeze in some Claude Sonnet 4.5 analysis. Maybe I should make my videos longer… https://80000hours.org/aiexplained AI Insiders ($9!): https://www.patreon.com/AIExplained Chapters: 00:00 - Introduction 00:40 - Two models? 01:15 - Rollout Details 01:43 - Versus Sora 1 / Veo 3 04:30 - Sora App / Social Media 06:40 - Masterplan 09:30 - Generalist Agent? Periodic Labs 12:05 - Claude Sonnet 4.5 13:42 - Future Outlook Announcement: https://openai.com/index/sora-2/ Launch Video: https://www.youtube.com/live/gzneGhpXwjU System Card: https://cdn.openai.com/pdf/50d5973c-c4ff-4c2d-986f-c72b5d0ff069/sora_2_system_card.pdf Sam Altman Blog Post on Sora App: https://blog.samaltman.com/sora-2 Most Intelligent Claim: https://x.com/willdepue/status/1973089331284681110 GTA: https://x.com/AndrewCurran_/status/1973298436536766666 Meta Vibes: https://x.com/alexandr_wang/status/1971295156411433228?s=46 Altman on Regulations: https://www.lesswrong.com/posts/5jjk4CDnj9tA7ugxr/openai-email-archives-from-musk-v-altman OpenAI Profit: https://www.theinformation.com/articles/openais-first-half-results-4-3-billion-sales-2-5-billion-cash-burn?rc=sy0ihq Periodic Labs: https://periodic.com/ https://www.nytimes.com/2025/09/30/technology/ai-meta-google-openai-periodic.html https://x.com/LiamFedus/status/1973055380193431965 https://baincapitalventures.com/insight/we-must-know-we-will-know/?s=09 Sonnet 4.5: https://www.anthropic.com/news/claude-sonnet-4-5 https://simple-bench.com/ Non-hype Newsletter: https://signaltonoise.beehiiv.com/ Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices

  • Yesterday · 23 min

    A Model Explosion: GPT 5.6 Sol, Grok 4.5 and Meta Muse Rewrite the Rules

    What a week in AI, for real. GPT 5.6 may actually beat Claude Fable, in what you get for your money, while the new Grok 4.5 and Meta Muse Spark 1.1 make the choice even harder. Uncovering a dozen nuggets of gold you may have missed from all the viral headlines, I can also assure you you’ll learn something you didn’t know before. For Exclusive Videos, go to AI Insiders (less than $9!): https://www.patreon.com/AIExplained Chapters: 00:00 - Introduction 01:03 - GPT 5.6 Sol Reveals 05:08 - Missing benches, plus Grok 4.5 07:17 - Gaming as the new frontier? 08:31 - Muse Spark 1.1 10:03 - SimpleBench Upgrade 11:17 - Ultra Sol + Self-Improvement 13:44 - well, this is awkward 15:41 - Why model improvement will not plateau anytime soon AI Consciousness: https://www.patreon.com/AIExplained/posts/anthropics-quite-163360718 I Smell Fear: https://x.com/thsottiaux/status/2075287108680601929 GPT 5.6: https://openai.com/index/gpt-5-6/ Grok 4.5: https://x.ai/news/grok-4-5?twclid=2ezs408o0z23pw07tmxcwbzibd Meta Muse Spark 1.1: https://ai.meta.com/blog/introducing-muse-spark-meta-model-api/ Proliferating GPT Toggles: https://x.com/rasbt/status/2075369179817902176/photo/1 Anthropic Call-out: https://x.com/Mononofu AI Security Institute Finding: https://x.com/alxndrdavies/status/2075279480331874306 Competitive Coding: https://x.com/FakePsyho/status/2075128093891801305/photo/1 Agents Last Exam: https://agents-last-exam.org/ Dawn Song: https://x.com/dawnsongtweets/status/2065095757988868190 https://simple-bench.com/ SWE-Marathon: https://www.swe-marathon.org/ https://www.frontierswe.com/ ARC-AGI 3: https://x.com/arcprize/status/2075270869992264003 Automation Bench: https://zapier.com/benchmarks VibeCode Bench: https://www.vals.ai/benchmarks/vibe-code ‘Post-Train Claim’: https://posttrainbench.com/ Redwall Game: https://redwall-bellmaker-7e03e4.surge.sh/ Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices

  • Yesterday · 27 min

    You Are Being Told Contradictory Things About AI

    With headlines of an imminent job apocalypse, code red for ChatGPT and recursive self-improvement, at the same time as Anthropic's CEO yesterday saying we know how to scale to AGI, and Gemini 3 DeepThink out today, it is easy to get lost among the narratives and counter-narratives. So here are both, plus the facts behind them, for you to decide. https://epoch.ai/ai-explained-datacenters Epoch AI is the sponsor of today’s video, and my views, and those expressed in this video, do not necessarily reflect Epoch AI’s views in any way. Chapters: 00:00 - Introduction 00:42 - Job Apocalypse? 01:45 - Scaling to AGI 04:15 - Recursive Self-Improvement Needed, or Not 09:57 - OpenAI Code Red vs Gemini 3 DeepThink vs Claude Opus 4.5 13:27 - DeepSeek Speciale vs Mistral Large v3 16:45 - Claude Soul Document https://lmcouncil.ai/ AI Insiders ($9!): https://www.patreon.com/AIExplained Guardian Interview: https://www.theguardian.com/technology/ng-interactive/2025/dec/02/jared-kaplan-artificial-intelligence-train-itself MIT Study on Jobs/Tasks: https://iceberg.mit.edu/report.pdf vs https://www.cnbc.com/2025/11/26/mit-study-finds-ai-can-already-replace-11point7percent-of-us-workforce.html Amodei on Scaling: https://www.youtube.com/watch?v=FEj7wAjwQIk Claude Soul Document: https://www.lesswrong.com/posts/vpNG99GhbBoLov9og/claude-4-5-opus-soul-document Capabilities Original Stance: https://www.anthropic.com/news/core-views-on-ai-safety Ilya Interview: https://www.dwarkesh.com/p/ilya-sutskever-2 Ricursive Intelligence: https://x.com/RicursiveAI/status/1995932204703346946 Economist Worker Usage of GenAI: https://www.economist.com/finance-and-economics/2025/11/26/investors-expect-ai-use-to-soar-thats-not-happening#selection-1409.94-1413.42 Mistral v3 Large: https://docs.mistral.ai/models/mistral-large-3-25-12 Compute Slowdown Paper: https://joel-becker.com/images/publications/forecasting_time_horizon_under_compute_slowdown.pdf https://x.com/joel_bkr/status/1993023436541903155 METR Chart: https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ https://www.theinformation.com/articles/openais-350-billion-computing-cost-problem?rc=sy0ihq OpenAI Code Red: https://www.anthropic.com/news/core-views-on-ai-safety Rocket Company: https://www.independent.co.uk/news/world/americas/sam-altman-rocket-elon-musk-spacex-b2878351.html DeepSeek Paper: https://arxiv.org/html/2512.02556v1 DeepSeek Crowdstrike CCP: https://www.crowdstrike.com/en-us/blog/crowdstrike-researchers-identify-hidden-vulnerabilities-ai-coded-software/ https://simple-bench.com/ Patreon Post: https://www.patreon.com/c/aiexplained/posts Robot: https://x.com/jloganolson/status/1985850115379351799 Non-hype Newsletter: https://signaltonoise.beehiiv.com/ Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices

  • Yesterday · 18 min

    Did AI Just Get Commoditized? Gemini 2.5, New DeepSeek V3, and Microsoft vs OpenAI

    Gemini 2.5 is out, on the same day as the new DeepSeek V3 (which should power Deepseek R2). Do both models prove AI is being commoditized? Let’s find out, on this blockbuster day of AI releases. Plus exclusives from the Information, Simple indications, Vista Bench, LM Arena and more… AI Insiders ($9!): https://www.patreon.com/AIExplained Chapters: 00:00 - Introduction 01:15 - Gemini 2.5 Benchmarks 05:46 - Long Context, Simple indication 07:08 - New Deepseek V3 -024 09:11 - Microsoft MAI 11:48 - 90% of code but new Claude jobs ‘World’s most powerful model’: https://x.com/OfficialLoganK/status/1904580368432586975 Gemini 2.5 Release Notes: https://blog.google/technology/google-deepmind/gemini-model-thinking-updates-march-2025/#gemini-2-5-thinking ‘Commoditized’: https://the-decoder.com/microsoft-ceo-satya-nadella-says-ai-models-are-getting-commoditized/ Microsoft Information report: https://www.theinformation.com/articles/microsofts-ai-guru-wants-independence-from-openai-thats-easier-said-than-done?rc=sy0ihq LMarena: https://x.com/lmarena_ai/status/1904581128746656099/photo/1 Free for now: https://x.com/btibor91/status/1904578053537476628 Vista Bench:https://scale.com/leaderboard/visual_language_understanding DeepSeek V3: https://huggingface.co/deepseek-ai/DeepSeek-V3-0324 Claude Plays Pokemon: https://www.twitch.tv/claudeplayspokemon Amodei: 100% Coding: https://www.youtube.com/watch?v=esCSpbDPJik&t=3017s Anthropic Jobs: https://job-boards.greenhouse.io/anthropic/jobs/4020717008 Microsoft Money from Onslaught: https://www.972mag.com/microsoft-azure-openai-israeli-army-cloud/ https://simple-bench.com/ Release Date Comments: https://x.com/zacharynado/status/1904647277861318979 Non-hype Newsletter: https://signaltonoise.beehiiv.com/ Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices

  • Yesterday · 17 min

    Manus AI - The Calm Before the Hypestorm … (vs Deep Research + Grok 3)

    Is Manus AI the memecoin of the AI world, or legit? I’ll compare it to OpenAI’s Deep Research, Operator, Grok 3 DeepSearch and more to find out. I’ll also let you in on some of the secrets of what makes a good hype campaign, the estimated costs of Manus AI, and where it is strong. Other news (yes, Gemini image editing and research hacking, I mean you), will have to wait for a few more hours, as millions enquire about Manus AI. https://app.grayswan.ai/arena AI Insiders ($9!): https://www.patreon.com/AIExplained Patreon Vid: https://www.patreon.com/posts/4-ai-trends-in-123857767 Chapters: 00:00 - Introduction 00:46 - Hype Campaign 02:40 - Single, Public Benchmark 03:12 - What is Manus AI? 04:22 - Test 1 05:12 - Cost and Rate Limits 06:15 - Test 2 vs Deep Research + Grok 3 DeepSearch 08:24 - Test 3 (not AGI) 11:10 - 4 Trends in AI in 2025 11:37 - Hype Works Manus AI: https://manus.im/app Xiao Hong Interview: https://www.chinatalk.media/p/manus-chinas-latest-ai-sensation Gaia Benchmark: https://openreview.net/pdf?id=fibxvahvs3 MIT Report: https://www.technologyreview.com/2025/03/11/1113133/manus-ai-review/ Information Report: https://www.theinformation.com/articles/anthropics-claude-drives-strong-revenue-growth-while-powering-manus-sensation?rc=sy0ihq Hype Examples: https://x.com/Saboo_Shubham_/status/1898425707401031940 https://x.com/EHuanglu/status/1899110687902978373 https://x.com/AJs_AI/status/1898756132384178291 Mistakes: https://x.com/TheXeophon/status/1898737178273829220 Tools and Code: https://x.com/peakji/status/1898994802194346408 https://operator.chatgpt.com/ Non-hype Newsletter: https://signaltonoise.beehiiv.com/ Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices

  • Yesterday · 23 min

    GPT 5.2: OpenAI Strikes Back

    Full GPT-5.2 breakdown - did OpenAI reclaim the crown? A story of tokens, time and cost, plus 9 details you wouldn’t get just from reading the headlines. https://www.youtube.com/@eightythousandhours AI Insiders ($9!): https://www.patreon.com/AIExplained https://lmcouncil.ai Chapters: 00:00 - Introduction 00:55 - Better than Human @ Professional Tasks? 04:42 - Test time Compute 07:05 - Benchmark Selection 09:32 - Simple Results + council comparison 13:01 - Long Context 13:52 - Self-Improvement 15:00 - 10 Years + New Models Release Page: https://openai.com/index/introducing-gpt-5-2/ GPT 5.2 Benchmark Comparison: https://www.reddit.com/r/singularity/comments/1pka1y9/gpt52_all_20_benchmarks_rankings_and_pricing/ https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini_3_table_final_HLE_Tools_on.gif https://lmcouncil.ai/benchmarks Charxiv: https://charxiv.github.io/#leaderboard GDPval: https://arxiv.org/pdf/2510.04374 My vid: https://www.youtube.com/watch?v=oK5LxMaROSA Kilpatrick: https://x.com/OfficialLoganK/status/1999270402712023158/photo/1 Noam Brown: https://x.com/polynoamial/status/1999189845164667132 New Model in New Year: https://www.theinformation.com/articles/openai-developing-garlic-model-counter-googles-recent-gains?rc=sy0ihq 10 Years of OpenAI: https://openai.com/index/ten-years/ GPQA: https://x.com/idavidrein/status/1841265634170278063 ARC-AGI 1-2: https://arcprize.org/arc-agi/2/ Sunday Robotics: https://x.com/tonyzzhao/status/1991204839578300813 Non-hype Newsletter: https://signaltonoise.beehiiv.com/ Podcast: https://aiexplainedopodcast.buzzsprout.com/ https://lmcouncil.ai Learn more about your ad choices. Visit megaphone.fm/adchoices

  • Yesterday · 22 min

    Phi-1: A 'Textbook' Model

    After a conversation with one of the 'Textbooks Are All You Need' authors, I can now bring you insights from the new phi-1 tiny language model. See if you agree with me that it tells us so much more than how to do good coding, it affects AGI timelines by telling us whether data will be a bottleneck. I cover 5 other papers, including WizardCoder, Data Constraints (how more epochs could be used), TinyStories, and more, to give context to the results and end with what I think timelines might be and how public messaging could be targeted. With extracts from Sarah Constantin in Asterisk and Carl Shulman on Dwarkesh Patel, Andrej Karpathy and Jack Clark (co-founder of Anthropic), as well as the Textbooks and TinyStories co-author himself, Ronen Eldan, I hope you get something from this one. And yes, the title of the paper isn't the best. Textbooks Paper: https://arxiv.org/pdf/2306.11644.pdf Karpathy Tweet: https://twitter.com/karpathy/status/1671587087542530049 TinyStories: https://arxiv.org/pdf/2305.07759.pdf GPT 4 Self-Repair: https://arxiv.org/pdf/2306.09896.pdf Yao Fu Tweet on Emergent Self-Repair: https://twitter.com/Francis_YAO_/status/1670618013089820674 WizardCoder: https://arxiv.org/pdf/2306.08568.pdf Evol-Instruct (WizardLM) paper: https://arxiv.org/pdf/2304.12244.pdf Scaling Data Constrained Language Models: https://arxiv.org/pdf/2305.16264.pdf Sarah Constantin, Asterisk Magazine: https://asteriskmag.com/issues/03/the-transistor-cliff Jack Clark Tweet: https://twitter.com/jackclarkSF/status/1673369486869811201 Carl Shulman, Intelligence Explosion, Dwarkesh Patel: https://www.youtube.com/watch?v=_kRg-ZP1vQc LLMs and BDTs, Oxford: https://arxiv.org/ftp/arxiv/papers/2306/2306.13952.pdf HumanEval: https://arxiv.org/pdf/2107.03374v2.pdf Decoder Piece (if anyone wants to know, I think George Hotz is super-naïve on safety): https://the-decoder.com/gpt-4-is-1-76-trillion-parameters-in-size-and-relies-on-30-year-old-technology/#google_vignette https://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices

  • Yesterday · 18 min

    9 of the Best Bing (GPT 4) Prompts

    Everyone knows by now how to prompt ChatGPT, but what about Bing? Take prompt engineering to a whole new level with these 9 game-changing Bing Chat prompts. Did you know you can get interviewed by Bing, time travel, force Bing to improve its output and so much more? These are the best 9 prompts that I could find, after analysing over 200 prompts and trying out dozens of examples personally. You might call them prompt hacks, but I prefer to think of them as simply creative exploration of Bing's capacities. Some prompts inspired by this post: https://github.com/f/awesome-chatgpt-prompts And by https://twitter.com/emollick https://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices

Showing 1–20 of 100 episodes