Skip to content
Artwork for SemiAnalysis Weekly
BusinessExplicit

SemiAnalysis Weekly

Jordan Nanos, Doug O'Laughlin

Everything semiconductors and AI

Covering the spectrum

Play
  • 22 episodes
  • weekly
  • Avg 43 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • S1 · E26
    August 20 · 43 min

    Ep. 026 - PJM's $12B Modeling Mistake Is Hitting Ratepayers Again (Datacenter, Energy) | Robert Boswall, Jordan Nanos

    PJM overpaid $12 billion across two capacity auctions, and the same modeling error is set to repeat in an upcoming emergency auction. Robert Boswall (@RobertBoswall) and Jordan Nanos (@JordanNanos) break down how the largest grid in America, 13 states and 66 million people, inflates demand while constraining supply. The result is scarcity pricing that ratepayers absorb, not the data centers driving the narrative. Of the $63 billion spent across four auctions, Boswell estimates $12 billion was avoidable, split $7 billion and $5 billion across the 2025-26 and 2026-27 auctions. This is the mechanics behind every headline blaming AI for rising power bills. Subscribe for weekly analysis on grid design, capacity markets, and AI energy demand! 0:00 Cold Open1:33 What Is PJM3:13 Capacity Auctions7:39 The $12 Billion12:33 How Plants Get Paid15:15 The Winterization Gap21:17 Data Center Demand26:41 The Emergency Auction30:00 Board Overrules Members40:32 Turning It Around Article: https://open.substack.com/pub/semianalysis/p/12b-of-us-ratepayers-money-wasted

  • S1 · E25
    August 17 · 38 min

    Ep. 25 - DYLAN IS HERE, LIVE! | Dylan Patel & Jordan Nanos

    Dylan joins the pod today, recorded live from our SF office. Jordan and Dylan discuss SemiAnalysis's AI spend, the model that escaped during training, AI rollups eating private equity, new accelerators vs NVIDIA, and whether the pod gives away too much information for free.0:00 Cold Open5:25 SA AI Spend9:29 AI Performance Reviews10:40 AI Rollups13:47 The Agent Moment15:28 The Escaped Model18:11 Model Exponential22:52 Compute and Chips30:18 Open Models32:00 ADHD and AI

  • S1 · E24
    August 9 · 50 min

    Ep. 024 - SpaceX's 10GW Plan Drives $300B ARR by 2027 (Datacenter, Energy) | Reyk Knuhtsen, Jeremie Eliahou Ontiveros, Jordan Nanos

    OpenAI and Anthropic are adding close to $30 billion of ARR per month, and the driver is gross margin expansion, not new compute. Jeremie Eliahou Ontiveros (@JeremieEO) and Reyk Knuhtsen (@robotknower) join Jordan Nanos (@JordanNanos) to trace the $100 million per megawatt per year figure from real inference workloads on GB200 and GB300 up through SemiAnalysis simulation. The Google deal prices GB300 capacity near $14 an hour against a $3 average, a premium justified by a 90 day cancellation clause and the ability to turn on megawatts immediately. From there the conversation moves to SpaceX's 10 gigawatt ambition, the sites and supply chain needed to hit it, the permitting playbook, and why Microsoft becomes the largest offtaker. The bull and bear cases both get airtime, followed by a Hugging Face security scare to close. Subscribe for weekly coverage of tokenomics, datacenter energy, and AI compute economics. CHAPTERS: 0:00 – Intro 0:47 – The $100M Thesis 3:59 – Pricing & Google's Deal 10:38 – Training vs Inference 13:21 – Sites & Supply Chain 20:55 – Permitting Playbook 25:40 – Microsoft's Role 33:31 – Paying For It 41:01 – Bull vs Bear 43:29 – Hugging Face Security Scare & Wrap-Up Referenced: SpaceX 10GW in 2027 – Why It's Real, Will Drive $300B ARR for SpaceX, and Why Microsoft Will Be the Largest Offtaker: https://newsletter.semianalysis.com/p/spacex-10gw-in-2027-why-its-real

  • S1 · E23
    August 7 · 44 min

    Ep. 023 - Everyone Leaves Google, Elon Forecasts 1T ARR, Reflecting On GPT-5, Building Personalized Software (Roundtable) | Jon Y, Doug O'Laughlin, Jordan Nanos

    Jeff Dean and the Gemini leads are leaving Google. Jon Y (@asianometry) makes his FIRST appearance to unpack the exodus with Doug O'Laughlin (@fabknowledge) and Jordan Nanos (@JordanNanos). The crew debates the Demis CEO question, whether Google acquired its innovation or invented it, and how the Bell Labs comparison reads once disruption arrives. First engineers who spent thirty-five years in one job now leaving for passion projects.The episode opens by revisiting the 2024 Dwarkesh clip where Jon questioned whether GPT-5 would be any good. Two years later the verdict is mixed: GPT-5 was a dud, 5.2 was ass, but 5.6 is strong and GPT-6, codenamed Doug, is rumored to write well. Then the Chinese transceiver ban, where the West depends on a supply chain China already owns. 00:00 Intro 01:02 Was GPT-5 Good? 05:00 The Transceiver Ban 08:57 Everyone Leaves Google 12:56 Google's L Culture 19:20 Do Legends Matter? 23:44 The Protestant Church 28:44 Elon's $1T Pull-In 30:55 Roll Your Own Software 38:17 Start With Memory

  • July 29 · 49 min

    Ep. 022 - Market Drawdown, Historic Bubbles, Funding The Buildout, AI Politics (Doug is Back)

    Doug is back this week! Timestamps: 00:00 Market Update 07:02 Comparing to Past Bubbles in Taiwan and Korea 10:08 Memory Prices, LTAs, and Market Cycles 18:11 Future Demand for AI and Model Usage 37:40 Scaling Laws and Supply Constraints 44:45 Financial and Capital Constraints in Tech Expansion 50:57 Geopolitical Risks, Policy Impact, Long Term Outlook

  • S1 · E21
    July 23 · 52 min

    Ep. 021 - The AI Project Trinity: Capital, Offtake, Data Center (Datacenter, Energy) | Dan Nishball, Jordan Nanos, Zane Fong, Kang Wen Cheang

    AI and datacenter CapEx hits $11 trillion cumulatively from 2024 to 2029, and $7.1 trillion of that needs funding, roughly 75% debt financed. Dan Nishball (@dnishball), Zane Fong (linkedin.com/in/zanefongzq), and Kang Wen Cheang (linkedin.com/in/cheangkangwen) sit with Jordan Nanos (@JordanNanos) to break down the AI project Trinity: capital, offtake, and data centers. Today only one deal reliably clears the lending bar, a five-year offtake from an investment-grade hyperscaler. Everything else struggles to finance.The crew explains how NVIDIA backstops are structured, how lenders price GPU loans against them, and what NVIDIA actually does with GPUs it takes back. AI debt financing is on track to become the second-largest US asset-backed market behind the $13 trillion mortgage market. Subscribe for weekly analysis on semiconductors and AI infrastructure. References: https://newsletter.semianalysis.com/p/nvidia-gpu-debt-backstop-unleashes CHAPTERS: 00:00 Intro| 01:23 The $11T Funding Problem 09:12 Startups Can't Get GPUs 11:32 How the Backstop Works 20:26 The Central Bank of AI 21:35 The Bullseye 26:23 CoreWeave Credit Spreads 36:58 APAC Examples 42:41 ClusterMax 48:58 Training vs Inference

  • S1 · E20
    July 18 · 50 min

    Ep. 020 - Anthropic vs OpenAI Usage, Margins, Meta Compute, Future of MSL (Tokenomics) | Crystual Huang, Max Kan, Joey Brookhart, Jordan Nanos

    Coding drives over 70% of lab API revenue, and token austerity policies mostly miss the point. Crystal (@crystalthegg), Max Kan (@maxkan), and Joey Brookhart (@SaasquatchC) break down why blocking teams from Opus saves nothing, while power users at the 99th percentile burn $100k per employee per year. Jordan (@Jordannanos) and the team run break-even math on the Max plans and Anthropic's margins. 00:00 Intro00:53 Token Budgeting: Maxing vs Austerity03:05 Coding Eats the Token Market04:49 The ROI Question06:42 Subscriptions vs API Pricing09:08 Break-Even Math on the Max Plans10:42 Anthropic's Profit Margins12:07 Consumer vs Enterprise Mix16:04 The Two-Horse Race18:45 Codex App vs CLI20:41 Who Comes in Third?23:14 Clawbacks and the SpaceX Playbook26:59 Meta's NeoCloud Backstop29:58 Token-as-a-Service Market Forecast32:03 Hyperscalers vs Inference Startups38:50 MSL and the RL Scaling Law41:58 How to Build a Five-Figure RL Task47:10 Vibe Checks48:13 The $400B Anthropic BetReferenced:TokenBudgeting: Our Conversations with Enterprises on Token Spend: https://newsletter.semianalysis.com/p/tokenbudgeting-our-conversationsMeta Compute: Everyone Wants To Be A Neocloud: https://newsletter.semianalysis.com/p/meta-compute-everyone-wants-to-beAnthropic 3Q26 Profit Over $1B: The Anthropic IPO Financials Sneak Peak: https://newsletter.semianalysis.com/p/anthropic-3q26-profit-over-1b-theThe Future of Meta Superintelligence: A 1 Year Progress Update: https://newsletter.semianalysis.com/p/the-future-of-meta-superintelligenceComplete Launch Kit

  • July 18 · 31 min

    [Emergency Episode] Moonshot’s Kimi K3 has Arrived! China has a Frontier Model

    A year ago, the big three was OpenAI, Anthropic, and Google. Things have changed.Moonshot's Kimi K3 sits above Gemini on every composite benchmark, and it's open source in 10 days.New episode: what K3 reveals about frontier margins, model sizes, and who's actually still in the game. 00:00 Intro 00:11 Is Kimi K3 the Third Best Model? 04:04 Why Delay the Weights? 05:30 2.8T Parameters and Serving Constraints 06:48 Frontier Margins and the 3x Price Hike 11:10 New Architecture, What Comes Next 14:09 Will Open Source Catch Closed? 19:51 Built for Chinese Accelerators 22:57 The Harness Is the Product 28:49 We're Still Early

  • S1 · E19
    July 16 · 36 min

    Ep. 019 - Inside the STEEL Lab: From Package to Transistor (Teardown Lab) | Afzal Ahmad, Andrew Wagner, Jordan Nanos

    SMIC's N+3 node shrank the M0 layer over 15% and cut SRAM area 10 to 20%, all without EUV. SemiAnalysis built a lab just for that; Andrew Wagner and Afzal Ahmad walk Jordan Nanos (@JordanNanos) through the STEEL teardown of Huawei's Kirin 9030, from package to transistor. They explain die shots, FIB and TEM cross sections, NPU discovery, cell height, and standard cell libraries. Then backside power, GAA, and where SMIC goes next. 00:00 Intro: The STEEL Teardown Lab01:02 What Is a Teardown?02:17 Who Uses Teardown Data03:22 SMIC N+3 and the Kirin 903005:02 Inside the Lab: Sourcing to Silicon09:10 Die Shots Explained12:35 The NPU Discovery15:42 Scaling Without EUV17:57 FIB, SEM, and TEM Cross Sections20:59 Cell Height and Transistor Shrink23:56 Standard Cell Libraries26:22 Export Bans and Huawei's Response27:40 What's Next: Backside Power and GAA32:43 Data Center GPUs and Logic Folding34:28 Closing Thoughts Read More: https://newsletter.semianalysis.com/p/steel-smic-n3-teardown

  • S1 · E18
    July 9 · 50 min

    Ep. 018 - Stop Saying Half of 2026 US Datacenter Capacity Is Canceled (Datacenter, Energy) | Jeremie Eliahou Ontiveros, Reyk Knuhtsen, Ellie Holbrook, Jordan Nanos

    Bloomberg said half of 2026 US data center capacity is delayed. The SemiAnalysis Data Center, Energy, and Industrials team pulled the underlying report and found a broken denominator. Amazon alone built 4GW in 2025 and is adding 5GW plus in 2026. CoreWeave adds a gigawatt, all under construction. Jeremie Eliahou Ontiveros (@JeremieEO), Reyk Knuhtsen (@robotknower), and Ellie Holbrook join Jordan Nanos (@JordanNanos) as they walk through why the number is wrong and what the real forecast shows. The team covers behind the meter power generation reaching 40GW by 2028, Oracle's New Mexico problem, and the gas turbine supply chain that is running toward peak. They break down the three types of data center delays, how OEMs are responding, and where solar, batteries, and nuclear fit. 00:00 Intro 00:51 The "half of capacity is canceled" myth 04:43 Why early stage projects get canceled 08:06 Three types of data center delays 08:40 Oracle's New Mexico problem 13:47 Behind the meter: 40GW by 2028 18:50 How OEMs are responding 20:24 Signed deal to powered GPUs 23:44 Hyperscaler market share 27:49 How SemiAnalysis tracks data centers 31:34 What behind the meter means 33:31 The gas turbine supply chain 38:22 Peak turbine 42:54 Solar, batteries, and nuclear 46:13 Final thoughts 48:20 Favorite projects Referenced: Stop Saying Half of 2026 US Datacenter Capacity Is Canceled: https://newsletter.semianalysis.com/p/stop-saying-half-of-2026-us-datacenter US Grid Constraints: Towards 40GW+ of Behind-The-Meter Datacenter by 2028?: https://newsletter.semianalysis.com/p/us-grid-constraints-towards-40gw

  • S1 · E17
    July 1 · 34 min

    Ep. 017 - DeepSeek V4 and Huawei Ascend NPU Performance (InferenceX) | Kimbo Chen, Cam Quilici, Bryan Shan, Jordan Nanos

    DeepSeek V4 claims a 100x KVcache reduction versus a standard MoE model, hitting 1M context length through compressed sparse attention and heavily compressed attention. Kimbo (@Kimbochen), Cam Quilici (@noslawextratost), Bryan Shan join Jordan Nanos (@JordanNanos) to break down what changed from V3, why the new MHC dimension tripped up NVIDIA on day zero, and how Mega MoE fuses compute and communication into a single kernel. The vLLM versus SGLang NDA access gap and the Huawei day zero optimization guide circulating on Twitter.The crew walks through the InferenceX article on going from day zero to day 43 support and what that grind actually looks like across different hardware. Subscribe for weekly semiconductor and AI infrastructure analysis from the SemiAnalysis team. Referenced:DeepSeekV4 1.6T Day 0 to Day 43 Performance Over Time - GB300 NVL72, Huawei, MI355X, B200: https://newsletter.semianalysis.com/p/deepseekv4-16t-day-0-to-day-43-performanceChapters: (00:00) DeepSeek V4 vs V3 changes(01:00) Sparse attention and KV cache reduction(03:04) Day zero runtime support challenges(05:34) What Mega MoE actually is(08:38) Downsides of fusing kernels(10:25) MegaKernel benchmark claims(12:59) AMD FP4 optimization gains(15:14) Compounding step by step improvements(17:59) vLLM versus SGLang competition(19:34) Open source vs vendor libraries

  • S1 · E16
    June 23 · 48 min

    Ep. 016 - What Unitree's Evolution Means For Robotics (Robotics) | Jordan Nanos, Reyk Knuhtsen, Niko Ciminelli

    Unitree is going public, boasting 67% gross margins on its humanoid robots. Jordan Nanos (@JordanNanos), Reyk Knuhtsen (@robotknower), and Niko Ciminelli discuss how the Chinese company achieves this through aggressive pricing, rapid iteration, and a focus on "good enough" hardware for the research and hobbyist markets. This strategy allows Unitree to dominate, much like DJI and BYD did in their respective fields.The discussion explores the reality of humanoid robot deployment versus market hype. While industrial applications are in their "baby days," Unitree's approach leverages economies of scale to create a significant moat, challenging US competitors to match their production volume and cost efficiency. The team analyzes if the US can truly compete with China's manufacturing might in the emerging robotics sector.Join SemiAnalysis Weekly for expert insights into the semiconductor, AI infrastructure, and robotics markets. Subscribe for deep dives into AI supply chain, chip economics, and market analysis.Article: https://newsletter.semianalysis.com/p/chinas-unitree-will-dominate-globalTimestamps:00:00 — Intro, deployment reality vs. hype02:39 — Unitree Business: why go public, margins, pricing05:48 — Parallels to DJI and BYD, economies of scale as China's moat16:52 — When is Claude Code moment for robotics18:38 — Real use cases and future demand shocks26:39 — Shenzhen and the humanoid BOM35:56 — The bear case: who actually buys them?42:17 — Can the US compete?

  • S1 · E15
    June 15 · 47 min

    Ep. 015 - DG Matrix Explains 800V DC vs Legacy AC Distribution (Datacenter, Energy) | Jordan Nanos, Jeremie Eliahou Ontiveros, Nicolas Bontigui, Haroon Inam

    NVIDIA's next architecture will demand 800V DC and datacenters can't wait to build the infrastructure. Haroon Inam, CEO and Co-Founder of DG Matrix, joins Jordan Nanos (@JordanNanos), Jeremie Eliahou Ontiveros (@JeremieEO), and Nicolas Bontigui to explain why megawatt GPU racks make 800V DC a necessity, not an option. Haroon details his journey from uncool power electronics to enabling superhuman intelligence infrastructure. Full Article Link: https://newsletter.semianalysis.com/p/inside-the-800vdc-revolution-part Chapters: 00:00 Haroon Inam's Background and Power Electronics Journey 00:29 Introduction to 800V DC Architecture and Its Significance 02:19 Why 800V? Economics, Semiconductor Ratings, and EV Influence 04:47 Physics and Economics Driving the 800V Revolution 08:27 DG Matrix's Multi-Port SST and Its Value Proposition 11:33 Design Challenges and Innovations in Multi-Port SSTs 13:53 Current State and Adoption of 800V DC in Data Centers 17:51 Future Data Center Architectures and Power Density Trends 22:08 Adoption Curve and Market Penetration of 800V DC 26:34 Risks, Challenges, and Future Proofing of Data Centers 30:36 Cybersecurity and Power System Resilience 38:04 The Broader Impact of Power Innovations on Society

  • S1 · E14
    June 4 · 23 min

    Ep. 014 - Finding Miscompiles For Fun, Not Profit (AI Infrastructure) | Justin Lebar & Jordan Nanos

    Justin Lebar (jlebar.com) recently spent $10,000 in an afternoon, uncovering critical miscompiles across NVIDIA's PTXAS, LLVM's AMD GPU, and X86 backends. He joins Jordan Nanos (@JordanNanos) to detail his methodology, which combined traditional fuzzing techniques with novel LLM-assisted bug finding. Their discussion highlights the unique challenges of detecting flaws in less-tested ML compilers compared to mature CPU environments.Lebar shares specific high-severity X86 findings, including an atomic operation bug that splits into two non-atomic operations. They explore the comparative efficacy of fuzzing versus LLM agents in identifying these elusive errors. This episode offers critical insights into compiler security and the burgeoning role of AI in automating rigorous code verification for AI infrastructure.FULL ARTICLE 00:00 Introduction and Content Overview00:25 Justin Lebar's Background and Recent Project00:59 Fuzzing Techniques for Compiler Bugs01:56 Motivation Behind the Project02:48 Challenges in Bug Detection in GPU and ML Compilers04:13 Bug Severity and Findings in AMD and x8605:38 Using LLMs to Read and Find Bugs in Code07:56 Impact of New Models and UltraCode Mode12:18 Estimating Time and Effort Without AI Assistance14:22 Limitations of Manual Code Review for Bugs15:03 Optimism About AI in Software Development16:17 Next Steps and Future Projects18:11 Key Takeaways for Developers and Researchers21:48 Call for Community Engagement and Scientific Approach

  • S1 · E13
    June 1 · 43 min

    Ep. 013 - AWS Margins Jump 10% While Azure and GCP Flatline (Tokenomics) | Jordan Nanos, Jeremie Eliahou Ontiveros, Joey Brookhart, Crystal Huang

    AWS operating margins jumped 10 percentage points while Microsoft Azure and Google Cloud stayed flat. The driver: Anthropic's Claude usage routing through Bedrock, Amazon's token-as-a-service platform. Jordan Nanos (@JordanNanos), Jeremie Eliahou Ontiveros (@JeremieEO), Joey Brookhart (@SaasquatchC), and Crystal Huang (@Egg1459) break down why stabilized token margins are fundamentally richer than GPU-as-a-service for hyperscalers. The crew analyzes Anthropic's recent $65B Series H raise, Claude Opus 4.8 release, and SpaceX partnership against the backdrop of 300+ neo clouds fragmenting the traditional cloud moat.The team forecasts how AWS's workload mix advantage creates sustainable returns while competitors struggle with asset-heavy GPU service models. They examine the $22.7T TAM question, earnings-before-training dynamics, and whether the 2026 AI infrastructure beat belongs to silicon vendors or platform integrators. Subscribe for weekly deep dives into semiconductor and AI infrastructure economics.00:00 Intro: Episode 13 and the AWS margins article00:56 What is Bedrock? The three hyperscaler buckets02:33 AWS margins rising while peers lag03:33 Cloud moats collapsing and the neo cloud explosion06:32 Why stabilized token-as-a-service margins are so rich09:54 Amazon's workload mix advantage12:41 Forecasting Anthropic and the 4.8 release16:33 The SpaceX deal and the $65B Series H raise19:30 Bullish or bearish? Demand becoming supply28:55 The $22.7T TAM and does the race even matter31:59 Earnings before training and open-ended TAM36:27 The 2026 beat is basically one company40:22 Who wins long term: silicon, partnerships, integration

  • S1 · E12
    May 15 · 44 min

    Ep. 012 - Cerebras IPO: 1,100 Tokens Per Second vs GPU's 500 (AI Cloud TCO) | Jordan Nanos, Howie , Myron Xie

    Jordan Nanos (@JordanNanos), Howie, and Myron Xie break down the economics of Cerebras's IPO on the eve of their public debut, examining their OpenAI and Amazon deals that have shifted the company away from Middle Eastern investor concentration toward frontier AI labs willing to pay exponentially more for speed.The discussion covers Cerebras's radical stitching innovation across a full wafer, creating compute density equivalent to an entire NVL72 rack without off-chip data movement. The hosts analyze whether businesses will accept these premium economics as fast tokens become the new standard for interactive AI applications.Subscribe for weekly deep dives into semiconductor economics and AI infrastructure developments. (00:00) Cerebras IPO Preview (18:58) Need for Speed Fast Tokens (21:19) Wafer Scale Engine Architecture (25:31) Radical Stitching Innovation (31:55) Power Delivery and Cooling (34:31) Bandwidth and IO Limitations (37:12) Scaling Beyond Wafer Size (40:05) Manufacturing and Assembly Bottlenecks (42:54) Data Center Service Model

  • S1 · E11
    May 6 · 45 min

    Ep. 011 - GPT 5.5 vs Claude 4.7: OpenAI's Comeback From the Brink (Tokenomics) | Jordan Nanos, Dylan Patel, Doug O'Laughlin, Max Kan

    OpenAI was in serious trouble at the beginning of this year. Anthropic's Claude Opus 4.5 release had triggered a wave of developers to start using Claude Code, pushing Anthropic's revenue past OpenAI's on a like-for-like basis by April. OpenAI's GPT 5.4 response was such an embarrassment they didn't even compare it to Claude in their model release card. Then came GPT 5.5 - finally back on the frontier, but is it enough to reclaim the crown? Jordan Nanos (@JordanNanos), Dylan Patel (@Dylan522p), Doug O'Laughlin (@FabricatedKnowledge), and Max Kan (@maxkan_) break down the latest AI model wars, from Claude 4.7's coding dominance to DeepSeek's long-delayed v4 release and what it reveals about China's AI capabilities. They analyze token efficiency, benchmark gaming, and why fast mode might be fake news. Subscribe for weekly deep dives into the semiconductor and AI infrastructure powering the future. The Coding Assistant Breakdown AI Value Capture Timestamps: 00:00 OpenAI's Comeback and the Latest AI Model Wars 04:05 The High Cost of AI Models and Fast Mode Effectiveness 08:16 When AI Tokens Become Too Expensive for Tasks 13:11 Why AI Model Quality Degrades and Benchmarks Fail 18:42 Deep Dive into Claude 4.7 Features and Tokenizer Changes 25:29 DeepSeek's Release and China's AI Compute Constraints 28:20 The Future of Context Windows and Agent Orchestration 30:47 The Great Debate: CLI vs. App for AI Interaction 36:33 Debunking AI Fake News and Context Window Limitations 40:51 The AI Race: China, Meta, and the Neo Cloud Vision 43:46 Final Thoughts and Listener Feedback Request

  • S1 · E10
    May 1 · 46 min

    Ep. 010 - How Much Do GPUs Really Cost, and Where Does the Value Go? (AI Cloud TCO) | Jordan Nanos, Dan Nishball, Kang Wen Cheang, Zane Fong

    This episode features Jordan Nanos (@JordanNanos) and Daniel Nishball (@dnishball) breaking down the economics of GPU clusters through real-world data and experience. Joined with Kang Wen Cheang and Zane Fong, the team discussed moving beyond theoretical TCO models as they examine how reliability differences between top-tier and lower-tier providers create significant cost disparities that aren't captured in simple per-GPU pricing. The discussion introduces practical frameworks for measuring goodput and understanding how system failures cascade through entire training jobs.Nanos walks through the mechanics of fault-tolerant frameworks including AWS's Checkpointless Training and explains why a single GPU failure can halt progress across hundreds of nodes. The conversation reveals how hyperscalers and NeoClouds price their services and why paying premium rates for reliable infrastructure often delivers better value than chasing the lowest per-hour costs. Subscribe to SemiAnalysis for in-depth analysis of AI hardware economics and infrastructure trends that impact the entire semiconductor ecosystem.

  • April 19 · 53 min

    Ep. 009 - Using Open Source Data To Drive Investment Decisions (ChipBook) | Chaim Eisenberg, Simi Sherman, Jordan Nanos

    This week the team from ChipBook (formerly Chips & Wafers) joins. Jordan Nanos talks with Chaim Eisenberg and Simi Sherman as they explore how they build the ChipBook with open source data, and how that drives investment decisions. This episode dives deep into the ChipBook itself, revealing how granular, historical data collection provides insights into supply chain dynamics, memory markets, wafer fabrication equipment trends and more. The guests also share compelling examples of how their data-driven approach has generated some viral social media recently.for more: SemiAnalysis.com/chipbook00:00 The Chipbook: Understanding Open Source Data07:56 Granularity in Data: The Key to Investment Insights09:11 Understanding the Semiconductor Supply Chain10:52 Memory Market Insights and Trends14:23 Tracking WFE and Its Impact on Production15:20 The Importance of Early Signals in Investment16:51 Geopolitical Implications on Semiconductor Supply20:24 The Impact of Tariffs and Regulations22:48 Granular Tracking for Investment Decisions27:07 Data-Driven Insights and Investment Strategies29:13 The Structure of the Chipbook33:02 Collaboration and Integration at Semi Analysis37:02 Geopolitical Analysis and Its Impact on TSMC43:48 Helium Supply Chain and Its Importance

  • S1 · E8
    April 10 · 34 min

    Ep. 008 Claude Code Psychosis: How SemiAnalysis Is Token Mogging Meta | Dan Nishball, Sam Harshe, Jordan Nanos

    Jordan Nanos (@jordannanos) Daniel Nishball (@dnishball) and Sam Harshe (@sharshe02) break down how SemiAnalysis is deploying Claude Code agents across its Singapore office at a scale that outpaces Meta on a per-employee basis. They cover the practical workflow changes, the trust and reliability questions that come with AI-generated analysis, and what it actually takes to build an agent swarm that does useful work. The conversation also gets into cybersecurity risks and where AI model development is headed next.00:00 - Introduction and Team Dynamics09:10 - The Evolution of Agent Utilization14:49 - Conference Insights and Research Efficiency15:58 - AI's Role in Learning and Analysis21:13 - Trust and Reliability in AI Outputs27:04 - Market Impact and Adoption of AI Tools32:14 - Cybersecurity and AI: Opportunities and Challenges39:32 - Future of AI Models and User Experience

Showing 1–20 of 22 episodes