Skip to content
Artwork for AI to ROI
BusinessManagementTechnology

AI to ROI

Ray Rike

AI to ROI is a podcast that shares how enterprises translate AI investments into measurable business value. Hosted by Ray Rike, Founder and CEO of Benchmarkit, the show features senior enterprise leaders and AI software executives who share how AI initiatives move from pilots to production, and how ROI is actually measured and achieved. In addition, each week, we publish a bonus episode with AI to ROI Newsletter co-author, Peter Buchanan to discuss the Big Story of the Week.

The AI to ROI podcast is the evolution of the original "Metrics to Measure Up" podcast.

Play
  • 22 episodes
  • Avg 33 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • Wednesday · 31 min

    AI Cost Governance to AI Cost Optimization with Sundeep Goel, CEO, Mavvrik

    Ray Rike sits down with Sundeep Goel, co-founder and CEO of Mavvrik, for a deep dive into the second annual AI cost governance research report from Maverick and Benchmarkit. With 30+ years in enterprise software, Sundeep brings a practitioner's view of why AI spend is proving far harder to forecast and govern than cloud spend ever was, and what finance and technology leaders need to do about it. Forecast accuracy is getting worse, not better. Only 11% of companies can forecast AI spend within plus or minus 10% in 2026, down from 15% the prior year, even as aggregate AI spend has climbed sharply. Unexpected AI costs are already changing business decisions. 49% of companies with AI-powered digital products have had to modify pricing, and 25% have stalled or shelved an AI project due to cost overruns. Traditional per-seat and percentage-of-spend pricing models are breaking down. Variable, high AI cost of goods sold is pushing companies toward consumption-based and outcome-based pricing. Cost tracking and cost allocation are two different problems. 98% of companies say they track AI costs at an aggregate level, but only about 5% can allocate that spend down to the customer, application, or agent level where it actually drives decisions. Cost governance is the precursor, cost optimization is the destination. Drawing a parallel to the 30%+ waste long documented in cloud spend, Sundeep argues a similar magnitude of waste exists in AI spend today, and that dynamic model selection, prompt caching, and rate optimization are the next frontier. ROI has a solvable half and a harder half. Cost is measurable today; value remains subjective and requires a company-specific framework to quantify. Ray and Sundeep close with three rapid-fire questions on who should own AI ROI measurement, the key variables for driving it, and advice for early-career professionals navigating an AI-disrupted job market. If you're finding value in the insights our guests share, subscribe to the AI to ROI podcast on your favorite platform and leave a five-star rating. Have a guest suggestion? Reach out to Ray Rike on LinkedIn. See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

  • Tuesday · 26 min

    Microsoft is Running an AI Marathon

    For 24 months, the market narrative held that Microsoft was losing the AI race to the lab it had funded. Then came the July 30th earnings call, and the stock added roughly $450 billion in market capitalization the next day, the largest single-day gain by any public company on any exchange in history. In this week's Big Story, Ray Rike and Peter Buchanan work through why the reassessment happened and what it says about where enterprise AI value is actually accruing. The thesis is not that Microsoft builds the best model. It does not. The thesis is that Microsoft has figured out how to monetize the gap between model quality and market value at the exact moment the novelty premium on frontier AI is wearing off and enterprise buyers are starting to ask price-performance questions instead of capability questions. What the episode covers: The numbers behind the trade. Azure grew 43% and crossed $100 billion in annual revenue for the first time. The AI-specific business, bundling AI consumption with Copilot and GitHub, hit a $37 billion annual run rate, up 123% year over year. Total quarterly revenue reached roughly $90 billion, up 18%, with net income up 31% to nearly $36 billion. The concentration question inside that number. Roughly $24 billion of the $37 billion AI run rate is hosted by OpenAI, about 7% of the company's total revenue. Peter frames the two ways to read it, as systemic customer concentration risk or as evidence that hundreds of thousands of workloads now run through Azure, and explains why he leans toward the second. The toll booth position. Microsoft is the only cloud provider that hosts the OpenAI, Anthropic, and Mistral models alongside its own MAI family on the same platform. As model orchestration becomes a real buying criterion, that means an Azure customer never has to leave the platform to route traffic across model families, and gets one contract and one bill for all of it. The equity stakes that make Microsoft indifferent to who wins. A 27% position in OpenAI now valued around $220 to $240 billion, plus an Anthropic stake that produced a $3.2 billion unrealized gain last quarter and added 33% to earnings per share. Seven years of OpenAI partnership history and how it changed. The original $1 billion investment in July 2019 in exchange for Azure exclusivity, the $13 billion follow-on in 2023 with exclusive commercial API rights, the mid-2025 strain when OpenAI signed separately with Oracle for compute, and the public benefit corporation restructuring that converted profit sharing rights into equity while Microsoft gave up hosting exclusivity, retaining a first mover window on new releases and non-exclusive licensing rights running through 2032. The MAI model family and Satya Nadella's frontier diffusion strategy. Seven models announced at Build in June with Microsoft owning the IP. MAI Thinking One is a mixture-of-experts reasoning model with about a trillion parameters, only 35 billion of which are active at any time, making it cheap to run at scale. MAI Code One Flash, a 5 billion parameter coding model, became the default in GitHub Copilot. Image, voice, transcription, and cybersecurity models fill out the set. What that routing actually saves. Microsoft reported an 84% reduction in GPU costs for image generation in PowerPoint and an 89% reduction in GPU costs for voice processing in the Dynamics 365 contact center by keeping high-volume, repetitive traffic on models it controls end-to-end rather than sending it to a frontier lab. Two production examples. Dragon Copilot, the clinical assistant built into Dragon Medical One, and DAX Copilot that drafts clinical notes from a patient visit, and Project Perception, an agentic cybersecurity platform running continuous vulnerability scanning and patching through coordinated agents rather than generating alerts for human triage. Both are high-volume, low-novelty work, which is exactly the profile these models were built for. Where the benchmark story does not match the marketing. On SWE Bench Pro, MAI Thinking scores around 53% against roughly 69% for the current Anthropic flagship on independent leaderboards. Microsoft's own AI chief has acknowledged a lead measured in months rather than weeks. Ray's point for buyers: MAI benchmark figures come from Microsoft's internal technical reports, and should be treated as vendor claims until independently validated. The distribution advantage that does not show up in any benchmark table. 70% of the Fortune 500 using Microsoft 365 Copilot, more than half a million companies in the Microsoft AI Cloud Partner Program, and 30 million paying Copilot seats, up from 20 million the prior quarter, plus more than 20 million developers and over 140,000 organizations on GitHub Copilot. The sovereignty play in Europe. A multi-billion dollar Mistral partnership funding European data center build out in exchange for prioritized capacity and distribution rights, which sidesteps the capital intensity of building EU capacity from scratch and answers the sovereignty question directly in a region where OpenAI and Anthropic are less established. Frontier Company, and the forward-deployed engineer story revisited. A $2.5 billion subsidiary staffed largely from existing engineers rather than new hires, with more than 6,000 forward-deployed engineers implementing inside client environments. Ray connects it back to the FDE economics the show covered previously, and Peter explains why pulling the large consultancies into the tent matters more than the headcount itself. What CFOs and GTM leaders should take away: Route by workload, not by vendor loyalty. Frontier quality for the problems that need it, cheaper controlled models for Excel formulas, transcription, and support tickets. The difference in margin between those two decisions is the whole story. Treat self-reported model benchmarks as vendor claims. Until independent evaluation catches up, keep MAI models off your most technically demanding workloads. Don't skip the security diligence. Microsoft shipped real vulnerabilities this year, including a Copilot flaw that could have exposed customer files. The new Microsoft and NVIDIA open weights security association, which grew from 20 founding members to 125, is a good leading indicator but not a substitute for your own testing. Recognize what a vendor with equity in multiple labs is actually incented to do. When the provider profits regardless of which model you select, the steering pressure toward a specific model family drops. Weigh the balance sheet underneath the AI investment. Microsoft is funding all of this from a core business that grew 18% and saw net income up 31%, a cushion the pure-play labs do not have. Ray's read: never bet against Microsoft, and the combination of a large profitable core business plus existing enterprise access is a genuine advantage. Peter's read: this is a price-performance play rather than a capability play, and whether it is a durable moat or a stopgap until the cost curve moves again remains an open question. For the full analysis behind this week's big story, subscribe to the AI to ROI newsletter at ai2roi.substack.com. See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

  • August 18 · 32 min

    Can U.S. Frontier AI Labs Survive a Price War with Open-Weight Models?

    When Kimi K3 landed, the headlines said Chinese open-weight models had caught the American labs, and the AI trade sold off from chipmakers to the labs themselves. Nobody stopped to ask whether cheaper and more profitable mean the same thing. In this week's AI to ROI Big Story, Ray Rike and Peter Buchanan run the actual business math on both sides of the fight and find that neither side has the balance sheet to fight a sustained price war. The setup is stark. Anthropic is projecting its first-ever quarterly operating profit of roughly $559 million in Q2, with an annualized revenue run rate near $47 billion, up from $9 billion at the end of last year. That profit disappears the moment it tries to match open-weight pricing. Gross margin is running around 40%, roughly 10 points below the internal forecast, against more than $350 billion in data center commitments coming due over the next three to five years. OpenAI's picture is even thinner: roughly $30 billion in ARR, a projected $14 billion operating loss this year, data center commitments approaching $1 trillion, and an advertising business off to a slow start that needs to reach $100 billion by the end of the decade to close the gap. What the episode covers: Why the Chinese open weight labs are not the subsidized price killers the coverage assumed, with Z.ai's gross margin falling from 41% to roughly 15%, DeepSeek near break-even on about $500 million of revenue and already back in market after a $7 billion round, and Moonshot and MiniMax raising at rising valuations rather than running toward profit The open weight versus open source distinction that changes the entire economic model, since every new customer requires more chips, power, and data center capacity, and Z.ai's own numbers show roughly 49 to 50% gross margin on customer hosted deployments versus about 19% when they host and serve via API The price war math itself: frontier models cost roughly $6 to $8 per million output tokens to serve, Anthropic's $25 per million on Opus produces about a 70% gross margin, and repricing down to the $4 to $6 range where Meta's Muse Spark sits flips that margin from positive 70% to negative 65% Why DeepSeek cut prices on a low-end model and then, two weeks later, told customers to prepare for substantial increases across the line, particularly on API access Google as the structural outlier, with 83% growth in its cloud and AI segment, Gemini embedded across fifteen products with more than a billion users each, and the Apple Siri deal extending reach toward two billion devices Where Kimi K3 actually fits, including the caveats nobody is pricing in: weights released only last week, no published large-scale production deployments, two to three times the token consumption on complex tasks, and infrastructure requirements around a 72 GPU rack that costs millions to install and millions a year to run The geopolitical wildcard, with Washington weighing sanctions or outright bans on Chinese open-weight models and distillation-related IP exposure still unresolved The metric that resets the argument: cost per completed task Price per token is the easiest unit to measure and the wrong one to buy on. Ray and Peter walk through a frontier lab evaluation that assumed a fully loaded remediation cost of $17 per failed attempt, roughly 10 minutes of a human operator's time. Claude Opus completed the task about 90% of the time at roughly $2.56. Meta's Muse Spark, priced at a quarter of Opus on tokens, succeeded 75% of the time and landed above $5.80 per completed task. The list price was 75% lower, and the delivered cost was more than double. A fifteen-point reliability gap did all the work. The formula Ray offers turns a squishy quality debate into something a CFO can actually evaluate: cost per attempt, plus failure probability times fully loaded remediation cost, divided by success rate. What CFOs and GTM leaders should take away: Build cost per completed task into vendor evaluation and make vendors compete on that number rather than on a token price list Segment AI workloads by what a failure actually costs, since a wrong answer in a regulated process or a customer support interaction carries a very different price than the model delta suggests, and standardizing on one model across both workload types to save on price is the common mistake Treat orchestration as a cost lever, not just the model choice, since Cursor's internal testing found coordination across multiple models delivered comparable code quality at a fraction of the cost of a single large model Run your own evaluations instead of trusting public leaderboards, since LMSYS Chatbot Arena measures human preference rather than task completion and says nothing about your workload Watch the switching trap, because a cheap model that attracts heavy traffic today can reprice two or three times higher in a quarter once the vendor needs margin Do not overreact to the open weight scare by standardizing on the cheapest option, since compute, power, and people costs are accelerating, and vendor durability still belongs in the evaluation Ray's read: on real-world performance and total cost of ownership, the closed-weight labs still hold the advantage, their prices keep coming down, and they have every incentive to avoid a price war. Peter's read: open weight economics favor whoever hosts the model more than whoever built it, which makes the hyperscalers the quiet winners regardless of how this resolves. For the full analysis behind this week's big story, subscribe to the AI to ROI newsletter at: ai2roi.substack.com See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

  • August 12 · 35 min

    Web Presence Intelligence with Stephan Bajaio, Co-Founder & CEO, VibeLogic

    On this week's AI to ROI podcast, Ray Rike is joined by Stephan Bajaio, co-founder and CEO of VibeLogic and a former co-founder of Conductor, the enterprise SEO platform he helped build over 14 years, culminating in a WeWork acquisition, a management buyback, and a valuation approaching half a billion dollars. With 25 years spanning e-commerce, Yahoo, Time Inc., and a stint as a software CMO, Stephan brings an operator's view of what has actually changed in the buyer journey and what has not. The conversation centers on a hard number. When Stephan asked the marketing team at a $3 billion company what share of their web traffic was being driven by LLMs, the estimates ranged from 20 percent to 50 percent. The measured answer was 1 percent, with a high conversion rate on that small base. That gap between perceived and measured contribution is the core problem for any executive being asked to fund an AEO, GEO, or AIO program this year. Stephan introduces Web Presence Intelligence, a supply-and-demand framing designed for executive conversations rather than channel specialists. Demand is what goes into the search bar or the prompt. Supply is everything that comes back, including publishers, Reddit threads, affiliates, partners, and competitors. The strategic question becomes whether you are influencing where your buyer's opinion is formed, before you decide where to place the bets. Ray pushes back directly, arguing that understanding where models source citations should now outrank owned web properties. Stephan holds his position, and the exchange gets to the heart of the allocation decision facing CMOs and CFOs. In this episode: Web Presence Intelligence defined, and why supply and demand is the right executive language for a channel conversation The four P's of owned content: point, paragraph, page, path, and the audit finding that most sites never actually name the problem they solve in the customer's own words Why owned assets matter more when the interpretation layer keeps changing, illustrated by a professional who lost a decade of LinkedIn equity overnight Attribution reframed as a consequence of measurement rather than a measurement itself, and how to work backward from the sale to the second and third best proxies The baseline requirement, or what Stephan calls the before photo, and why AI deployed against a process you do not already understand produces outcomes you cannot judge Control groups in practice, including a market share test on terminology that showed one enterprise software company exactly what its brand guidelines were costing it in visibility Why the gold rush money went to the people selling pans, and where the actual near-term return sits: auditing your own sales calls, renewal conversations, and customer logs From funnel to hourglass, and why channel ownership models are breaking down as generalists get more capable with AI Augment, automate, or build net new: three different projects that require three different budgeting methods and three different attribution models The operator takeaway: control what you can control, establish a baseline before you fund the initiative, and structure AI marketing investments so the return can be defended against assets you own rather than placements you do not own. See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

  • August 12 · 34 min

    Will the AI Data Center Backlash Really Make a Difference?

    For three years, the AI infrastructure story has been about chips, power, and capital. In 2026, a fourth variable arrived that the hyperscalers were not prepared for: organized, well-funded, and increasingly successful local opposition to data center development. Ray and Peter walk through the numbers behind the fight. Data Center Watch tracking shows at least $130 billion in US projects blocked or delayed in the first quarter of this year alone, with Carbon Direct putting the cumulative figure at $170 billion since 2024. Set against roughly $725 billion in projected 2026 capex across Alphabet, Amazon, Meta, and Microsoft, the question for operators is whether this is a genuine constraint or noise around the edges of an inevitable buildout. The bigger shift is who carries the risk. With digital infrastructure funds raising $157.6 billion in 2025 according to PitchBook, and private capital taking positions like Blue Owl's 80 percent interest in Meta's $27 billion Hyperion campus, a permitting fight in a single county is now underwriting risk for pension funds and insurers nationwide. Delay no longer just moves a launch date. It extends the payback period and compresses return on invested capital. In this episode: The scale of the buildout: 12 gigawatts of national compute capacity today, with the industry targeting 60 gigawatts by the end of the decade Who is writing the checks: OpenAI's $1.1 trillion in total infrastructure commitments, Anthropic's roughly $350 billion in compute commitments, and the rise of third-party capital as the load-bearing wall Why projects slip: 75 percent of capacity under construction already pre-leased, 1.6 percent vacancy, five-year grid interconnection backlogs, and 3-5 year transformer lead times What is driving residents into council meetings: $29.4 billion in added PJM customer costs, utility bills up as much as 267 percent in some markets, Google's water consumption up 34 percent year over year, and a Gallup finding that 71 percent of Americans oppose a data center in their own community, a higher share than opposes living near a nuclear plant How communities win: 833 active opposition groups across 49 states, more than 300 municipal bans and moratoriums since 2023, and Monterey Park's permanent ban approved with 88.34 percent of the vote Fifty different rulebooks: New York's statewide permitting pause, Ohio's 85 percent capacity payment requirement, Virginia's energy consumption tax, and West Virginia moving in the opposite direction The three playbooks that get projects built: Meta's community investment model in Louisiana, the legal route in Michigan, and Microsoft's quieter university and water reuse partnership in Wisconsin The operator takeaway: The buildout is getting slower and more negotiated, not stopped. For anyone modeling enterprise AI unit economics, the bottleneck has moved from chips to concrete, copper, and zoning boards. Cost models built on an assumption of falling compute prices need a second look this quarter. See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

  • August 4 · 36 min

    Forward-Deployed Engineers (FDEs) - AI's New ROI Battleground

    AI spending continues to accelerate, but the ROI story has not kept pace. This week, Ray Rike and Peter Buchanan dig into the proposed fix that has the market talking: turning frontier AI labs into professional services firms through the forward-deployed engineer (FDE) model. The role is not new. Palantir built the FDE function in 2005 to embed technical teams directly within customer environments, and by 2016, it had more FDEs than software engineers. What is new is the capital. Five ventures from the major AI labs and top three hyperscalers have committed more than $10 billion, betting that the bottleneck to enterprise AI value is not the model; it is getting that model wired into a customer's data, processes, workflows, and compliance requirements. Ray and Peter connect this moment back to the ERP era, when SAP and Oracle needed four to five dollars of services for every dollar of software, and explain why the same people, process, and services reality is playing out again with agentic AI. What the episode covers: Why OpenAI's $4 billion DeployCo, with a guaranteed 17.5% investor return, is the most aggressive and most financially puzzling bet of the five How Anthropic (Ode), AWS, Microsoft, and Google each took a different structural path, from joint ventures to capital-light partner ecosystem plays Why the MIT 95% pilot failure stat and McKinsey's finding that two-thirds of organizations have not started scaling AI make this a real problem, not a fringe one How incumbent consulting firms are playing every side at once to protect their AI practices Alex Karp's argument that the AI industry broke its own business model, and the irony of him making it What CFOs and GTM leaders should take away: Ask any FDE partner for two or three production use cases with measurable outcomes before expanding scope Start narrow with one win, not four or five simultaneous projects Price for outcomes up front, before mid-project renegotiation Model the ongoing maintenance cost, since 20 to 40% of the initial investment often goes to keeping it running, and an outside FDE team at $300 to $800 an hour is an expensive long-term maintenance line item Weigh the neutrality trade-off honestly, since lab-backed ventures get you the deepest model roadmap access but also the deepest lock-in Lawyer up on data privacy and IP, because these engagements can feed your proprietary workflows back into the next model release Ray's read: independent consulting firms are the likely long-term winners, with hyperscalers close behind. The AI labs are, in Peter's words, still leaving the Shire without their full posse ready to go. Listen to the full breakdown, and for the deeper analysis behind each big story, subscribe to the AI to ROI newsletter at ai2roi.substack.com. See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

  • July 30 · 36 min

    Chinese Open Weight Models Overtake US Frontier AI: What Every Enterprise Executive Needs to Know

    Twelve months ago, US frontier models controlled roughly 70 percent of AI traffic. Today, Chinese open-weight providers, led by DeepSeek, Z.ai, Moonshot, and Minimax, account for 45 to 61 percent of top-tier model traffic on OpenRouter, with DeepSeek alone processing more tokens than Google, Anthropic, or OpenAI individually. In this Big Story edition, Ray Rike and Peter Buchanan unpack how this shift happened, how US labs and regulators are responding, and three scenarios for how the closed versus open weight competition plays out for enterprise AI buyers. Key topics discussed: The pricing collapse driving enterprise migration. DeepSeek made a 75 percent price cut permanent in May, bringing its V4 Pro model to a fraction of a cent per million tokens versus $2.50 per million for GPT 5. Minimax delivers GPT 5.5 class coding performance at 5 to 10 percent of the cost. This is why Uber, Microsoft, and Walmart are now implementing formal usage governance on frontier models rather than treating cost control as temporary. The Mythos and Fable shutdown as a trust event. The 18-day suspension of Anthropic's top models over export control concerns spooked global enterprise buyers who realized mission-critical workloads could be cut off without warning. This single event accelerated the adoption of open-weight alternatives and pushed allied governments to invest in sovereign AI capacity. Distillation attacks and the IP leakage problem. Anthropic accused Alibaba's Qwen lab of running a large-scale adversarial distillation campaign, using tens of thousands of accounts and tens of millions of exchanges to extract agentic reasoning capability from Claude. This reframes the security conversation from model safety to unauthorized technology transfer, which is a distinct and arguably bigger risk for any enterprise relying on proprietary model capability as a moat. Cybersecurity parity is closing faster than expected. Multiple Asian labs, including Z.ai's GLM 5.2, Beijing based 360 Security, and Japan's Sakana AI, now claim benchmark performance approaching Anthropic's Mythos model on vulnerability detection and both offensive and defensive cyber tasks, often at significantly lower compute cost. This weakens the safety and capability gap argument that has justified restricting access to frontier models. Real deployments have moved from theory to production. Coinbase cut AI spend in half after migrating to Z.ai and Moonshot's Kimi models, even as token usage grew. Cursor shipped a coding tool built on Kimi, with a Grok-based version reportedly imminent. Andreessen Horowitz estimates 80 percent of its portfolio companies already use open-weight models in production AI products. Three scenarios for how this settles, and why the decision belongs at the board level. The hosts outline bifurcation (premium closed models for regulated use cases, open weight for commodity workloads), export control entrenchment (Washington treats the Fable ban as a template rather than a one-off), and capability convergence (the rationale for unilateral bans erodes as the performance gap closes). All three are already visible simultaneously, which means enterprise AI architecture decisions, including primary and backup model orchestration, are becoming strategic decisions that belong with the CEO and board, not just the technical team. Why does this podcast episode matter for enterprise executives selecting models, especially for agentic AI deployments? The vendor you choose today may not be the vendor you can use tomorrow, for reasons that have nothing to do with model quality. Regulatory risk, geopolitical exposure, and pricing volatility are now first-order variables in model selection, alongside capability and cost. Any agentic AI architecture built on a single model provider carries concentration risk that didn't exist a year ago. Building orchestration flexibility with primary and backup models, and understanding the true economics behind token pricing, is quickly becoming a board-level governance question rather than a procurement detail. For the metrics and framework to instrument and report ROI on your AI investment, get the Big Book of AI Metrics at benchmarkit.ai under Media. See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

  • July 28 · 34 min

    AI is Hot - AI Regulation is a Hot Mess

    On June 12th, the US Department of Commerce ordered Anthropic to suspend global access to Fable 5 and Mythos 5, its most powerful frontier models, with no advance warning and criminal penalties attached. In this week's Big Story episode, Ray Rike and Peter Buchanan walk through the first use of export controls to take a deployed frontier model offline, why it backfired for security rather than strengthening it, and what a coherent federal AI regulatory framework would actually need to look like. The shutdown mechanics. Because Anthropic could not verify user nationality in real time, the directive knocked out access for every user globally, including Anthropic's own non US employees and more than 150 companies across 15 countries running critical infrastructure workloads. Three reaction tracks. Industry, the developer community, and allied governments each responded differently. OpenAI pushed back on talent restrictions while its legal team blocked coordination with Anthropic on antitrust grounds, and a coalition of 150 cybersecurity leaders published an open letter asking not for less regulation but for a transparent, science based process. The regulatory vacuum, in five forces. Ray and Peter unpack the drivers behind what they call a hair on fire crisis: capability jumps outpacing any statutory framework, a state level policy tsunami of over 1,500 AI bills across 45 states, data center community opposition, the unresolved Anthropic and DOD dispute, and a bipartisan Congressional letter questioning why comparable models were treated differently. A six part framework. The episode lays out proposed solutions modeled on existing regulatory precedent: global market access agreements similar to military sales processes, mandatory pre release testing run by NIST modeled on FAA certification, clear federal versus state jurisdictional lines modeled on pharmaceutical regulation, expanded child safety authority, FERC fast tracking for data center grid access, and coordinated environmental standards. What enterprise leaders should do now. Ray closes with practical guidance: review AI vendor contracts for shutdown protection since force majeure clauses were never written with export controls in mind, build fallback infrastructure for critical AI dependent workflows, delay rushing into brand new model releases, and automate tracking of the fast growing state regulatory landscape. Full details are in the AI to ROI newsletter at ai2roi.substack.com. Subscribe, and consider reaching out to your representatives on this one. See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

  • July 22 · 34 min

    Building the 100x Org - The CFO as AI Architect with Dan Zhang, CFO & CBO at ClickUp

    Most companies treat AI as a layer they add on top of how work already gets done. Dan Zhang, Chief Business Officer and CFO at ClickUp, argues that it is exactly backward. In this episode, Ray sits down with Dan to unpack the "100x Org," ClickUp's framework for rebuilding the business around AI rather than sprinkling tools and tokens on top of a human-driven workflow, and why that distinction determines whether an AI initiative shows up as activity or as income statement impact. The conversation covers: Why ClickUp expanded the CFO's charter to own AI transformation end to end, after both a top-down mandate and a bottoms-up experimentation push failed to produce results that made it into production The "jobs to be done" framework Dan uses to separate primary work that actually moves the business from secondary work that just generates busy AI activity, and why most companies have a work redesign problem before they have an AI problem Dan's psychological test for AI ROI (would you pay for it with your own money) and why ARR per headcount is the right metric, but a lagging one that plays out over years, not weeks How ClickUp instrumented daily, not monthly, visibility into AI cost and token consumption, and why over half of enterprise companies in Ray's own research have blown their AI budget by 25% or more Why the real anxiety CFOs have about AI ROI isn't the return, it's confidence in cost management, and how a governance layer turns that anxiety into a guardrail instead of a monthly surprise Dan's rapid-fire advice on who should own AI ROI measurement, the two-to-three variables every CFO needs in place, and how early and mid-career professionals protect their relevance by owning the question, not just the answer If your organization has AI activity but can't yet point to AI impact, or your CFO isn't sure whether AI spend is a competitive advantage or a leak, this episode gives you the operating model and the financial discipline to tell the difference. See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

  • July 21 · 33 min

    AI is a Compensation Scale Expense

    Token prices have fallen 98 percent since GPT-4, but enterprise AI bills are up 320 percent. In this week's Big Story episode, Ray Rike and Peter Buchanan trace where that money is actually coming from, and the answer is not the software budget. Drawing on Gartner, Oxford Economics, Zylo, Challenger Gray and Christmas, and Goldman Sachs data, the two lay out why labor, not IT, is becoming the primary funding source for AI at scale, and why almost no company has the measurement infrastructure to manage it. The price paradox. Per token costs have collapsed, but usage has grown faster than costs have fallen. Ray walks through the math behind average enterprise AI budgets rising from $1.2 million to $7 million in two years. Three cautionary tales. Uber consumed its entire annual Claude Code budget in under four months, Microsoft revoked thousands of Claude Code licenses over cost, and one unnamed enterprise ran up a $500 million bill in a single month. Ray and Peter break down why each was a governance failure rather than a technology failure. Only two budget pools are big enough. The IT and software budget represents just 3 to 4 percent of revenue, while labor represents 25 to 40 percent depending on industry. Ray makes the case that labor is the only pool large enough to absorb the AI spending trajectory Gartner and Oxford Economics are projecting. The attrition lever. Ray and Peter unpack how not backfilling open roles has quietly become the primary way enterprises are funding AI investment, supported by data showing over 113,000 tech layoffs in 2026 with 48 percent explicitly attributed to AI. Revenue per FTE as the tell. Ray shares benchmark data showing SaaS company revenue per employee up 25 to 35 percent over the last twelve quarters, and explains why this metric will be the clearest signal of whether the AI budget transfer is actually working. Five metrics every CFO needs now. The episode closes with a practical starting list: AI spend as a percent of revenue, AI spend per employee, inference spend as a percent of opex, inference cost as a percent of COGS for AI enabled products, and revenue per FTE tracked against labor cost and agent cost as a percent of OPEX. AI to ROI is always looking for guests with real-world examples of measuring AI budget impact. Reach out Ray on LinkedIn (@rayrike). Subscribe to AI to ROI at ai2roi.substack.com for the full June 9th edition, and leave a review wherever you listen See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

  • July 16 · 36 min

    AI-Native Services: The $100B Disruption of Professional Services

    Professional services firms have billed by the hour for over 200 years. That model is now under direct attack from a new category of company: the AI Native Services firm. Ray Rike and Peter Buchanan break down exactly what makes these companies structurally different from traditional professional services firms and AI-augmented incumbents, profile four companies proving the model at scale, and lay out the six critical success factors that will separate the winners from the well-funded failures. Episode Highlights: Defining the category. Drawing on the Emergence Capital AI Native Services Playbook, Ray and Peter establish a clear three-part taxonomy: AI Native Services companies (AI does 80 to 90% of the work, a licensed human reviews and is accountable for the outcome, and the client pays for results), AI Augmented Services companies (humans still do most of the work with AI as a productivity layer), and traditional SaaS tools (the customer's team does the work, and the vendor takes no accountability for the outcome). The distinction matters enormously for enterprise buyers evaluating contracts and liability. How the operating model actually works. The AI Native Services delivery model runs in four phases: client intake and data ingestion, AI-driven execution of the primary service work, licensed human review and approval, and outcome delivery back to the client. Critically, that fourth phase is where the model compounds, because every accepted output becomes training data that makes the system smarter and harder to displace over time. Four companies are proving the model. Ray and Peter profile four AI Native Services companies at different stages of scale, each dominating a regulated vertical wedge. Top AI-Native Service companies covered include: 1) Field Guide is automating audit workflows for nearly half of the top 100 US accounting firms, including KPMG and RSM, and recently raised a $75 million Series C at a $700 million valuation; 2) Even Up has built a proprietary PI AI model trained on hundreds of thousands of personal injury cases, processing 10,000 cases per week for over 2,000 law firm clients, following a $385 million funding round; 3) A-Bridge converts physician-patient conversations into structured clinical notes integrated with Epic and other EMR platforms, serving over 150 enterprise health systems including Kaiser Permanente, Mayo Clinic, and Johns Hopkins, with $100 million in ARR and a $5.3 billion valuation; 4) Harper is a licensed commercial insurance broker, not a software tool, that processes applications across 160+ carriers simultaneously and delivers final coverage in 24 to 48 hours versus the industry standard of five to seven days. Six critical success factors. The hosts lay out what separates durable AI Native Services companies from those that will stall: genuine domain expertise on day one, a proprietary data flywheel that compounds with every case resolved, a clear migration path from labor-based to outcome-based pricing, honest gross margin accounting that properly classifies LLM inference and human labor as cost of goods sold, narrow vertical focus on a specific wedge rather than broad horizontal expansion, and distribution through regulated industry incumbents who provide both credibility and enterprise access. Gross margin as the early warning signal. If gross margins are declining as an AI Native Services company scales, that is a signal the human-in-the-loop is becoming the bottleneck rather than the leverage point. The financial goal is to continuously increase gross margins as AI does more of the work, moving from a 35 to 45% gross margin profile toward 50 to 60% over time. What enterprise buyers should ask. Ray closes with a direct call to action for executive buyers: if an AI vendor is pricing by the seat, by the partial FTE, or by the hour, push hard on how much of that is human supervision of AI versus a truly AI-native delivery model. You are no longer contracting a resource to do the work. You are contracting an organization to deliver the outcome you need. See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

  • July 14 · 33 min

    AI Math is not Adding Up - Where is the ROI?

    Token spend is exploding across the enterprise, but the value it creates remains largely invisible on corporate dashboards. In this week's Big Story episode, Ray Rike and Peter Buchanan unpack why AI investment and ROI visibility are moving in opposite directions, and what enterprises need to do about it. Drawing on Ramp data, Exponential View, Semianalysis, and Ray's recent conversation with Russ Frayden, CEO of Lariden, the two dig into the measurement infrastructure gap that is turning individual AI productivity gains into an unmeasured expense line. The productivity paradox. Individual output is up across engineering, sales, and research functions, but those gains are not translating into company-level financial impact. Ray connects this to Parkinson's Law and explains why more productive workers do not automatically produce more profitable companies. AI dark output. Peter introduces the concept from Semianalysis: real economic value created by AI that never registers on a P&L, using the example of a legal document that drops from $400 to $5 to produce, where the savings disappear while the token expense shows up in plain sight. The cost to compensation shift. Ray walks through why token spend approaching 50 to 100 percent of engineering compensation changes the entire calculus for measurement, contrasted against IT's historical 3.5 to 6 percent share of revenue. Case studies in good and bad. The episode breaks down three real examples: Uber's Claude Code rollout that ran out of budget without measurable output gains, Lowe's cross-functional agent deployment that built proper context tracking, and Petrobras's narrow tax compliance pilot that identified $120 million in savings and is scaling toward $1 billion. A four-stage framework for real measurement. Ray and Peter lay out the progression from cost visibility to utilization, proficiency, and business impact, and explain why almost every enterprise is stuck at stage one. Six tactical takeaways. The episode closes with concrete actions: measure outcomes not activity, design for the middle 70 percent of users rather than power users, give CFOs real budget ownership, start narrow and expand from proof points, redesign decision rights as AI scales, and treat organizational role changes as a planned program rather than a side effect. Subscribe to the AI to ROI newsletter at ai2roi.substack.com for the full breakdown, and leave a review wherever you listen See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

  • July 7 · 34 min

    AI Governance, Ethics, and the Stewardship Framework: A Conversation with Evan Schwartz, Chief Innovation Officer at AMCS Group

    Most companies jumped straight to headcount reduction when AI arrived. AMCS Group went the other direction, starting with governance, ethics, and a question most companies never ask: not what can AI do, but what should it do? In this episode, Ray talks with Evan Schwartz, Chief Innovation Officer at AMCS Group, a platform serving resource-intensive industries across 80 countries, about how that single question reframed their entire AI strategy and produced results measured in multiples rather than percentage points. Topics we discussed include" Why "person plus AI" beats "AI replaces person" from day one Companies that moved quickly on headcount reduction before AI left the lab found themselves rehiring at a higher cost when the technology did not perform in the wild as it did in controlled conditions. AMCS took a different path. Rather than treating headcount reduction as the goal (a finite game with a ceiling of zero), they pursued asymmetric growth: amplifying the capabilities of experienced employees so the business could scale revenue without incurring proportional costs. The result was output multiples, not efficiency percentage points. The governance and ethics framework that drove better AI decisions Operating across 80 countries with GDPR, SOC 1, SOC 2, and a range of regional regulatory requirements, AMCS could not afford to move fast and fix things later. They codified existing governance frameworks (including the EU AI Act and NIST standards) into a use-case design framework that forced a structured question before any deployment: what should this AI do? That question filtered out low-value applications, surfaced the high-impact ones, and created the foundation for what Evan calls the stewardship model. What an AI steward actually does, and why the role is human As AMCS built out orchestrator agents and sub-agents, they needed a clear accountability structure. The steward is always a human. Effective AI stewards share three skills: they communicate tasks clearly to orchestrators, they understand what data context the agent needs to do the job well, and they know what good output looks like even without knowing how the system produced it. That last skill, the ability to look at a result and say "that number is wrong," is what keeps agentic systems on the rails and prevents AI sprawl from becoming unmanageable. Two external agentic AI use cases with hard ROI numbers The dispatch management agent now monitors 700,000+ trucks globally, dynamically reroutes based on real-time events (blocked containers, missed pickups), and automatically notifies customers through their preferred channel, including rescheduling VIP accounts before they can call in a complaint. The result: 17 gallons of diesel saved per truck per month in fuel optimization, plus a $650,000 pull-forward of aged receivables (from 90-day to 30-day collection cycles) in just the first month at one customer. The customer service agent enables CSRs to double or triple their customer-touch volume by having AI handle all post-call documentation, action items, scheduling, and follow-up. That increased coverage cut AMCS's own churn rate from 6% to 3%. How AMCS justifies AI investments internally, and why it starts with board-level metrics AMCS is targeting ISO 42001 compliance (the AI management system standard) by year-end, which requires registering every AI tool, documenting bias risks and mitigations, and tying each use case to measurable outcomes. Evan's framework for approval is straightforward: identify your current baseline, set a target, and trace the expected return all the way to a board-level financial metric, EBITDA, free cash flow, or SG&A. Stopping at "we saved three hours" is what he calls lazy intellectualism. The real question is what those three hours produce when redirected to high-value work. Career advice for the AI era: stop valuing yourself by the output. Evan's message to early-career professionals is direct. If AI can produce the output, the output itself has an approaching-zero value. What has value is the ability to get AI to produce it, to steward the system, to know what good looks like, and to course-correct when it does not. The leaders of the next decade will be those who can direct a digital workforce of agents toward outcomes that matter, not those who were best at producing the deliverables themselves. See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

  • July 7 · 34 min

    The AI Coding Wars, Inflection Point and the Cursor-SpaceX Deal | AI to ROI: The Big Story

    The April 2026 announcement that SpaceX may acquire Cursor for $60 billion, or alternatively pay $10 billion for a compute partnership, stopped the enterprise tech world in its tracks. A four-year-old company founded by four MIT students with $2.7 billion in annualized revenue but nearly $900 million in losses on $700 million in actual revenue. This deal is not primarily a valuation story. It is a signal and a cautionary tale about the economics of the AI coding tool market. In this Big Story edition, Ray Rike and Peter Buchanan break down what is really happening in the AI coding wars, why Cursor ended up at SpaceX's door, and where this market goes from here. Key topics covered in this episode: From copilot to autonomous agent: how the AI coding market structurally shifted. Four years ago, AI coding tools suggested your next line of code. Today, they read entire codebases, plan multi-step tasks, edit files across a project, run tests, and submit pull requests with minimal human direction. Claude Code reached $1 billion in annualized revenue six months after launch, the fastest of any enterprise software product in history, and crossed $2.5 billion by February 2026. Meanwhile, 90% of enterprise developers now use at least one AI coding tool, and nearly half of all GitHub code is AI-generated or AI-assisted. The productivity gains are real but uneven, and the risks are underappreciated. JPMorgan deployed AI coding agents to 40,000 engineers and reported 10 to 20% productivity gains in code creation and conversion, along with a 70% increase in code deployments. But CodeRabbit's research found 1.7 times as many defects in AI-authored pull requests as in human-authored code. Meta's brief "token maxing" leaderboard experiment, designed to spotlight power users, had to be taken down within two weeks after producing high token consumption and limited usable code. Senior developers are shifting toward architecture and review roles while junior developer pipelines are shrinking, even as total software developer job postings are up 5 to 10% year over year. A tour of the seven major players and where the structural tension lives. Ray and Peter profile Anthropic Claude Code, GitHub Copilot, Cursor, OpenAI Codex, Google Gemini Code Assist, Replit Agent, Lovable, and Cognition's Devin across revenue, differentiation, and risk. The common thread: most point-solution coding agents run on Anthropic or OpenAI models, and those same model companies have now launched their own competing coding products. The Oracle database-to-applications parallel is not subtle. Why the Cursor-SpaceX deal happened and what it actually reveals. Cursor had $2.7 billion in annualized revenue, negative 23% gross margins, and was losing money faster than it was growing. Even with a $2 billion funding round in process from Andreessen Horowitz, Thrive Capital, NVIDIA, and Battery Ventures, Cursor's leadership concluded they would need to raise billions more by year-end to fund compute costs. SpaceX's acquisition offer, or the $10 billion partnership payment that Ray reads as a very generous breakup fee, solved that problem while giving XAI a revenue base three times its current size ahead of a $1.75 trillion IPO valuation push. The Chinese open-source threat and three scenarios for where this market goes. Kimi, DeepSeek, and Qwen models are improving rapidly and are significantly cheaper. They are, as Peter puts it, lurkers haunting every company on the list. Ray and Peter then lay out three scenarios: model makers consolidate, and IDE players get marginalized; a durable multi-tool ecosystem persists because different tools serve different workflow stages and buyer profiles; or a compute-native player builds a fully autonomous coding agent that eliminates the need for an IDE entirely. Claude Code already resolves 64.3% of real-world GitHub issues, and full autonomy for defined-scope tasks may be 18 to 36 months away. The AI coding market is projected to reach $49-$50 billion by 2030, at a 38% CAGR. The speed gains are real. So are the defect rates, the governance gaps, the model-dependency risks, and the token-budget surprises landing on CFO desks. If you are a CIO, CFO, engineering leader, or investor trying to make sense of who wins and what it costs, this is the episode to start with. Read the full story at ai2roi.substack.com. See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

  • July 2 · 34 min

    Big Book of AI Metrics

    The AI to ROI team, Ray Rike and Peter Buchanan, mark the official launch of The Big Book of AI Metrics, an 180-page, 81-metric operator's reference guide built to close the gap between AI adoption and AI ROI. Twenty-seven percent of executives say AI has met their ROI expectations, enterprise AI token spend is up 13x since last year, and most companies still can't explain what they got for the investment. Ray and Peter break down why that gap exists and what to do about it. Topics covered: Why adoption, utilization, and outcomes are three different things, and why most companies stop measuring at adoption The five layer causal chain framework: input signals, leading indicators, operational KPIs, financial outcomes, and strategic value Why establishing a baseline before deployment is the single most skipped step, and why skipping it turns results into opinion instead of evidence Four real world case studies: Petrobras ($120M in tax savings), Stocks Insurance (83% reduction in claims processing time), Uber's cautionary token budget blowout, and Klarna's revenue per employee gains Three actions operators should take this week: define the outcome metric, establish a baseline, and build a measurement cadence before and after deployment Key quote: "Adoption still is not ROI. Outcomes are ROI. And outcomes that translate into better financial performance, that's true ROI that a CFO, investor, and a board of directors can get behind." - Ray Rike The Big Book of AI Metrics is organized into 13 functional roles, covering both operating executives investing in AI to improve their functions and B2B software executives whose product economics now depend on token consumption, inference costs, and gross margin impact. Get the Big Book of AI Metrics: https://www.benchmarkit.ai/ai-big-book-of-metrics See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

  • June 9 · 33 min

    Measuring the costs, utilization, proficiency and impact of AI - with Russ Fradin, Founder and CEO, Larridin

    Most enterprises have deployed AI broadly. Far fewer know what they are actually getting from it. Russ Fradin, Co-Founder and CEO of Larridin, has spent his career building measurement infrastructure at inflection points in technology adoption, from early days at ComScore measuring internet advertising to founding Larridin with backing from Andreessen Horowitz and Google's Gradient fund. In this episode, Russ makes the case that AI spend is on a trajectory to become the number-one or number-two driver of enterprise OpEx, and that most organizations still lack the basic visibility needed to manage it. Topics covered: The AI visibility gap: Why AI adoption moved faster than measurement infrastructure, and why enterprises are only now scrambling to answer fundamental questions about what they are spending, where, and by whom Utilization vs. proficiency vs. business impact :Why these three dimensions require separate measurement, and why the 1,800 heavy users at a 30,000-person company are not a success story on their own Token spend as a new category of OpEx risk: How consumption-based pricing turns every employee into a cost endpoint, with real examples of runaway agent spend and blown budgets that no one turned off CFO ownership of AI investment: Why AI spend is the first technology cost category large enough to pull the CFO into governance conversations that historically belonged to the CIO and department heads Change management as the bottleneck: Why the hard work is not experimentation but operationalizing what works, scaling proven behaviors from the top 5% of users to the full organization Career advice for AI-era professionals: Work harder than the room, achieve deep tool mastery, and invest in relationships, the same fundamentals that applied before AI, now with higher stakes for the people who act on them Russ closes with a memorable framing: "Companies have committed to a fitness journey but have not yet bought a scale; Larridin is building that scale." See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

  • June 2 · 30 min

    Leveraging AI to Reduce Churn and Increase NRR - with Dan Harmeson, Co-Founder and Co-CEO at QuadSci

    Most B2B software companies are sitting on one of the most powerful and underutilized data assets in their business: product telemetry. Every click, API call, and feature interaction is a signal. The question is whether your go-to-market organization knows how to read it. In this episode, Ray Rike is joined by Dan Harmeson, co-founder and co-CEO of QuadSci, to explore how machine learning applied to telemetry data is changing how software companies predict churn, protect the base, and accelerate expansion revenue. Key topics covered in this episode: Why telemetry data is the largest untapped GTM asset in B2B software. Dan defines telemetry data, from front-end product analytics events to back-end observability metrics, and explains why these trillions of usage signals are the single biggest data set B2B software companies generate but rarely use to make go-to-market smarter. QuadSci deploys AI locally inside the customer environment so sensitive data never moves to a third party. How QuadSci builds trust before the sale. Rather than asking customers to take predictions on faith, QuadSci runs a retrospective exercise: predicting churn and growth events that already happened, including data the model never trained on. Customers consistently see 90%+ accuracy, which becomes the foundation for acting on forward-looking risk signals. Gross revenue retention is under pressure and the data is clear. Per Benchmarkit's not-yet-published 2026 benchmarking data, GRR has declined four percentage points to 84% as an industry benchmark. For companies above $100M in ARR, roughly 95% of revenue comes from renewals and expansion, which means a two-point GRR drop cannot be offset by new logo acquisition within a 12-month window. Expansion revenue is a precision play, not just a CS motion. Dan walks through how QuadSci identifies Goldilocks-zone consumption patterns, surfaces cross-sell opportunities aligned to actual usage behavior, and helps account teams build nine-to-twelve month consumption forecasts that customers can actually plan around. The result is expansion conversations grounded in data, not intuition. Token consumption is the next frontier. As agentic AI deployments scale, CIOs and CFOs are facing unpredictable inference costs. Dan explains why the same telemetry-based approach that protects software GRR today is directly applicable to governing AI token spend inside Fortune 5,000 enterprises, a market QuadSci is beginning to address. Rapid fire: ROI measurement, ownership, and career advice. Dan ties AI ROI to trust and verifiability rather than vanity metrics, identifies StratOps as the emerging owner of go-to-market performance measurement, and offers practical guidance for early-career professionals on why deep business process expertise paired with AI fluency is the highest-value combination in the market right now. If your company is facing pressure on retention, trying to build a more systematic expansion motion, or wrestling with unpredictable AI infrastructure costs, this episode delivers both the framework and the evidence behind it. Subscribe to AI to ROI on your favorite podcast app, leave a five-star rating, and connect with Ray at Ray Rike on LinkedIn to suggest a future guest. See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

  • May 27 · 33 min

    The AI Agent Outcome-Based Pricing Journey - with Kunal Agarwal, CFO Gorgias

    What does it actually look like when a CFO drives the strategic, pricing, and financial decisions behind an AI-first product transformation? Kunal Agarwal, CFO at Gorgias, the leading e-commerce customer experience platform for Shopify merchants, joins our host, Ray Rike to share the unfiltered story of how Gorgias built, priced, and operationalized its AI agent product from the ground up. This episode goes well beyond theory, covering the real decisions, real numbers, and real lessons learned from a company that has roughly half its customer base already using its AI agent product. Episode Highlights: The build decision: re-architect, don't bolt on. In early 2024, Gorgias made the deliberate choice to re-architect its platform around an agentic future rather than layering AI on top of an existing help desk product. The first AI agent focused exclusively on email support, shipped in July/August 2024, and expanded from there into chat and shopping assistance. Kunal explains why starting with a single, high-confidence use case was critical to earning early adoption and trust from merchants. The North Star metric: full resolution rate, not deflection. Gorgias intentionally moved away from deflection rate as its primary success metric, which can mask frustrated customers who simply abandon a conversation, and anchored instead on end-to-end AI resolution rate. That metric started with a target of 20 to 25% and has scaled to 60 to 80% for their largest enterprise customers. Why outcome-based pricing was the only intellectually honest answer. Seat-based pricing misaligns incentives, and per-ticket pricing creates the wrong incentive to grow ticket volume rather than resolve issues. Gorgias charges per resolution, meaning it only gets paid when the AI agent delivers a measurable outcome. Kunal explains how that pricing model forces the company to stand behind product quality and why keeping it simple, at the cost of short-term revenue maximization, was the right call to accelerate adoption. Gross margin reality: AI-native economics are structurally different from SaaS. Kunal is candid that AI agent gross margins are lower than traditional SaaS and that denying that fact is living in an alternate reality. With LLM inference costs running approximately 55 to 60% of fully loaded cost per interaction, and infrastructure as the fastest-growing expense line, Gorgias built real-time cost instrumentation by feature, a rolling 28-day average LLM cost per interaction, and a CFO-led governance model with weekly to bi-weekly engineering check-ins to stay ahead of cost drift. The shopping agent and the attribution problem. Gorgias expanded its AI platform from post-sale support into pre-sale shopping assistance, helping Shopify merchants drive incremental AOV and repeat purchases. The challenge is attribution: when a customer engages with a product recommendation but converts two to three days later, did the AI agent drive that sale? Kunal describes the approach of co-creating attribution logic with customers, which is the only way to make the ROI story believable and defensible. The CFO as owner of AI ROI, internally and externally. On measuring the return on internal AI investments, Kunal's view is clear: the Office of the CFO owns AI ROI measurement across every function, including product, marketing, and sales. Product and engineering teams are important stakeholders but have inherent incentives to measure outcomes favorably. Independent, finance-led measurement is what gives the numbers credibility with the board. See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

  • May 19 · 38 min

    AI to ROI: OpenAI - The Most Important AI Company in the World, and the Most Fragile

    OpenAI built $25 billion in annualized revenue and 910 million weekly active users in three and a half years. It also has 33% gross margins, a projected $14 billion loss, a CFO who was reportedly demoted for saying the company is not ready to go public, and an investor presentation that told its software partners it plans to replace them. In this episode, Ray and Peter work through six documented challenges facing OpenAI, six specific actions that could right the ship, and what enterprise leaders should actually do with their AI strategy given all of it. What we covered in this episode: The model is not the moat, and ChatGPT's market share is eroding Analyst Benedict Evans has noted that the six leading large language model companies are now roughly equivalent in capability, with no proprietary data advantage or network effect allowing any one to pull decisively ahead. ChatGPT's share of enterprise and developer usage has fallen from roughly 80% two and a half years ago to around 60% today, growing at just 4% while Claude grew 14% and Gemini 12%. OpenAI is a consumer-first product trying to pivot to enterprise at a moment when Anthropic is already the preferred first purchase for 73% of enterprise buyers according to Ramp data. Leadership integrity and financial credibility are both under pressure A 16,000-word New Yorker profile drawing from over 100 interviews raised serious questions about Sam Altman's management behavior and integrity. The Wall Street Journal followed with reporting on his personal investment conflicts. The CFO, Sarah Friar, was reportedly demoted after privately advising colleagues the company is not ready for an IPO. At a $852 billion valuation (roughly 28x projected 2026 revenue) with 33% gross margins and a $14 billion projected loss, institutional investors interviewed by The Information said they would not buy the stock and some indicated they would short it. The partner ecosystem problem could be existential In a February investor presentation, OpenAI stated it intends to build products that replace Salesforce, Workday, Adobe, Slack, and Atlassian, companies with whom it has active revenue-generating partnerships. Every systems integrator and enterprise software company building on top of OpenAI's models is now evaluating whether that is a safe long-term bet. Bill Gates defined a platform as something that creates more value for partners than for itself. OpenAI's current stated strategy is the opposite. Six actions that could change the trajectory Ray and Peter walk through a specific set of recommendations: launch a structured enterprise customer evidence program with named deployments and quantifiable outcomes; stop the public sniping at competitors and replace it with product and customer communication; fund an independent AI governance and safety board with real veto authority; impose IPO-grade communications discipline and treat major leaks as firing offenses; commit credibly to a partner ecosystem with defined product boundaries that give integrators a durable business case; and operate as a mature growth company, not a startup, because $30 billion in revenue demands the leadership behaviors that go with it. What enterprise leaders should watch and do right now Three signals will tell the real story over the next 12 months: whether Sarah Friar stays or exits, whether the IPO timeline slips to 2027, and whether enterprise case studies with quantifiable outcomes start appearing in volume. In the meantime, the strategic prescription is straightforward. Do not build single-model dependency into your AI architecture. Require the same evidence from OpenAI you would from any other vendor: verified outcomes, clear product roadmap, and accountability. And build API portability into your application design so you can move if you need to. The closing question: if you had to pick one LLM company to invest a million dollars in, where does it go? Peter picks Google, citing distribution advantages, DeepMind's research depth, and full control over its own financial destiny. Ray picks Anthropic, citing a lower revenue base with larger upside, near-universal goodwill across hyperscalers and enterprise buyers, and a safety-first positioning that is proving to be a genuine competitive differentiator. They agree on the conclusion: OpenAI is the defining company of the AI generation, but Netscape, Lotus, and BlackBerry were all category leaders too. See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

  • May 13 · 33 min

    NVIDIA – The Full-Stack Maestro

    Five months ago, Ray and Peter called NVIDIA the maestro of the AI economy. Since then, NVIDIA has not just conducted the orchestra. It has rewritten the music and may be building the entire concert hall. In this episode, Ray and Peter revisit their October thesis, walk through everything NVIDIA unveiled at GTC, and break down what it all means for enterprise AI buyers navigating infrastructure, inference costs, and procurement strategy. What we covered in this episode: From GPU maker to full-stack AI platform: the transformation is complete NVIDIA's strategic intent is no longer just selling chips. It is embedding its technology across the entire AI stack and becoming the foundational layer on which the rest of the AI economy rests. Ray draws the only historical parallel he can find: what IBM was to enterprise technology from the 1960s through the 1980s. The difference is NVIDIA is moving faster, with more cash, and with a software flywheel IBM never had. GTC was not a product launch, it was a platform declaration NVIDIA unveiled the Vera Rubin platform, a fully integrated AI supercomputer with liquid cooling and a two-hour installation window. They licensed Groq's LPU architecture in a $20 billion deal that combines GPU and LPU chips to deliver 35x token throughput over current Blackwell systems. They launched NemoClaw (an enterprise-grade agent framework already partnered with Adobe, Salesforce, and SAP), Dynamo (an open-source inference operating system), and the Nemotron family of open-source frontier models. Jensen committed $26 billion over five years in free cash flow to build best-in-class frontier models with no outside funding required. The financial performance is in a category by itself Fiscal year 2026 revenue came in at $215.9 billion, up 65% year over year and 8x since 2022. Data center revenue exceeded $190 billion. Free cash flow hit $97 billion, translating to a 47% free cash flow margin. Combined with 65% growth, that is a Rule of 40 score of 109. Ray notes he has never seen anything like it at scale, and NVIDIA is a hardware company running 80% gross margins. CFO Colette Kress described their inference position as: "right now, we are the king of inference." The moat is not hardware. It is ecosystem lock-in Since 2022, NVIDIA has committed over $50 billion across 170 venture deals, with corporate deal volume growing from 12 deals in 2022 to 67 deals in 2025. Portfolio companies include OpenAI, Anthropic, xAI, CoreWeave, and Lambda. Sovereign AI contracts signed since October total $30 billion across France, the Netherlands, Canada, Singapore, and the Middle East. Hyperscalers still represent roughly 50% of revenue, but the faster-growing segments are sovereign entities, enterprise verticals, and NeoCloud providers, which is exactly the diversification NVIDIA needs as hyperscaler CapEx normalizes. The risks are real but manageable from where NVIDIA sits today Custom ASICs from Google, Amazon, Meta, and Microsoft represent the most credible competitive threat, though those chips are optimized for internal platforms and do not solve multi-cloud or on-premise deployment needs. Export control escalation remains a live risk, with NVIDIA restarting NH200 production for China. TSMC concentration is a structural vulnerability, especially given geopolitical risk around Taiwan. And three hyperscalers account for over half of NVIDIA's receivables, some of whom are actively building competing chips. What enterprise AI buyers should do right now. Ray and Peter close with four concrete takeaways for enterprise buyers: evaluate the full infrastructure stack, not just GPU cost; model inference economics carefully before deciding which models to run and where; pursue a strategic partnership with NVIDIA rather than transactional procurement, because partnership creates supply access standard customers do not get; and do not assume custom silicon from hyperscalers solves your problem, because data residency and on-premise requirements often mean NVIDIA needs to be part of the solution regardless. See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

Showing 1–20 of 22 episodes