Skip to content
Artwork for The Experimentation Edge
BusinessTechnology

The Experimentation Edge

Growthbook

How do product teams decide what to build and what not to? The Experimentation Edge is the podcast where product, growth, and engineering leaders share how A/B testing, feature flags, and experimentation drive real business outcomes — backed by named companies and real numbers. From DoorDash's 12,000 A/B tests a year to Atlassian's experimentation-led product win to UPS's $500M experimentation team, each episode goes deep with operators running experimentation programs at scale.

Hosted by Ashley Stirrup, CMO at GrowthBook and a 25-year executive in data and experimentation. For product managers, engineers, data scientists, and growth leaders at B2B tech companies who care about experimentation culture, statistical rigor, and shipping with confidence. No marketing speak. Just operators explaining what they shipped, what moved the needle, and how experimentation reshaped their teams.

Topics: A/B testing, experimentation, growth experimentation, product experimentation, tech experimentation, feature flags, experimentation culture, statistical significance, marketplace experimentation, conversion rate optimization, experimentation at scale.

Play
  • 24 episodes
  • Avg 26 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • #40
    Yesterday · 20 min

    Realtor.com on using your AI as a junior data scientist

    Summary What do you do when your biggest experiment win turns out to be a loss? Whitney Perez, Director of Product Management at Realtor.com, joins host Ashley Stirrup to share the checkout bundling test that posted a 300% attach rate and still lost revenue, the 30/30/30 rule she uses to set expectations for a new experimentation team, and how AI is turning an English major into an aspirational data scientist. This episode is for product managers, engineers, and data scientists building experimentation programs from the ground up. Chapters 00:00 Cold open and welcome 01:10 From growth hacker to Realtor.com 02:50 Three foundations for a new experimentation team 04:20 The 30/30/30 rule 05:05 The 300% bundling win that lost revenue 07:10 You don't need a stats degree to experiment 09:20 Cascading North Star metrics 12:50 Do the homework before the experiment 13:50 The wishlist: instrumentation, embedded knowledge, culture 15:50 AI as an aspirational data scientist 18:45 Keeping a human in the loop Takeaways -A winning decision metric is not enough. Realtor.com's bundling test hit a 300% attach rate, but funnel fallout from the extra step made it a net revenue loser. Set secondary metrics and their thresholds before launch. -Expect the 30/30/30 rule: roughly a third of tests win, a third are inconclusive, and a third lose. The math is the math, and the losers carry most of the learning. -Start a new team on foundations: what a clean test and an A/A test look like, which surfaces should not be tested, and a peer review program that lets people graduate to more complex experiments. -You don't need a stats background to run good experiments. Teach the simplest definition of a good test, then let people learn by doing. -AI can make anyone an aspirational data scientist for analyzing results and spotting opportunities, but it can be confidently wrong. Keep a human in the loop and sanity check output the way you'd peek at a freshly launched test. Connect with the Guest LinkedIn: https://www.linkedin.com/in/whitneykperez/ Website: https://www.realtor.com Sponsor GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    • Transcript
  • #39
    Wednesday · 20 min

    How Zalando connects every experiment to its North Star

    Summary How do you keep 1,000 experiments a year pointed at one North Star? In this episode of The Experimentation Edge, host Ashley Stirrup talks with Mi Tian, Head of Applied Science at Zalando, about running experimentation inside a central economics org that reports to the CFO. Mi shares how Zalando balances safe confirmatory tests with game-changing bets, how the team measured the discovery feeds homepage launch when success had no established metric, and how a KPI tree cascades the company North Star down to the controllable inputs teams ship every day. She also looks ahead to LLM-based agents as a simulation layer for screening hypotheses. A practical conversation for anyone building an experimentation program that wants both rigor and ambition. Chapters 00:45 About Zalando and its marketplace model 02:00 Mi's path from engineering to experimentation 03:10 Economists and data scientists in one decision-making org 04:15 Running over 1,000 experiments a year 05:55 What makes an experiment high risk 07:10 Sharing learnings through standardization and champions 09:10 The discovery feeds launch and its measurement plan 11:15 Balancing the experimentation portfolio 12:35 Growing a KPI tree from the North Star 16:05 LLM agents and the future of experimentation at Zalando 18:35 New missions for a longstanding business Takeaways -Treat experimentation as a portfolio: balance confirmatory tests that protect the business with game-changing bets that can win big. -Assess risk tiers when building the roadmap so measurement rigor scales with the stakes instead of slowing every decision down. -When a launch is too new to have a success metric, pair short term A/B tests with long term holdouts from day one. -Connect every experiment to the company North Star by cascading it down to sensitive proxy metrics and controllable inputs. -Use LLM based agents as a cheap simulation layer to screen hypotheses, not as a replacement for real A/B tests. Connect with the Guest LinkedIn: https://www.linkedin.com/in/mi-tian-941b4767/ Website: https://www.zalando.com Sponsor GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    • Transcript
  • #38
    Tuesday · 21 min

    Learneo on testing the opposite of every hypothesis

    Summary Rich Liebling, senior director of engineering at Learneo, joins host Ashley Stirrup to explain the practice that came out of growing Shop It To Me from 50,000 subscribers to one million in nine months: test your hypothesis, and test its opposite. Rich covers the page where cutting text lost and adding text won, why the inverse wins more often than teams expect, and why small focused tests are the only ones where "the opposite" means anything. He also walks through translating a million subscriber goal into a target of 300 A/B tests, making a new engineer's second merge request their own experiment, and the different constraints at Course Hero, where competing team metrics were resolved with an early lifetime value model and three-day SQL analyses quietly capped testing velocity. This episode is for product managers, engineers, data scientists, and growth leaders building or scaling an experimentation program. Chapters 00:00 Cold open and introduction 01:50 Shop It To Me and the first engineering hire 02:55 300 A/B tests as the path to one million subscribers 04:00 Testing as a core value and the second merge request 07:20 Learneo, Course Hero, and two ways to get access 08:50 Modeling lifetime value to stop teams competing 11:10 Testing the opposite of the hypothesis 14:35 A portfolio of small, medium, and large tests 17:15 Multi armed bandits and seasonal traffic 20:35 The flywheel that keeps copycats behind Takeaways -Test the opposite of every hypothesis. At Shop It To Me the inverse won surprisingly often, and even when it lost it proved the variable mattered. -Keep tests small and focused. Redesign a whole page and lose, and you learn that version failed but not what to change next. -Translate a growth goal into an execution count. One million subscribers is not actionable. 300 A/B tests by year end is, and everyone can influence it. -Make experimentation part of hiring and onboarding. Every new engineer's second merge request was their own test idea. -Analysis friction sets the ceiling on testing velocity. Three days of ad hoc SQL per test quietly discourages teams from running more. Connect with the Guest LinkedIn: https://www.linkedin.com/in/richliebling/ Website: https://www.learneo.com Sponsor GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    • Transcript
  • #37
    September 3 · 21 min

    The four questions Early Warning asks before any A/B test

    Summary What separates a valid A/B test from an expensive guess? Priya Singhee, VP of Enterprise Analytics & Data Science at Early Warning — the bank-owned consortium that fights payment fraud and operates Zelle, which processed a trillion dollars last year — joins host Ashley Stirrup to share the experimentation playbook she built leading storefront analytics at Wayfair. She walks through the four questions to ask before launching any A/B test, why 85 to 90% of tests are supposed to fail, how pre-registration and kill criteria stop p-hacking before it starts, the pitfalls that fake wins (novelty effects, hidden heterogeneity, multiple comparisons), and how to roll out winners with gradual ramps and long-running holdouts. A practical episode for product managers, engineers, data scientists, and growth leaders building rigorous experimentation programs. Chapters 00:45 Meet Early Warning: fraud detection, Zelle, and a trillion dollars in payments 02:00 Wayfair and optimizing every step of the storefront funnel 03:05 The four questions to ask before any A/B test 05:15 Test setup best practices: hypotheses, guardrails, power, and pre-registration 07:40 Why 85 to 90% of tests fail and why that's a learning agenda 09:25 Novelty effects, hidden heterogeneity, and the multiple comparisons problem 12:55 Pre-registration, kill criteria, and stopping p-hacking 15:10 Rolling out winners: gradual ramps and long-running holdouts 17:15 Causal inference when you can't A/B test 20:05 The case for more A/B testing, not less Takeaways Run the four-question framework before any test: clean randomization, a plausible effect size for your traffic, a reversible and cheap change, and a falsifiable hypothesis. Treat A/B testing as a learning agenda: 85 to 90% of tests are supposed to fail, and a suspiciously high win rate is a red flag, not a trophy. Pre-register the full analysis plan, including hypothesis, mechanism, primary metric, exact statistical test, and subgroups, so p-hacking can't creep in when a test goes sideways. Define kill criteria and success, failure, and guardrail-dip actions before launch, with leadership sign-off, so nobody chases a loss into a fake win. Log every test and its learnings where the whole organization can see them; that reinforcement loop is what separates world-class experimentation programs. Connect with the Guest LinkedIn: https://www.linkedin.com/in/priya-singhee/ Website: https://www.earlywarning.com Sponsor GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    • Transcript
  • #36
    September 1 · 23 min

    How Supercell A/B tests 300 million players without breaking trust

    Summary On this episode of The Experimentation Edge, Ashley Stirrup talks with Shan Huang, data scientist on the central experimentation team at Supercell, the Helsinki mobile game company behind Clash of Clans, Clash Royale, Brawl Stars, Hay Day, and Boom Beach. Shan explains how a famously decentralized, creative-first company with 300 million monthly active users runs fewer than 100 A/B tests a quarter and why that restraint is deliberate, how Supercell announces experiments to players in advance and promises make-up events to keep testing fair, and how importable AI skills now let anyone at the company analyze their own experiments, making quality consistency the next big challenge. It's for product managers, data scientists, and growth leaders balancing creative conviction with experimental rigor. Chapters 00:00 Intro 01:10 About Supercell and 300 million players 02:30 The central team and a decentralized culture 04:25 Fewer than 100 tests a quarter 05:40 Sharing learnings across independent game teams 07:35 Retention as the North Star 08:35 Onboarding experiments with gems and tutorials 12:45 Telling players about A/B tests 15:15 Hypotheses and proxy metrics 20:45 AI and the future of experiment analysis Takeaways -Supercell runs fewer than 100 A/B tests a quarter for 300 million monthly players, because the goal is to become more hypothesis driven while staying creative, not to maximize volume. -In a decentralized company, a central experimentation team earns its impact by providing the platform, partnering on rigor, and making sure learnings travel across independent game teams. -Announce experiments to users before they run; Supercell's community managers tell players what is being tested and why, which turns a skeptical community into a research partner. -Promise fairness, not just transparency; players who get the worse variant always receive a make-up event later, because game players come to have fun, not to be disadvantaged. -AI-powered self-serve analysis means everyone can now run and analyze experiments, so the next challenge is making the quality of AI analysis consistent across the whole company. Connect with the Guest LinkedIn: https://www.linkedin.com/in/cnshanhuang/ Website: https://supercell.com Sponsor GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    • Transcript
  • #35
    August 27 · 30 min

    How Clover experiments when billions of dollars flow through daily

    Summary How do you run an experimentation program when classic A/B testing is off the table? Ben Schein, Director of Product Management at Clover, joins host Ashley Stirrup to explain how a platform serving 300,000+ merchants and processing billions of dollars daily proves every feature through pilots and ground-level testing before rollout, why uncertainty and downside — not feature visibility — decide testing depth, and what his years leading product at Shake Shack taught him about turning a checkout funnel into a brand channel. This episode is for product managers, engineers, and data scientists building experimentation programs where the stakes are real. Chapters 00:00 Cold open and welcome 01:30 The Clover business model and its scale 03:45 Ben's role and the metrics that matter 07:45 Deciding what gets tested: uncertainty and downside 09:05 Why Clover can't test in production 12:45 Testing the Shake Shack checkout experience 19:35 Advice for PMs new to experimentation 23:45 Context over personalization 25:45 The future of experimentation and the human element Takeaways -At Clover's scale, testing happens before rollout: pilots and detailed go-to-market plans replace in-production A/B tests, because a merchant's work tool can never change overnight without warning. -Uncertainty and downside set the testing depth. High-risk changes like payment authorization flows earn deep, ground-level experimentation, while table stakes features like Apple Pay earn a monitored rollout. -Testing is the evidence that justifies rollout investment: if the data doesn't show a feature will succeed, the go-to-market dollars never get spent. -A checkout funnel can carry the brand. At Shake Shack, Ben's team tested prep-time expectations, fixed wrong-location orders, and used loading screens to deliver hospitality digitally. -Structure experiments for durable business value, not pass-fail verdicts: pair headline metrics with counter metrics and anchor on ground-level measurements like items per check that resist marketing noise. Connect with the Guest LinkedIn: https://www.linkedin.com/in/benschein/ Website: https://www.clover.com Sponsor GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    • Transcript
  • #34
    August 26 · 26 min

    Why JobLeads says one test won't move you, but 100 will

    Summary What happens when the most logical feature you've ever built has zero impact? In this episode of The Experimentation Edge, host Ashley Stirrup, CMO of GrowthBook, sits down with Edd Saunders, product experimentation manager at JobLeads, to unpack the pizza personalization experiment that cut an ordering flow from 22 clicks to 5 and changed nothing. Edd shares the problem mapping framework he uses to move new experimenters from solution space to problem space thinking, how JobLeads grew from 0.3 to 2.8 experiments per month, and why democratizing experimentation across a whole company comes down to habit change rather than education. A practical conversation for product managers, data scientists, engineers, and growth leaders building experimentation cultures. Chapters 00:00 Introduction 01:40 Meet Edd Saunders and JobLeads 03:12 How JobLeads uses AI for prototyping 04:05 Building experimentation operations and a knowledge base 05:35 Learning over winning and compounding growth 08:10 The pizza personalization experiment 12:50 Moving from solution space to problem space 17:05 Problem mapping on a 2x2 matrix 22:20 Velocity and democratizing experimentation 24:20 AI automation for the unsexy work Takeaways -A one click reorder feature that cut a pizza ordering flow from 22 inputs to 5 had zero impact on purchases, proving that removing friction can also remove the customer's sense of control. -Exploration is part of the customer's delight; returning customers wanted to browse the menu even though they ordered the same thing every week. -Moving new experimenters from solution space to problem space thinking raises win rates and produces learnings the whole organization can use. -Problem mapping on a 2x2 matrix of evidence versus impact turns customer research into a prioritized experiment roadmap, and one validated problem can spring a whole tree of testable ideas. -Scaling experimentation from 0.3 to 2.8 tests per month is less about education and more about habit change, shared learnings, and giving non specialists the tools to launch their own experiments. Connect with the Guest LinkedIn: https://www.linkedin.com/in/eddsaunders/ Website: https://www.jobleads.com Read Blog: https://www.growthbook.io/podcast/episode/1-34 Sponsor GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    • Transcript
  • #33
    August 20 · 23 min

    Why The Aspen Group targets a 25% win rate

    Summary What does a winning A/B test mean when your website serves 1,100 dentist owned offices? In this episode of The Experimentation Edge, host Ashley Stirrup talks with Arie Polycarpou, Senior Manager of Test and Learn at Aspen Dental, about building experimentation programs at Kohl's, Marriott, Total Wine, and now the largest company in The Aspen Group's healthcare retail portfolio. Arie explains why appointment bookings are the North Star but never the whole story, why he deliberately targets a 25% win rate and expects it to fall as the program matures, and why he applies a haircut to every stacked win before it reaches leadership. A practical conversation for product managers, engineers, data scientists, and growth leaders building test and learn programs at multi-location businesses. Chapters 00:00 introduction and Arie's path into experimentation 02:05 building programs at Kohl's, Marriott, and Total Wine 03:15 healthcare retail and 1,100 dentist owned offices 04:30 the test and learn team and 100 tests a year 07:15 building a culture of shared wins and explained losses 09:50 what a losing navigation redesign revealed 12:30 testing your way into big redesigns 13:55 the case for a 25% win rate 15:50 haircuts, holdouts, and honest math on stacked wins 17:55 metrics beyond conversion and where AI fits next Takeaways -Aspen Dental runs experimentation as healthcare retail: with 1,100 dentist owned offices, the office, not just the website visitor, is the real unit of analysis. -A mature program should target a true win rate around 25%; a 40% win rate usually signals a young program still picking off low-hanging fruit. -Apply a haircut to stacked wins: lifts depreciate as customers acclimate, and two 5% wins never add up to 10%. -Losing tests are valuable when they're designed to isolate why: test your way into big redesigns instead of shipping them whole. -Experimentation culture grows from sharing wins and explaining losses: keep the statistical rigor in the back end and the communication simple. Connect with the Guest LinkedIn: https://www.linkedin.com/in/ariepolycarpou/ Website: https://www.aspendental.com Sponsor GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    • Transcript
  • #32
    August 18 · 31 min

    How Principal A/B tests on customers who don't exist

    Summary What if you could A/B test on customers who don't exist before spending a single live impression? Erika Dunn, assistant director of data science at Principal Financial Group, joins host Ashley Stirrup, CMO at GrowthBook, to share how she built synthetic digital audiences entirely in house: profiles shaped by real data that rank content by likelihood of engagement, whose first live A/B test selection just beat the control. They also dig into the looping metric her team built in SQL to find where customers get stuck without heat-mapping tools, why testing gets watered down into "let's try something" at large companies, and how a center of excellence that shares wins and losses defeats the "we tried that years ago" reflex. This episode is for product managers, data scientists, marketers, and experimentation leaders, especially those working inside large, risk-averse organizations. Chapters 00:00 Cold open and welcome 01:40 Erika's path from quantitative psychology to experimentation 03:35 When testing gets watered down 06:50 Experimentation in Principal's marketing space 09:55 The looping metric that finds stuck users 12:45 Building synthetic digital audiences 15:15 The first synthetic audience A/B test wins 22:30 The PDF lesson and meeting customers on mobile 25:15 North star metrics and the right contact cadence 28:05 Where experimentation goes next with AI agents Takeaways - Synthetic digital audiences let teams rank 20 content options by predicted engagement before spending a single live impression, and Principal's first synthetic selection beat the control in a real A/B test. - A looping metric built from web behavior data can find where customers get stuck without heat-mapping tools: watch how often users cycle back to the same page within tight time windows. - Testing gets watered down when "let's try something" replaces a control group; a little pre-planning gets far more out of every experiment. - A center of excellence that shares wins and losses turns tribal knowledge into shared knowledge and stops "we tried that years ago" from killing valuable retests. - An experimentation mindset requires that people can't get punished for mistakes; give teams guardrails and a safe playground and they'll stop running the same test forever. Connect with the Guest LinkedIn: https://www.linkedin.com/in/erikadunn/ Website: https://www.principal.com Sponsor GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    • Transcript
  • #31
    August 11 · 22 min

    Why US Bank considers missing even 1% of customers unacceptable

    Summary How does a major bank scale experimentation when even one percent of customers missing an experience is unacceptable? Vijay Lal, Lead Product Manager for Experimentation at US Bank, joins host Ashley Stirrup, CMO at GrowthBook, to share how his team made their experimentation platform self serve for non technical marketers, how a login widget experiment led to a two second fallback that accounted for every customer, and why metrics should be driven by hypotheses instead of handed down by leadership. They also dig into where AI genuinely saves time in experiment analysis, why a human in the loop is non negotiable, and what real time personalization means for the future of testing. This episode is for product managers, data scientists, and experimentation leaders, especially those working in regulated industries. Chapters 00:00 Cold open and welcome 00:40 Vijay's path from Comcast to financial services 03:16 Making the experimentation platform self serve 05:08 The login widget experiment and the two second fallback 08:59 Documenting learnings from every experiment 10:38 AI in experimentation and the human in the loop 12:39 Advice for new product managers 15:24 Hypothesis driven metrics 18:25 Real time personalization and agentic AI 20:11 Democratizing experimentation with responsibility Takeaways -Self serve experimentation lets a small central team support a huge testing volume, but it only works with continuous training and guardrail metrics attached. -In a regulated industry, every customer must be accounted for. Even one to two percent of users missing an experience is unacceptable. -A simple fallback, like a two second load rule, can save an ambitious experiment without sacrificing coverage or security. -Metrics should be driven by the experiment's hypothesis, not chosen by leadership in a silo. Pair a primary KPI with secondary KPIs for return behavior. -AI saves real time in experiment analysis, but a human in the loop must validate anything AI produces before it goes live. Connect with the Guest LinkedIn: https://www.linkedin.com/in/vijay-lal/ Website: https://www.usbank.com Sponsor GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    • Transcript
  • #30
    August 5 · 40 min

    Why Farfetch manages by learning rate, not win rate

    Summary Luis Trindade, Principal Product Manager of Experimentation at Farfetch, joins host Ashley Stirrup to explain how one of the world's largest luxury marketplaces built its own experimentation platform and the culture around it. Luis covers the move from a hybrid setup with an external testing vendor to Fabs 2.0, the in house system where a feature toggle is the single entry point for every experiment, why Farfetch manages by learning rate instead of win rate, and the two year Inspire experiment that replaced the world's leading recommendation engine vendor. He also shares how a deliberately shrinking center of excellence supports hundreds of experiments a month through clinics, shared templates, and open learning sessions. This episode is for product managers, engineers, data scientists, and growth leaders building or scaling an experimentation program. Chapters 00:00 Cold open and introduction 01:45 Inside Farfetch, the global marketplace for luxury fashion 08:00 From startup validation to an experimentation mindset 09:45 A center of excellence that enables instead of executes 12:45 Fabs, build versus buy, and dropping the external vendor 16:45 One feature toggle as the entry point for every experiment 20:45 Learning rate over win rate 23:15 The two year experiment that replaced the recommendation vendor 29:45 Onboarding new product managers into experimentation 33:15 AI, corporate knowledge, and what comes next for experimentation Takeaways -Manage by learning rate, not win rate. The only failed test is one that was badly designed, with wrong metrics or sampling biases. Every other test produces a learning. -Route every experiment through a single entry point. Farfetch's feature toggling system connects segmentation, user systems, CMS, and messaging so every team tests in the same language. -External JavaScript injection tools carry hidden costs: broken pages, inconsistent results, and rework to reclaim your own data for deep dives. -Strategic bets deserve a longer clock than fail fast allows. Farfetch iterated on its Inspire engine for two years before it beat and replaced the market leader. -A center of excellence should enable, not execute. Farfetch's central team shrank while experiment volume grew because its job is ceremonies, templates, and coaching. Connect with the Guest LinkedIn: https://www.linkedin.com/in/ltrindade/ Website: https://www.farfetch.com Sponsor GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    • Transcript
  • #29
    July 23 · 33 min

    How Cogniteer Built an Experimentation Engine From Scratch

    Summary In this episode of The Experimentation Edge, host Ashley Stirrup talks with Fabian Hans, founder and behavioral psychologist at Cogniteer, a consultancy that helps enterprises build in-house experimentation programs and raise both test velocity and win rate. Drawing on fifteen years in conversion rate optimization, Fabian explains why mass-producing the same A/B tests across clients quietly kills learning, why most ecommerce drop-offs are structural rather than your fault, and how matching the interface to how people actually buy, new versus returning, B2C versus B2B, can move conversion far more than another button. It is a practical, psychology-grounded conversation for product managers, engineers, data scientists, and growth leaders who want their experimentation programs to compound understanding, not just volume. Chapters 00:00 Introduction 01:25 From agency mass production to in house deep dives 04:05 Why some products resist selling online 06:35 The drop offs every ecommerce shop shares 07:45 The 50% win rate test Cogniteer reused 09:25 Why alignment beats developer resources 12:45 Two teams, two goals, one broken checkout 17:05 Selling water dispensers without a product catalog 22:15 Designing every experiment to lose 26:15 Personalizing buyers and where AI takes experimentation Takeaways -Deep dives beat mass produced tests, because understanding one business's users uncovers bigger levers than reusing the same test across many clients. -Many ecommerce drop offs are structural, since the basket and product page leak in roughly 80% of shops because it is ecommerce, not because of your product. -Product to channel fit decides what sells online, so books and fashion judge well on a screen while perfume and washing machines need cues the interface cannot fully provide. -The real bottleneck is alignment, not developer resources, so agree on the problem and its hierarchy before anyone builds a variation. -Match the interface to how people actually buy, because new buyers need information, returning buyers want speed, and B2B buyers often want a solution and an offer instead of a product catalog. Connect with the Guest LinkedIn: https://www.linkedin.com/in/fabianhans-cogniteer/ Website: https://www.cogniteer.de/ Sponsor GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    • Transcript
  • #28
    July 21 · 45 min

    How Fin A/B Tests Millions of Samples in Days

    Summary On this episode of The Experimentation Edge, host Ashley Stirrup talks with Pedro Tabacof, Principal Machine Learning Scientist at Fin (formerly Intercom), about how one of the most advanced AI customer support agents in the world is built on relentless experimentation. Pedro explains why unit tests don't work on non-deterministic AI, how Fin runs up to two dozen concurrent A/B tests pulling millions of samples in days, and shares two counterintuitive experiments: one where slowing the agent down improved every metric, and one where adding more context made Fin more helpful and more prone to fake promises until a targeted prompt fix kept the upside without the hallucinations. It's a candid look for product managers, engineers, and data scientists at how a $100M ARR AI product actually ships improvements. Chapters 00:00 Welcome and what Fin actually does 02:00 How Fin became Anthropic's first line of support 02:30 Why Fin sells resolutions not deflections 06:00 Owning the stack with custom models 10:40 Pedro's path from fuzzy logic to AI 12:55 Why A/B testing is the only gold standard for AI 15:50 Do no harm testing on every change 18:00 The latency experiment that shocked the team 27:30 When more context made Fin hallucinate 30:15 Win rates and the future of AI driven experimentation Takeaways -Faster is not always better. Fin increased latency artificially and positive feedback went up, likely because a small delay makes an AI feel like it is doing real work. -You cannot unit test a non-deterministic AI. A/B testing at scale, millions of samples in days, is the only reliable way to know a change actually helped. -Adding more conversation history made Fin more helpful and more prone to fake promises, until a targeted prompt fix removed the hallucinations and kept most of the gain. -A losing experiment is often a winner with one broken part. Diagnose which element hurts the experience, fix only that, and rerun. -Fin A/B tests everything, even one-character prompt changes and many bug fixes, and treats a 20 to 30 percent win rate as a healthy sign of a real experimentation program. Connect with the Guest LinkedIn: https://www.linkedin.com/in/tabacof/ Website: https://fin.ai Sponsor GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    • Transcript
  • #27
    July 14 · 22 min

    How Kargo turns losing experiments into competitive edges

    Summary In this episode of The Experimentation Edge, host Ashley Stirrup, CMO of GrowthBook, sits down with James Falzone, Director of Product Management at Kargo, to unpack how a high scale ad tech marketplace turns failure into its biggest advantage. James explains how Kargo connects advertisers to publishers through real time auctions that resolve in milliseconds across up to 10 billion ad requests a day, why experimentation is embedded in the company's culture rather than siloed in a team, and what happened when a winning click optimization model failed completely after being copied to a new customer type. The conversation is built for product managers, data scientists, engineers, and growth leaders who want a practical, honest view of running experiments at scale, learning from losses, and keeping AI grounded in solid infrastructure. Chapters 00:00 Welcome and introducing James Falzone 01:45 What Kargo does and how real time ad auctions work 04:45 Why experimentation is embedded in Kargo's culture 07:45 The three things every marketplace has to deliver 10:15 The experiment that failed: click optimization on third party demand 12:15 A bad result versus a bad experiment 13:45 Why different customer types need different signals 15:30 Putting "where did you fail?" on every retro 18:45 How experimentation evolves with AI 21:15 Better not bigger: the closing takeaway Takeaways -A bad result is not a bad experiment. If you're not failing, you're probably not trying anything new. -The same metrics and signals don't apply to every customer type. Bad results often come from a lack of context, not bad tech. -Metrics and signals you test against should always be business driven, not ported from the last thing that worked. -Put failure on the agenda. A biweekly "where did you fail?" retro turns one person's dead end into the whole team's shortcut. -AI's biggest unlock is access. More people can run experiments, but it has to be built on solid ML and infrastructure. Better, not bigger. Connect with the Guest LinkedIn: https://www.linkedin.com/in/jamesafalzone/ Website: https://kargo.com Sponsor GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    • Transcript
  • #26
    July 9 · 27 min

    The 'wine effect' and other surprises that reshaped how Box runs e-commerce experiments

    Summary In this episode of The Experimentation Edge, host Ashley Stirrup talks with Danielle Olean, Director of E-commerce at Box, about what it really takes to build a culture of experimentation inside a B2B company. Drawing on 15 years across B2C and B2B at Wayfair, Drizly, Zoom, and now Box, Danielle explains why experimentation belongs to every product team and not just e-commerce, walks through a pricing page saga of one win and two losses that exposed the limits of simplification, and shares the "wine effect" test that won for a reason no one predicted. It's a practical, story rich conversation for product managers, growth leaders, and anyone trying to make better decisions with data. Chapters 00:45 Meet Danielle Olean and Box's reinvention 02:45 Owning the entire customer life cycle 04:45 Why experimentation matters even without a checkout 07:45 The feature that's used but hidden 11:45 Proving ROI with a scrappy manual test 12:45 Building a culture that shares wins and losses 16:45 The pyramid strategy for prioritizing tests 18:45 The simplification tightrope on the pricing page 24:45 When a test wins for the wrong reason 27:45 Where experimentation at Box goes next Takeaways - Experimentation isn't only for e-commerce. Any product with a funnel, even an AI chatbot, can be measured and improved through testing. - Simplification has a limit. Removing too much can strip away the cues and context buyers actually need to decide. - Share losses as openly as wins. Wins build credibility, and losses build the psychological safety a testing culture runs on. - Prioritize like a pyramid. Fix the widest-impact experiences first, then optimize down into smaller cohorts. - Surprising results are the point. A test can win for a reason you never hypothesized, like the "wine effect," and that's where the real learning lives. Connect with the Guest LinkedIn: https://www.linkedin.com/in/dolean1/ Website: https://www.box.com Sponsor GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    • Transcript
  • #25
    July 7 · 21 min

    Dilligent explains why moving on from an experiment might cost you

    Summary Dan Layfield, Director of Product Management at Diligent, joins host Ashley Stirrup on The Experimentation Edge to trace what fifteen years of A/B testing across Codecademy, Uber Eats, and the Fortune 1000 boardroom actually taught him. He breaks down the Codecademy trial-model rebuild that took four months and several rounds to deliver a 35% conversion lift, why moving on from a losing experiment too early is one of a PM's costliest mistakes, how to escape the B2B feature factory with metrics that genuinely ladder up, why retention should ride a product's natural use case instead of fighting it, and where AI is already replacing weeks of research and analysis. It's a practitioner's guide for product managers, growth leaders, data scientists, and engineers bringing experimentation rigor to both B2C and B2B. Chapters 00:45 Meet Dan Layfield and Diligent 01:45 Two worlds of experimentation, Codecademy and Uber 03:45 The trial model that lifted conversion 35% 06:20 What to do with a losing experiment 08:50 Two flavors of experimentation 09:45 Reading forty metrics at Uber Eats 13:10 Escaping the B2B feature factory 16:45 Anchoring the North Star to real usage 19:15 Where AI fits in research and analysis Takeaways A losing experiment is often inconclusive, not negative; treat it as a map of the funnel rather than a verdict, and know when a big problem is worth another round. Persistence paid off at Codecademy: four months and three to four rounds of trial-model testing produced a 35% conversion increase. Separate your two experimentation modes; high-volume CRO chases many small wins, while big, uncertain bets are worth taking multiple shots to de-risk. Most B2B product teams are feature factories; the fix is a top-down OKR system, and planning usually breaks in the connections between layers, not inside them. Anchor retention and engagement to the product's natural use case, and use AI to synthesize research and simple A/B analysis in hours instead of weeks. Connect with the Guest LinkedIn: https://www.linkedin.com/in/layfield/ Website: https://www.diligent.com Sponsor GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    • Transcript
  • #24
    July 2 · 20 min

    The metric Stitch Fix says every experimenter should chase

    Summary In this episode of The Experimentation Edge, GrowthBook CMO Ashley Stirrup sits down with Nick Beyler, data science manager at Stitch Fix, where he leads the decision and insights team and owns the company's internal experimentation platform. Nick shares why the metric he most wants is the one he can't measure yet, a North Star that predicts a client's long-term value from their earliest behaviors, and why the most impactful experiment learnings tend to come from adoption friction rather than product bugs. He makes the case that if you're only testing winners you're not taking enough risks, explains how guardrails make that risk safe, and looks ahead to a new in-house platform and the promise of agentic AI. It's a practical, statistician's-eye view of experimentation for product managers, data scientists, and engineers building serious testing programs. Chapters 00:00 Cold open and welcome to the show 01:45 What Stitch Fix actually does 04:15 Balancing AI with the human stylist 05:15 From public policy to the A/B testing adrenaline rush 07:15 Inside the weekly experimentation review group 08:45 The AI style assistant and listening to qualitative feedback 10:45 Why adoption friction beats product bugs 13:45 Testing for losers and building guardrails 15:45 Keep rate, successful fixes, and the holy grail metric 18:15 The new platform and the promise of agentic AI Takeaways The most impactful experiment learnings usually come from adoption friction, not product bugs. By the time a big feature reaches A/B testing, it's often already a winner, so the open question is how and where to introduce it. A losing test is a finding, not a failure. If every experiment wins, you're not taking enough risk to learn anything new. Guardrails and stopping criteria are what make risk-taking safe, especially when the experience is as personal as shopping. The most valuable North Star metric is the one you can't measure yet, long-term client value, and causal-inference modeling helps predict it from short-term behavior. Quantitative results are only half the story. Direct, qualitative client feedback inside an experiment often reshapes the rollout more than the numbers do. Connect with the Guest LinkedIn: https://www.linkedin.com/in/nick-beyler-381864119/ Website: https://www.stitchfix.com Sponsor GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    • Transcript
  • #23
    July 1 · 28 min

    What the Expedia Group cannot measure, it cannot ship

    Summary Amir Moghaddam, Director of Software Engineering at Expedia Group, joins host Ashley Stirrup on The Experimentation Edge to make the case that measurement is not a reporting step but a gate: what you cannot measure, you cannot ship. Drawing on nearly four years at DoorDash and his current work leading Expedia's air booking platform, Amir explains why he refuses to label experiments winners or losers, how a "failed" pricing test pushed his team toward full personalization, and why a three sided marketplace forces hard trade-offs between competing metrics. The conversation closes on how the same experimentation discipline now applies to shipping and measuring AI. Built for product managers, engineers, data scientists, and growth leaders who care about rigor over opinion. Chapters 00:00 Cold open 00:50 Meet Amir and the air booking platform at Expedia 03:10 DoorDash, growth, and a 70 experiment year 04:20 Three kinds of experimentation at Expedia 06:30 AI velocity and the new frontier model pace 08:30 What you cannot measure, you cannot ship 10:45 The DoorDash carousel and the price experiment 12:45 The three sided marketplace and competing metrics 16:55 There are no losing experiments 20:45 Predictability, LLMs, and Expedia's road ahead Takeaways "What you cannot measure, you cannot ship" — if you can't measure an outcome, you can't decide whether it's better, so you're just debating opinions. Measurement spans three live dimensions: spend (more with less), speed (sprints instead of quarters), and quality, with guardrail "do no harm" metrics on top. There are no losing experiments. A flat result is a signal to either refine the hypothesis or step back and look from a completely different angle. DoorDash's price experiment proved price by itself doesn't predict orders. Different customers want different things at different times, which pushed the team toward personalization. A three sided marketplace (buyers, merchants, Dashers) makes metrics compete. Running the test is easy; deciding what to optimize when goals conflict is the real work. Connect with the Guest LinkedIn: https://www.linkedin.com/in/amirmoghaddam Website: https://www.expediagroup.com Sponsor GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to growthbook.io

    • Transcript
  • #22
    June 30 · 23 min

    How Fin went from weeks to hours of analysis using AI

    Summary In this episode of The Experimentation Edge, host Ashley Stirrup sits down with Raunak Kumar, senior manager of GTM analytics at Fin (formerly Intercom), to unpack how experimentation actually works when the data is messy and the traffic is thin. Drawing on nearly 12 years in marketing analytics across Atlassian, Stripe, and Fin, Raunak explains how AI tools like Claude Code have collapsed analysis from weeks to hours and freed his team to clear its experiment backlog, why declining organic search traffic and a 5x jump in untagged ChatGPT referrals are forcing teams to rethink attribution, and how the most valuable experiments are often the ones that "lose." From a Jira Service Desk bundling test that won on trials but had to be rolled back, to a Stripe contact form that was quietly blocking real buyers, this conversation is a practical guide for product managers, engineers, data scientists, and growth marketers who want to learn more from every test they run. Chapters 0:45 Welcome and what the show is about 1:45 Raunak's role and 12 years in marketing analytics 2:45 How AI and Claude Code changed the analyst's day 4:15 LLMs, declining organic traffic, and the 5x ChatGPT jump 5:15 Two kinds of experiments at Fin: on page and off page 7:15 The Jira Service Desk bundling experiment 10:45 Why the trial winner became a rollback 11:45 Contextual onboarding turns the loser into a winner 14:45 Reading an experiment that loses 18:45 What's next: incrementality, connected TV, and testing creative Takeaways AI has collapsed marketing analysis from weeks to hours, and the real payoff is a cleared experiment backlog plus analysts who compete on the questions they ask, not the speed they query. Organic search traffic is declining as ChatGPT, Gemini's AI mode, and Claude answer buyers in place; Fin saw a 5x rise in ChatGPT referrals, but LLMs don't tag that traffic, so attribution has to be proven through experiments. A guardrail metric saved Atlassian from a costly mistake: bundling Jira Service Desk lifted trials more than 50 percent but tanked activation and paid conversion, forcing a rollback. A failed test can hold the real winner; contextual onboarding matched to user intent roughly doubled activation and became the default variant after the bundling experiment was rolled back. In low-volume B2B, read losing experiments for sub-segment signal; a "failed" Stripe form simplification revealed the form was blocking legitimate small-business buyers using Gmail. Connect with the Guest LinkedIn: http://linkedin.com/in/raunakkumar1991 Website: https://fin.ai Sponsor Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts. Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse. With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction. See a demo at https://www.growthbook.io/

    • Transcript
  • #21
    June 29 · 11 min

    Inside The Home Depot's experimentation at a $25B scale

    Summary What does experimentation look like inside a $150 billion retailer? In this episode of The Experimentation Edge, host Ashley Stirrup talks with Kim Ting Li, Senior Manager of Experimentation at The Home Depot, where one centralized team tests every major change to a $25 billion online business. Kim explains how 40 people serve 40–50 business teams, why executives join test readouts and ping analysts directly, how every result since 2020 lives in a searchable library, and why scaling beyond hundreds of experiments per year depends on server-side testing capabilities more than AI. For product, data, and engineering leaders building or scaling experimentation programs. Chapters 00:00 Intro 00:45 From neuroscience research to Home Depot 01:45 A $150B enterprise, a $25B online business 02:45 The centralized experimentation model 03:45 Inside the 40-person team 04:30 Readouts, blast emails, and the experiment library 05:40 Executive visibility and the golden rule 06:15 "If you won't act on a bad result, don't run the test" 11:15 Learning from losing tests 12:30 Scaling up: AI, server-side testing, and what's next Takeaways One centralized team of about 40 people tests every major change to Home Depot's $25B online business, serving 40–50 business teams with consistent hypothesis and analysis standards. Executive engagement is real at Home Depot: leaders join 30-minute readouts, search the experiment library, and ping analysts directly because they treat A/B testing as the golden rule for measuring incrementality. Institutional memory is infrastructure — every test result since 2020 lives in a centralized, searchable archive so no one re-runs a question the company already answered. Kim's stakeholder filter: if you wouldn't do anything differently after a bad result, don't run the test. Scaling past low hundreds of experiments per year is a capabilities problem before it's an AI problem — Home Depot is moving from client-side to server-side testing so winners release quickly, end to end. Connect with the Guest LinkedIn: https://www.linkedin.com/in/kimtingli Website: https://www.homedepot.com Sponsor Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts. Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse. With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction. See a demo at https://www.growthbook.io/

    • Transcript
Showing 1–20 of 24 episodes