Skip to content
Artwork for The Test Set by Posit

The Test Set by Posit

Posit, PBC

A Posit podcast for data science junkies, anomaly hunters, and those who play outside the confidence interval. Hosted by Michael Chow, with co-hosts Wes McKinney & Hadley Wickham.

Play
  • 21 episodes
  • fortnightly
  • Avg 1 hr 3 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • S1 · E29
    Monday · 1 hr 1 min

    Nobody Remembers Your 20 Charts — with Ruth Milligan

    Ruth Milligan has coached hundreds of conference speakers, curates TEDx Columbus, and wrote The Motivated Speaker. She's here to tell you that speaking is habitual, not natural. Michael, Wes, and Hadley get a live fifteen-second filler word hack, an uncomfortable truth about listening back to your own voice, and the three kinds of pitches Ruth has been reading for seventeen years. What's inside Why a great idea beats a great speaker Speaking is habitual, not natural, and that’s good The fifteen-second breathing trick to kill fillers How many charts do you need? Fewer than you think The three pitches Ruth has read for seventeen years You can't get better until you listen to yourself Hadley contends people who send voice memos are monsters Mentioned in this episode Patsy Rodenburg's "Three Circles of Energy" — a summary of the framework Rébecca Kleinberger, "Why you don't like the sound of your own voice" (TED)

    • Transcript
  • S1 · E28
    August 10 · 1 hr 10 min

    Let the Agent Cook — with Trevor Manz

    Trevor Manz went from measuring plant apertures by hand in a wet lab to building the notebook that lets coding agents take the wheel. The creator of anywidget and founding engineer at marimo (marimo.io/pair) popped into The Test Set to spill on reactive notebooks, why marimo pair threw out every MCP tool but one, and what agents really want out of a data environment. This conversation also features a jacket bouncer, a hidden Python API, and Michael's slow-motion war with the word "marimo." What's inside: Cell order doesn't matter in a reactive Python notebook Wet-lab pipettes and Harvard's visualization group, via Raspberry Pi The problem with building beautiful tools nobody actually uses The reason marimo pair deleted every agent tool but one Code mode: the hidden API humans aren't supposed to touch What happens when you ship the API your LLM hallucinated

    • Transcript
  • S1 · E27
    July 27 · 1 hr 14 min

    The Answer Was Never Us — with Leilani Battle

    Leilani Battle studies how software shapes what we see and believe. The University of Washington professor and co-director of the UW Interactive Data Lab talks with Michael, Hadley, and Wes about an experiment that manipulated people using nothing but loading speed, and why AI models don't seem to recommend charts the way the community that studies charts actually does. Other highlights: rationality's blind spots and a thorough disc golf origin story. What's inside: Loading spinners that quietly change what people find AI models that don't recommend charts like humans do The blurry line between databases and human factors A behavior-change experiment for building fairer models Rationality's role in producing unethical outcomes Disc golf origin story A vegan pancake recipe

    • Transcript
  • S1 · E26
    July 13 · 1 hr 5 min

    Curiosity, duty, and existential dread — with Joe Cheng

    Joe Cheng is the CTO of Posit and the creator of Shiny. He joins Michael and Hadley to talk about why he almost walked away from AI work entirely over ethics concerns and what it takes to lead a team that didn't necessarily choose you. Plus, why saying yes to everyone is a worse strategy than it sounds. Bonus: Hadley calls out Joe's people-pleasing in real time. What's inside: Joe's 2012 self-doubt spiral that accidentally created Shiny Why Joe almost quit working on AI entirely The "loaded guns" problem with releasing AI tools Hadley's blunt leadership style vs. Joe's people-pleasing Nobody actually wanted to make Joe CTO? Joe's take on curiosity, duty, and fear as motivators

    • Transcript
  • S1 · E25
    June 29 · 1 hr

    Confidently Incorrect — with Caitlin Colgrove

    Caitlin Colgrove is the CTO of Hex, the data workspace for building and sharing data projects using SQL and Python that somehow counts a Sweetgreen chef as a power user. She joins Michael, Hadley, and Isabel to talk about what AI agents actually get wrong in data work (it's not the hallucinations, it's supreme overconfidence), why data teams aren't going anywhere, and how she thinks about building products for humans and agents at the same time. What's inside What Hex's Context Studio does, and why it's a data team's new job More code is now written in Hex by agents than by humans "My job is to vouch for the correctness of the answer" — redefining the data team The vibe-coded CEO PR is coming for your data team (if it hasn’t already) Soulsborne games as couple's therapy, aka, the Elden Ring co-op report

    • Transcript
  • S1 · E24
    June 15 · 1 hr 14 min

    The Bothness of It — with Alex Hillman

    Alex Hillman built one of America's first co-working spaces, wrote a business book in tweets, and recently handed his inbox to a Claude Code agent — not to draft emails, but to notice when a friendship is going cold. In this episode, Alex, Michael, Wes, and Hadley dig into marketing for people who hate marketing, what 20 years of email reveals about your relationships, and why the hardest part of AI-assisted coding was always before you wrote a single line. What's inside: Marketing is really just listening at scale Building a 20-year relationship database from your sent folder "Hot rod vs. plumbing" — the two kinds of software you build now What early internet and the AI boom have in common The case for reading 20-year-old engineering books with a coding agent Karaoke philosophy as a framework for community building

  • S1 · E23
    June 1 · 1 hr 8 min

    The Code Doesn't Lie — with Mike Bostock

    Mike Bostock made D3 when the browser was still a joke. He built bl.ocks when people needed somewhere to share their work. Now he's building Observable — reactive notebooks with an AI that actually looks at what it made. In this episode: the three-GIF bar chart that launched 25 years of viz, why open source needs both intrinsic and extrinsic motivation, and why an agent that can't see its own output is likely to be confidently wrong. What's Inside The 1998 visualization library that could only make bar charts Why D3 hit #3 on GitHub, and what killed the gallery What spreadsheets got right that notebooks ignored for years "The agent can lie with text, but not with code" Why Observable scrapped canvases and went back to notebooks The penguin dataset that exposes AI Strength training, tennis mind games, and a resurrected Stanford game

    • Transcript
  • S1 · E22
    May 18 · 45 min

    The Wonder-Driven Builder — with Paige Bailey

    Paige Bailey is a developer relations engineering lead at Google DeepMind. She's a geophysicist-turned-AI-engineer who was once told by her professors that building open-source libraries was a waste of time. We talk about her path from planetary science to TensorFlow, why statisticians have a hidden edge in the age of AI, and what it means to be a curious generalist when the cost of building software is approaching zero. Bonus: installing solar-powered silent-film birdhouses as street art in San Francisco. What's inside From planetary science to TensorFlow, before it was GPU-capable Geophysicists as early GPU adopters The professors who said open-source wasn’t “real science” Building silent-film birdhouses as San Francisco street art Hiding Gemini API tests inside whimsical side projects The right-tool-for-the-job case for mixing AI models Why “taste” is the skill that matters when code costs nothing

    • Transcript
  • S1 · E21
    May 4 · 1 hr 15 min

    Widgets Are Lego Bricks (and Other Things People Are Sleeping On) — with Vincent Warmerdam

    Vincent Warmerdam has been the first full-time hire at a startup, a spacey punster who accidentally got himself a job, a bartender at an Amsterdam comedy theater, and a Dutch bike tour guide — and he'll tell you all of it was career development. Now doing DevRel at Marimo, Vincent makes the case for reactive notebooks, Lego-brick widgets, and why "number go up" is not a data science strategy. Also: chickens die. The model doesn't know. This matters more than you think. What's inside How a spacey pun accidentally launched Vincent's career Why Marimo's constraints make it better for LLMs, not just humans The gorilla hiding in your dataset — and why the model missed it Vibe coding vs. notebooks: three cells at a time as a discipline Widgets as Lego bricks: reusable, composable, criminally underused Cognitive debt, confirmation bias, and sycophantic data science Why natural intelligence is still, actually, a pretty good idea

    • Transcript
  • S1 · E20
    April 20 · 1 hr 35 min

    Everything's a Fad (Including This Podcast) — with Benn Stancil

    Benn Stancil built Mode Analytics, spent a decade in the data trenches, and now writes some of the sharpest, funniest essays in the data world. On The Test Set, he talks about the cultural shift from Nate Silver to Rick Rubin why AI might kill the analytics dashboard, and what happens when a thousand startups all build the same thing. Plus: boy bands as a model for collaboration, and why the best creative work starts with cheating. What's inside: Why the modern data stack was basically big data 2.0 The cultural flip from Nate Silver to Rick Rubin Gas Town, tar pits, and the AI startup zero-sum game Software is becoming content, and that changes things Benn's creative process: Lorde lyrics, Codenames, and cheating The boy band as a model for small-team collaboration BI is (mostly) dead, and vibes might replace SQL

    • Transcript
  • S1 · E19
    April 6 · 1 hr 5 min

    Deeply Unsexy: SQL's Redemption Arc — with Tristan Handy

    dbt Labs CEO Tristan Handy drops into The Test Set to map the fault lines between the data science world and the enterprise data world — and explain why analytics engineers are basically pissed-off data analysts who decided to organize the bookshelf. We get into SQL's glow-up, the community magic of dbt Slack, what AI agents mean for data warehouses, and why everyone's building iOS apps with Claude now. What's inside: What analytics engineers *actually* do SQL's journey from deeply unsexy to indispensable How dbt turned source control into a source of truth Building a tech community without the RTFM energy AI agents on your data lake: permissions get personal Will LLMs kill the open-source package ecosystem? Edible gardening, welding dreams, and digital dysphoria

    • Transcript
  • S1 · E18
    March 23 · 1 hr 22 min

    Your VP Is Doing a Rogue Analysis in Cursor Right Now — with Nell Thomas

    Nell Thomas has spent two decades in data — from equity research to the DNC to Facebook to leading a 400-person data org at Shopify. She walks Michael and Wes through the modern data stack role by role, gets honest about what AI is and isn't changing about data work, and admits the semantic layer has been her greatest leadership failure. Plus: Sneakers gets the respect it deserves. Episode Notes What does it actually look like to run data infrastructure for millions of merchants while the entire industry reinvents itself in real time? Nell Thomas (VP of Data, Shopify) talks vibe-coded dashboards, political campaign data scarcity, blameless postmortems, and why no one should be locking in on an AI strategy just yet. Recorded live in Times Square. What’s Inside Mapping the modern data stack, role by role Why data quality is still the #1 problem What "good scrutiny" looks like on a data team Vibe coded dashboards and the trust problem Shopify's MCP for their data warehouse The throwaway tech problem in political campaigns Why the semantic layer is so damn hard Sneakers!

    • Transcript
  • S1 · E17
    March 9 · 56 min

    Sleeping Rats and Sociopathic Agents — with Phillip Cloud

    Phillip Cloud has been shaping the Python data ecosystem since the early pandas days — and he has *opinions*. Now a principal engineer at NVIDIA leading the Ibis project, Phillip talks about how he stumbled into open source via an eye movement lab, why he prefers his coding agents cold and emotionless, and what happens when you ask an LLM for woodworking trig. Plus: terminal user interfaces, the file hierarchy standard hot take nobody asked for, and the pineapple-on-pizza hill he's willing to die on. Episode Notes Phillip Cloud (NVIDIA, Ibis project) joins Michael Chow, Wes McKinney, and Hadley Wickham to talk about his path from eye movement labs to pandas core team, why developer productivity tools have quietly gotten amazing, his brutally honest take on coding agents, and what it would actually take to impress him. Also: VisiData love, NixOS evangelism, and yard work as therapy. What’s Inside From eye movement labs and MATLAB to pandas core team Column multi-indexes: the feature nobody likes but Phillip needed Why your command line tool better have a sweet TUI VisiData: the terminal data tool you're sleeping on Are we writing code for humans or for agents now? Phillip's AI skepticism journey: Cursor, Claude Code, and frustration The Numba CUDA test suite port that would finally impress him

  • S1 · E16
    February 23 · 1 hr 35 min

    More productive but a lot less fun — with Charlie Marsh

    Charlie Marsh built Ruff, uv, and Ty — the tools that mass-fixed Python's worst pain points. Now he's grappling with what happens when agents start writing most of the code. In this episode, Charlie gets real about his team trusting his PRs less, the gnarly middle of coding with agents, and whether Python is even the right language for an agentic future. It's honest, a wee existential, and deeply relatable if you ship code for a living. Episode Notes Charlie Marsh is the founder and CEO of Astral — the company behind Ruff, uv, and Ty. He sits down with Michael and Wes to talk about what it's actually like building with coding agents every day, why his team's code review dynamics completely changed, and big open questions about code quality, open source community, and Python's future nobody has answers to yet. What’s Inside How Ruff convinced mass adoption when switching tools is painful Why "just uv run it" became the killer feature Hiring outside Python's ecosystem to build tools for it His team said "we trust your PRs less now" Engineers screen-sharing their actual agent workflows at Astral The "Lisp Curse" reborn: cheap code breaking open sourceIs Python the wrong language for an agentic world?

  • S1 · E15
    February 9 · 1 hr 1 min

    Alenka Frim: What yoga teaches us about discipline and collaboration in data science

    Alenka Frim went from teaching yoga full-time to becoming a committer and PMC Member on Apache Arrow. In this episode, Alenka joins The Test Set hosts to talk about how Arrow grew from spec to critical infrastructure, and why she started contributing to a project she had never even used. She reflects on imposter syndrome, the discipline of showing up (on the mat and in GitHub), and how agents are changing what it means to write code. Plus: managing 4,000 open issues without losing your mind. Episode Notes Alenka's path into Arrow is unconventional: Sshe wasn't looking for a job, she wasn't using the tool, and she'd spent the previous five years focusing on mind-body fitness. But open source felt like the right place to learn, have fun, and figure things out, so she jumped in. What followed was a journey from her first R bindings to becoming a PMC member on one of the most critical pieces of data infrastructure in the world. What’s Inside Alenka's journey, from yogi to Arrow committer Signs of a healthy open source community: people, dialogue, and turnover Arrow as critical infrastructure: DuckDB, Polars, Pandas, and the spec that unifies them Managing 4,000 open issues without losing your mindImposter syndrome in open source What the yoga mat teaches you about discipline and collaborationAI and the future of programming: 100x more software or 10x better software?

  • S1 · E14
    January 26 · 58 min

    Emily Riederer: Column selectors, data quality, and learning in public

    Emily Riederer writes Python with an R accent, and we’re all comfortable with it. In this episode, Emily reflects on her journey through R, Python, and SQL — from lessons learned in averaging default values (oops, we're not all rich!) to discovering that column selectors are way cooler than they sound. She weighs in on the delicate art of learning in public, why frustration often makes the best teacher, and how to find your niche by solving the boring problems. Oh, Oh, and the crew casually drops that she's keynoting posit::conf 2026! Episode Notes Emily’s had a wild ride through modeling, data engineering, machine learning, and back again, and she knows a thing or three about the evolution of SQL tooling (from nightmare multi-page scripts to the dbt renaissance). She reveals how building internal packages became her gateway to making work enjoyable. Plus: the surprising Stata origins of column selectors, the eternal struggle of naming packages across R and Python, and why watching people code teaches you more than any tutorial ever could. The conversation gets real about imposter syndrome and the magic of tacit knowledge. What’s Inside Why real-world data is chaos, not truthThe path from modeling to data engineering (and back) What a data pipeline really is (extract, load, transform) and why organization matters How dbt changed the SQL game Learning by watching: Tacit knowledge and coding over the shoulder Imposter syndrome and learning in public Building internal tools to escape busyworkposit::conf 2025 keynote preview

  • S1 · E13
    January 12 · 56 min

    Rebecca Barter: Persistent learning, tool building, and ‘Will code even exist?’

    Rebecca Barter, senior data scientist at Arine and adjunct assistant professor at the University of Utah, refuses to work on things she doesn’t care about. Lucky for us, she cares about a lot, most of all impact. In this episode, Rebecca joins The Test Set to talk about learning fast, building better tools, and staying motivated and adaptable. She shares how moving between R, Python, SQL, and dashboards reshaped how she thinks about expertise. Plus a reflection on her recent posit::conf talk, “AI: Hype, Help, or Hindrance.” Episode Notes Rebecca digs into what it’s really like to work with AI every day and why humans still rule, especially in exploratory data analysis. She explains how tool building can be the fastest way out of busywork and how teaching beginners sharpened her ability to communicate clearly. The conversation circles a bigger question too: As AI keeps improving, are we headed toward a future where code looks completely different … or maybe disappears altogether? What’s Inside Why motivation matters even more than productivity Escaping busywork by building better tools From R to Python to dashboards: Learning fast as a survival skill Reality check on AI in the IDE Why exploratory analysis still needs human intuition The 80/20 of coding: Automate the boring, protect the judgment Teaching beginners and earning trust The uncertain future of code

  • S1 · E12
    Dec 15, 2025 · 51 min

    Marco Gorelli: Narwhals, ecosystem glue, and the value of boring work

    You’ve probably used Narwhals without realizing it. It’s the compatibility layer helping apps and libraries like Plotly play nice with Pandas, Polars, Arrow, and more — while keeping computation native instead of converting everything to Pandas. In this episode, Marco Gorelli explains how his weekend experiment turned into essential ecosystem infrastructure and why data types, not APIs, are where interoperability gets tricky. Plus what it takes to build trust and community around an open-source project. Episode Notes Marco shares the Narwhals origin story (including the meme-powered name), the hard edge cases that live in data types and null semantics, and why he’s cautious about using AI for code generation when correctness hinges on tiny details. We also jam on proactive “GitHub surfing,” conference talks as trust-building exercises, celebrating contributors, and how early commit messages capture the genuine excitement of building something new. What’s Inside Narwhals 101: You’ve probably used it (even if you didn’t know it) The real interoperability traps: data types, null semantics, and “looks-the-same” operations Why expression systems won, and how they shaped Marco’s approach — with nods to Ibis, Polars, and Pandas Open source as social work: proactive outreach, trust-building, and a Discord-powered community Extending Narwhals to new engines, starting with the Daft plugin

  • S1 · E11
    Dec 1, 2025 · 51 min

    Kelly Bodwin — Quarto hacks, AI in the classroom, and why R should stay weird

    In this episode, we’re joined by Kelly Bodwin — candy corn defender, board game enthusiast, and Associate Professor of Statistics and Data Science at Cal Poly. We discuss her path from English and French to statistics, how she builds teaching tools and navigates AI in the classroom, and what it takes to keep a programming community weird in the best possible way. Episode notes Kelly is curious, collaborative, and unafraid to lean in on quirky. Kelly shares how she balances teaching three courses with master's student supervision, applied research projects spanning Polish history and beyond, and her belief that the best part of academia is the people. We also dive into the practical and philosophical challenges of staying current in a field that reinvents itself every few years. What's inside Breakfast mixology Building Quarto extensions with JavaScript and AI When ChatGPT helps students learn (and when it doesn't) Applied stats meets history: analyzing social networks from the Polish Revolution Why remarkable, welcoming communities matter more than perfect code

  • S1 · E10
    Nov 17, 2025 · 42 min

    James Blair: Part 2 — Solutions engineering, critical thinking, and staying human

    This episode is Part 2 of our conversation with James Blair. He explains how he found his “accidental perfect fit” as a solutions engineer and how that role became a pipeline into product management. Get a peek into the AI-powered tooling he’s now building for the Posit ecosystem, and hear how he’s using Claude Code, Positron Assistant, and DataBot to generate synthetic, industry-specific demos on the fly — plus, why the real magic is keeping humans firmly in the loop. Episode notes This is a story about listening deeply to users and using AI to make that listening scale. James explains what solutions engineers actually do, how that work shaped Posit’s product team, and how synthetic data plus agents are changing the way they build demos and teach data science. What’s inside What a solutions engineer really is and why the role was such a good fit for James How solutions engineering became a natural pathway into product management at Posit Multi-agent “bot posse” workflows and why context management matters Using AI the right way and why code literacy, critical thinking, and staying human are the real superpowers in an AI-saturated world

Showing 1–20 of 21 episodes