Skip to content
Artwork for Just Now Possible
TechnologyBusinessEntrepreneurship

Just Now Possible

Teresa Torres

How AI products come to life—straight from the builders themselves. In each episode, we dive deep into how teams spotted a customer problem, experimented with AI, prototyped solutions, and shipped real features. We dig into everything from workflows and agents to RAG and evaluation strategies, and explore how their products keep evolving. If you’re building with AI, these are the stories for you.

Play
  • 20 episodes
  • fortnightly
  • Avg 1 hr 4 min
  • English
  • #30
    July 23 · 1 hr 4 min

    Building AI for Women's Health: How Hertility Combined Bayesian Diagnosis and Scan Automation

    Guests Tulsi Patel, Director of Product and Technology, Hertility Lorna Brightmore, Head of Data and AI, Hertility Jack Pickard, Head of Engineering, Hertility In this episode What makes Hertility's data set unique: seven years of linked symptoms, blood tests, and pelvic scans from over a million women How Gyn.AI uses a Bayesian network to give clinicians probability-based diagnoses instead of binary yes/no calls Why showing clinicians the reasoning behind a diagnosis—not just the label—builds trust and speeds up triage Guarding against automation bias with holdout sets and independent, fresh-eyes review Inside the scan automation pipeline: classifying ultrasound images, detecting follicles, and measuring ovarian volume more precisely than manual methods Using an agentic loop to check AI-drafted clinical letters against patient data and catch hallucinations before a human sees them The infrastructure challenge of securely piping DICOM ultrasound images from third-party scan providers into Hertility's systems How Hertility handles PII and PHI: pseudonymization, data minimization, and running models in-house on AWS Bedrock Why treating healthcare regulation as a product requirement from day one makes AI products more scalable, not slower Key Takeaways Probabilistic, transparent AI outputs build more clinician trust than binary classifications. Guardrails against automation bias are as important as the model itself. Data minimization and in-house infrastructure make it possible to build AI responsibly with sensitive health data. Treating regulation as a design constraint from day one makes AI products more defensible and scalable, not slower. Resources & Links Hertility — At-home hormone testing and reproductive health diagnostics for women in the UK and Ireland AWS Bedrock — The platform Hertility uses to run LLMs in-house under its own governance and regulatory controls PyTorch — The foundation for Hertility's in-house image classification and contouring models Chapters 00:00 Meet the Team 00:13 What Hertility Does 01:51 How Customers Access It 04:06 A Unique Women’s Health Dataset 07:03 Mission and Efficiency with AI 10:03 Why Long Assessments Convert 13:52 Before AI Workflows 16:52 Research Publications and Impact 18:48 GynAI Reducing Time to Diagnosis 21:21 Triage and Clinician Support 24:37 Keeping Patient UX the Same 26:12 Bayesian Network and Explainability 30:19 Multiple Diagnoses and Probabilities 32:37 Probabilistic Diagnosis Shift 33:50 Clinician Adoption and Workflow Fit 34:58 Communicating Medical Uncertainty 36:43 Scan Automation Overview 40:30 In House Image Analysis 44:25 DICOM Pipeline Engineering 47:30 Evals and Automation Bias 50:31 LLM Letter Guardrails 56:47 PHI Handling and Regulations 01:00:43 Infrastructure Choices and Wrap Up

  • #29
    July 9 · 1 hr 14 min

    From COVID Pivot to AI World Building: How Snapbar Reinvented the Photo Experience

    Guests Sam Eitzen, Co-founder & CEO, Snapbar Joe Eitzen, Co-founder & CPO, Snapbar Patrick Ellis, CTO, Snapbar You'll hear how they: Pivoted from physical photo booths to a cross-platform virtual product in spring 2020 using WebRTC—built from first principles out of necessity Integrated Stable Diffusion 1.5 as their first generative AI model and ran custom LoRA fine-tunes on H100/H200 GPUs to produce brand-quality outputs nobody else in their space could match Evolved from negative prompts to reasoning model long-form prompts, giving brands more precise creative and safety control Built a meta-prompting pre-processing pipeline to ensure user likenesses—including non-obvious details like disabilities—are accurately represented in generated images Designed an experiential marketing platform that lets brands "world build" at conferences, trade shows, and live events by bringing fans into branded creative worlds Added participatory user inputs through Mad Lib-style prompts and prompt injection, turning photo experiences into co-creation moments between brands and their audiences Used Claude Code and Codex to build and ship features rapidly as a small bootstrap team, and developed a four-pillar agent orchestration framework: context, tools, verification, and workflows Are building customer-facing "vibe coding" using the Claude Agent SDK so brands can configure and create experiences themselves within Snapbar's platform Resources & Links Snapbar — AI photo booth and experiential marketing platform Patrick Ellis on YouTube — Patrick's channel on AI engineering and agent orchestration Sam Eitzen on LinkedIn — Follow Sam for experiential marketing insights Claude Code — The AI coding tool Patrick and Teresa both use heavily Stable Diffusion — One of the generative image models in the mix Snapbar evaluates for its AI photo experiences Black Forest Labs / Flux — Another model family Snapbar tests and can route activations to, depending on the creative style needed Chapters 00:00 Meet the SnapBar Team 00:59 From Photo Booth Side Hustle 02:31 Pandemic Pivot to Virtual 08:06 What AI Photo Booth Does 11:03 Why It Works at Events 13:15 When AI Entered the Product 14:26 Early GenAI Tech Stack 17:43 Patrick’s Learning Path 20:44 Why They Bet on AI 28:25 Building with Necessity 34:25 Applied AI Beats Models 36:01 Prompting Got Easier 37:53 Arcades to Home Shift 38:58 Commoditization vs Application Layer 40:58 Experiential Platform Overview 42:33 World Cup Activation Examples 45:34 QR Flow and WebRTC 47:51 Capture to Display Pipeline 49:24 Dark Factory and Vibe Coding 55:38 Brand Safety Prompting 01:00:23 Meta Prompting and Representation 01:03:22 Experiential Marketing Momentum 01:06:12 Agent Orchestration Framework 01:10:02 Claude Code and Fable Rant 01:12:48 Wrap Up and Where to Find Them

  • #28
    June 25 · 55 min

    Is This Okay? How Override Labs Built a Safety-First AI Consent Coach for Teen Boys

    Guests Priya Nakra, Founder and Product Lead, Override Labs Olivia Rowley, AI Advisor and Board Member, Override Labs In this episode Why Priya left a 10-year tech career to found a nonprofit focused on gender-based violence prevention How scraping 2,000 Reddit posts per subreddit validated demand for a consent reflection tool Why Override Labs defined a "South star" — the worst-case outcome — and designed the product to avoid it How a licensed therapist and positive masculinity coaches shaped the product's tone and eval rubric Why the product never gives a "green flag" response — and what it offers instead How risk classification runs deterministically before Claude is ever called, then tailors the response by tier The three-part response structure grounded in motivational interviewing: validate, reflect, invite reflection Why "privacy by design" meant no accounts, no cookies, no cross-session tracking — and how that became a feature The challenge of measuring prevention when success means something didn't happen How Override Labs is building a product ecosystem: a women's tool, a web game as a top-of-funnel, and a future institutional API layer Resources & Links Override Labs — Building technology to prevent harm from gender-based violence The Fund to Prevent Sexual Assault — The philanthropic incubator that funded Override Labs' first cohort You can find Override Lab's products here: Is this okay? Game Boi Chapters 00:00 Meet Priya and Olivia 00:50 Why Override Labs Exists 02:11 From Tech to Mission 04:51 Is Technology Good 07:49 Narrowing to Teen Consent 11:07 Researching Teen Scenarios 14:33 Clinical Guidance and Tone 16:48 Designing for Reflection 20:51 Prototype to Product Flow 24:46 Risk Tiers and Evals 28:10 Defining the Conversation Goal 29:35 Motivational Interviewing Framework 31:37 Measuring Real World Impact 35:39 Privacy First Data Practices 40:29 Reaching Teens With Ads 41:56 Web Game Funnel Strategy 44:30 Funding and Institutional Pathways 46:50 UX for Sensitive Moments 49:34 Matching the Older Brother Tone 52:09 Building Safe AI for Prevention 54:35 Closing Thoughts and Thanks

  • #27
    June 11 · 1 hr 12 min

    Beyond Black Box Scores: How Musubi Trains Custom AI for Trust and Safety Teams

    Guests Nikki Marinsek, Data Scientist, Musubi Brian McCaffrey, Software Engineer, Musubi Dan Means, Machine Learning Engineer, Musubi In this episode Why off-the-shelf moderation scores fail and how custom-trained models fix that How Musubi combines traditional ML with LLMs for different moderation tasks The discovery that AI can outperform human moderators—and how to communicate that to clients Using AI as a judge to referee disagreements between AI and human decisions How Musubi onboards new customers with "reverse demos" What custom model training actually means: fine-tuning, feature engineering, and reusable deployment pipelines The policy optimizer: an agentic flow that helps customers iterate on their LLM moderation policies Why pushing eval tools directly to customers is a core product strategy How Musubi is building flexible orchestration workflows for non-technical trust and safety teams Resources & Links Musubi — AI-powered trust and safety toolkit for content platforms Maven AI Evals Course — The course Teresa took to learn about evals (get 35% off with Teresa's affiliate link) Chapters 00:00 Meet the Team 01:18 Why Everyone Wears Product 02:32 What Musubi Builds 04:51 AI for Human Moderation 09:59 Adversaries and Asymmetry 11:48 Early Days and Low Latency 13:35 First Prototype Slice 15:33 Traditional ML Meets LLMs 19:52 Benchmarking Against Humans 23:09 LLM as Judge and Policy Gaps 29:53 From Prototype to Platform 31:15 Customer Onboarding Reverse Demos 36:08 Custom Models Per Customer 38:05 Fine Tuning vs Training 39:14 Embedding Driven Classification 40:04 Cost and Latency Tradeoffs 43:21 Productizing Customization 49:16 Scaling Prototypes to Production 51:58 Golden Sets and Policy Loops 56:17 Coaching Customers Safely 01:02:06 Gamified Feedback Signals 01:06:19 Agentic Toolkit Roadmap 01:09:05 Workflow Orchestration Future 01:12:06 Wrap Up and Thanks

  • #26
    May 28 · 1 hr 8 min

    Building Lorikeet: How AI Humility and a Dual-Agent Architecture Are Redefining Customer Support

    Guests: Jamie Hall, Co-founder & CTO, Lorikeet Xharmagne Carandang, Product Engineer, Lorikeet Rona Wang, Product Engineer, Lorikeet In this episode: How Lorikeet evolved from failed ops tools to a full AI customer support concierge The dual-agent architecture: Concierge for customer tickets, Coach for configuration and ongoing improvement Why "AI humility" — defaulting to human handoff when uncertain — is a core design principle How Lorikeet integrates with Zendesk and Intercom instead of replacing them The UX evolution from workflow builder to conversational interface — and why the blank chat box is still hard "Resolution in the loop": how human agents unblock the AI without taking over a ticket Why guardrails need to be domain-specific — the cannabis company story How customers define their own evals and guardrails through the Coach interface Using AI to diagnose failure modes in traces and automatically suggest fixes Lorikeet's product engineering culture: every engineer asks weekly what they learned from a customer Resources & Links: Lorikeet — AI customer support concierge for enterprises in regulated industries Gradient Labs on Just Now Possible — another AI agent team in regulated financial services Neople on Just Now Possible — AI digital coworkers with a similar training-by-conversation approach Incident.io on Just Now Possible — AI SRE with multi-agent hypothesis investigation Chapters 00:00 Meet the Team 01:05 What Lorikeet Builds 02:34 Origin Story and Early Missteps 06:42 Finding the Real Support Pain 07:37 Why AI Fits Support Work 11:16 First Prototype and Early Evals 14:42 Design Partners and Selling the CLI 16:30 Product Mindset and the Real Moat 19:47 Rona Joins and Scaling Up 21:02 Milestones Voice Actions Escalation 23:48 Integrations with Zendesk Intercom 25:59 How the Agent Works Today 28:30 Coach Agent and Configuration UX 32:58 SOPs to Test Cases 34:35 Refund Flow Setup 36:12 Coach Conversational UI 38:12 Hybrid UX Guidance 40:46 Resolution in Loop 43:17 Collaboration Middle Ground 49:40 Process Maturity Limits 53:30 Confidence and Guardrails 55:59 Customer Defined Guardrails 01:01:14 Trace Diagnosis Agent 01:03:14 Product Engineers Culture 01:07:46 Closing Thoughts

  • #25
    May 14 · 1 hr 10 min

    Building Rhea's Factory: How AI-Designed Enzymes Could Finally Solve Plastic Recycling

    Guests Arzu Sandıkçı, Co-founder & CEO, Rhea's Factory Mert Topcu, Co-founder, Rhea's Factory In this episode: Why only 10% of plastic gets recycled—and why mechanical and chemical methods hit a ceiling How enzymatic recycling breaks plastic all the way back to its original monomers, unlike traditional methods that just shorten polymer chains Why enzymes are selective: they can target specific plastic types even in mixed waste streams The discovery of a plastic-eating bacteria in Japan that opened the door to enzymatic recycling How AlphaFold and the Nobel Prize in Chemistry transformed what's possible in enzyme engineering How Rhea's Factory uses protein language models (PLMs) and multi-step AI pipelines to design novel enzymes computationally The evolution from a human-orchestrated pipeline to an agentic AI scientist How guardrails at each pipeline step keep the AI pointed in the right direction without limiting exploration Why wet lab data—even just hundreds of proprietary data points—can be enough to train a powerful domain-specific prediction model Why Mert sometimes wants the model to hallucinate (and how high temperature settings help explore the full enzyme design space) The business constraint: enzymatic recycling must compete economically with cheap, oil-based plastic production What's next: a process agent, a 5,000-ton demo plant in California, and enzymes for new plastic types Resources & Links Rhea's Factory — Enzymatic plastic recycling technology AlphaFold — DeepMind's AI system for protein structure prediction (inspiration for the Nobel Prize in Chemistry) Maven AI Evals Course — The course Teresa took to learn about evals (35% off with Teresa's affiliate link) Chapters 00:00 Meet the Founders 01:50 Why Plastic Circularity 03:19 Mechanical vs True Recycling 04:52 Biology as the New Tool 07:20 Necklace and Pearls Analogy 13:22 Low Energy Reactor Process 17:33 Origin Story and PET Enzyme 22:52 Protein Folding and AlphaFold 28:32 AI Designed Enzymes 34:28 Protein Language Models Stack 37:14 Multi Step Protein Generation 39:00 Building on Foundation Models 40:50 Lab First Success Metrics 43:10 From Human to Agentic Orchestration 43:59 Problem Statements as Inputs 46:18 Guardrails at Every Stage 47:48 Prediction Models and Data Limits 50:03 Industrial Reality and Cost 52:30 Agentic Parallels and Orchestrators 57:45 Impact on Timelines and Diversity 01:03:23 When Hallucination Helps 01:04:09 Scaling Up and Process Agents 01:06:56 Enzyme Blends for Mixed Plastics 01:07:49 Why Clamshells Aren't Recyclable 01:09:34 Closing Thoughts and Thanks

  • #24
    April 30 · 1 hr 7 min

    Building AI Employees for Hospitality: How AITropos Takes Orders Where Customers Already Are

    Guests Santi Marchiori, CEO, AITropos Juan Haedo, CTO, AITropos You'll hear how they Spent two years exploring hundreds of startup ideas before finding the specific niche of AI-powered order taking in hospitality Went through three product iterations — hardware for waiters, a waiter app, and finally a customer-facing WhatsApp agent — before landing on the right form factor Identified order item identification accuracy as their single most important KPI Chose a tools-based agent architecture over MCP or pipelines to hit real-time response speed requirements Built a parallelized pipeline that searches for multiple products simultaneously and pre-fetches product context before the agent even calls a tool Use smaller, fast sub-agents to build an "immediate system prompt" that injects relevant data into each turn without extra tool calls Test with thousands of agent-simulated customer conversations run overnight before deploying to new venues Reduced new customer onboarding from three months to a few weeks — and continue to shrink it as they build domain templates Resources & Links AITropos Chapters: 00:00 Meet the Founders 00:59 What Tropos Builds 01:51 AI vs Human Touch 06:17 Restaurant Use Cases 08:16 Why Hospitality 10:47 Finding the Wedge 16:00 Early Prototypes 16:46 Hard Parts of Ordering 18:03 Speed and Channels 21:15 Iteration and Model Jumps 30:50 Customer Order Flow 35:48 Menu Discovery Question 36:07 Menus Inside WhatsApp 36:50 Finding the Chat Entry 37:37 Why Text Ordering Wins 38:30 Under the Hood Pipeline 40:54 Tools Over Workflows 45:05 Tooling and Prompt Composer 49:29 Preloading Context Fast 54:02 Founder Learning Mindset 57:21 Evaluating Order Accuracy 01:00:03 Testing and Human Takeover 01:03:56 Onboarding and Scaling Up 01:06:10 Whats Next and Wrap

  • #23
    April 16 · 1 hr

    Building Todoist Ramble: How Doist Turned Voice Braindumps into Real-Time Task Capture

    Guests Ernesto Garcia, Front-end Product Engineer, Doist Thomas Jost, Backend Software Engineer, Doist Hugo Fauquenoi, Product Manager, Doist In this episode How Doist's 2-3 month AI exploration phase led to Ramble — and why voice-to-task emerged as the top contender The user research insight behind Ramble: people using pen and paper or ChatGPT voice to brainstorm tasks before committing them to Todoist Why Ramble skips transcription entirely and processes raw audio directly with a Gemini live audio model How the model makes tool calls (add task, edit task, delete task) in real time while the user is still speaking — no text output at all Designing for the driving use case: sound effects as audio confirmation cues alongside visual task cards The challenge of teaching an LLM to capture tasks literally without over-interpreting or doing them — and how temperature tuning played a role Date handling complexity: injecting the current date, normalizing to days vs. months, and always outputting dates in English for the natural language parser Building an LLM-judge eval system with 20+ language recordings from 100+ employees across 35 countries to catch prompt regressions Why Doist chose to inject the full project/label list into the system prompt instead of building a RAG pipeline — and why it worked How easy correction beats perfect first-time accuracy in natural language interfaces What's next: multimodal task capture from images and text blobs, Apple Watch support, and automation integrations Resources & Links Todoist Doist Google Vertex AI (Gemini) Chapters: 00:00 Meet the Doist Team 01:40 What Doist Builds 02:27 Ramble Voice to Tasks 04:16 Why Voice Matters 07:42 Brain Dump Insight 09:46 Prototyping With LLMs 11:08 Live Audio Workflow 14:32 Driving Friendly UX 18:47 Tool Only Architecture 26:06 Evals and Multilingual Testing 28:41 Taming Dates and Time 33:28 Fixing Date Confusion 33:43 Defining Task Boundaries 34:34 Capture Versus Do 37:17 Tuning Creativity Levels 39:01 Evals Across Languages 41:23 Feedback and Regressions 44:09 Model Upgrades Over Time 46:33 Projects Labels Context 51:40 Handling Ambiguous Names 54:23 Whats Next Multimodal 58:48 From Capture to Execution 59:46 Closing Thoughts

  • #22
    April 2 · 1 hr 10 min

    Building Banani: How a Canvas-First AI Designer Is Raising the Floor on Product Design

    Guests Vlad Solomakha, CEO & Co-founder, Banani Vova Parkhomchuk, CTO & Co-founder, Banani Vlad Ostapovats, Founding Growth, Banani In this episode Why Banani started as a Figma plugin and what they learned from early organic distribution The canvas-first approach: why Banani is built around a design canvas rather than a chat interface How their agent architecture splits prompts into surgical edits instead of regenerating full screens The "gulf of specification" problem and what Banani is building to help agents and designers speak the same visual language Managing context across canvases with hundreds of screens — per-screen history with shared project context Why Banani doesn't compile running applications — just HTML/CSS mockups — and how that shapes everything How they evaluate design quality without traditional evals: spinning up 10 screens from one prompt to compare models Their approach to building at the edge of what's possible: identifying which model limitations to work around vs. wait out The role of context engineering and specialized agent tools in producing tasteful, high-quality design Resources & Links Banani TL Draw Chapters 00:00 Meet the Founders 01:12 What Bonani Builds 02:18 Why an AI Designer 03:40 Raising the Design Floor 06:23 Why AI Was Finally Ready 10:48 First Prototype Figma Plugin 14:10 Early Growth and Distribution 15:25 Standing Out in a Crowded Market 20:13 Product Tour Canvas First AI 23:40 Autopilot vs Manual Control 27:07 Tech Behind High Quality Design 32:08 Craft Beyond 80 Percent 33:40 Gulf of Specification 36:44 Proactive Agent Interviews 38:40 Canvas First UX Choices 42:54 Agent Architecture Under Hood 48:48 State History Context Tricks 52:32 Tooling Context Engineering 56:04 Navigating Busy Canvases 01:00:13 Betting on Model Progress 01:03:47 Shipping Around Imperfections 01:07:20 Try Banani and Next Steps 01:07:52 Building the Banani MCP 01:09:19 Final Thanks and Wrap

  • #21
    March 19 · 1 hr 6 min

    Building Agent Studio: How Medable Is Using Agentic AI to Accelerate Clinical Trials

    Guests Luke Bates, Product Leader (Agent Studio), Medable Jen Brown, Product Manager, Medable Matt Schoolfield, Product Designer, Medable Fiachra Matthews, Principal Architect, Medable What we cover in this episode: What Medable does: enabling global clinical trials across 100+ languages and accelerating drug-to-market timelines The two agents built on Agent Studio—ETMF (document classification) and CRA (clinical data monitoring)—and the problems they solve Why Medable chose a platform approach to agents instead of one-off builds How Agent Studio works: models, skills, knowledge bases, MCP connectors, versioning, and trigger types Three deployment models: Medable-built products, services-led custom builds, and self-serve platform access RAG approaches at scale: embeddings vs. markdown hierarchies vs. just-in-time MCP retrieval How they built a unified ontology layer to map terminology across 13 different clinical data systems Why they built custom MCPs with an authentication and credentialing wrapper Context window management with sub-agents and automatic tool filtering Evaluation design in a GXP-regulated environment: golden datasets, production monitoring, and the challenge of human feedback as ground truth How they document agent intent → specification → test evidence to satisfy regulatory bodies The "full self-driving" vision for clinical trials and what it would take to get there Resources & Links Medable - Clinical trial platform powering Agent Studio Chapters 00:00 Meet The Medable Team 01:14 Medable Mission And Scope 03:27 Agent Studio Platform Overview 06:29 ETMF Document Automation 08:47 CRA Agent For Monitoring 10:40 Clinical Trial Workflow Primer 14:34 Why Build A Platform 17:51 Learning AI As A Team 21:47 Early Days Of Agent Studio 23:15 How Agents Are Built 25:15 Customer Adoption And UX 30:00 Skills And MCP Standards 31:15 Scaling Context Retrieval 33:07 RAG Patterns And Tradeoffs 34:48 Ontology Data Layer Explained 38:01 Customer Friendly Agent Setup 42:19 MCP Security And Connectors 44:36 Tool Bloat And Subagents 50:44 Evals For Reliable Agents 54:40 Human Feedback Isn’t Truth 57:43 GXP Compliance For Agents 01:03:34 Full Self Driving Trials

  • #20
    March 5 · 1 hr 4 min

    Building GitHub for Product Management: How Momental Uses AI to Find Merge Conflicts in Strategy

    Guests Matthias Kleverud - Co-Founder, Momental Charlotte Kleverud - Co-Founder, Momental What we cover in this episode: What "GitHub for product management" means: finding merge conflicts in strategy, not code The product chain: signals → learnings → decisions → principles, and how AI maps it Three trees that model an organization: the product tree (OKRs to epics), the wisdom tree (decisions and their reasoning), and the people/time tree How a document processing agent uses OODA-loop thinking to extract and connect context across documents Why traditional chunking and RAG breaks down at scale and what Momental does instead The origin story: building a team of AI agents in 2024, only to discover agents hit the same alignment problems as humans Starting in 2022 with DaVinci 002 and learning that the market wasn't ready for AI-assisted product thinking How conflicts are detected, auto-resolved, or escalated to humans with merge options Why metadata—who said it, when, and in what context—is critical to preventing hallucinations The self-improving agent: collecting user feedback weekly and rewriting its own prompts Moving from chat-first to UI-first to proactive agents as an AI product design pattern Design partner strategy and what's next for Momental's public launch Resources & Links Momental - GitHub for product management Spotify - Where both founders started their PM careers Claude Code - AI coding tool discussed in the conversation Perk episode on Just Now Possible - Referenced episode about eliminating shadow work Chapters 00:00 Meet The Founders 01:14 GitHub For PMs Explained 03:19 Strategy Merge Conflicts 06:49 Product Chain Model 09:49 Capturing Context Fast 12:17 Context Graph And RAG Limits 16:52 Origin Story Since 2022 20:01 From Agent Team To Foundation 25:26 Three Trees Of Context 28:42 Two Agents Secret Sauce 31:37 How Document Processing Works 34:55 Agents Ask Better Questions 35:41 Human In The Loop Context 36:38 Data Models And Graphs 39:21 Beyond Documents And Vectors 42:25 Specialized Tools Win 44:50 Quarterly Planning At Scale 49:38 Discovery Versus Vibe Coding 51:00 Tree Building And Conflicts 53:49 UI Over Chat Interfaces 56:01 Proactive Agents In Practice 58:00 Quality Evals And Feedback 01:00:22 Launch Plans And Mission 01:03:19 Eliminating Shadow Work 01:03:56 Closing Thanks

  • #19
    February 19 · 1 hr 3 min

    Building AI Sales Reps: How ShowMe Orchestrates Voice, Video, and Multi-Agent Workflows to Close Deals

    Guests Yuri Vela Tulopov -- Co-Founder and CEO, ShowMe Quique Gomez -- Co-Founder and Lead Product Engineering, ShowMe What we cover in this episode: How ShowMe builds AI digital workers that function as inbound sales reps The origin story: spotting the conversion gap at a previous company and realizing AI could fill it Why the first MVP was a voice agent with product videos and a simple RAG knowledge base Adding a realistic avatar via HeyGen and how it changed user engagement through better affordances Decomposing a single sales conversation into multiple specialized sub-agents (greetings, qualifying, pitching) The three agent types: conversation agents, evaluator agents, and creator agents How deterministic workflows manage the lead-to-close journey across days Building toward a smart orchestrator agent that breaks out of rigid workflow paths Ingesting sales transcripts and training materials to teach agents company-specific sales skills Customer-driven evaluation loops that start at 100% review and taper to ~5% over time Creating automated tests from customer feedback to prevent prompt regression Confidence scoring and frustration detection for real-time human handoff decisions Treating the agent as a coworker: onboarding via Slack, weekly reporting, CRM integration Future plans: self-serve PLG motion, smart orchestration, and expanding to customer success Resources & Links ShowMe - AI digital sales reps for inbound teams HeyGen - AI avatar platform used for ShowMe's video calls Chapters 00:00 Meet the Founders: Juri & Kike Introduce Show Me 00:45 What Show Me Builds: AI Sales Reps as Digital Coworkers 02:17 Why Inbound-First Sales Agents (and Not Outbound Spam) 03:51 Origin Story: The Website Conversion Problem That Sparked the Idea 08:11 MVP Launch: Voice + Video Product Demos in Two Weeks 10:45 Bootstrapping the Knowledge Base: Videos, Docs, Scraping & RAG 11:54 Beyond Demos: Multi-Stage Buyer Journeys, Follow-Ups & Orchestration 14:34 Building Trust: Avatars, Video-Call UX, and AI “Affordances” 20:18 Whiteboard Architecture: Agents, Workflows, and the Orchestrator Layer 29:43 Where Agents Run: Creators, Evaluators, and Breaking Deterministic Flows 32:18 Sales Is High-Stakes: Personalization vs. Hallucinations & Revenue Risk 33:35 Conversation Agent Evolution: From Q&A Bot to Guided Sales Discovery 34:10 Why One Agent Becomes Many: Decomposing Stages for Latency & Memory 36:36 Orchestrator + Tooling: Routing Between Greeting, Qualify, Pitch, Next Steps 38:46 Teaching Sales Skills: Generic Prompting vs Company-Specific Playbooks 42:52 Ingesting Real Calls & Onboarding Like a Teammate (Transcripts, Training Docs) 45:05 Real-Time Voice + Avatar Demos: Latency Tricks and Video Clip Libraries 47:33 Creator vs Evaluator Agents: Data Cleaning, Custom Fields, Sentiment & Confidence 49:15 Human Handoff Guardrails: When Confidence Drops or Users Get Frustrated 50:15 Proving Quality in Production: POCs, A/B Rollouts, Dashboards, and CRM Logging 53:21 Evals at Scale: Customer Feedback Loops, Regression Tests, and the 5% Review Set 58:19 What’s Next: Smarter Orchestration, Self-Serve Setup, and More Digital Workers 01:02:07 Closing Thoughts: Customer Insight as the Moat in a Fast-Moving AI World

  • #18
    February 5 · 1 hr

    Building Earmark: How a Two-Person Team Turned Meetings into Finished Work

    Guests Mark Barbir – CEO, Earmark Sanden Gocka – Co-Founder, Earmark What we cover in this episode: How Earmark differs from generic AI notetakers by producing finished work, not just summaries The pivot from Apple Vision Pro presentation coaching to a web-based meeting assistant Running multiple agents in parallel during live meetings Template-based agents: Engineering Translator, Make Me Look Smart, Acronym Explainer Personas that simulate absent team members (security architect, legal, accessibility) Why ephemeral mode (no data storage) became a selling point for enterprise Reducing AI costs from $70/meeting to under $1 through prompt caching Why GPT 4.1 still beats newer models for prose quality in their use case The limits of vector search for analysis questions across meetings Building agentic search with multiple retrieval tools (RAG, BM25, metadata queries, bespoke summaries) Designing for product managers as the extreme user to solve for everyone Their vision for an AI chief of staff that goes beyond automating deliverables Resources & Links Earmark — Productivity suite where the work completes itself ProductPlan — Roadmapping tool where both founders previously worked Granola — AI notetaker mentioned for comparison Assembly AI — Speech-to-text service used by Earmark OpenAI API — LLM provider with prompt caching support Cursor — AI code editor with build integration in Earmark V0 by Vercel — AI prototyping tool with build integration in Earmark Chapters 00:00 Introduction to Earmark Founders 00:28 Background and Experience 01:05 What Does Earmark Do? 01:23 AI and Productivity 03:09 Comparing Earmark to Competitors 03:41 Earmark's Unique Features 05:53 Templates and Personas 10:06 Technical Details and Development 17:12 Early Product Versions and Challenges 28:44 Understanding Prompt Caching 29:49 Managing Multiple Tools and Costs 30:59 Optimizing Transcript Summarization 35:11 Challenges with Context and Reasoning Models 38:10 Innovative Search and Retrieval Techniques 44:06 Creating Actionable Artifacts from Meetings 48:30 Ensuring Quality and Managing Hallucinations 58:20 Future Vision for AI Chief of Staff

  • #17
    January 22 · 49 min

    When Trust Is Everything: Building AI for Physicians at Healio

    Guests Jennifer Deal – SVP of Product Development, Healio Casey Utley – Senior UX Designer, Healio Matthew Skepner – VP of Technology, Healio What we cover in this episode: Why physicians need AI at the point of care—and how they actually use it (hint: it's preparation, not bedside) The surprising discovery that physicians wanted help with patient communication and empathy, not just clinical answers Building a working prototype in a weekend with Cursor after starting with Figma mockups How Healio's RAG system combines lexical search, vector search, and semantic search across multiple trusted sources Why "just use PubMed" isn't simple—five different ways to access the same data, each with trade-offs Designing citations that physicians trust: subscripts, hover states, and progressive disclosure Serving contextual ads while the LLM processes queries—a practical monetization approach HIPAA compliance and input guardrails for masking personal health information Eight LLM judges for evals: safety, medical accuracy, faithfulness, relevancy, completeness, reasoning, clarity, and overall quality Why physician feedback trumps LLM-as-judge feedback in high-stakes medical contexts The role of the Healio Innovation Partners in ongoing discovery and validation Resources & Links Healio — Medical news, education, and clinical guidance for healthcare professionals PubMed — Database of biomedical literature Cursor — AI-powered code editor used to build the prototype Chapters 00:00 Introduction to Healio Team 01:00 Overview of Healio's Services 01:57 Introducing Healio AI 03:39 Addressing Physician Needs with AI 05:45 Building Trust in AI Solutions 13:56 Prototyping and Testing Healio AI 18:02 Refining the AI Product 21:48 Technical Architecture and Advertising Integration 25:16 Balancing Speed and Accuracy in AI Responses 26:30 Ensuring Credible and Trustworthy Content 27:41 Challenges in Data Integration and Web Crawling 29:00 Optimizing Search Strategies for Different Data Types 31:09 User Interface and Trust Building 34:31 Human Feedback and Continuous Improvement 35:41 Guardrails and Evaluations for Reliable AI 39:11 Experimenting with LLM as Judges 45:13 Future Directions and User-Centric Design

  • #16
    January 15 · 1 hr 5 min

    Building Tendos AI: How an Agent Swarm Turns Construction Emails into Quotes

    Guests Daniel Kappler — CPO (Product & Design), Tendos AI Matthias Hilscher — CTO (Engineering), Tendos AI Key Takeaways Start narrow to prove value: Tendos AI began with just radiators for one design partner before expanding to all building products Own the interface: building a web application (vs. integrating into legacy systems) gave them control over UX and the ability to iterate toward full automation Evaluate each agent, not just the chain: per-agent evals make debugging tractable and show exactly where performance changed Use review agents: a separate agent that checks work (like code review) catches errors before they reach humans Let customers pull you: customers asked Tendos to replace their CPQ software—strong signals of product-market fit Topics Covered The tendering chain in construction and why it's ripe for automation How domain expertise (CEO's construction background) helped identify and validate the opportunity Entity extraction from PDFs ranging from 1 page to 1,800+ pages Planning patterns in agentic systems—creating and updating plans based on findings How agents evaluate product fit against customer requirements Building custom tracing and observability tools for complex agent chains The path toward self-learning systems through human feedback loops Links & Resources Tendos AI Chapters 00:00 Introduction to Tendo and Key Roles 01:01 Understanding the Tendering Chain 02:26 Real-World Construction Analogy 03:34 Challenges in the Construction Industry 04:48 AI's Role in Tendo's Product 12:59 Early Prototypes and AI Integration 18:31 Expanding Product Capabilities 28:56 Customer Collaboration and Workflow Automation 33:15 Strategic Partnerships and Technical Groundwork 34:20 Focusing on Specific Customer Segments 36:03 Product Evolution and Current Capabilities 38:17 Technical Workflow and Automation 40:12 Evaluating and Matching Product Requests 47:00 Dynamic Agent Architecture 55:29 Quality Measures and Evaluation 01:02:59 Future Directions and Customer-Centric Development

  • #15
    January 8 · 1 hr 9 min

    Building a Career Co-pilot for Disadvantaged Students: How Zero Gravity Bridges Knowing and Doing

    Guests Elliot Little, Product Manager, Zero Gravity Dan St. Paul, Software Engineer, Zero Gravity What we cover in this episode Zero Gravity's mission: breaking down barriers to elite careers for disadvantaged UK students The "knowing-doing gap"—why students struggle to act even when they know what to do Why their first prototype (a job suitability summary) didn't create the "wow moment" they expected The decision to use text chat over voice input and why guided prompts beat empty text boxes Context management techniques: removing stale tool calls, summarizing history, exposing tools conditionally Using different models for different tasks (GPT-5 Nano for structured outputs, lighter models for quick replies) Safeguarding architecture: moderation endpoints plus external verification with Unitary Building a failure taxonomy through internal red team/green team exercises What's next: long-term memory management for multi-year student journeys Links & References Zero Gravity Unitary – AI-powered content moderation Blue Dot Impact AI Safety Course – free AI safety course Elliot recommended Chapters 00:00 Introduction to Dan and Elliot 00:45 Zero Gravity's Mission and Impact 02:14 Introducing the AI Career Co-Pilot 04:01 Challenges Faced by Disadvantaged Students 06:49 Zero Gravity's Mentorship Program 09:14 Building the AI Career Co-Pilot 12:01 Early Prototypes and User Feedback 17:05 Refining the AI Career Co-Pilot 37:36 Introduction to Career Co-Pilot 38:02 Current Student Interactions 40:22 Technical Deep Dive 42:14 Context Management Challenges 44:43 Tool Call Optimization 51:48 Safeguarding and Moderation 57:52 Evaluating AI Performance 01:04:09 Future Directions for Career Co-Pilot 01:07:52 Concluding Thoughts

  • #14
    Dec 18, 2025 · 1 hr 1 min

    Automating the Full Customer Support Iceberg: How Gradient Labs Built a Multi-Agent Platform

    Guests** Jack Taylor, Product Engineer, Gradient Labs Ibrahim Faruqi, AI Engineer, Gradient Labs In this episode The iceberg metaphor: why frontline support is only the tip of automation potential How three agent types (inbound, back office, outbound) coordinate on complex tasks like fraud disputes Natural language procedures that let subject matter experts train agents without engineering bottlenecks The "turn" architecture: state machines that orchestrate agent logic across async, multi-day conversations Skills as modular agent capabilities—and how they're scoped deterministically per turn Defining "done" for outbound agents when the customer isn't the one ending the conversation Guardrails as classification problems: balancing recall and precision for regulatory compliance Ask a Human: a tool call that brings humans into the loop for approvals or missing APIs Auto-eval pipelines that flag conversations for manual review and feed labeled datasets Links & References Gradient Labs Incident.io episode – Referenced in the conversation Chapters 00:00 Meet the Engineers: Jack and Ibrahim 00:39 The Role of Product Engineers in Tech 01:21 Introduction to Gradient Labs 02:11 The Three Pillars of Customer Support Automation 04:32 The Evolution and Growth of Gradient Labs 05:29 Building and Refining AI Agents 06:39 Outbound Agent: Addressing Customer Problems 09:12 Defining Success in Outbound Procedures 17:08 Ensuring Compliance and Guardrails 30:17 Understanding Agent Guardrails 31:54 Complexities of Natural Language Input 36:21 Skill Design and Management 39:53 Deterministic Skill Execution 41:54 Customer-Specific Guardrails 44:21 APIs and Customer Tools Integration 46:02 Ask A Human Tool 48:24 Guardrails as Classification Problems 57:12 Auto Eval System 59:12 Future of Multi-Agent Systems

  • #13
    Dec 11, 2025 · 1 hr 7 min

    Building Mowie: How a Concierge Service Became an AI Marketing Platform

    Guests Chris O'Connor – CEO, Mowie Jessica Valenzuela – Co-Founder, Mowie What we cover in this episode How Mowie evolved from a concierge marketing service to an AI-powered platform The "document hierarchy" architecture: how Mowie builds and maintains context about each business Why they moved from structured schemas to loosely structured markdown for intermediate processing Using Simon Sinek's Golden Circle framework to validate early product-market fit How Mowie generates quarterly content calendars and weekly posts across email and social media The three mini-calendars: public events, business-specific events, and recommended campaigns Building traceability so customers can see which context documents influenced their content Using customer approvals, edits, and regeneration requests as lightweight evals Connecting marketing performance back to point-of-sale data for attribution What's next: deeper attribution, omnichannel expansion, and digital out-of-home displays Resources & Links Mowie AI — AI marketing platform for SMBs Simon Sinek's Golden Circle — The framework Mowie used for early validation Chapters 00:00 Introduction to Mowie AI 00:06 Meet the Founders: Chris and Jessica 00:45 Understanding the Target Customers 01:35 Challenges Faced by SMBs in Marketing 04:20 The Evolution of Mowie AI 05:40 How Mowie AI Works for Businesses 08:16 Onboarding and Data Collection 14:56 Content Strategy and Calendar Creation 22:29 Technical Challenges and Solutions 29:41 Iterative Development and Feedback 37:23 Automated Content Calendar Approval 39:19 Content Calendar Creation Process 40:14 Marketing Pillars and Campaigns 42:11 Transparency and Traceability in Recommendations 45:23 Generating Weekly and Quarterly Calendars 48:24 Customizing Campaigns for Specific Events 54:15 Evaluating Campaign Performance 58:48 Customer Feedback and Document Hierarchy

  • #12
    Dec 4, 2025 · 54 min

    From Prototype to Production: How Perk Built a Voice AI Agent That Makes 10,000 Calls a Week

    Guests Steven Payne, Product Manager, Perk Gabriel Stock, Senior Engineering Manager, Perk Philipe Steiff, Senior Software Engineer, Perk What we cover in this episode How Perk's team identified an AI use case by connecting prior experimentation with a real operational problem Why they chose Make.com for prototyping—and shipped to production without touching backend code The evolution from a single prompt to structured conversation stages (IVR handling, booking confirmation, payment request) How breaking up the agent's task dramatically improved reliability Building two eval systems: classification for success rates and LLM-as-judge for conversational behavior Why the team still listens to calls manually even with automated metrics The challenge of prompt engineering for voice: numbers, booking references, and text-to-speech markup Lessons learned from expanding to German (prompts in native language improve results) How this project uncovered other operational problems they didn't know existed Resources & Links Perk Make.com – No-code automation platform used for the prototype Twilio – Voice/telephony provider 11 Labs – Text-to-speech provider (used in early experiments) Chapters 00:00 Introduction to the Team 01:54 Understanding PERK's Mission 02:59 Challenges in Travel Booking 07:27 AI Solutions for Customer Care 09:52 Prototyping with AI and Voice 17:00 Implementing AI in Production 25:51 Learning Through Trial and Error 26:40 Prompting Challenges and Solutions 27:58 Iterating on Prompts and Evaluations 30:08 Scaling and Production Challenges 32:43 Advanced Evaluation Techniques 35:32 Real-World Applications and Success 49:07 Future Directions and Expansion 53:53 Conclusion and Team Reflections

  • #11
    Nov 20, 2025 · 1 hr 6 min

    Building an AI Sleep Coach: How Rest is Making CBTI Principles Accessible to DIY Sleep Hackers

    Guests Martin Siniawski, CEO and co-founder, Rest Ignacio, CTO, Rest You'll hear how they: Discovered the sleep use case from podcast app user behavior (10% of users, but high willingness to pay) Used jobs-to-be-done research to identify "DIY sleep hackers" as an underserved segment Chose CBTI (Cognitive Behavioral Therapy for Insomnia) as their foundation—a clinically proven approach with 80% efficacy Evolved from text chatbot to voice-first AI using Vapi for voice and OpenAI for reasoning Built a memory system that remembers user context (like traveling, having a dog) with time-based relevance Created dynamic agendas that drive daily conversations based on sleep data, program stage, and user compliance Managed parallel development paths (text via OpenAI Assistants and voice via Vapi) Moved from massive system prompts to RAG for general sleep knowledge, keeping user data in prompts Navigated wellness vs. medical product positioning with clear guardrails against diagnosis and medication advice Used weekly error analysis with domain experts (sleep therapists) to drive product iterations Built LLM-powered evals for safety boundaries and experimented with Hamming for voice testing Resources & Links Rest – AI sleep coach app Vapi – Voice agent platform Rest uses Langfuse – Observability and evals platform Hamming – Voice testing platform AI Evals Maven Course by Hamel Husain and Shreya Shankar (Get 35% off with Teresa's affiliate link) Chapters 00:00 Introduction to Rest and Its Founders 00:33 The Origin Story of the AI Sleep Coach 02:07 Exploring the Podcast App and Sleep Use Case 03:35 Transitioning to a Dedicated Sleep Audio App 05:47 Understanding User Segments and Sleep Challenges 07:45 Introduction to the AI Sleep Coach 13:14 The Role of Voice in the AI Sleep Coach 18:46 Daily User Interaction and Features 21:30 Prototyping and Early Learnings 28:09 Navigating Ethical and Regulatory Concerns 30:39 Navigating the Line Between Health and Wellness Apps 31:00 Incorporating Adjacent Disciplines into the App 32:15 The Power of 24/7 Availability 32:53 Evolution of the Chatbot and Error Analysis 34:49 User Experience Improvements and Voice Integration 46:49 Implementing Memory and Personalization 50:18 Dynamic Agenda and User-Centric Conversations 57:37 Evaluation and Guardrails 01:00:05 Future Roadmap and Enhancements 01:03:38 Combining Data Layers for Enhanced AI 01:06:00 Conclusion and Final Thoughts

Showing 1–20 of 20 episodes