Skip to content
Artwork for Two Voice Devs

Two Voice Devs

Mark and Allen

Mark and Allen talk about the latest news in the VoiceFirst world from a developer point of view.

Play
  • 20 episodes
  • Avg 26 min
  • English
  • S1 · E279
    August 13 · 26 min

    From Google Glass to the Next Wearables Wave

    In this episode, Allen Firstenberg is joined by guest host Cecilia Abadie, a computing pioneer who has been at the forefront of every major tech wave—from personal computers to mobile, wearables, and now AI. They look back on their shared roots at the 2012 Google Glass Foundry, trace Cecilia's journey through enterprise eyewear at Tesla and Boeing, her time at Google X/Android XR, and her transition to founding 33 Labs to research the frontier of AI and smart glasses. Cecilia shares her recent experiences presenting at the Meta Wearables Summit, discussing the electric, yet cautious, developer vibe and the distinct communities—enterprise and accessibility—leading the charge. They dive deep into the debate between immersive headsets and lightweight, assistive smart glasses, and discuss how conversational AI (such as OpenAI's live voice mode and Google's Gemini) is redefining voice as a primary interface. They explore the critical challenges facing today's developers, from the lack of a cohesive developer story and monetization models to Cecilia's project "Halos," designed as mini-apps or skills tailored specifically for voice-first interactions. Whether you are a developer excited about Google's Android XR, Meta's wearables ecosystem, or interested in the future of human-AI collaboration, this episode is packed with invaluable, real-world perspective on the past, present, and future of intelligent eyewear. Learn More: * https://33labs.org/haloField * https://www.youtube.com/@MultithreadedReality [00:00:00] Welcoming guest host Cecilia Abadie [00:00:40] Cecilia's background: From Uruguay to early technology waves [00:03:50] Life after Google Glass: Genie, Lynxfit, and enterprise eyewear [00:07:06] Launching 33 Labs and presenting at the Meta Wearables Summit [00:08:56] The developer vibe at the Meta Wearables Summit [00:10:40] Immersive vs. Assistive eyewear and the resurgence of voice [00:13:05] The missing developer story and monetizing conversational AI [00:15:34] Introducing "Halos" as mini-apps for conversational interfaces [00:19:39] Highlights from the summit: Carbon tracking and accessibility [00:22:05] Comparing SDKs: Google's Android XR vs. Meta's DAT Wearables Kit [00:24:32] Outro and where to connect with Cecilia #VoiceFirst #AIAgents #GenerativeAI #SmartGlasses #IntelligentEyewear #VoiceUX #GoogleGlass #AndroidXR #MetaWearables #33Labs #OpenAI #WearableTech Episode 279

  • S1 · E278
    August 6 · 26 min

    AI Agents: Six Lessons from Six Years of Two Voice Devs

    Happy 6th Anniversary to Two Voice Devs! In this milestone episode, Mark Tucker and Allen Firstenberg look back at six years of podcasting and discuss how the industry is coming full circle. We started in the era of hardware assistants like Alexa and Google Assistant, shifted into text-based LLMs, and are now witnessing the return of voice-first interfaces through smart glasses and conversational LLM voice modes. As developers rush to build the next generation of AI agents, are we repeating the painful mistakes of the past? Mark and Allen share six critical lessons today's agent developers must learn, covering the abysmally poor developer-to-consumer discovery experience, the puzzle of monetization for independent creators, the true meaning of "voice first, not voice only," the absolute necessity of concise responses, the lost art of crafted entertainment over just-in-time generation, and why we need asynchronous interactions modeled after the Star Trek computer. If you're thinking about the next wave of agents in all sorts of form factors and modalities, this is the episode for you to watch. And we'd love to hear your take on what we've learned and what we still need to learn. [00:00:00] Celebrating six years of Two Voice Devs! [00:01:00] The full circle return of voice-first LLM interfaces [00:03:59] Lesson 1: The discovery and installation bottleneck [00:08:59] Lesson 2: The monetization puzzle for indie developers [00:10:48] Platform plays: Android's advantage vs. Amazon's closed beta [00:16:50] Lesson 3: Designing for "voice first, not voice only" [00:18:45] Lesson 4: LLMs are too verbose—managing output conciseness [00:20:09] Lesson 5: Tailored entertainment and crafted storytelling [00:23:13] Lesson 6: Latency, response times, and the Star Trek computer [00:24:42] Outro and looking ahead to another year #VoiceFirst #AIAgents #GenerativeAI #SmartGlasses #IntelligentEyewear #VoiceUX #AmazonAlexa #GoogleAssistant #Gemini #AndroidXR #OpenAI Episode 278

  • S1 · E277
    July 16 · 22 min

    The Intersection of AI, Fashion, Design, and Development

    In this episode, Allen Firstenberg welcomes GDE Margaret Maynard-Reid as guest host to discuss the exciting intersection of AI, art, design, and fashion. Margaret shares her hands-on experiences testing Google's Gemini Omni Flash and Veo models, highlighting the game-changing capabilities of conversational video editing. We dive into her creative workflows, such as composing initial visuals using NanoBanana before adding motion, and storyboarding for short videos. Margaret also showcases several of her open-source projects (from fashion mood boards to agentic design workflows using Antigravity 2), and provides essential advice for developers prototyping with AI. Find Margaret's work and open-source projects at: https://margaretmz.me/ [00:00:00] Welcome and Introduction [00:00:22] Margaret's journey: AI, art, and fashion design [00:03:38] Hands-on with Gemini Omni Flash and Google Veo [00:05:58] Conversational editing and fashion use cases [00:07:41] Image-to-video workflow & storyboarding [00:12:43] UI platforms vs. writing custom developer code [00:13:49] Open source projects: Mood boards, Met Museum, Antigravity [00:17:14] Tips for developers: Choosing models & solving real problems [00:21:00] Outro and contact information #FashionAI #GenMedia #GeminiOmni #Veo #NanoBanana #AIArt #AIDesign #Antigravity #GenerativeAI #TwoVoiceDevs Episode 277

  • S1 · E276
    July 10 · 29 min

    What is an Agent Harness?

    In this episode, Allen Firstenberg and Sam Witteveen dive into one of the newest and most discussed concepts in the developer community: the "Agent Harness." What exactly is a harness, and how does it transform a simple, one-shot AI model into a truly capable, autonomous agent? Sam and Allen demystify this new paradigm, tracing its evolution from early frameworks like LangChain and LangGraph to the modern, bespoke, "off-the-rails" architectures powering tools like Claude Code, Hermes Agent, and Antigravity. They explore essential best practices—including sandboxing, persistent file access, agent loops, planning tools, and fine-grained security—and weigh the critical trade-offs between highly customizable single-tenant deployments and scalable, cloud-managed multi-tenant agent infrastructures. Whether you are prototyping with managed APIs or writing custom scaffolding in Python, Go, or Rust, this episode provides a clear map of the shifting landscape of agent engineering. [00:00:00] Catching Up & the Summer of AI Evolution [00:01:07] What is an Agent Harness? Scaffolding the Model [00:02:03] The Shift from Rigid Frameworks to Bespoke Code [00:04:12] Best Practices of Modern Agent Harnesses [00:05:26] Autonomous Agents and 'Off the Rails' Evolution [00:09:48] Security, Sandboxing, and Environment Lockdown [00:13:04] The Core Agent Loop and Formalization [00:14:53] Exploring Modern Harnesses: Hermes, Claude Code, and Antigravity [00:17:04] Managed Agents, SDKs, and Cloud Infrastructure [00:23:56] Single-Tenant Customization vs. Multi-Tenant Scaling [00:28:08] Wrap-Up & Where to Find Sam #AIAgents #AgentHarness #ClaudeCode #Antigravity #SoftwareEngineering #LangChain #OpenClaw #HermesAgent #AIProgramming #TechPodcast #TwoVoiceDevs Episode 276

  • S1 · E271
    June 25 · 14 min

    Set the scene with Gemini TTS

    Roll tape and prompt! In this episode of Two Voice Devs, Allen and Mark explore how Google’s new advanced prompting guidelines turn developers into voice directors for Gemini Text-to-Speech. Instead of coding rigid SSML tags, you can now establish a scene, write stage directions, and give "director's notes" to shape a base voice's gender, accent, style, and pacing. Allen showcases a web app where he directs a single base voice—to play two entirely different characters: a rough Brooklyn cab driver and a classic Southern belle. The hosts discuss using natural language audio tags as cues for laughter, sighs, gasps, and more, and how these theatrical controls are coming alive in real-time with Gemini Live and Gemini 3.1 Flash TTS. Learn more: * https://ai.google.dev/gemini-api/docs/speech-generation [00:00:05] Welcome to Two Voice Devs [00:00:27] Intro to Gemini Text-to-Speech and Advanced Prompting [00:01:57] Moving Beyond SSML to Flexible Base Voices [00:03:07] Prompting Genders and Accents (The Storytelling Analogy) [00:04:40] Web App Demo: Zephyr as a Brooklyn Cab Driver vs. Southern Belle [00:06:50] Building Multi-Voice Conversations with Stage Directions [00:08:41] Using Natural Language Audio Tags for Expressive Cues [00:11:02] Gemini Live Integration and Dynamic Tone Selection [00:12:27] Model Details: Gemini 3.1 Flash TTS Preview and Release Info [00:13:53] Wrap-up and Call for Feedback Hashtags: #GeminiTTS #TextToSpeech #GenerativeAI #GoogleDeepMind #GeminiLive #GeminiFlash #AIStudio #DeveloperTools #SpeechSynthesis #VoiceFirst #AdvancedPrompting Episode 275

  • S1 · E274
    June 11 · 15 min

    Project Solara: Welcome to Agent-First Hardware

    After months of conferences and busy schedules, Mark Tucker and Allen Firstenberg return to discuss Microsoft’s surprising Build conference announcement: Project Solara. Moving from the legacy voice-first consumer world of Amazon Alexa and Google Assistant, Microsoft is pioneering a secure, business-focused "Agent-first" platform. In this episode, we unpack Microsoft's two new concept devices, a desktop smart display and a wearable camera-equipped badge, and explore the Android Open Source Project (AOSP)-based platform behind them: the Microsoft Device Ecosystem Platform (MDEP). We discuss how Project Solara integrates enterprise security standards like Intune, Windows Hello for Business, and Entra ID to allow agents to act on behalf of authenticated users. We also dive into the future-proof promise of "Just In Time UI" (Generative UI) which dynamically adapts interfaces to any form factor, and explore how these agentic tools could liberate deskless workers from being "slaves to a slab of glass." More Info: * https://commandline.microsoft.com/project-solara-build-2026/ Timestamps: [00:00:00] Intro & Catching Up [00:00:49] Transitioning from Voice-First (Alexa/Assistant) to Agent-First [00:01:35] Designing for Echo Show and Google Assistant vs. GenAI [00:02:37] Project Solara: Custom Agentic Devices for Business [00:03:09] Google Glass & the Early Spark for Enterprise Use Cases [00:04:30] Smart Displays and Wearable Badge Concept Hardware [00:05:12] Built on Android (AOSP) vs. Google's Android XR [00:05:46] Security: Microsoft MDEP, Intune, and Alexa for Business [00:07:10] Bring Your Own Agent (BYOA) on Azure [00:08:41] Just-In-Time UI & Generative UI [00:12:09] Developer Availability and Future Outlook [00:13:26] Rethinking Computers: Lessons from Google Glass & Assistant [00:14:32] Wrap Up and Future Form Factors (Watches, Rings, Glasses) #ProjectSolara #MicrosoftBuild #AgentFirst #VoiceFirst #MDEP #GenerativeUI #GenUI #AOSP #BYOA #EnterpriseTech #TwoVoiceDevs Episode 274

  • S1 · E273
    June 4 · 19 min

    New Horizons for Android: XR, MCP, and Agents

    Allen and Mike record live from Google I/O in the Builders podcast space. They discuss their impressions of this year's conference, the evolution of I/O over the years, and the big announcements from the keynote. Key topics include Gemini's "any output from any input" vision, how the new NanoBanana and Omni models are different than Imagen and Veo, the state of Android XR development, and the introduction of App Functions (Android MCP) for better AI agent integration. They also share their thoughts on the new Gemini app UI and what they hope to see in the world of wearables by next year. More info: * Android XR Developer Program: https://developer.android.com/develop/xr/catalyst [00:00:11] Live from Google I/O Builders Podcast Space [00:00:37] Reflections on I/O over the years [00:02:01] Gemini's "Any Input to Any Output" Vision [00:02:54] What's the big deal with NanoBanana and Omni? [00:03:41] Android XR and the future of intelligent eyewear [00:06:02] New Android developer tools and AI coding agents [00:08:29] App Functions and Android MCP [00:13:08] Spark, Halo, and AI agents on Android [00:15:07] The new Gemini app UI and design feedback [00:17:26] Looking ahead: Hopes for I/O 2027 and wearables #GoogleIO #GeminiAI #AndroidXR #AndroidMCP #AppFunctions #GoogleGlass #TwoVoiceDevs #AIAgents #AndroidDev #Wearables #AppFunctions #NanoBanana #GeminiOmni

  • S1 · E272
    May 28 · 19 min

    Google I/O 2026: GenUI, Glass, and Android XR

    Allen and Noble are live from Google I/O! This episode breaks down the biggest keynote news: agentic coding in Search, the power of Generative UI, and the future of "intelligent eyewear." They share what these changes mean for the venerable Google Search, what works (and what doesn't) with the new Google Glass, and how Android XR fits into the picture. From wearable AI to interactive search, find out what's here and what's coming this fall. More info: * Agentic Coding in Search: https://blog.google/products-and-platforms/products/search/search-io-2026/#agentic-coding * Android XR Developer Program: https://developer.android.com/develop/xr/catalyst [00:00:00] Introduction from Google I/O [00:01:32] Agentic Coding in Google Search [00:04:00] Generative UI: Beyond the Chatbot [00:08:19] The Three Pillars: Models, Coding, and Agents [00:11:00] Intelligent Eyewear and the Return of Glass [00:13:09] Hands-on with the AI Sandbox [00:15:44] The Human Impact of Real-Time Translation [00:16:47] Android XR and the Developer Experience [00:18:36] Developer Opportunities and Early Access #GoogleIO #IO26 #AndroidXR #GeminiAI #GenerativeUI #GoogleGlass #IntelligentEyewear #GoogleSearch #AgenticAI #TechPodcast #TwoVoiceDevs #AI #IOCreatorStudio #GoogleForDevelopers Episode 272

  • S1 · E271
    May 14 · 18 min

    Live from Next 2026: The Year of the Agent

    Allen and Alice are on the ground at Google Cloud Next, breaking down the biggest shifts in the AI landscape. This episode explores the transition from focusing on models to building agents with the launch of the Gemini Enterprise Agent Platform. They discuss the new TPU v8 hardware, the power of the Model Context Protocol (MCP) for Workspace integration, and how tools like Workspace Studio are making agent development accessible to everyone. Plus, a look at the incredible AI-powered Wizard of Oz experience at the Sphere! Timestamps: [00:00:12] Live from Day Two of Google Cloud Next [00:01:13] New Hardware: TPU v8 for Training and Inference [00:02:53] Gemini's Current State and Future Models [00:04:27] Vertex AI Rebrands as Gemini Enterprise Agent Platform [00:06:14] Building Reliable Agents: Identity, Registry, and Observability [00:07:18] Powering Agents with Model Context Protocol (MCP) [00:11:06] Workspace Studio: Automation for Everyone [00:15:00] Immersive Experiences at the Sphere [00:17:12] Final Thoughts and Where to Follow Hashtags: #GoogleCloudNext #Gemini #AIAgents #VertexAI #TPU #MCP #WorkspaceStudio #TwoVoiceDevs #GenAI #QueenOfSpreadsheets

  • S1 · E270
    March 5 · 49 min

    Episode 270 - Beyond the Big Three: Open Models, Agents, & the Future of Devs

    In part two of this insightful conversation, Allen and Sam Witteveen dive deep into the rapidly expanding world of AI models beyond the "big three." They explore the impact of open-weight and Chinese models like DeepSeek, Mistral, and Qwen, discussing their impressive efficiency and coding capabilities. The conversation shifts to the rise of agentic workflows and how tools like Claude Code are fundamentally changing the day-to-day lives of developers. They also tackle the tough questions: Are junior developers being replaced? Is AI just the next level of abstraction in programming? Finally, they cover the enterprise side of AI, from on-premise deployments to the evolving landscape of prompt engineering and observability frameworks like LangChain. Timestamps: [00:00:00] Introduction [00:00:49] Exploring Open Weights and Chinese Models [00:03:41] The Value of "Thinking" Models and Distillation [00:06:41] Running Models Locally [00:08:34] The Shift Towards Agentic Workflows [00:12:17] How AI is Changing the Role of Developers [00:29:04] AI as the Next Level of Abstraction [00:35:00] Best Models for Tool Calling and Coding [00:39:04] On-Premise Models and Enterprise Solutions [00:44:49] The Future of Prompt Engineering and LangChain [00:48:37] Outro and Where to Find Sam Hashtags: #TwoVoiceDevs #AI #OpenWeights #DeepSeek #Mistral #Qwen #ClaudeCode #Gemini #LangChain #SoftwareEngineering #AgenticAI #MachineLearning

  • S1 · E269
    March 3 · 37 min

    Episode 269 - The "Big Three" AI Models and Training Evolution

    In Part 1 of a two-part series, guest host Sam Witteveen joins Allen to catch up and dive deep into the rapidly evolving world of AI models. Sam shares his fascinating journey from being a successful pop songwriter to becoming a Machine Learning Google Developer Expert (GDE) and running the massive Machine Learning Singapore meetup. The conversation shifts to the latest AI developments, exploring the "Big Three" model builders—Anthropic, OpenAI, and Google. Sam and Allen discuss the frenetic pace of new model releases, changes to the Gemini 3 API, and how developers navigate the trade-offs between intelligence, latency, and cost. Finally, they pull back the curtain on how these models are actually trained today. Discover why models are no longer trying to be "fact machines" and how post-training breakthroughs, code execution sandboxes, and Reinforcement Learning (RL) environments are dramatically improving AI capabilities. Stay tuned for the end of the episode, where they hint at what's coming in Part 2! Timestamps: [00:00:00] Introduction and catching up [00:01:33] Sam's fascinating journey from pop music to machine learning [00:05:23] Running the massive Machine Learning Singapore meetup [00:07:42] Stumbling into YouTube and teaching AI with Google Colab [00:12:38] Analyzing the "Big Three" AI models and rapid release cycles [00:17:52] Gemini 3 API updates, Flash models, and thinking levels [00:22:00] Tool use, knowledge cutoffs, and why LLMs aren't fact machines [00:26:00] How post-training and code sandboxes revolutionized AI [00:32:00] Scaling Reinforcement Learning (RL) environments for design [00:34:04] Structured outputs and the return to predictable rules [00:36:43] Tune in next time for more! And where to find Sam online Hashtags: #TwoVoiceDevs #AI #MachineLearning #DeepLearning #LLM #GoogleGemini #Gemini #OpenAI #ChatGPT #Anthropic #Claude #ReinforcementLearning #RAG #Developers #SamWitteveen

  • S1 · E264
    February 19 · 18 min

    Episode 268 - The New @langchain/google Package

    Allen has been busy! This week, he unveils the new `@langchain/google` package for LangChain JS. This major update consolidates five previous libraries into a single, standardized, and powerful tool for developers working with Gemini and Vertex AI. Allen walks Mark through the motivation behind the change, the focus on backward compatibility, and the exciting new features like simplified multimodal input/output and text-to-speech support. If you're building with Google AI and JavaScript, this is the update you've been waiting for. [00:00:57] The confusion of previous packages [00:02:52] Creating a unified package [00:03:45] Introducing @langchain/google [00:04:35] Backward compatibility [00:06:48] Multimodal inputs [00:07:54] Standardizing output and image generation [00:08:58] Text-to-Speech support [00:11:29] Simplifying parameters and reasoning [00:14:55] Future roadmap #LangChain #Gemini #NanoBanana #TextToSpeech #GoogleAI #JavaScript #TypeScript #VertexAI #OpenSource #AI #WebDevelopment #TwoVoiceDevs

  • S1 · E267
    February 6 · 36 min

    Episode 267 - Behind the Scenes: How We Use AI to Build Two Voice Devs

    Ever wonder how "Two Voice Devs" goes from a raw recording to a finished episode? In this episode, Allen Firstenberg takes Mark Tucker on a deep dive into his production workflow. They discuss how Descript’s text-based editing revolutionized their process, how Allen uses a custom Gemini CLI agent to automate show notes and descriptions, and the technical (and ethical!) journey of creating AI-generated thumbnails using Google's Nano Banana. It’s a candid look at how AI can act as a force multiplier for creators while keeping the "human in the loop." [00:00:01] Introduction and Check-in [00:01:27] Behind the Scenes: Why We Use AI [00:03:42] Descript: Text-Based Video Editing [00:05:24] Building a Knowledge Database from Transcripts [00:08:13] Editing Video Like a Document [00:12:34] Exploring Descript's AI [00:13:36] Automating Show Notes with Gemini CLI [00:14:10] The Power of System Instructions (GEMINI.md) [00:19:30] AI Thumbnail Generation with Nano Banana [00:26:10] The Ethics of Synthetic Media and Artistic Style [00:28:40] Keeping the Human in the Loop [00:33:00] Evolution of the Two Voice Devs Workflow #TwoVoiceDevs #PodcastProduction #GeminiAI #Descript #GeminiCLI #NanoBanana #Automation #ContentCreation #Ethics #GenerativeAI #AIWorkflow #PodcastEditing

  • S1 · E266
    January 29 · 43 min

    Episode 266 - Supercharging Your AI Agent with Skills

    Mark and Allen dive into the emerging world of Agent Skills, an open standard for extending the capabilities of AI coding assistants like GitHub Copilot, Claude Code, and Gemini CLI. They explore how these skills work, how they compare to the Model Context Protocol (MCP), and walk through creating and installing a custom skill using the `skills` CLI. They also discuss the skills.sh website by Vercel, which acts as a registry and leaderboard for the ecosystem. The conversation touches on the potential for standardization, the current fragmentation in the ecosystem, and critical security considerations for these powerful new tools. More Info: * https://agentskills.io * https://skills.sh * https://cra.mr/mcp-skills-and-agents [00:00:00] Introduction & Context: AI Agents and Tools [00:02:18] Getting Information into Context (Instructions files) [00:06:50] What are Agent Skills? (AgentSkills.io) [00:09:55] Agent Skills vs. MCP Servers [00:16:35] How Skills Work: Progressive Disclosure [00:19:50] Mark's Example: List Global NPM Skill [00:22:56] Installing Skills with skill.sh and the Skills CLI [00:26:55] Demo: Installing on GitHub Copilot [00:30:58] Demo: Installing on Gemini CLI [00:37:37] Discussion: Discovery, Standardization, and Security [00:43:05] Conclusion #AgentSkills #AI #GitHubCopilot #GeminiCLI #CodingAssistants #MCP #ModelContextProtocol #DeveloperTools #TwoVoiceDevs

  • S1 · E265
    January 23 · 24 min

    Episode 265 - Gemini's New Personal Intelligence: A Second Brain?

    Allen and Mike discuss Google's new "Personal Intelligence" feature for Gemini. They explore how it connects to your personal data like Photos, Gmail, and Docs to provide context-aware answers. The conversation covers real-world use cases, privacy concerns regarding training data, and the importance of transparency and granular control in AI systems. They also touch on the "blackmail" scenario found in other AI research and what developers can learn from Google's implementation. More Info: * https://blog.google/innovation-and-ai/products/gemini-app/personal-intelligence/ [00:00:30] Google's Gemini Personal Intelligence announcement [00:01:48] Connecting personal data sources to Gemini [00:03:45] Google's unique advantage with user data [00:06:40] Real-world use case: Tracking travel history [00:07:30] Potential use case: Combining health data sources [00:09:15] Privacy: Is your data used for training? [00:12:40] The debate: Opting in vs. privacy concerns [00:16:30] AI safety and the "blackmail" scenario [00:18:50] Lessons for developers: Granular permissions and transparency [00:20:30] Verifiability and user trust [00:23:50] Conclusion #Gemini #GoogleAI #PersonalIntelligence #Privacy #MachineLearning #Developer #TechPodcast #AI #TwoVoiceDevs

  • S1 · E264
    January 20 · 24 min

    Episode 264 - AI, Context, and the "No-UI" Future

    Allen Firstenberg welcomes back guest host Mike Wolfson, an Android Google Developer Expert, to discuss the shifting landscape of User Experience (UX) in the age of Artificial Intelligence. As we move toward autonomous agents and multimodal interactions—incorporating voice, haptics, and environmental data—the reliance on traditional screens and touch interfaces is set to diminish. Mike shares insights on why "context" is the new "queen," illustrating the challenges of current AI with real-world examples (like his Meta glasses failing to find a decent breakfast burrito). The conversation tackles the critical question: "Who does the agent truly work for?" and explores how developers can avoid "notification fatigue" while ensuring users remain informed and in control. From the "creepy" factor of hyper-personalized data to the "Beverage Butler" concept, this episode dives deep into designing for a future where the best UI might be no UI at all. [00:00:00] Welcome and Introduction [00:01:21] The Evolution of UX: Beyond Touchscreens [00:03:52] The Importance of Context: A Meta Glasses Example [00:06:21] Privacy, Creepiness, and Agent Loyalty [00:08:43] Application Development in an AI World [00:11:34] Avoiding the "Seinfeld Assistant" of Notification Fatigue [00:17:15] Feedback Modalities: Tones vs. Speech [00:19:39] Discoverability of Features in Non-Visual Interfaces [00:23:00] The "Beverage Butler" and Future Outlook #AI #UX #UserExperience #ArtificialIntelligence #AndroidDev #ContextAwareness #VoiceUI #NoUI #TechPodcast #TwoVoiceDevs #GDE #GoogleDeveloperExpert #BeverageButler

  • S1 · E263
    January 16 · 37 min

    Episode 263 - Exploring the Parlant Agent Framework

    In this episode, Mark introduces Allen to Parlant, an open-source framework for building agentic AI. They explore how Parlant differs from other frameworks like LangChain and LangGraph by using concepts like "journeys" for flexible conversation flows and "guidelines" for conditional rule application. Mark walks through the key features, including the ability to define glossaries, tools, and the engine's matching process. They also discuss the recent version 3.1 updates, such as linked journeys and behavior criticality levels. Finally, they dive into "Emcie," Parlant's managed NLP service that utilizes a teacher-student model architecture to optimize performance and cost using Small Language Models (SLMs). [00:00:00] Welcome and Introduction [00:00:43] Introduction to Parlant [00:02:03] Journeys: Flexible Conversation Flows [00:03:36] Guidelines: Conditional Rules [00:06:27] Motivations and Compliance [00:08:44] NLP Services and Providers [00:11:42] The Balance Between Rigid and Loose Conversations [00:18:43] Parlant 3.1 Updates: Linked Journeys and Behavior Levels [00:22:20] Custom Matchers [00:23:14] Emcie: Parlant's Managed NLP Service [00:25:05] Model Tiers: Jackal and Bison [00:26:14] Teacher-Student Architecture and SLMs [00:30:26] Cost and Optimization with Student Models [00:35:19] Conclusion and Wrap-up #Parlant #AI #AgenticAI #LLM #SLM #OpenSource #SoftwareDevelopment #TwoVoiceDevs #Emcie #NLP

  • S1 · E262
    January 2 · 28 min

    Epsiode 262 - 2025 Wrap-Up: The Great Agent Takeover & 2026 Vibe Check

    Happy New Year! Allen and Mark kick off 2026 by looking back at the whirlwind of AI developments in 2025. From the explosion of agentic frameworks like LangGraph, Google's Agent Development Kit, and the Microsoft Agent Framework to the emergence of protocols like MCP and A2A, it was a year of rapid evolution. They discuss the rise of "vibe coding," the state of voice assistants like Alexa Plus and Gemini, and the challenges of monetization and discovery in a model-centric world. What lies ahead in 2026? The duo shares their big predictions! What do we need on the hardware side? Perhaps something new from Google, Microsoft, OpenAI, Amazon, or someone else? What companies are we rooting for? What's next for agents? Tune in and find out! Timestamps: [00:00:00] Introduction and Happy New Year [00:01:13] Reflecting on a busy 2025: Google's weekly announcements [00:02:05] New terms: MCP (Model Context Protocol) and A2A (Agent-to-Agent) [00:03:52] The shift to Agents and Agentic solutions [00:05:40] Framework evolution: Microsoft Agent Framework and LangGraph [00:07:54] The rise of Coding Assistants and "Vibe Coding" [00:09:25] State of Voice: Alexa Plus and Gemini on smart devices [00:11:15] Smart Glasses and the future of ambient AI [00:12:40] MCP: Challenges with discovery, monetization, and security [00:15:10] Microsoft Foundry and low-code agent building [00:20:25] Live Streaming Models: Video, Audio, and Text [00:22:00] 2026 Predictions: Voice Flow acquisition [00:23:05] Prediction: Moving from Chatbots to Autonomous Agents [00:25:20] Prediction: The role of Small Language Models (SLMs) [00:27:10] Closing thoughts and 2026 outlook Hashtags: #AI #GenerativeAI #Agents #AutonomousAgents #MCP #A2A #LangGraph #Gemini #VoiceAssistant #SmartGlasses #VoiceFirst #SLM #VoiceFlow #TwoVoiceDevs #YearInReview #2026Predictions #VibeCoding

  • S1 · E261
    Dec 26, 2025 · 20 min

    Episode 261 - The Great Holid-AI Rebus Battle

    Get ready for the ULTIMATE SHOWDOWN of holiday cheer and artificial intelligence! In this SPECIAL HOLIDAY EPISODE, Mark and Allen aren't just exchanging pleasantries—they're exchanging MIND-BENDING REBUS PUZZLES generated by some AI models themselves! It's a battle of wits, a clash of code, and a festive face-off as Microsoft Copilot takes on Google's Gemini (and the famous "Nano Banana Pro") to solve visual riddles that will have you shouting at your screen. Can our hosts decipher the scribbles of silicon brains? Or will the AI stump the humans once and for all? Grab your eggnog, put on your thinking cap, and play along! It's Two Voice Devs like you've never seen (or puzzled) them before! HAPPY HOLIDAYS! [00:00:00] Intro: The Rules of Engagement [00:02:48] Puzzle 1: A Nipping Chill [00:03:33] Puzzle 2: Going for Gold [00:07:00] Puzzle 3: Escaping the Cage [00:10:00] Puzzle 4: The Silent Mouse [00:11:00] Puzzle 5: A Knightly Gesture [00:12:15] Puzzle 6: Sweet Ballerina [00:13:30] Puzzle 7: Listen Closely to the Animal [00:15:30] Puzzle 8: A Holiday Wish [00:16:15] Puzzle 9: The Grand Finale Challenge [00:19:30] Happy Holidays from Two Voice Devs! #TwoVoiceDevs #HolidaySpecial #RebusPuzzles #LLMBattle #AIShowdown #Copilot #Gemini #NanoBananaPro #ChatGPT #ArtificialIntelligence #MachineLearning #JackFrost #WinterWonderland #Freeze #SugarPlumFairy #Nutcracker #NewYears #Gnu #MerryChristmas #FestivalOfLights #Hanukkah #Kwanzaa #Unity #HolidayFun #Games #Puzzles #TechHumor #DevLife #HappyHolidays #SeasonsGreetings #Fun #Creative #OverTheTop #Podcast #Developers #SoftwareEngineering

  • S1 · E260
    Dec 24, 2025 · 19 min

    Episode 260 - Turn Your AI Agent into a Voice Agent With Microsoft Foundry

    Mark Tucker explores Microsoft Azure's "Voice Live" feature within the newly rebranded Microsoft Foundry. He demonstrates how to take a standard text-based AI agent—in this case, one that talks like a pirate—and instantly give it a voice using WebSockets to bridge speech-to-text and text-to-speech. Mark walks through the differences between the "Old" Foundry (V1) and the "New" Foundry (V2), shows the configuration steps, and dives into a Python code example to connect it all together. Learn more: * https://github.com/rmtuckerphx/voicelive-agents-quickstart [00:00:00] Intro and Holiday Plans [00:00:46] Introducing Microsoft Azure Voice Live [00:02:00] Microsoft Foundry Overview [00:02:46] Creating a Pirate Agent in Foundry [00:04:46] Enabling Voice Live in the Playground [00:07:46] Demo: Speaking with the Pirate Agent [00:08:46] Comparing Old and New Foundry [00:13:46] Code Walkthrough: Voice Live Quick Start [00:15:46] Connecting Version 2 Agents [00:18:46] Conclusion #MicrosoftAzure #VoiceLive #MicrosoftFoundry #AI #VoiceFirst #GenerativeAI #SpeechToText #TextToSpeech #Python #Coding

Showing 1–20 of 20 episodes