Skip to content
Artwork for The Data Engineering Show
TechnologyBusinessManagement

The Data Engineering Show

The Firebolt Data Bros

The Data Engineering Show is a podcast for data engineering and BI practitioners to go beyond theory. Learn from the biggest influencers in tech about their practical day-to-day data challenges and solutions in a casual and fun setting.

SEASON 1 DATA BROS
Eldad and Boaz Farkash shared the same stuffed toys growing up as well as a big passion for data. After founding Sisense and building it to become a high-growth analytics unicorn, they moved on to their next venture, Firebolt, a leading high-performance cloud data warehouse.

SEASON 2 DATA BROS
In season 2 Eldad adopted a brilliant new little brother, and with their shared love for query processing, the connection was immediate. After excelling in his MS, Computer Science degree, Benjamin Wagner joined Firebolt to lead its query processing team and is a rising star in the data space.

For inquiries contact tamar@firebolt.io
Website: https://www.firebolt.io

Play
  • 21 episodes
  • fortnightly
  • Avg 22 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • #60
    August 18 · 17 min

    How AI Is Reshaping Modern Data Teams and the Future of Platforms with Xavier Gumara Rigol

    What if AI could transform how your entire data team works—without replacing them? In this episode, Benjamin Wagner explores with Xavier Gumara Rigol, Head of Data at Manychat, how natural language-to-SQL tools are reshaping data analyst and data scientist roles, why building a strong data platform is essential for AI-powered self-serve analytics, and the critical strategies for evaluating text-to-insight solutions in 2026. Whether you're leading a data organization or building your next analyti

  • #59
    August 6 · 19 min

    Building Modern Data Platforms Without Legacy Bottlenecks ft Andrew Jones

    What if you could transform your data team from a bottleneck into a self-serve platform that empowers the entire organization? In this episode, host Benjamin Wagner sits down with Andrew Jones, Staff Data Engineer at LocalStack, to explore how to balance building new data infrastructure while supporting legacy systems, why reliability and data contracts are becoming non-negotiable, and how AI agents are accelerating platform development. Whether you're scaling a data organization or rethinking y

  • #58
    July 21 · 20 min

    Why 99% of BI Tools Get Embedded Analytics Wrong And How Omni Fixed It ft. Chris Merrick

    What if the future of analytics wasn't about building another middleware layer, but about getting closer to the actual business user? In this episode, Benjamin sits down with Chris Merrick, CTO and cofounder of Omni, to explore why semantic layers matter more than ever in an agentic world, how AI is reshaping embedded analytics and customer-facing data experiences, and the key strategies for keeping complex data models aligned across federated sources. Whether you're building analytics platforms

  • #57
    June 16 · 19 min

    AI for Data and Data for AI: The Dual Frontier of Modern Data Engineering with Pranav Motarwar

    What if the data engineering skills you have today become obsolete in five years? In this episode, host Benjamin Wagner sits down with Pranav Motarwar, a data engineer who's witnessed the industry's transformation from traditional ETL to AI-powered pipelines, to explore how AI is fundamentally reshaping data engineering roles, why you need to master both "AI for data" and "data for AI" to stay relevant, and the emerging infrastructure required to handle multimodal data at scale. Whether you're a

  • #56
    May 7 · 18 min

    AI Won't Replace Engineers, But This Framework Will Change How They Build with Rohit Girme

    What if you could build AI features with confidence while moving at the pace of innovation? In this episode, Benjamin Wagner sits down with Rohit Girma, Staff Software Engineer at Airbnb, to explore how to evaluate generative AI in production, why breaking down complex problems into smaller chunks accelerates development, and the key strategies for scaling AI-powered products beyond zero-to-one. Whether you're shipping AI features or transforming your engineering workflow, this conversation offe

  • #55
    April 28 · 22 min

    The Framework Canva Uses for 200M+ Designers with Paul Tune

    In this episode of The Data Engineering Show, Benjamin sits down with Paul Tune, Staff Research Scientist at Canva, to explore the advancement of machine learning at one of the world's leading design platforms. Learn how Canva is transitioning from traditional ML like recommendation engines for templates to cutting-edge agentic workflows that allow users and AI to collaborate on complex design tasks. Whether you're interested in the infrastructure behind distributed training or the nuances of po

  • #54
    April 8 · 22 min

    Llama 2 & 3 Safety: Soumya Batra on Agentic AI Training

    What if the expertise that built foundation models could reshape how you think about AI's future? In this episode, Benjamin sits down with Soumya Batra, founder and CEO of WisePort AI and former safety lead on Llama 2 and Llama 3 at Meta, to explore how foundation models evolved from traditional NLP, why post-training holds the highest leverage for safety and controllability, and what natively agentic AI means for the next frontier of AI development. Whether you're curious about the model traini

  • #53
    March 24 · 18 min

    The Data Fusion Secret & Why Custom Query Engines Fail with Nikita Lapkov

    What if building a distributed SQL engine meant rethinking everything about how query execution works at scale? In this episode, Benjamin sits down with Nikita, Senior Software Engineer at Cloudflare, to explore how R2 SQL leverages object storage and distributed computing to power analytics across 300 global locations, why backward compatibility becomes critical when you can't control infrastructure rollouts, and the key strategies for handling joins and adaptive query execution in a stateless,

  • #52
    March 10 · 24 min

    How Zipline AI Turns Weeks of Engineering Into Minutes of SQL Queries ft. Nikhil Simha

    What if you could deploy ML features and real-time data pipelines without building complex infrastructure from scratch? In this episode, host Benjamin sits down with Nikhil Simha, CTO at Zipline AI and co-author of Chronon AI, to explore how Chronon, an open-source system that generates data infrastructure from simple queries, is transforming feature engineering at companies like OpenAI and Airbnb. Learn why iteration speed matters for fraud detection, how to serve thousands of signals at a ma

  • #51
    February 19 · 16 min

    The Geo-Data Problem Nobody Talks About And How Voi Solved It ft. Magnus Dahlbäck

    What if your data platform could power both critical business decisions and real-time product features at scale? In this episode, host Benjamin sits down with Magnus Dahlbäck, Senior Director of Data and Platform at Voi, to explore how a metrics-first approach and semantic layers transform data accessibility, why traditional ML and LLMs require different strategies for different problems, and how to balance FinOps costs while processing billions of IoT events daily. Whether you're building data

  • #50
    February 3 · 29 min

    Why 99% of Data Teams Give Up on Real-Time And How Artie Changes That

    What happens when a team of seven engineers spends a year trying to build a production-ready CDC connector and fails? For Artie CTO and co-founder Robin Tang, it was the spark needed to build a platform that makes data streaming accessible. In this episode, Robin joins Benjamin to discuss the "DFS" (Deep First Search) approach to data sources, the engineering hurdles of real-time Postgres-to-Snowflake pipelines, and why "theoretically correct" architectures often fail in practice.

  • #49
    Dec 16, 2025 · 25 min

    The $100M Problem: How Lyft's Data Platform Prevents ML Failures with Ritesh Varyani at Lyft

    What if your data platform could serve AI-native workloads while scaling reliably across your entire organization? In this episode, Benjamin sits down with Ritesh, Staff Engineer at Lyft, to explore how to build a unified data stack with Spark, Trino, and ClickHouse, why AI is reshaping infrastructure decisions, and the strategies powering one of the industry's most sophisticated data platforms. Whether you're architecting data systems at scale or integrating AI into your analytics workflow, thi

  • #48
    Nov 19, 2025 · 19 min

    60 Billion Predictions Daily: Inside Credit Karma’s Agentic Data Layer with Maddie Daianu

    What does MLOps look like when you are deploying 60 billion machine learning predictions a day? Maddie Daianu, Head of Data and AI at Intuit Credit Karma, joins the Data Bros to pull back the curtain on one of the most high-volume data environments in FinTech. With a 100-person team serving 140 million members, standard data practices break down. Maddie shares how her team manages terabytes of daily data on Google Cloud and explains the massive strategic pivot they are undertaking right now: Th

  • #47
    Oct 7, 2025 · 20 min

    Block Bad Data Before the Write with Nike’s Ashok Singamaneni

    Nike’s Principal Data Engineer Ashok Singamaneni joins Benjamin and Eldad to discuss his open-source data quality framework, Spark Expectations. Ashok explains how the tool, which was inspired by Databricks DLT Expectations, shifts data quality checks to before the data is written to a final table. This proactive approach uses row-level, aggregation-level, and query data quality checks to fail jobs, drop bad records, or alert teams - ultimately saving huge costs on recompute and engineering effo

  • #46
    Sep 17, 2025 · 21 min

    Postgres vs. Elasticsearch: The Unexpected Winner in High-Stakes Search for Instacart with Ankit Mittal

    Modernizing Search Infrastructure: How Instacart Transitioned from Elasticsearch to PostgreSQL for Enhanced Performance and Simplicity. In this episode of The Data Engineering Show, host Benjamin Wagner speaks with Ankit Mittal, former senior engineer at Instacart, about the company's innovative approach to modernizing their search infrastructure by transitioning from Elasticsearch to PostgreSQL for single-retailer search functionality.

  • #45
    Aug 28, 2025 · 21 min

    Is Self-Service BI a False Promise? Lei Tang of Fabi.ai Thinks So

    AI is reshaping business intelligence by enabling true self-service analytics and transforming how organizations interact with their data through natural language processing. In this episode of The Data Engineering Show, host Benjamin interviews Lei, Co-founder and CTO of Fabi.ai, to explore how AI-native BI platforms are reshaping data analytics and empowering non-technical users to derive meaningful insights from complex datasets.

  • #43
    Jun 10, 2025 · 22 min

    From Zero to 100M Users: Inside Notion’s Data Stack and AI Strategy with Sumit Gupta

    Dive into the future of data engineering with Sumit Gupta, Lead BI Engineer at Notion, as he shares insights with the bros on navigating the AI revolution in modern data stacks. From leveraging tools like Snowflake and dbt to automating content creation with AI, discover how traditional technical skills are evolving alongside the rise of AI. Whether you're a seasoned data professional or just starting your journey, learn why embracing AI isn't optional and how to balance technical expertise with

  • #42
    May 7, 2025 · 31 min

    How Rising Wave Is Redefining Real-Time Data with Postgres Power

    In this episode of The Data Engineering Show, the bros sit with Yingjun Wu, founder and CEO of Rising Wave, to explore the innovative world of stream processing systems. Yingjun shares his journey from academic research to creating a Postgres-compatible streaming system that drastically reduces resource usage. They discuss how Rising Wave's S3-based architecture and Postgres compatibility provide advantages over traditional systems like Flink, and explore the increasing role of Apache Iceberg in

  • #41
    Apr 8, 2025 · 23 min

    Revolutionizing Data Governance with DataStrato’s Unified Open Source Approach

    In this episode of The Data Engineering Show, the bros sit with Lisa Cao, Product Manager at DataStrato, to explore data catalogs and Apache Gravitino, a unified metadata lake used to manage access and perform data governance for all data sources. They discuss data catalogs and how they refine the data management process.

Showing 1–20 of 21 episodes