Skip to content
Artwork for The Data Science Podcast with Fexingo: Analytics, Machine Learning, and Data-Driven Conversations
The Data Science Podcast with Fexingo: Analytics, Machine Learning, and Data-Driven Conversations · August 9 · 8 min

How Data Scientists Use Data Lineage to Debug Pipelines

When a model's predictions start drifting from reality, the first question every data scientist asks is: where did the data go wrong? In this episode, Lucas and Luna explore a subtle but powerful debugging technique: data lineage. They walk through a real example at a mid-sized e-commerce company where a sudden drop in recommendation click-through rates turned out to trace back to a silent upstream schema change in a product catalog feed. Lucas explains how lineage graphs — the network of data sources, transformations, and downstream consumers — let teams trace a model's output all the way back to the raw inputs, and why most debugging setups only track half the story. The conversation also covers how lineage differs from data provenance, why it's often buried in orchestration logs and metadata catalogs, and how a simple lineage query can cut a multi-day investigation down to minutes. If you've ever stared at a stale model output and wondered what changed, this episode gives you a practical mental model for finding the answer. #DataLineage #DataDebugging #DataPipelines #MachineLearning #DataScience #DataGovernance #Metadata #DataProvenance #ModelDrift #DataQuality #RecommendationSystems #Ecommerce #DataEngineering #Tech #Business #FexingoBusiness #BusinessPodcast #DataSciencePodcast Keep every episode free: buymeacoffee.com/fexingo

0:00-8:47

transcript

No transcript — this publisher did not publish one.

show notes

When a model's predictions start drifting from reality, the first question every data scientist asks is: where did the data go wrong? In this episode, Lucas and Luna explore a subtle but powerful debugging technique: data lineage. They walk through a real example at a mid-sized e-commerce company where a sudden drop in recommendation click-through rates turned out to trace back to a silent upstream schema change in a product catalog feed. Lucas explains how lineage graphs — the network of data sources, transformations, and downstream consumers — let teams trace a model's output all the way back to the raw inputs, and why most debugging setups only track half the story. The conversation also covers how lineage differs from data provenance, why it's often buried in orchestration logs and metadata catalogs, and how a simple lineage query can cut a multi-day investigation down to minutes. If you've ever stared at a stale model output and wondered what changed, this episode gives you a practical mental model for finding the answer.

#DataLineage #DataDebugging #DataPipelines #MachineLearning #DataScience #DataGovernance #Metadata #DataProvenance #ModelDrift #DataQuality #RecommendationSystems #Ecommerce #DataEngineering #Tech #Business #FexingoBusiness #BusinessPodcast #DataSciencePodcast

Keep every episode free: buymeacoffee.com/fexingo

links1