
How Data Scientists Use Data Lineage to Debug Pipelines
transcript
show notes
When a model's predictions start drifting from reality, the first question every data scientist asks is: where did the data go wrong? In this episode, Lucas and Luna explore a subtle but powerful debugging technique: data lineage. They walk through a real example at a mid-sized e-commerce company where a sudden drop in recommendation click-through rates turned out to trace back to a silent upstream schema change in a product catalog feed. Lucas explains how lineage graphs — the network of data sources, transformations, and downstream consumers — let teams trace a model's output all the way back to the raw inputs, and why most debugging setups only track half the story. The conversation also covers how lineage differs from data provenance, why it's often buried in orchestration logs and metadata catalogs, and how a simple lineage query can cut a multi-day investigation down to minutes. If you've ever stared at a stale model output and wondered what changed, this episode gives you a practical mental model for finding the answer.
#DataLineage #DataDebugging #DataPipelines #MachineLearning #DataScience #DataGovernance #Metadata #DataProvenance #ModelDrift #DataQuality #RecommendationSystems #Ecommerce #DataEngineering #Tech #Business #FexingoBusiness #BusinessPodcast #DataSciencePodcast
