
How Data Scientists Use Knowledge Distillation to Compress Models
transcript
show notes
In episode 170 of The Data Science Podcast, Lucas and Luna explore knowledge distillation—a technique for compressing large, powerful machine learning models into smaller, faster ones. They use the concrete example of a large language model distilled into a compact version that runs on a smartphone, cutting latency and energy use while preserving most of the accuracy. The hosts explain the core idea: a 'teacher' model trains a 'student' model to mimic its predictions, transferring knowledge through soft labels. They also discuss real-world trade-offs, including a case where a distilled model lost nuance in edge cases, and practical advice like when distillation beats pruning or quantization. Listeners will learn how this approach is changing model deployment in production and what it means for the next wave of on-device AI.
#KnowledgeDistillation #ModelCompression #MachineLearning #DataScience #AI #DeepLearning #OnDeviceAI #TechPodcast #FexingoBusiness #BusinessPodcast #DataSciencePodcast #TeacherStudent #SoftLabels #ModelEfficiency #EdgeAI #NeuralNetworks #ModelDeployment #AICompression
