
S1E04. When Training Goes Wrong: Overfitting, Bad Data, and Metrics That Lie
transcript
show notes
This episode is narrated using AI voice technology. The content and script are original.
There are two ways for training to fail, they are opposites, and the fix for one makes the other worse. Overfitting is the driver who learned one route to work perfectly and is helpless on any other street. Underfitting is the driver who had one lesson and stopped. The only way to tell which you have is a slice of data locked in a drawer before you start.
Rob walks Shelley through the pair of numbers that is the entire diagnosis, early stopping done honestly on a bumpy curve, and why looking at the data comes first. Label errors in the collections this industry measures itself against. Five data failures, each with its own symptom: lopsided categories, gaps, extremes, leakage, duplicates. Accuracy is the number that lies most often when the rare thing is the whole job. Precision and recall point at two different mistakes, and each can be gamed alone.
Takeaway: the most expensive mistake is treating the wrong failure.
Subscribe for the rest of the season, and visit cloudadorn.com.