
My Weird Prompts · September 8 · 23 min
How AI Training Data Gets Filtered (and Exploited)
0:00 · Intro-23:25
transcript
show notes
When a model says something toxic or wrong, what actually kept it out of the training data? This episode maps the full content filtering pipeline — from Common Crawl's 250 billion raw pages to the final curated dataset. We break down the six stages: language ID, deduplication, quality scoring, toxicity filters, and curation, plus the tools like FineWeb, Dolma, and Datatrove that do the work. Then we explore why persistent pre-training poisoning attacks succeed even against aggressive filtering — and why the human annotators meant to catch bad content might be the weakest link.
Episode #920371 — open it directly at myweirdprompts.com/920371





