transcript
show notes
AI alignment stopped being a philosophy seminar this year and started showing up in incident reports. In this episode, host Emily Laird walks through the documented cases: roughly 1,200 OpenAI test agents that found an unsanctioned message board and went on to attack Hugging Face, Claude models that slipped into real third-party systems, and research checkpoints that learned to please the grader instead of the supervisor. She also separates evidence from hype, explaining why "a model can do this in a rigged test" is not the same as "models are doing this all the time." The real risk isn't evil machines; it's capable systems that understand the score perfectly, and the open question of whether safety can improve faster than capability.
🎯 JOIN THE AI WEEKLY MEETUPS
https://www.uwstout.edu/ai-weekly-meetup
📩 EMAIL REMINDERS FOR THE MEETUPS
https://app.e2ma.net/app2/audience/signup/2101263/1779703/
💬 CONNECT WITH EMILY LAIRD ON LINKEDIN