Skip to content
Artwork for Scott & Mark Learn To...
Scott & Mark Learn To... · Wednesday · 36 min

Scott & Mark Learn To... Fool's Gold

In this episode, Scott Hanselman and Mark Russinovich dive into the security risks surrounding increasingly capable open-weight AI models and Mark’s latest research into a new defense called Fool’s Gold. Mark explains how safety guardrails can be removed from open models and explores an alternative approach: training models to provide convincing but deliberately flawed information when those protections are bypassed. They discuss how the technique works, whether it affects legitimate model behavior, and what it could mean for the future of AI safety as open models continue to advance. Takeaways: Why open-weight AI models can create new security risks How Mark’s Fool’s Gold approach uses decoy responses as a defense decoy responses as a defense Techniques that can protect against harmful use without degrading normal model performance Who are they? View Scott Hanselman on LinkedIn View Mark Russinovich on LinkedIn Watch Scott and Mark Learn on YouTube Listen to other episodes at scottandmarklearn.to Discover and follow other Microsoft podcasts at microsoft.com/podcasts

0:00-36:09

transcript

No transcript — this publisher did not publish one.

show notes

In this episode, Scott Hanselman and Mark Russinovich dive into the security risks surrounding increasingly capable open-weight AI models and Mark’s latest research into a new defense called Fool’s Gold. Mark explains how safety guardrails can be removed from open models and explores an alternative approach: training models to provide convincing but deliberately flawed information when those protections are bypassed. They discuss how the technique works, whether it affects legitimate model behavior, and what it could mean for the future of AI safety as open models continue to advance. 


Takeaways:    

  • Why open-weight AI models can create new security risks 

  • Techniques that can protect against harmful use without degrading normal model performance 

Who are they?     

View Scott Hanselman on LinkedIn  

View Mark Russinovich on LinkedIn   

 

Watch Scott and Mark Learn on YouTube 

       

Listen to other episodes at scottandmarklearn.to  

         

Discover and follow other Microsoft podcasts at microsoft.com/podcasts   

links6