
0:00-1:52:22
Streams straight from the publisher. podnod never proxies or re-hosts episode audio.
Intro topic: Grills
News/Links:
- You can’t call yourself a senior until you’ve worked on a legacy project
- Recraft might be the most powerful AI image platform I’ve ever used — here’s why
- NASA has a list of 10 rules for software development
- AMD Radeon RX 9070 XT performance estimates leaked: 42% to 66% faster than Radeon RX 7900 GRE
Book of the Show
- Patrick:
- The Player of Games (Ian M Banks)
- https://a.co/d/1ZpUhGl (non-affiliate)
- The Player of Games (Ian M Banks)
- Jason:
- Basic Roleplaying Universal Game Engine
Patreon Plug https://www.patreon.com/programmingthrowdown?ty=h
Tool of the Show
- Patrick:
- Pokemon Sword and Shield
- Jason:
- Features and Labels ( https://fal.ai )
Topic: Reinforcement Learning
- Three types of AI
- Supervised Learning
- Unsupervised Learning
- Reinforcement Learning
- Online vs Offline RL
- Optimization algorithms
- Value optimization
- SARSA
- Q-Learning
- Policy optimization
- Policy Gradients
- Actor-Critic
- Proximal Policy Optimization
- Value optimization
- Value vs Policy Optimization
- Value optimization is more intuitive (Value loss)
- Policy optimization is less intuitive at first (policy gradients)
- Converting values to policies in deep learning is difficult
- Imitation Learning
- Supervised policy learning
- Often used to bootstrap reinforcement learning
- Policy Evaluation
- Propensity scoring versus model-based
- Challenges to training RL model
- Two optimization loops
- Collecting feedback vs updating the model
- Difficult optimization target
- Policy evaluation
- Two optimization loops
- RLHF & GRPO
https://amzn.to/3ES4p5i
amzn.tohttps://fal.ai
fal.ai★ Support this podcast on Patreon ★
patreon.com
- 0:00PhD Throwdown
- 1:36An AI Explains The Meaning of Hood
- 2:12Grills vs Pizza Ovens
- 5:29What is a Senior Engineer? (
- 11:30The Art of Refactoring Code
- 12:51Recraft Is the Most Powerful AI Image Platform I've Ever Used
- 18:12Open Source vs. Closed Source
- 18:24NASA's list of 10 rules for software development
- 21:59AMD Radeon RX 9070 XT performance estimates leaked
- 29:42What Kind of Jobs Are You Happy With?
- 32:24Book of the Week
- 33:44Basic Role Playing: The Universal Game Engine
- 38:33Pandas vs. Pokemon
- 43:46Foul AI: The Middleman Between Your AI Models and the
- 48:13Reinforcement Learning
- 55:22Reinforcement Learning vs Unsupervised Learning
- 1:00:23Value Based Reinforcement Learning vs Policy Based Learning
- 1:05:32Policy Algorithms for Complexity
- 1:13:41Reward Shaping in Machine Learning
- 1:14:41AlphaGo 2.8: Trust Region and Probability
- 1:19:25AlphaGo: Inferring and Substitution
- 1:20:51Model-based Reinforcement Learning
- 1:25:10Post-supervised learning: What's the big deal?
- 1:26:23Using Reinforcement Learning in the Store
- 1:30:39ChatGPT: RLHF and Reinforcement Learning
- 1:40:08The Deep Sequile Leap
- 1:41:19DeepSeam Sniped OpenAI's AI
- 1:46:10Facebook's Reinforcement Learning Explained
- 1:48:44A Future of AI Is Not Scary
- 1:51:10Reno on the Reno Hotel