transcript
show notes
Welcome back to Builders Learn ML with Professor Andy! In this episode we're talking about training your own AI models.
Should you train your own model? Maybe. But probably not yet. If you haven't nailed your prompts, workflows and evals, training won't fix your agent, and it only starts to pay off once you're running millions of tokens a day. In this episode, you'll learn when it's worth it and how to do it.
There are two main techniques for training your own AI models: supervised fine-tuning (SFT) and reinforcement learning (RL). SFT teaches a smaller, cheaper model to copy a stronger one using thousands of good examples from your agent. RL gives an open-weight model like Qwen, GLM or Kimi a task and rewards it when it gets a good result, so it figures out its own way there.
The hard part of RL is the reward. Score it badly, and you get reward hacking, where the model learns that doing nothing is the safest way to avoid a penalty.
They also talk about RL environments and Harbor, unit tests vs LLM judges as verifiers, how much data you need, why good evals are worth having even if you never train, and where to start learning on your own.
Andy Lyu is the co-founder and CTO of Osmosis (YC W25), a platform that helps companies train their own models with reinforcement learning. Before Osmosis he was a tech lead on TikTok's real-time recommendations and data infrastructure team.
Recorded live on AI Agents Hour, a weekly livestream by Mastra CPO Shane Thomas and CTO Abhi Aiyer. Mondays 12PM Pacific.
📚 ABOUT OSMOSIS
Osmosis: https://osmosis.ai
The Hugging Face RL handbook: https://huggingface.co/learn
Unsloth: https://unsloth.ai
📚 MASTRA RESOURCES
Mastra: https://mastra.ai
Mastra on X: https://x.com/mastra_ai
Mastra Discord: https://mastra.ai/community/discord
Mastra GitHub: https://github.com/mastra-ai
Learn Mastra: https://mastra.ai/learn
Principles of Building AI Agents (Book): https://mastra.ai/books/principles-of-building-ai-agents
Patterns for Building AI Agents (New Book): https://mastra.ai/books/patterns-of-building-ai-agents
WHAT IS MASTRA?
Mastra is an open-source TypeScript framework designed for building and shipping AI-powered applications and agents with minimal friction. It supports the full lifecycle of agent development—from prototype to production. You can integrate it with frontend and backend stacks (e.g., React, Next.js, Node) or run agents as standalone services. If you're a JavaScript or TypeScript developer looking to build an agentic or AI-powered product without starting from first principles, Mastra provides the scaffolding, tools, and integrations to accelerate that process.
⏱️ CHAPTERS
00:00 Intro
00:56 Meet Professor Andy (Osmosis)
01:34 What "training a model" means
01:43 Supervised fine-tuning (SFT)
02:04 Reinforcement learning (RL)
03:54 Scale & cost
04:46 How much data do you need
06:51 Data quality & distribution
08:20 Why RL beats SFT
09:08 Base models: why open source
09:53 Inside Osmosis: the compute crunch
11:32 RL environments & Harbor
12:58 Reward functions & verifiers
16:51 Reward hacking
18:18 Designing the reward (verifiable + rubric)
20:49 Why evals matter, even without training
23:15 Where to start