
AI Agents in the Enterprise | Sierra, Mercor, Intercom, Turing | $2.8B+ Raised
transcript
show notes
(We know the audio quality isn't great on this one :( But the conversation is still well worth it!)
Last week I hosted a fireside chat on what it actually takes to build AI agents in the enterprise with Natalie Meurer (Head of Agent Eng, Sierra), Harsh Trivedi (founding engineer, Mercor), Juhi Parekh (GM, Turing), and Kevin Lynch (Senior FDE, Fin).
We get into why new models aren't always better (and why you can't just swap in the latest release and assume your agent improves), how the data-labeling/RL environment business might only have a couple years left, why real-time voice-to-voice models still aren't production-ready, how cheaper inference is still causing prices to go up, how baking in a constellation of models into enterprise agents is so important for reliability,
and much, much, more!
Chapters below
00:00 Intro
00:21 Meet the panel
01:57 What everyone's actually using agents for day to day
06:10 The reality of forward deployed work
09:28 What agents couldn't do a year ago that they can now
12:29 Why you have to tell agents what NOT to do
16:26 What a harness actually is
22:03 RL environments explained
28:46 Does the data-labeling and RL environment business even last?
37:02 Why benchmarks don't tell you what works in production
38:12 Agent engineering vs forward deployed engineering
41:38 Deploying into 100-year-old enterprise systems
44:25 Why AI adoption is an org problem, not a tech problem
45:36 Hiring for judgment when engineers aren't really coding anymore
48:16 Why agents are a new kind of software
50:46 The first 90 days of an enterprise deployment
53:20 Why compliance environments break normal testing
56:58 Layering AI on AI to get to 99% accuracy
01:00:56 New models aren't always better — the swap problem
01:02:35 Improving agents without waiting for a new model
01:06:36 Does agent performance secretly degrade over time?
01:09:40 Why one model is never enough: the constellation approach
01:11:14 Building resilience when inference providers go down
01:13:48 When fine-tuning actually makes sense
01:14:53 Why voice-to-voice still isn't production-ready
01:16:25 The cascaded pipeline that real voice agents use
01:21:45 Audience Q&A: managing change inside the enterprise
01:23:24 Why inference getting cheaper makes things more expensive
01:26:54 Charging for outcomes instead of conversations
01:30:19 What the real moat is when everyone uses the same models
01:37:04 Synthetic data and where the data wall actually is
01:38:50 Closing thoughts





