Skip to content
Artwork for Forward Deployed
Forward Deployed · June 23 · 1 hr 41 min

AI Agents in the Enterprise | Sierra, Mercor, Intercom, Turing | $2.8B+ Raised

(We know the audio quality isn't great on this one :( But the conversation is still well worth it!) Last week I hosted a fireside chat on what it actually takes to build AI agents in the enterprise with Natalie Meurer (Head of Agent Eng, Sierra), Harsh Trivedi (founding engineer, Mercor), Juhi Parekh (GM, Turing), and Kevin Lynch (Senior FDE, Fin). We get into why new models aren't always better (and why you can't just swap in the latest release and assume your agent improves), how the data-labeling/RL environment business might only have a couple years left, why real-time voice-to-voice models still aren't production-ready, how cheaper inference is still causing prices to go up, how baking in a constellation of models into enterprise agents is so important for reliability, and much, much, more! Chapters below 00:00 Intro 00:21 Meet the panel 01:57 What everyone's actually using agents for day to day 06:10 The reality of forward deployed work 09:28 What agents couldn't do a year ago that they can now 12:29 Why you have to tell agents what NOT to do 16:26 What a harness actually is 22:03 RL environments explained 28:46 Does the data-labeling and RL environment business even last? 37:02 Why benchmarks don't tell you what works in production 38:12 Agent engineering vs forward deployed engineering 41:38 Deploying into 100-year-old enterprise systems 44:25 Why AI adoption is an org problem, not a tech problem 45:36 Hiring for judgment when engineers aren't really coding anymore 48:16 Why agents are a new kind of software 50:46 The first 90 days of an enterprise deployment 53:20 Why compliance environments break normal testing 56:58 Layering AI on AI to get to 99% accuracy 01:00:56 New models aren't always better — the swap problem 01:02:35 Improving agents without waiting for a new model 01:06:36 Does agent performance secretly degrade over time? 01:09:40 Why one model is never enough: the constellation approach 01:11:14 Building resilience when inference providers go down 01:13:48 When fine-tuning actually makes sense 01:14:53 Why voice-to-voice still isn't production-ready 01:16:25 The cascaded pipeline that real voice agents use 01:21:45 Audience Q&A: managing change inside the enterprise 01:23:24 Why inference getting cheaper makes things more expensive 01:26:54 Charging for outcomes instead of conversations 01:30:19 What the real moat is when everyone uses the same models 01:37:04 Synthetic data and where the data wall actually is 01:38:50 Closing thoughts

0:00-1:41:36

transcript

No transcript — this publisher did not publish one.

show notes

(We know the audio quality isn't great on this one :( But the conversation is still well worth it!)

Last week I hosted a fireside chat on what it actually takes to build AI agents in the enterprise with Natalie Meurer (Head of Agent Eng, Sierra), Harsh Trivedi (founding engineer, Mercor), Juhi Parekh (GM, Turing), and Kevin Lynch (Senior FDE, Fin).

We get into why new models aren't always better (and why you can't just swap in the latest release and assume your agent improves), how the data-labeling/RL environment business might only have a couple years left, why real-time voice-to-voice models still aren't production-ready, how cheaper inference is still causing prices to go up, how baking in a constellation of models into enterprise agents is so important for reliability,

and much, much, more!

Chapters below

00:00 Intro

00:21 Meet the panel

01:57 What everyone's actually using agents for day to day

06:10 The reality of forward deployed work

09:28 What agents couldn't do a year ago that they can now

12:29 Why you have to tell agents what NOT to do

16:26 What a harness actually is

22:03 RL environments explained

28:46 Does the data-labeling and RL environment business even last?

37:02 Why benchmarks don't tell you what works in production

38:12 Agent engineering vs forward deployed engineering

41:38 Deploying into 100-year-old enterprise systems

44:25 Why AI adoption is an org problem, not a tech problem

45:36 Hiring for judgment when engineers aren't really coding anymore

48:16 Why agents are a new kind of software

50:46 The first 90 days of an enterprise deployment

53:20 Why compliance environments break normal testing

56:58 Layering AI on AI to get to 99% accuracy

01:00:56 New models aren't always better — the swap problem

01:02:35 Improving agents without waiting for a new model

01:06:36 Does agent performance secretly degrade over time?

01:09:40 Why one model is never enough: the constellation approach

01:11:14 Building resilience when inference providers go down

01:13:48 When fine-tuning actually makes sense

01:14:53 Why voice-to-voice still isn't production-ready

01:16:25 The cascaded pipeline that real voice agents use

01:21:45 Audience Q&A: managing change inside the enterprise

01:23:24 Why inference getting cheaper makes things more expensive

01:26:54 Charging for outcomes instead of conversations

01:30:19 What the real moat is when everyone uses the same models

01:37:04 Synthetic data and where the data wall actually is

01:38:50 Closing thoughts