
How a Logistics Giant Keeps AI Data Locked Down
transcript
show notes
Picture a startup with ten engineers hammering away on AI and one person lying awake over a $20,000 bill that hasn't arrived yet. That's the scenario we put to Jason Ward, who handles FinOps for AI at C.H. Robinson and recently joined the FinOps Foundation's AI working group.
Jason's first answer is not glamorous: tag every AI resource so you know who owns it. The rest of the episode is what that makes possible. He walks through how C.H. Robinson runs AI across order entry, quoting, booking, and tracking, and why the team routes Anthropic models through Vertex AI to keep data locked down.
Then come the metrics. Jason uses AI to dig through his own observability platform for signals he didn't know were there. One of them is how chatty a model is. That signal turned a prompt bloat alert into a bug in the code that kept retrying and burning tokens. Its opposite, context starvation, burns tokens too: a model with too little context keeps failing and trying again.
We also cover why agentic and conversational workloads need separate baselines. Jason explains why cost per order is the easy win, and why most of the real work doesn't fit into neat discrete tasks. That's where his experimental cost per thought metric comes in, with reasoning ratio and cache hit rate alongside it. His advice is simple: your AI is the best tool you have for understanding your AI.
Alex Salkever: https://www.linkedin.com/in/alexsalkever
Jason Ward: https://www.linkedin.com/in/jward2
Timestamps:[0:00] Cold open[0:33] Meet Jason Ward from C.H. Robinson[1:18] AI use cases in logistics[1:52] Azure OpenAI and Vertex AI[2:19] Why keeping data in-house matters[2:44] How developers use AI day to day[3:40] Using AI to hack AI observability[4:21] Measuring how chatty a model is[5:45] Advice for a startup afraid of its AI bill[6:33] Step one is tag every AI resource[7:05] The Copilot billing blind spot[7:44] Why spend is only half the story[8:12] Break down cost by app and by model[9:09] The prompt bloat that exposed a bug[10:03] Context starvation[11:28] What goes into the AI spend report[12:21] Discrete tasks are low-hanging fruit[12:58] Agentic vs conversational workloads[14:12] Know the workflow before you report on it[15:02] The CTO's end goal[16:08] Cost per thought[17:59] Reasoning ratio and cache hit rate[19:06] Dynamic model routing[20:04] Use your AI to improve your AI