Skip to content
YPO Technology Network AI Brief

Your Agents Need a Spending Limit

Yesterday · 9 min · Season 1 · Episode 124 · 8.8 MB
0:00-9:14

Streams straight from the publisher. podnod never proxies or re-hosts episode audio.

For two years, "is your company good at AI" was a question about models and vendors. This episode goes where the answers actually live now: the engineers and operators publishing what works in production, in their own words, with their own numbers. What they have converged on looks nothing like the vendor decks. It looks like treasury management. One operator posted his AI bill and found 84 percent of it was cache traffic, then cut costs roughly in half by restructuring sessions. A SaaS company named Manifest built a four-tier model-routing system, ran it across 7,000 users for four months, and shut it down, because simple prompt caching saved more money more reliably. Sierra, which runs customer-facing agents for other businesses, published an architecture in which agents never hold live credentials at all. Zendesk disclosed an incident in which its AI agents looped for two hours because an unrelated database cleanup job held locks, the kind of boring ticket nobody review-gates. Ramp graded its bookkeeping agent against a 237-task suite and found that cutting a prompt 64 percent improved accuracy. Brex's engineers wrote the line of the year: upgrading the model improved investigation quality less than writing better runbooks. And Box put "AI model evaluator" on its payroll. Stephen Forte on the spending limit your agents do not have, the four-column controls one-pager to ask your team for, and why the frontier of AI management is not technical at all.