
transcript
show notes
Most budget conversations about private AI infrastructure start and end with GPU sticker prices — and that's exactly where the planning goes wrong. This episode of LLM.co delivers a rigorous, component-by-component total cost of ownership analysis for a private large language model cluster serving 200 knowledge workers at a regulated organization, drawing on this detailed GPU cluster cost breakdown current to late 2026. Whether you're at a mid-sized bank, a hospital network, or a defense contractor, the math is more tractable — and more illuminating — than most teams expect.
The episode walks through every major budget line over a three-year horizon, explaining not just what things cost but why each variable moves the total the way it does:
- Concurrency shapes hardware, not headcount. At 10–15% peak concurrency, 200 users typically require just two eight-GPU nodes — not a data center — and the right GPU choice (H100 vs. L40S) depends entirely on model size and inference workload.
- Capital expenditure for a two-node H100 build runs roughly $820K–$1.4M over three years, covering GPUs, high-speed networking, storage for weights and vector indexes, and rack infrastructure; an L40S equivalent comes in at about half.
- Power and cooling are a hidden budget line. Two nodes drawing ~22 kW of IT load, factoring in a realistic facility PUE, translate to $280K–$580K in colocation and power costs alone over three years.
- Software is free to download but expensive to run. Open-weight runtimes cost nothing upfront, but commercial support across 16 GPUs adds $210K–$300K over the life of the system — a line item many initial estimates omit entirely.
- Staffing is the largest non-hardware cost. A realistic 1.5–2 FTE model (platform engineer plus partial ML, security, and data engineering support) runs $400K–$700K over three years, dwarfing networking and storage budgets combined. For organizations exploring broader deployment, custom LLM deployment services can shift some of that operational burden off internal teams.
- All-in, the three-year TCO for a production-grade H100 cluster with high availability, software support, and staffing lands at $1.9M–$3.2M; an L40S build for smaller models runs $1.2M–$1.9M. For teams in regulated sectors like healthcare or finance — where private LLM infrastructure for financial services carries additional compliance requirements — understanding these ranges before procurement is essential.
The episode also covers a practical headroom rule: clusters running at 90% utilization leave no room for the fine-tuning jobs that inevitably surface next quarter, making 60% steady-state utilization the smarter design target. If you're preparing for procurement conversations, the related episode What to Put in a Private LLM RFP Before You Sign is the logical next listen.