Skip to content
Artwork for cloud2030
cloud2030 · May 29 · 28 min

Back After a Break

In this episode, we discuss the rising cost of using AI and how usage-based pricing, model changes, and capacity limits are affecting daily work as AI moves from experimentation into operational use. We also talk about multi-model workflows, hybrid infrastructure, and examples of using hosted models alongside open models locally for tasks such as writing and named entity resolution. We get into the need for enterprises to run their own AI infrastructure, including questions around GPU pooling, routing, reservation, data sovereignty, and service levels. Transcript: https://otter.ai/u/pihJkUzDWWcqBnWyM24CIyxX4Qs?utm_source=copy_url

0:00-28:40

transcript

No transcript — this publisher did not publish one.

show notes

In this episode, we discuss the rising cost of using AI and how usage-based pricing, model changes, and capacity limits are affecting daily work as AI moves from experimentation into operational use. We also talk about multi-model workflows, hybrid infrastructure, and examples of using hosted models alongside open models locally for tasks such as writing and named entity resolution. We get into the need for enterprises to run their own AI infrastructure, including questions around GPU pooling, routing, reservation, data sovereignty, and service levels.

Transcript: https://otter.ai/u/pihJkUzDWWcqBnWyM24CIyxX4Qs?utm_source=copy_url

more episodes

All episodes