

Solving the $2,000 Per Hour GPU Problem With Lumen
What happens when an enterprise invests heavily in AI compute but cannot move data quickly enough to keep those expensive processors working? In this episode of Infrastructure as a Conversation, I speak with Jim Fowler, Chief Technology and Product Officer at Lumen, about why network performance is becoming one of the defining factors in enterprise AI ROI. Jim brings experience from both sides of the relationship. He previously led technology programs inside large enterprises and served on Lumen’s board before joining its management team. His description of Lumen as a logistics company for information provides a useful way to understand the network’s role. Models can process information, but the network must deliver the right data to the right location at the required speed. Our conversation begins with the $2,000 per hour GPU problem. Jim recalls speaking with a Fortune 100 CIO whose company was training its own model using data stored across multiple clouds and regions. The company had GPU availability, but moving petabytes of data to those processors had become its largest bottleneck. Expensive compute was waiting for information. We discuss why an AI workload can perform well inside a controlled development environment before slowing down when users, data sources, and cloud services become distributed. Jim says many inference workloads require response times between 10 and 20 milliseconds. Factory automation can require latency below 10 milliseconds, while natural voice conversations may need responses within five to 10 milliseconds. Poor data movement also affects employees. Jim describes spending time with data scientists who began their mornings by moving information from three global regions. The transfers took between two and two and a half hours, leaving skilled employees waiting, attending meetings, or drinking coffee before they could begin modeling. In his words, these were “idle humans” waiting on the network. The problem becomes harder as AI workloads span public clouds, private data centers, colocation facilities, and edge environments. Jim says bandwidth growth between clouds and data centers is rising far faster than traffic between premises and cloud environments. Enterprises are also managing multiple carriers, tools, and operating systems, making performance and predictability harder to control. We examine data gravity and the decision to move compute closer to data rather than transporting large data sets over long distances. Security, resilience, sovereignty, compliance, and cost must be considered according to the requirements of each workload. Jim also explains why networking is beginning to behave like a cloud service. Traditional capacity could take weeks or months to provision. A software controlled network can increase or reduce capacity as demand changes. Lumen describes this model as Cloud 2.0, with automation, APIs, and real time control replacing static capacity planning. Our discussion closes with ownership. Network, cloud, application, data, and AI teams may each measure success differently. Jim believes a business owner should oversee the complete path from data source to AI outcome because a bottleneck can appear anywhere along the way. Could the biggest constraint on your AI investment be the network connecting your data and compute? Listen to the episode and share your thoughts with me.


















