As artificial intelligence reshapes the technological landscape, the demands on data center infrastructure are evolving at an unprecedented pace. In this exclusive VMblog Q&A, Aniket Khosla, Vice President of Wireline Product Management at Spirent Communications, explores the intricate world of AI workloads and their impact on data center networks.
Khosla unveils the hidden complexities behind AI data centers, from the staggering costs of specialized hardware to the critical importance of network fabric optimization. With insights gained from real-world scenarios and conversations with industry leaders, he illuminates the path forward for organizations looking to harness the full potential of their AI investments.
As we stand on the brink of an AI revolution, this interview offers invaluable guidance for tech leaders navigating the challenges of building and maintaining cutting-edge AI infrastructure. Discover why thorough pre-deployment testing could be the key to unlocking billions in GPU investments and learn how the data center landscape must transform to meet the demands of next-generation AI workloads.
VMblog: What are the challenges of AI workloads in the data center?
Aniket Khosla: The AI data center landscape differs significantly from traditional data centers managed by major cloud providers, with most AI data centers existing as separate entities either physically or logically. Setting up these specialized AI facilities comes at a high cost, primarily driven by NVIDIA’s dominance in the market with their premium-priced GPUs, switches, and InfiniBand technology.
The expenses associated with establishing and operating AI data centers are substantial, covering operational aspects like power consumption, cooling systems, and the need for specialized expertise to ensure efficient operation. An inherent challenge in effectively managing these AI data centers lies in the Ethernet fabric that connects them, as issues such as packet loss and latency can impede optimal GPU utilization.
Certain forward-thinking customers are conducting initial tests on the Ethernet fabric of AI data centers to maximize the value of their GPU investments before full deployment. However, acquiring and configuring GPUs for testing presents complexities and resource-intensive challenges, adding another layer of difficulty for customers looking to conduct comprehensive fabric testing for their AI data centers.
Increasing the number of GPUs is not the solution. The key lies in optimizing the network for maximum efficiency. This requires a continuous, proactive testing and verification process to ensure organizations can fully leverage their infrastructure as they expand.
VMblog: How do data center networks need to transform to support new AI workloads?
Khosla: As a data center network operator, the initial decision involves choosing between the Ethernet and InfiniBand pathways, each with its own advantages and disadvantages. Let’s assume you’re going down the Ethernet route. There are multiple bottlenecks in an AI data center. The Ethernet fabric is one of them, and that’s the one we can help people test. Organizations need to test their fabric as scale before deploying. They’re spending billions of dollars buying GPUs, only to find out that 20-30% of the time, they’re sitting idle because the Ethernet fabric has high network latency.
So, make sure you test your AI fabric before you deploy. Make sure that you’ve identified and resolved potential choke points and congestion scenarios before deploying in a live production network. Know what your design needs to look like to either make your fabric lossless, or minimize these disruptions. Operators need to significantly transform the way they think about what kind of early testing they do before they actually deploy.
I’ve spoken with a large hyperscaler who didn’t do a lot of testing in pre-production and when they deployed a cluster, there were major issues. It’s very important to understand how the topology, scale, and fabric configuration will perform under real-world scenarios prior to deploying in a live network
VMblog: What is unique about testing AI and why is it so complex?
Khosla: AI workloads are different from traditional workloads due to their sheer scale. In a traditional data center, you have this notion of high-performance computing where you can throw a lot of compute power at a single job and get it done a lot quicker. In the AI world, these workloads are so massive that the computation of them has to be distributed across multiple systems. There’s no single system that can put or process all these AI workloads.
They are, by their very nature, distributed. You take a large AI job, break it into much smaller chunks, and then spread it across thousands or tens of thousands of GPUs for it to be “processed.” That comes with a lot of complexity. How do you optimize sending this data out to all these thousands of nodes running in an AI cluster? How do you train these GPUs? How do you make sure that once these things are distributed across thousands of GUPs, that every packet is getting to every destination it is intended to go to, on time, with low latency, and very low packet loss?
Unlike conventional high-performance computing setups where significant compute power can speed up tasks, AI workloads are too massive to be handled by a single system alone. The nature of AI workloads demands distribution across multiple systems, making them inherently complex.
VMblog: How have you seen AI evolve over the past 12 months?
Khosla: Just from a market perspective, the models have gotten smarter and more complex. There’s ChatGPT-4o that just came out. It’s pretty incredible. For example, it’s shown a mathematical problem, and it actually recognizes what the problem is when it looks at it, and it guides the student through how to solve the problem with voice prompts.
There’s been a lot of evolution in the past months, in what we’ve seen out there in the market. We’re just at the tip of the iceberg.
##






