Kubernetes has moved from experimentation to expectation. Enterprises are no longer just deploying containers; they’re running production workloads at scale, layering AI on top of cloud-native stacks and asking platform teams to deliver consistent, repeatable infrastructure across environments.
This shift introduces new challenges. Organizations now need to run containerized and VM-based workloads side by side, while managing clusters at massive scale. They must also power GPU-driven AI, all without getting trapped in proprietary ecosystems and ensuring tight security with efficient operations.
VMware vSphere Kubernetes Service (VKS) delivers a CNCF certified Kubernetes runtime as an integral part of VMware Cloud Foundation (VCF). It combines integrated lifecycle management, GPU enablement for AI workloads, and strong substrate-level isolation into a single platform.
To break it all down, VMblog sat down with Timmy Carr, Product Manager in the VCF Division at Broadcom, to discuss how VKS and VCF are helping enterprises run production Kubernetes at scale, enable AI workloads without lock-in, and supporting platform engineering teams.
VMblog: Let’s get technical—what specific architectural challenges or operational bottlenecks are you addressing for cloud-native teams in 2026?
Timmy Carr: The biggest bottlenecks we see for cloud-native teams are less about “VMs vs. containers” and more about how you run both, consistently, at scale. Architecturally, our customers want to platform their new applications on Kubernetes, while still operating alongside their existing, VM-based workloads that remain business-critical. We’re directly addressing the friction by delivering our Kubernetes offering and all services in VCF, including VMs, with a standard Kubernetes API so customers can run their VMs and containers together with a consistent security forward operating model.
Operationally, running Kubernetes at scale is difficult. Being able to stand up that many clusters in your environment, lifecycle them, and consistently manage them can create bottlenecks that affect day-to-day work. Because we’ve included management tooling as part of our stack and because networking, storage, load balancing, and performance are integrated into VCF out of the box, it’s easier to run Kubernetes infrastructure at scale.
VMblog: The cloud-native landscape is maturing rapidly. How has your technology evolved to meet the sophisticated demands of enterprises running production Kubernetes at scale?
Carr: We’ve evolved by validating and hardening our Kubernetes stack. We were one of the select launch partners for CNCF Kubernetes AI Conformance program and VKS is a CNCF-certified Kubernetes AI Platform. We also strengthened GPU and accelerator enablement, including tight integration with NVIDIA accelerators, so accelerators can be passed directly through to VKS. That lets enterprises run GPU-driven workloads within the same platform they’re already operating at scale.
Lastly, we are constantly expanding VKS support for the open source ecosystem. One solution to highlight here is Cosmonic in the WebAssembly (Wasm) space, enabling fast, optimized application experiences via wasmCloud on top of our platform. In 2026, we’re delivering more curated CNCF “design libraries” to help enterprises deploy production-ready architectures on VKS using their favorite open source and commercial tooling.
VMblog: Where do Broadcom’s offerings – both VKS and VMware Cloud Foundation – fit within the modern cloud-native stack? How do you integrate with and complement other CNCF projects and ecosystem tools?
Carr: VCF is the full-stack infrastructure platform that delivers a complete private cloud in your data center. It turns hardware into virtualized infrastructure. From that foundation, we build virtualized Kubernetes clusters. VKS is an integral part of VCF, and sits squarely at the Kubernetes layer, specifically as a certified Kubernetes runtime. That certification matters because we validate with the CNCF as we ship new Kubernetes versions, giving enterprises confidence that their Kubernetes workloads and CNCF projects that expect standard Kubernetes APIs will run consistently.
We continue to actively contribute upstream. A key example is Cluster API, which we helped create and underpins cluster lifecycle management in our stack, and has become a de facto standard across major on-prem distributions. We also contribute to projects such as Harbor, Velero, Contour, Antrea, etcd, and many others.
VMblog: Platform engineering continues reshaping how organizations approach cloud-native infrastructure. What’s your perspective on this shift, and how does your technology enable effective platform teams?
Carr: Platform engineering is becoming core to how modern software gets delivered because today’s software is both the applications and the infrastructure it runs on. Platform teams aren’t just standing up clusters; they’re now increasingly responsible for the full system, giving application teams a clear, repeatable path into production, whether that’s on-prem or across multiple clouds.
That’s where VKS fits in. We give platform teams a consistent Kubernetes substrate with the underlying infrastructure capabilities—compute/VMs, networking, storage, and load balancing—integrated into the VCF stack, so they can standardize once, manage lifecycle reliably, and securely provide those golden paths to production.
VMblog: Are you launching any new products, unveiling significant features, or announcing partnerships at KubeCon EU 2026?
Carr: At KubeCon Europe 2026, we are announcing that Velero has been accepted as a Sandbox project by the CNCF. Velero is a Kubernetes-native backup, restore and migration project that has grown through years of collaborations across maintainers, contributors and users. Velero’s move toward CNCF governance strengthens community stewardship, encourages broader participation and ensures long-term sustainability beyond any organization.
We’ve also been working with the etcd community to improve operational visibility and recovery for Kubernetes control planes, helping operators better understand cluster health and recover from failures.
On the partnership front, we are excited to share that our collaborations with cloud-native leaders across the Kubernetes ecosystem continue to expand, with newest validations on VKS being announced with F5, Kong and Tigera. These add to the existing VKS validations with Nvidia Run:ai, Mulesoft, Canonical, Cohesity, Redis, Vectara, and many others.
VMblog: Can you share a customer example that illustrates how VKS addressed a critical infrastructure or operational challenge—and what measurable outcomes the organization achieved as a result?
Carr: We’ve worked with several large enterprises that were historically running Kubernetes on bare metal infrastructure. There’s nothing inherently wrong with that, but over time, they began to see where efficiency challenges exist at scale. When Kubernetes runs directly on bare metal, you’re relying on the Linux kernel to schedule containers across CPU and memory. Our hypervisor is designed to context switch across virtual machines far more efficiently, which drives better utilization.
In some cases, customers increased infrastructure utilization from roughly 60% to 90%—a 30% gain in usable capacity or equivalent cost savings. Considering that a single server can cost tens of thousands of dollars before factoring in cooling, rack space, and networking, that improvement is significant. Then add in AI workloads, increasing data center demand, and footprint constraints tightening, the question becomes, “How do you get the most out of the smallest footprint?”
That’s where virtualization delivers. It’s a value proposition VMware has had for decades, fitting more workloads into fewer boxes. This is something we’re now bringing directly to the Kubernetes community.
VMblog: What clearly differentiates Broadcom’s offerings in an increasingly crowded Kubernetes market—and why should CTOs view your approach as strategically distinct from other vendors offering similar capabilities?
Carr: A key differentiator is that we are a virtualization company at our foundation. Virtualization enhances the security, performance and economics of running Kubernetes in the enterprise.
On bare metal, Kubernetes relies on a shared Linux kernel. If it’s compromised, everything on the cluster is exposed. With virtualization, each component runs in its own VM with its own kernel, creating stronger isolation. Now, think about if the kernel gets compromised, that’s a serious problem, and that should matter to every CTO.
Performance is equally important. Running on a hypervisor substrate allows our intelligent scheduling algorithms to dynamically allocate CPU across VMs. One VM can receive more resources when it needs them, and the workloads don’t even notice the shift. That level of control and predictability becomes critical at scale.
From a CIO perspective, there’s also the benefit that you don’t need to purchase additional licenses or stitch together separate vendors.
VMblog: Platform engineers are increasingly responsible for enabling AI teams. How are you helping platform teams support AI workloads at scale?
Carr: Most modern AI workloads run on Kubernetes, so we focus on ensuring that foundation is ready. We tightly integrate with vendors like NVIDIA, enabling GPUs and accelerators to be attached directly to Kubernetes clusters. And, we offer a CNCF certified Kubernetes AI platform, so teams can run the tools they need from the broader ecosystem with confidence.
For platform teams that don’t want to assemble everything themselves, VCF also includes VCF Private AI services out of the box, which include a model runtime, model gallery, agent builder service, data indexing and retrieval service, etc. If you don’t know how to build inferencing pipelines or manage those components, we can provide that capability as part of the stack.
At scale, AI models are large, memory-intensive, and maintain context while they’re running. In most environments, restarting the workload to perform maintenance means losing context and disrupting users. By running on our vSphere virtualization platform, we can use vMotion to move a running workload in real time to another machine without stopping it. Never having to restart, never losing context with minimal interruption.
VMblog: Enterprises want AI capabilities without locking themselves into proprietary ecosystems. How does your commitment as a top CNCF contributor influence your approach to AI on Kubernetes?
Carr: Open standards guide our approach.We helped bootstrap the CNCF Kubernetes AI Conformance effort as a launch partner because we believe enterprises should be able to run AI workloads on a Kubernetes substrate that’s validated by the community. When it comes to executing AI workloads, we support all the capabilities teams need for AI–GPU, scheduling, and the surrounding operational tooling–in a way that remains consistent with CNCF standards.
There’s often an assumption that running on VMware means locking into the hypervisor. The reality is entirely different. If you’re using VKS, you’re running CNCF-certified Kubernetes. If you run your AI workloads on VKS, you can run on any CNCF-conformant Kubernetes distribution. At the same time, customers choose to stay because of what we deliver under Kubernetes: efficiency, performance, and security.
##






