A recent survey by Edge Delta and the Cloud Native Computing Foundation (CNCF) found that more than 60% of enterprises use Kubernetes, and 96% of organizations are using or evaluating it. This coupled with the rising use of AI – which a 2025 McKinsey & Company State of AI report placed at 20% in 2017 increasing to 78% in 2024 – means organizations need to rethink/evolve their approach to running and managing workloads. To address this, platform engineering teams can benefit from taking a closer look at bare metal servers which can help ensure high performance and efficient resource management for computationally intensive tasks.
To learn more about the proliferation of AI and the benefits of bare metal for Kubernetes and GPUs, VMblog talked with Lukas Gentele, CEO of vCluster.
++
VMblog: The use of Kubernetes is now so widespread, and having just passed its tenth birthday, the use cases have matured. Now that more organizations are leveraging the benefits of AI, the game is changing. What are some best practices or considerations that you can share?
Lukas Gentele: We have achieved significant performance improvements and efficiencies running Kubernetes on bare metal. This move has helped significantly for AI/ML workloads, but moving to bare metal requires careful planning and attention to several best practices to ensure optimal performance, reliability, and security. Key areas include hardware selection, infrastructure-as-code, high availability, security, and resource management. Other considerations include scalability, AI/ML pipelines and local development.
VMblog: How does multi-tenancy impact AI workloads, and how does vCluster address it?
Gentele: Multi-tenancy can significantly impact AI workloads, both positively and negatively. While it can improve resource utilization and reduce costs by sharing infrastructure, it can also introduce challenges related to performance isolation, security, and operational complexity so organizations need to focus on these areas with the right solutions.
Positive impacts include: improved resource utilization enabled by multi-tenancy allowing multiple AI workloads to share the same underlying infrastructure (GPUs, CPUs, storage), leading to better utilization of expensive hardware and potentially lower costs; reduced costs by sharing resources resulting in fewer physical machines being needed (ie, lowering Capex and Opex); simplified management since managing a single, shared infrastructure is often easier than managing multiple isolated environments, particularly for large organizations; and, faster experimentation because multi-tenancy can facilitate quicker experimentation with AI models by providing a readily available shared environment.
Challenges related to performance isolation (aka “the noisy neighbor”), security, and operational complexity are the reasons we built vCluster and vNode, which create virtual clusters and virtual nodes within a host cluster, enhancing isolation and resource management for multiple tenants. These solutions address the traditional drawbacks related to multi-tenancy, and ensure organizations only experience the positive benefits.
VMblog: What do these changes mean for platform engineering teams?
Gentele: Increased AI workloads will bring about several major changes for platform engineers, transforming their roles and responsibilities in a significant way. A few areas to consider and prepare for are:
- Managing Specialized Infrastructure – AI workloads demand specialized hardware like GPUs and TPUs for computationally intensive tasks like model training and execution. Platform engineers will need to manage the allocation, scaling, and cost control of these resources effectively.
- Scaling Compute Resources – As AI applications grow in complexity and data volumes increase, the underlying infrastructure must be able to scale up or out, including storage and networking capabilities. Platform engineers will be crucial in implementing solutions like automated scaling based on traffic patterns and resource slicing to manage these resources efficiently.
- Infrastructure Fragmentation – AI workloads often span multiple environments (on-premise, public clouds, edge locations). Platform engineers will face the challenge of maintaining consistency, observability, and security across these fragmented environments.
- Developer Self-Service – Data scientists and AI engineers will increasingly desire self-service access to the infrastructure and tools they need. Platform engineers will be responsible for creating platforms that enable this, while ensuring proper governance through role-based access controls and policies.
VMblog: What else should organizations consider when preparing for future workloads like AI and LLMs?
Gentele: The first step – if they haven’t already – is to invest in AI infrastructure: This includes specialized hardware like GPUs to handle the high computational demands of AI and LLM workloads. Similar to bare metal Kubernetes, bare metal GPUs offer several advantages, particularly for tasks requiring high performance and consistent access to hardware. They provide dedicated access to the GPU’s resources, eliminating the overhead of hypervisors and virtualization, leading to increased speed and consistency for applications like AI training, model inference, and scientific simulations.
VMblog: What’s next for vCluster? Anything else our readers should know?
Gentele: Our product roadmap is focused on features that support the future of Kubernetes tenancy, as we believe providing multi-tenancy will address the broadest set of challenges with managing Kubernetes implementations, by providing more control over isolation performance and cost. The significant use of AI also has dictated how we are helping customers future-proof their infrastructure, so our roadmap is heavily focused on features that increase GPU utilization for AI workloads; running Kubernetes on bare metal to eliminate the cost and overhead of VMs; and, everything platform engineering teams need to build secure, scalable, multi-tenant Kubernetes environments. To address this, we recently introduced Private Nodes and Autoscaling powered by Karpenter, and soon we will introduce Standalone vCluster.
##






