Opens in a new tab
vmblog logo 2024 wht (updated)

Unlocking the Power of Cloud Native Platforms for AI Workloads

Share: 

David Marshall | Published: October 9, 2024

Introduction

In the rapidly evolving landscape of artificial intelligence (AI), cloud-native platforms have emerged as the backbone for deploying, managing, and scaling AI solutions. These platforms offer a plethora of features that cater to the unique demands of AI workloads, ensuring high performance, resilience, and ease of management. This article delves into the essential features that a cloud-native AI solution must possess, focusing on the power of GPUs, high-speed storage, advanced networking, resilience, multitenancy, deployment flexibility, and full observability.

Unlock the Power of GPUs

GPUs (Graphics Processing Units) are the workhorses of AI, providing the computational power needed for training and inference of complex models. A cloud-native AI solution must support a variety of GPU types and models, as different AI workloads have different requirements. The platform should be able to discover and categorize these GPUs, ensuring that the right type of GPU is allocated to the right workload. This GPU awareness extends beyond mere recognition; it involves automated workload placement that takes into account the specific capabilities and performance characteristics of each GPU.

Supply Data with High-Speed Storage

AI workloads are data-intensive, requiring high-speed storage solutions to feed data to GPUs efficiently. A cloud-native AI platform should support various types of storage, including block, file, and object storage, to cater to different data needs from core to edge. Advanced features like data compression, thin provisioning, and per-application replication policies are crucial for optimizing storage usage and ensuring data availability. The platform should also offer near bare-metal performance for storage, ensuring minimal latency and high throughput.

Provide Advanced Networking

Networking is another critical aspect of cloud-native AI solutions. The platform should support multiple networks, including overlay and underlay networks, and provide multiple IP addresses per pod or network function (NF). Technologies like SR-IOV and Open vSwitch can be used to ensure high throughput, low jitter, and redundancy. Customizable CNI plugins, support for IPv4/v6, and built-in load balancers like MetalLB are essential for robust and flexible networking.

Resilience with Advanced Workload and Storage Placement

Resilience is paramount in AI workloads, which often run for extended periods and require high availability. Resilient application and storage-level disaster recovery (DR) for AI applications ensures continuous availability and app+data integrity by implementing advanced object tracking, for app+data DR failover mechanisms. Automated recovery processes across distributed environments further enhance reliability, minimizing downtime and data loss. 

A cloud-native AI platform should offer advanced workload and storage placement strategies, taking into account NUMA awareness down to every physical node. This high granularity ensures optimal resource utilization and prevents configuration failures. The platform should also support CPU and CPU sibling isolations, NIC bonding for redundancy, and built-in resource pooling to maintain high availability even under node failure. 

Policy-Driven Operations

Every operation in a cloud-native AI platform should be policy-driven, ensuring consistency and reducing the risk of errors. This involves modeling and discovering all resources and configurations, eliminating the need for manual intervention. Policies can govern various aspects of the platform, from resource allocation and workload placement to security and compliance. This approach ensures that the platform can adapt to changing requirements and workloads dynamically.

Deploy Anywhere: Far Edge, Intermediate, and Core-Cloud

Flexibility in deployment is crucial for distributed AI solutions, which may need to run in diverse environments ranging from far edge to core-cloud. A cloud-native AI platform should support deployment across private, public, hybrid, and multi-cloud environments. This flexibility ensures that AI workloads can be placed close to the data source, reducing latency and improving performance. The platform should also support bare metal server lifecycle management (LCM), NF and supporting application LCM, and service stitching and LCM for seamless deployment and management.

Full Observability

Observability is critical for monitoring and managing AI workloads. A cloud-native AI platform should offer a multitenant and RBAC and multi-tenant observability suite for every node and workload. This includes detailed metrics, logs, and traces that provide insights into the performance and health of the system. Full observability ensures that issues can be detected and resolved quickly, maintaining the reliability and performance of AI applications.

Easy-to-Use App Store Experience for Deployment and Lifecycle Management

Ease of use is a significant factor in the adoption of cloud-native AI solutions. The platform should offer an app store experience for deployment and lifecycle management, allowing users to deploy entire multi-domain workflows at the click of a button. This includes rolling upgrades across NFs, services, hardware, and software components, as well as the deployment of appliances, switches, routers, and more. The platform should eliminate the need for hunting and hardcoding, CLI experience, or manual configuration, making it accessible to a broader range of users.

Rich Multitenancy and Chargeback

Multitenancy is a key feature for cloud-native AI platforms, allowing multiple users or teams to share the same infrastructure while maintaining isolation and security. The platform should support a potent combination of separation based on tenant, cluster, pod, container/VM, namespace, and cloud resource pools. Resource pools can provide isolation at both the user and application levels, with automated application placement ensuring that workloads are allocated to nodes with the required resources. Personalized chargeback for resources like GPU, CPU, memory, and storage is also essential for effective resource management.

Conclusion

A cloud-native AI solution must be designed with a comprehensive set of features to support the unique demands of AI workloads. From unlocking the power of GPUs and providing high-speed storage to advanced networking, resilience, flexible deployment, full observability, and ease of use, these platforms are essential for the successful deployment and management of AI applications. By leveraging policy-driven operations and ensuring high granularity in resource management, cloud-native AI platforms can deliver the performance, reliability, and flexibility needed to unlock the full potential of AI, anywhere on the planet.

To learn more about Kubernetes and the cloud native ecosystem, join us at KubeCon + CloudNativeCon North America, in Salt Lake City, Utah, on November 12-15, 2024.

##

ABOUT THE AUTHOR

W. Brooke Frischemeier, Sr Director, Head Of Product Management, Rakuten Cloud BU, Rakuten Symphony

Brooke-Frsichemeier 

Brooke has spent his last 10+ years building cloud platforms and services.  His overall experience is a strong mix of small start-ups, large companies and systems integrators. Currently he leads product management for the Rakuten Symphony cloud team, where he manages the cloud-native platform, cloud-native storage and orchestration solutions.