By Prabh Simran Singh, Sr. Director, Kubernetes Data Protection, NetApp and Ashwin Palani, Principal Engineer, Kubernetes Data Protection, NetApp
Kubernetes is on track to be the infrastructure of choice for modernizing applications, especially AI workloads, with an 85% adoption rate by global organizations anticipated by 2027, according to Gartner. This highlights the need for storage solutions that are container-centric and cloud-native. These solutions should provide seamless integration and persistent, scalable storage for containers and microservices.
Data is a pivotal asset for enterprises, driving strategy and necessitating rigorous protection. The recent cyberattack on Maersk, which led to a $300 million loss, starkly illustrates the devastating consequences of compromised security and the need for a comprehensive data protection strategy. Businesses must ensure data integrity and the portability of applications to enable recovery and continuity under any conditions.
Keys to a successful data protection strategy
In Kubernetes, managing stateful applications introduces complexity. Ensuring that these applications can scale seamlessly, manage persistent storage effectively, and maintain data integrity and consistency is challenging. Proper data management is crucial because stateful applications require persistent storage that endures pod restarts and rescheduling. Preventing data loss during failures or network issues is critical for reliability. Additionally, applications handling sensitive information must comply with regulations like General Data Protection Regulation (GDPR) and Health Insurance Portability and Accountability Act (HIPAA), necessitating stringent data protection measures.
To address these challenges, organizations need to think through the following to have a robust data protection strategy:
- 3-2-1 backups. This strategy advocates having three copies of data stored on two different types of media, with at least one copy kept off site. Using object storage for off-site/off-cluster backups offers a durable, scalable solution resistant to infrastructure disruptions.
- Comprehensive application backups. Backups should encompass both the application data and its metadata. This holistic approach guarantees that not only the essential data is preserved but also the configurations and settings necessary to restore the application to its operational state seamlessly.
- Recovery point objectives (RPOs). Establish clear RPOs to define the maximum acceptable amount of data loss. To meet these RPOs, implement regular, application-consistent backups that are performed frequently. This strategy safeguards against data loss by enabling restoration to the most recent stable state whenever issues arise.
- Disaster recovery. Cross-cluster replication enhances disaster recovery by ensuring that data copies exist in multiple locations, providing resilience against regional failures.
- Encryption. Encryption at rest and in transit is important for policy compliance.
- Recovery time objectives (RTOs). Efficient recovery processes enable swift restoration during disruptions, minimizing downtime and maintaining operational continuity.
- Role-based access control (RBAC). This control is vital in multitenant environments for ensuring resource isolation.
Data protection for new workloads
One of the key aspects of having a robust data protection strategy is the ability to support all existing and new workloads. Over the past 12 months, we have seen an increase in the adoption of AI and virtualization workloads running on Kubernetes, and including those as part of the unified strategy is critical.
Artificial intelligence/machine learning
In the realm of AI model development, data protection is vital, threading through each phase from data preparation to model deployment. It begins with securing data integrity, forming a reliable base for AI construction. As the AI model matures, it’s crucial to protect training checkpoints and configurations to preserve progress and flexibility. Secure backups at the deployment stage ensure that AI can consistently perform, even during disruptions. This protection is key to maintaining the accuracy of training datasets, ensuring compliance with governance standards, and facilitating swift disaster recovery. Together, these practices lead to AI applications that are both trustworthy and strong, capable of delivering precise outcomes in a dynamic technological environment.
Modern virtualization
As enterprises modernize IT, many are transitioning from virtual machines (VMs) to containerized environments. KubeVirt bridges this gap by enabling VMs to run within Kubernetes, thereby unifying platforms and allowing a gradual migration to container workloads without the need for immediate refactoring.
However, this transition brings its own set of challenges. Managing VM lifecycles alongside containers requires careful orchestration, and data protection is critical. VMs need block device support for high-performance storage, quality of service (QoS) for tiered storage management, and continuous replication to maintain data consistency. Ensuring data consistency during backup is vital, and it can be achieved through VM-level quiescing, pausing operations for consistent snapshots. Leveraging features like snapshots and rapid cloning helps organizations manage VMs effectively within Kubernetes.
Introducing advanced data protection
Recognizing the crucial role of data protection in maintaining data integrity and enabling disaster recovery, NetApp® TridentTM, our proven data provisioning solution, is broadening its capabilities. Trusted by numerous customers to manage vast amounts of storage, Trident is poised to enhance its offering with integrated data protection features.
This enhancement brings a new layer of resilience to your Kubernetes deployments, offering advanced features like high-speed data replication and cross-cluster disaster recovery. With the added convenience of Kubernetes API extensions, developers can now safeguard their environments more effectively within the familiar confines of their Kubernetes workflow. Keep an eye out for these new features in Trident’s upcoming release, designed to bolster your data against the unexpected while maintaining a seamless and empowering developer experience.
NETAPP, the NETAPP logo, and the marks listed at http://www.netapp.com/TM are trademarks of NetApp, Inc. Other company and product names may be trademarks of their respective owners.
To learn more about Kubernetes and the cloud native ecosystem, join us at KubeCon + CloudNativeCon North America, in Salt Lake City, Utah, on November 12-15, 2024.
##
ABOUT THE AUTHOR
Prabh Singh is an engineering leader at NetApp, where he leads the Kubernetes provisioning and data protection teams with a strategic focus on elevating NetApp storage as the premier choice for containerized workloads. His extensive experience in guiding enterprises through Kubernetes and cloud-native transformations has been essential in shaping the future of containerization and cloud adoption across the industry.
++
Ashwin Palani is an accomplished Kubernetes and microservices developer with expertise in storage and data management. With a passion for innovation, he has contributed to several features across NetApp’s suite of storage solutions. Ashwin is a firm believer in continuous learning and is always excited by the opportunity to explore new technologies.





