
Industry executives and experts share their predictions for 2019. Read them in this 11th annual VMblog.com series exclusive.
Contributed by Govind Rangasamy, CEO, Appranix, The Site Reliability Automation Company
Container Data Management, SREs, AIOps for Reliability Automation
Container Data Management with Kubernetes
As cloud-native architectures take off, Kubernetes has spawned a completely new ecosystem that rival ecosystems created by VMware and AWS. For those that have been in the technology industry for a long time and seen the emergence of many ecosystems, the speed with which Kubernetes has emerged as the de-facto standard is nothing but astonishing and simply makes other container ecosystems irrelevant. Without having a standard, solving thorny issues like data management in the container space is very difficult, especially considering the need to co-exist with existing enterprise data and network infrastructures. We predict 2019 will bring some of the much-needed collaborations and toolsets to solve storage and data infrastructure around the Kubernetes ecosystem that could lead to better application reliability and movement of application containers between the cloud zones, regions and between two clouds.
The Rise of Site Reliability Engineers
Along with Kubernetes, another Google service/site/app management concept, Site Reliability Engineering, is taking hold across organizations that need to manage distributed systems and large always-on services. Service management will have to be data-driven, codified and continually automated. As infrastructure-as-code gets standardized for managing a cloud platform or multiple cloud platform services together, the old IT operations methodology and service management, in particular, will go through a sea change. The emergence of Site Reliability Engineering and the need to manage application reliability has started attracting IT operations teams to collaborate with developers more and particularly some of the infrastructure aware developers are becoming the core SRE team members. We have been hearing more and more technology CxOs getting rid of their old IT ops teams and forming SRE teams along with the movement to the cloud and containers. Just run a LinkedIn search for Site Reliability Engineers to see how many open positions that offer attractive salaries pop up.
Availability and uptime are keys to running any digital enterprise. Old approaches to disaster recovery and business continuity will have to be thrown out the window in order to provide the necessary SLAs for an always-on enterprise. Service Level Objectives (SLO) based on RTO and RPO requirements need to be data-driven with Service Level Indicators (SLI).
Culture and Resilient Digital Infrastructures
Organizational culture is difficult to change. Think about the overall DevOps culture, it’s great on paper but it takes a village to actually put to work. However, we could point to respected analysts in the Infrastructure and Operations space, for instance, Mark Jaggers, Senior Director Analyst at Gartner “I&O leaders planning and delivering resilient digital infrastructure must realize that people are just as important as infrastructure and processes,”. Gartner recommends a culture of continuous improvement to create a resilient digital infrastructure by focusing on the four key principles of Site Reliability Engineering,
- Eliminate Toil Through Automation
- Set Service-Level Indicators and Objectives
- Perform Blameless Root Cause Analyses
- Create an SRE Community of Practice
I&O leaders need to shift focus from rapidly fixing problems to continuous process improvement.
Chris Gartner at Forrester says Forty-Six Percent Of Google’s SRE Principles Apply Directly To Your Enterprise. Focusing on the core principles of Site Reliability Engineering such as continuous automation with proper Service Level Indicators that are application and infrastructure performance driven is key.
AIOps for Reliability Automation
IT infrastructure operations, disaster recovery automation, data analytics, predictive maintenance, and machine learning have all existed before AIOps. It is about data collection, analytics, machine learning, and artificial intelligence used together to satisfy several I&O use cases such as anomaly detection, causal analysis, prediction, alarm management, and intelligent remediation. These AI Operations concepts will be expanded to covering the service management to improve applications reliability in 2019 and beyond for always-on enterprises.
##
About the Author
Govind Rangasamy, a serial entrepreneur, is founder and CEO of Appranix. With extensive experience in building products in the cloud automation and enterprise IT management space, Govind founded Appranix with a belief that existing infrastructure centric cloud and IT automation solutions are completely inadequate to handle application resiliency. Prior to Appranix, Govind was the CEO of FogPanel, a multi-cloud service management company that was sold to UST Global. Before starting FogPanel, Govind led Actifio’s cloud and resiliency solutions group. He also led products at Eucalyptus Systems, an AWS compatible open source private cloud leader. Prior to Eucalyptus, he successfully transformed HP’s Storage Management product line to be a leader in the Gartner Magic Quadrant.





