By Anant Adya, EVP, Service Offering Head and Head of Americas Delivery, Infosys
Up to 45 minutes. That’s how long it could take for reactive cloud operations to scale up to meet a surge in demand. No wonder then that reactive cloud operations are failing under real-time, data-intensive AI workloads, causing severe lags and service disruptions, to say nothing of the costs. Reactively scaling infrastructure, while requiring lower upfront investment, can cause substantial financial damage due to unplanned outages: this is because by the time traditional autoscaling kicks in (typically, after capacity utilization exceeds a defined threshold), it’s already too late.
Latency apart, a major concern with reactive cloud operations is that it troubleshoots issues after they occur rather than preventing them in the first place. It also assumes traffic is static and predictable, when in fact, AI workloads can swing hugely and unpredictably in a moment.
That’s not all. As infrastructure expands into vast multi-hybrid clouds, a single failure can rapidly propagate across the environment, causing an alert flood and failure chain reaction that legacy monitoring and manual systems find difficult to manage effectively. The October 2025 AWS outage is an example of this; an automation bug in the domain name system (DNS) sparked a cascading failure and disabled the control plane, defeating the reactive cloud operations.
Also, when multi-agent AI workflows scattered across cloud providers overload the servers unevenly, they create hardware bottlenecks and fritter away GPU resources, which reactive ops often struggle to predict and prevent.
Enter proactive AIOps
When cloud ops fail at AI scale, the damage can run up to millions of dollars per incident. Besides financial losses, enterprises may have to contend with alert fatigue, delayed remediation and critical security vulnerabilities. The good news is that AI also provides the solution: with artificial intelligence for IT operations (AIOps), enterprises can progress from rectifying issues after they have occurred to predicting and even preventing them in advance. Should a problem still occur, it is automatically taken care of by AIOps’ self-healing capabilities.
What AIOps is and what it does
In an AIOps model, artificial intelligence is embedded throughout the operational cycle:
- In the observation phase, AI processes data (real-time logs, metrics, etc.) to create dynamic baselines in place of static thresholds.
- During the diagnostic phase, it performs root-cause analysis across disparate data ecosystems.
- In the remediation phase, embedded AI agents repair issues by launching corrective measures independently.
- Finally, AI learns from all these outcomes to constantly hone its predictive abilities.
The model creates three strategic shifts in cloud operations:
- In place of regular monitoring, AIOps offers autonomous actions at every step: root-cause analysis, dynamic scaling and self-healing. For instance, upon detecting a problem, AI agents may launch remediation workflows such as restarting failed Kubernetes pods, clearing memory exhaustion or reversing bad deployments, without any instruction.
- Next, AIOps frameworks automate cloud governance. For example, upon seeing a configuration drift, they activate security policies to ensure compliance throughout the environment. Further, AIOps platforms correlate incidents across systems to identify the location and source of a problem within no time.
- Finally, these models optimize operationsat AI-scale by analyzing usage patterns across all cloud environments in real-time, continuously right-sizing resources and triggering cost-efficiency actions within the prescribed guardrails.
What’s driving its growth?
Globally, North America is the largest AIOps market, and it is rapidly advancing towards a projected 38.5 percent share of the AIOps market by 2035. The region’s advanced IT infrastructure and high cloud adoption are key factors of growth. In the United States, in particular, besides a mature IT ecosystem and enterprise cloud adoption, the government’s push for digital infrastructure is providing impetus to the market.
One of the top applications of AIOps is FinOps operationalization. Hybrid cloud has many benefits, but often runs the risk of overprovisioning. FinOps prevents this type of resource wastage by forcing different functions, including finance and engineering, to collaborate and share ownership for cloud usage. AIOps supports this goal by making decisions that balance cloud spends with performance.
The importance of governance
AIOps creates a self-healing cloud environment that is capable of detecting a problem, deciding and implementing the right course of action, verifying the outcome and improving (by itself) over time. This environment relies on analysis rather than set rules for making decisions, and autonomously resolves problems, typically even before human teams realize the issue.
Unfortunately, this also applies to AI-related risk. By the time people figure out that the model underlying AIOps suffers from automation bias or a lack of explainability, the damage is already done. Enterprises need to address this with responsible AI practices and AIOps governance, clearly defining which activities may be performed autonomously and which should be supervised. While enterprises should leverage AIOps to run cloud operations, they must also ensure that overall control is retained by humans.
##
ABOUT THE AUTHOR
Anant Adya is Executive Vice President, Service Offering Head, and Head of Americas Delivery at Infosys. He leads the company’s cloud, infrastructure, cybersecurity, AI, and digital transformation business across the Americas, helping global enterprises accelerate innovation, strengthen resilience, and modernize their technology landscapes. He also serves as a Trustee of Infosys Foundation USA, supporting initiatives that expand access to computer science and digital skills education.





