Opens in a new tab
vmblog logo 2024 wht (updated)

KubeCon NA 2025: Ciroos CEO Ronak Desai on How AI SRE Teammates Are Transforming Site Reliability Engineering and Kubernetes Operations – VMblog QA

Share: 

David Marshall | Published: November 5, 2025

     

As organizations grapple with increasingly complex hybrid and multi-cloud environments, the traditional approach to site reliability engineering is reaching its breaking point. Manual dashboard investigations, cross-domain complexity, and alert fatigue are pushing SRE teams to their limits, particularly as AI-native workloads built on Kubernetes demand faster, more intelligent operational responses. At KubeCon + CloudNativeCon North America 2025 in Atlanta, Ciroos is showcasing a fundamentally different approach: an AI SRE Teammate that uses multi-agentic AI reasoning to automate root cause analysis, proactively detect anomalies, and enable autonomous operations across multiple domains.

In this exclusiveVMblog pre-show Q&A, Ronak Desai, co-founder and CEO of Ciroos, explains how his company is reimagining observability and production operations for the modern cloud-native era. From achieving 20x faster incident resolution to delivering 10x productivity gains in production deployments, Desai discusses the measurable business impact that’s convincing CTOs and platform engineering teams to rethink their operational strategies. Attendees can experience interactive demos at booth #1552 or join an exclusive evening event on November 12 to explore the future of AI in SRE alongside industry leaders. Read on to discover how Ciroos is helping enterprises transform from reactive firefighting to proactive, intelligent operations. 

++

VMblog: Can you give us your elevator pitch? What key message will attendees hear from you at KubeCon NA 2025, and what actionable insights will they take back to influence their management teams?

Ronak Desai:  Ciroos is reimagining observability and production operations. Our AI SRE Teammate acts as a digital co-worker for operations teams to automate, augment, and drive autonomous operations across hybrid, multi-cloud, and multi-domain environments. Our platform uses multi-agentic AI reasoning to identify root causes, surface insights, and take corrective action.

Attendees will see how this approach transforms operational practices from reactive to proactive, reduces SRE toil, and creates measurable productivity gains for engineering, operations, and platform teams.

VMblog: Where can attendees find you at the event? What hands-on demos, interactive experiences, or booth activities have you designed to showcase your technology?

Desai:  You’ll find us at booth #1552, where we’re running interactive demos of our AI SRE Teammate. Attendees can simulate real incidents – Kubernetes events, cloud misconfigurations, cross-domain issues, and other anomalies – and watch our AI SRE Teammate reason through them in real time. Attendees can also explore how to apply learnings from the analysis to increase overall system reliability.

Attendees interested in learning more can visit ciroos.ai to request a demo or connect with our team before the conference.

VMblog: Are you hosting any exclusive events, networking sessions, or after-hours meetups during KubeCon? How can attendees participate?

Desai:  Yes, we are hosting an exclusive event on the evening of Wednesday, Nov. 12, called “An Evening with Ciroos: Exploring the Future of AI in SRE.” It’ll be a small, high-caliber gathering – founders, engineering heads, and reliability leaders exchanging perspectives on how AI is transforming operational intelligence and resilience. To reserve a spot, interested attendees can register at https://luma.com/vfptfgrj (space is limited and spots are filling up fast; so they are subject to availability).

VMblog: Can you dive deeper into your company’s core technologies? What specific challenges do you solve for KubeCon attendees in their day-to-day operations?

Desai:  Ciroos is a multi-agentic AI platform that mimics how expert SREs reason through complex issues. Each AI agent carries domain-specific skills – spanning Kubernetes, cloud, networking, security, and AI stacks – and collaborates to diagnose anomalies and identify the root cause.

We help teams overcome four major challenges:

  • Manual, dashboard-driven investigations
  • Cross-domain complexity and knowledge silos
  • Alert fatigue and escalating MTTR
  • Limited bandwidth to address non-critical tickets 

We address these challenges by surgically analyzing telemetry across a range of enterprise observability, cloud, incident response, collaboration, CI/CD tools, and more. With just read-only access to a subset of these tools, Ciroos dynamically builds a knowledge graph with minimal human intervention. Then, based on an incoming service level objective alert or other event, Ciroos applies its patent-pending Behavior Patterns technology to dynamically reason like a human expert to pinpoint root causes, answer difficult questions during investigations, proactively detect anomalies before they become incidents, and remediate issues that cross multiple domains. Participants can interact with our AI SRE Teammate to gain a real understanding of how to combine the best of AI and human expertise. 

The end result: slash time to root cause discovery by more than 95%. With a range of deployment options – including SaaS or self-hosted, cloud or on-premises – and advanced role-based access controls per user, Ciroos offers the broadest flexibility to meet enterprise-grade security requirements, delivering value in minutes. At all times, humans are in control, choosing their desired level of augmentation and autonomous operations throughout their AI journeys. 

VMblog: With GenAI workloads and LLM deployments reshaping cloud-native architectures, how does your solution address these AI infrastructure demands?

Desai:  A majority of GenAI workloads are built on Kubernetes as a foundation. SRE teams can’t rely on legacy, dashboard-based “click operations” that are inherently manual and too slow for such workloads. Ciroos’s AI SRE Teammate is built for this environment – using reasoning-based AI to analyze telemetry, apply relevant context, correlate cross-domain dependencies, detect early signals of failure, and recommend corrective actions. This enables enterprises to maintain reliability and performance objectives even as AI workloads grow more complex and dynamic. 

VMblog: What’s your executive pitch for CTOs and CIOs? How do you demonstrate measurable business impact and ROI?

Desai:  Ciroos turns operational efficiency and reliability into a strategic advantage. Our production deployments in large enterprises show up to 20x faster resolution, 10x productivity gains, and over 80% RoI. These metrics translate to reduced downtime, improved developer velocity, and superior reliability. By teaming AI with their operational staff, Ciroos allows leaders to scale site reliability operations while keeping costs and resource usage in check. 

VMblog: What are the biggest obstacles to Kubernetes adoption and scaling in 2025? How does your solution help organizations overcome these barriers?

Desai:  Today, Kubernetes is a mainstream architectural choice for modern applications, including AI-native workloads. However, Kubernetes does not work in isolation – it depends on the underlying infrastructure, managed database services, third-party services, services based on VM- or monolithic-based systems. Platform engineering teams have to wrestle with huge spikes in demand from their tenants. Rather than viewing Kubernetes as an isolated domain, Kubernetes platform owners need to consider the system as a whole, including the aforementioned interactions, even as they strive for a higher service-level objective. Without this expanded horizon, we see platform teams struggling to understand the “why” behind failures. Ciroos provides a unified reasoning layer across these integrations and interactions, connecting symptoms across multiple domains to root causes before they become incidents, and reducing investigation cycles from hours to minutes. Our reasoning goes deep to uncover hidden issues beyond the superficial, symptomatic level. 

VMblog: With platform engineering gaining momentum, how do you support organizations building internal developer platforms and improving developer experience?

Desai:  Ciroos integrates directly into platform teams’ workflows via Slack/Teams/WebEx, Jira/ServiceNow, CI/CD systems such as ArgoCD/Flux etc. We provide a reasoning layer for operational data across Kubernetes, cloud, network, and other domains. For platform engineering teams, this means visibility and automated insights are embedded directly into their workflows. The result is faster troubleshooting, higher developer confidence, and a smoother path toward reliable, AI-assisted internal developer platforms. 

VMblog: What emerging technologies or industry shifts is your company monitoring most closely as we head into 2026?

Desai:  In the last twelve months, the foundations of interoperable agentic AI systems have been laid with the advent of Model Context Protocol (MCP), Agent-to-Agent (A2A), and AGNTCY framework. These innovations set the stage for the next era of enterprise operations. We are particularly excited about three specific areas heading into 2026 as they pertain to production operations:

  • Our forward-looking customers deploying AI SRE solutions view this as a journey where the focus evolves from rapid, reactive response on day one to proactive response on day two to predictive and autonomous self-healing operations on day three.
  • We are inspired by the potential of thought experiments, i.e., what-if scenarios that SREs can do when armed with an intelligent AI system. What would be the impact on the overall system if certain input variables were to change?
  • We believe there is a larger opportunity for “outer feedback loops”, where data from investigations over a certain (longer) time horizon can be used to identify brittle system elements of digital infrastructure/applications to increase overall reliability.
##