Opens in a new tab
vmblog logo 2024 wht (updated)

VMblog Expert Interview: Moving Beyond Service Level Objectives: A Single Source of Truth for Software Reliability

Share: 

David Marshall | Published: September 25, 2023

 

Today, more than ever, customers expect applications to be reliable, fast, and available on demand, and that demand is only increasing. Software reliability engineering (SRE) is critical to successfully managing and delivering these applications and ensuring smooth operations. To understand the importance of moving beyond Service Level Objectives towards a single source of truth for software reliability, I sat down with Kit Merker, Chief Growth Officer of Nobl9. Nobl9 recently launched its Reliability Center, a comprehensive platform allowing engineers and managers to understand the reliability of their vast software systems to identify weak or risky areas that need attention and improve overall business functions while streamlining costs. 

VMblog: It’s been a while since we last spoke about Service Level Objectives (SLOs). What has Nobl9 been up to?

Kit Merker: I think the biggest trend we’ve seen recently is increasing recognition of site reliability engineering and service level objectives as a mainstream and necessary approach to ensuring reliability for organizations and achieving reliability goals efficiently and scalably within those organizations. In fact, in the Gartner® Hype CycleTM for Site Reliability Engineering, 2023, Gartner predicted that by 2027, a substantial 75 percent of enterprises will have embraced SRE practices across their entire organizations. This widespread adoption is expected to drive significant enhancements in product design, cost management, and overall operational efficiency, all aimed at meeting and exceeding customer expectations. This represents an increase from the relatively modest 10 percent adoption rate observed in 2022.

Our own recent State of SLOs 2023 report, which we conducted earlier this year with Dimensional Research, found that 72 percent of the responding companies use six or more monitoring and observability tools, and 80 percent indicate an increased focus on system reliability due to the pandemic driving cloud adoption, remote workers, and supply chain issues.

We’re seeing more customers adopting SLOs – particularly customers like Ticketmaster, Software AG, Procore, Cisco and ServiceNow- and many financial services, retail, and e-commerce platforms. And that’s very exciting because we’re seeing this technique that used to be almost exclusively implemented by big technology companies becoming something that pretty much anybody can do. And it’s expected that more and more organizations will adopt it over the coming years. The key takeaway is that the SRE market is going mainstream. 

VMblog: What was announced last week?

Merker: Last Thursday, we introduced our Nobl9 Reliability Center, the next generation of the Nobl9 SLO Platform. It incorporates new features to become the single source of truth for the reliability of an organization’s internal, customer-facing, and mission-critical software. The Reliability Center introduces some exciting new features, including dashboards and reports designed to let executives and engineers understand a Reliability Score calculated from all the SLOs they want to manage. The key new capabilities are an enhanced Reliability Experience (RX), helping engineers and teams become more productive in identifying targets, prescribing SLOs and policies, and automating runbooks for reliability risks; SLO-Backed Operations for continuous monitoring and management systems using SLOs and the ability to receive timely error-budget-backed alerts; and Reliability Insights, giving instant visibility into the overall health of an organization’s systems, and ability to align technology investment with business needs.

When we started Nobl9, we wanted to be able to show executives and senior leaders the view of reliability across their organizations. But it’s actually pretty hard if you start from the goal of creating this dashboard because, in the beginning, you don’t have accurate data. So, instead, we built a tool that practitioners would use daily to improve the accuracy of their data. When we roll it up across the organization, we are confident that the data is actionable and provides real insight to management so they can trust it. It’s not just some selective, curated dashboard. Rather, Nobl9 Reliability Center represents the state of the systems themselves because the practitioners use it day in and day out.

VMblog: There has always been discussion of five 9s, but reliability is becoming a more critical topic. What are you seeing?

Merker: We see that the number of nines is not the only way to measure reliability. In fact, what is important is to separate which parts of your systems need high availability or five nines and which can get away with best-effort availability and take on more risk – maybe even two or three nines. By segmenting your system into different reliability goals and setting very targeted goals based on customer behavior and expectations, you better utilize your resources to ensure that you’re getting the reliability where it’s critical. It could be your payment system or movie streaming, or it could be retailers on Black Friday or trading software while the stock market’s open. These are critical services and critical times, and you want to ensure the highest reliability. And since every company has limited resources, they need to make some trade-offs. Lowering the reliability goals in some system that wouldn’t impact customers as much, is one of those trade-offs. 

That really is the nuance people are getting with service level objectives. Every organization is having more and more mission-critical services that are digital, so they have to make important choices because no one can have it all. And we help them do that. Nobl9 Reliability Center gives you a single source of truth, not only of the performance of systems, but the performance against your goals – and it lets you question those goals. If you see a goal on a mission-critical service set too low, you can ask questions about why it’s not set higher. Conversely, if you see five nines everywhere, you can start to question which services could have more degrees of freedom that could free up resources to tackle the more important, pressing customer journeys that drive business value.

VMblog: Monitoring and observability tools are increasing in most organizations, and can be up to more than a dozen. Any best practices?

Merker: There are tons of observability tools. They all have a different purpose, and using them together gives you a better view of your overall system. Recently, there was huge news that Cisco is acquiring Splunk. This is in addition to the big focus Cisco has on Full Stack Observability including AppDynamics, ThousandEyes, and several others. This shows that several different observability tools exist even within the Cisco ecosystem. Across the broader industry, a proliferation of tools continues, organizations use multiple tools to gain deeper insights into security, compliance, AI/ML, and other applications. However, even with all of these solutions, less than half of the companies surveyed in our State of SLOs report have visibility into all their IT environments, and the expanding use of the hybrid cloud is compounding this challenge. Given the increasing adoption of containers and microservices, it was staggering to see just 45% and 35% have visibility into those systems, respectively. 

Our point of view is that customers don’t necessarily try to consolidate and tell every team “Thou shalt use a particular solution.” They want them to have the right tool for the job, for the detailed metrics, events, tracing, logs, etc. Nobl9 Reliability Center embraces that fragmentation and allows teams to use their tools of choice while simultaneously consolidating and aggregating into a centralized hub that clearly displays what’s going on across services. 

Nobl9 Reliability Center lets teams drill down into different services and how they’re working, link back to the original systems, click in to get more detail, and also gives users a system-agnostic view so they don’t have to know the granular details of each tool. For example, maybe the payment system is being monitored by New Relic, and Datadog is monitoring the sign up experience. The organization can get all of that data and see it consistently across these different implementations.

VMblog: How can managing monitoring and observability tools differently help optimize costs?

Merker: There are two parts to the cost story for monitoring and observability. One part is the efficiency of the system you’re monitoring, meaning, is it over-provisioned? Are you spending more on capacity than you need? This is an important question, because if the service is hitting its reliability goals, you may have an opportunity to optimize the cost of the infrastructure. 

The second part is the cost of storing data in the monitoring system and the related costs, which can often be quite shocking. I’ve heard horror stories of people blowing their entire monitoring budget for the year in a single month because they didn’t realize how much data they were collecting. One of the important rules for cost saving is making sure that you will take action on whatever data you store.

Observability without action is just storage. And this is something that Nobl9 Reliability Center can help you with as well because you’ll understand when there are interesting times versus more boring times. You can choose to selectively keep more data, add more context when things are most intriguing, and get rid of it otherwise. Crucially, we don’t charge for storage at Nobl9. We give you two years of data retention on all of your SLO data, plus you can export it to Snowflake or Amazon S3 whenever you want. We ensure that storage issues and associated costs go away. By setting clear reliability and performance targets for all of your services, you can figure out which ones are over-provisioned relative to the performance you see in those services.

VMblog: How can Reliability Center help support better business decisions, and where are you seeing the most effective use cases?

Merker: Executives are often quite detached from the implementation details of their vast IT systems. Meanwhile, they have to make sense of lots of conflicting information and make resourcing decisions when there’s often no great answer. They’re trying to manage the uptime reliability of the service, deliver new features, and keep employees happy without really knowing the technical details of the underlying systems. The most important way that Reliability Center helps with this decision-making is by giving people clarity and accuracy. We clearly display the system’s performance and aggregate data from all tools and services, but we also compare this to the organization’s goals, which helps teams rationalize what the goals should be relative to performance.

Then, they can decide if they’ve set their goals too high or too low, or perhaps they’re not measuring the right things about their service. It forces them to make some explicit decisions, which is precisely what you need when you’re building and running a large-scale enterprise to serve customers. Above all, organizations need to avoid surprises and emergencies – to turn bad days into normal days, where teams can accept the fact that things are going to go wrong, but they have a plan to keep operations running smoothly and most of all, keep customers happy. Reliability Center is really about being proactive, getting ahead of issues, and properly managing operations. We are not trying to avoid occasional inevitable failures. Instead, we put guardrails around the system and add trip wires so that organizations are notified early and can take action before customers start to complain.

VMblog: Is there anything else VMblog readers should know? And, what can we expect from Nobl9?

Merker: If you plan to attend AWS re:Invent in Las Vegas, please come see us there. We’ll have a booth that says “Smooth Operations” because we’re trying to help our customers achieve that lofty goal. As for what you can expect, we will keep investing in making reliability engineering easier, making it more of a team sport, and making it easier for executives and non-technical people to speak to IT operators, developers, and technical folks. They all work from the same data and share the same goals and information. And if you want to try Nobl9’s Reliability Center yourself, there is a perpetual free edition with lots of introductory documentation. For those looking to dive deeper, we provide access to the SLO Academy and offer support to help you get started.

##