Opens in a new tab
vmblog logo 2024 wht (updated)

The New Face of High Availability: IT Admins on the Frontline

Share: 

David Marshall | Published: October 28, 2025

Organizations rely on a growing array of software and services to carry out their missions and achieve their objectives. Increasingly, services and applications critical to operations are moving out of traditional on-premises data centers to cloud and hybrid environments. As such, the role of IT administrators has evolved to adapt to the dynamic landscape of emerging challenges, advanced technologies, and expanding responsibilities.

No longer limited to routine maintenance, IT admins are now on the frontlines of ensuring continuous availability and resilience-including for resources they may not own. That creates new challenges for IT managers who, while comfortable with what it takes to keep the lights on with their on-premises resources, may not yet grasp what is required to achieve high availability for assets running on cloud infrastructure.

Flexible Infrastructure

The flexibility and cost advantages offered by shifting to a cloud or hybrid environment was an attractive incentive for these organizations, but high availability and disaster recovery remain a top priority no matter who owns the infrastructure. Whether those resources are on-premises or in the cloud, IT staff and leadership face the critical task of ensuring the services and applications they depend on are safeguarded against the significant risk of “going dark” during an outage, which could result in severe disruptions to operations and business continuity.

Yes, your contract with a cloud services provider includes a high availability service level agreement (SLA) for their infrastructure, but it doesn’t cover the applications running on that infrastructure. And an SLA isn’t a guarantee that your cloud services won’t fail, such as when Azure was knocked offline in 2024 following a distributed denial of service (DDoS) attack, or when an AWS region went offline during planned maintenance. What often goes unnoticed is that you can have application downtime on operational servers for a variety of reasons. In that case, the cloud vendor’s high availability (HA) SLA is met, but your systems are still offline.

4-Nines Availability

Often an HA SLA will promise 99.99% (four-9s) availability per year. The customer is responsible for keeping their mission-critical applications running at a level that prevents a complete and indefinite pause in operations. Good news: it is possible to design for operational resilience and high-availability even when you are dependent on cloud-based services and infrastructure.

There are several strategies that can be implemented to protect application availability in the cloud. You can run your application on a primary node in the cloud and cluster it with advanced failover clustering software. You can provision local storage on each cluster node and use purpose-built replication software to synchronize each storage node. Clustering software will monitor the health of your application and move operation of the application to a secondary node in the event of a failure.

If you are concerned about the availability of a single cloud, some clustering software will allow you to create cluster nodes in two different clouds and failover from one to the other if necessary. Distribution of resources and workloads in a multi-cloud environment using SANless clusters can reduce or eliminate single points of failure and keep service running on secondary resources until your primary service is restored.

Three Key Objectives

SANless clusters also have an advantage over traditional SAN-based clusters in that they are more cost efficient, draw fewer computing resources, and are more flexible than their hardware counterparts. For that reason, a multi-cloud approach makes a lot of sense, supporting three key objectives for the enterprise:

  • Aligning high availability with operational priorities
  • Achieving cost-efficient resilience in hybrid environments
  • Influencing the burgeoning effects of AI on IT operations

Alignment with Priorities

When your goal is to align an IT operations strategy with an organization’s business objectives, multi-cloud earns favor with leadership. That’s because it is the best way for a cloud-dependent enterprise to reduce operational risk. The number one priority for IT is always to ensure services are available for its users, including customers, partners, and employees. If services are unavailable, nothing else matters.

Taking a multi-cloud approach means you can build in the redundancy needed to achieve high availability-essential for enabling seamless failover for critical applications across clouds. In other words, if Cloud A goes down, operations can automatically switch over to redundant services hosted on Cloud B. Even when operating at a lower level of performance, this ensures a level of business continuity that can avert a catastrophe while demonstrating a commitment of reliability to your customers.

Cost Efficiency

The idea that using two different cloud services providers can be cost efficient may seem counter-intuitive, but it makes sense when you take a thoughtful approach. First, some cloud services excel at different things, and so if you choose one cloud for its strength in available business services, you don’t have to settle for sub-optimal infrastructure performance as a tradeoff.

When you work with two or more cloud service providers, you also gain leverage in negotiating a better deal. (And for those times when a cloud service is knocked offline, the ability to failover to a secondary provider will be regarded as money well-spent.)

AI and IT Operations

Artificial intelligence is driving wonderful innovation in IT operations, and that role will only increase over time. But AI can’t eliminate risk. You still need a reliable and robust infrastructure to maintain efficient performance. And as AI-driven tools are deployed to support operations, they will rely on more-and more timely-data, further emphasizing the importance of reliable, highly available infrastructure.

AI-driven tools operating in conjunction with good high availability and disaster recovery solutions can strengthen operational resilience, ensure business continuity, and serve as the basis of effective application management well into the future.

Too Much at Risk

When shifting from traditional on-premises infrastructure to a cloud-based posture, there is too much money and reputation at risk to not also invest in building a highly available architecture to ensure operational continuity. IT administrators can increase their value to the organization by evolving their approach in alignment with the priorities of leadership. And with the right strategy and the right tools, it is possible to build resilient, highly available services even when you don’t own the infrastructure.

##

ABOUT THE AUTHOR 

Greg Tucker, Senior Windows Product Support Engineer, SIOS Technology

 

Greg Tucker is a seasoned Senior Windows Product Support Engineer at SIOS Technology Corp., where he has been delivering expert technical support and solutions for over 13 years. With extensive expertise in advanced high availability technologies as well as Hyper-V and Amazon EC2, Greg specializes in helping organizations protect their critical applications for maximum reliability in on-premises, hybrid, and cloud environments.

A proud graduate of South Carolina State University, Greg combines technical acumen with a customer-centric approach, ensuring businesses achieve robust IT infrastructures that meet evolving demands. Based in Columbia, South Carolina, Greg is actively engaged in the Cloud Computing, SaaS, Data Center, and Virtualization communities, where he shares insights and fosters innovation in the rapidly advancing tech landscape.