By Don Boxley, CEO and Co-Founder, DH2i (www.dh2i.com)
Somebody wants to talk about AI infrastructure everywhere I go lately. Usually that means GPUs. Or language models. Or, who’s spending billions building the next AI data center. That’s ok… Those are important conversations.
But I keep finding myself asking a different question: What happens the first time something breaks?
Because it will.
I’ve been around enterprise infrastructure long enough to know that hardware fails, software has bugs, people make mistakes, storage fills up, networks go sideways, and eventually somebody has to patch something at two o’clock in the morning.
That hasn’t changed because we started calling applications “AI.”
The stakes are higher now. The funny thing is… we’ve spent years talking about making AI smarter. The infrastructure underneath meanwhile has quietly become more important than ever. Not less.
One thing I think people miss is that AI isn’t replacing enterprise infrastructure. It’s piling on top of it. The database that feeds your AI? It still has to stay online. The storage holding years of business data? Still has to stay online. Authentication services? Networking? Virtual machines (VMs)? Replication? They’re all still there doing the same jobs they’ve always done. Just because chatGPT showed up, we didn’t rip all of that out.
Most organizations I talk to are running AI alongside hundreds, sometimes thousands of existing workloads. Some live in VMs. Some are containerized. Some are on-premises. Some are in the cloud. Some are at the edge.
That’s reality. Which means the challenge isn’t building AI. It’s keeping all those moving pieces available at the same time. I’ve always thought our industry has a funny habit. We get excited about the new thing and assume the old problems disappeared.
AI is changing plenty, but the fundamentals don’t change.
Applications still need data. Data still needs to be available.
And, here’s something I don’t think gets enough attention. When people say an AI application “went down,” that’s rarely what actually happened. Usually the model is fine. It’s waiting on a database. Or it lost access to storage. Or a VM hosting a critical service disappeared. Or somebody was performing maintenance.
The AI didn’t fail. The infrastructure around it did.
That’s an important distinction because it changes where IT teams should be spending their time.
I’ve also noticed something else. Hybrid infrastructure isn’t really optional anymore. Some data has to stay on-premises because of regulations. Some AI services make sense in the cloud. Inference might happen at the edge because latency matters.
But unfortunately, many organizations are still treating availability differently depending on where the workload happens to live.
Whether it’s a VM… container… or on-prem, the expectation from the business is exactly the same: Don’t let it go down.
Years ago, everybody accepted maintenance windows. You’d send an email. “Systems will be unavailable Saturday between midnight and 4 a.m.” Nobody liked it, but everyone understood. I’m not sure that’s true anymore. Businesses operate around the clock. Customers don’t know or care that IT is patching a hypervisor. They just know something stopped working.
That’s why I think HA has become less about DR and more about operational freedom.
Can you patch infrastructure without interrupting applications?
Can you move workloads without anybody noticing?
Can you modernize without scheduling downtime?
Those are becoming the questions that matter.
The AI conversation will keep changing.
Next year, it will be different models, different hardware, different benchmarks.
That’s the easy part.
The hard part is making sure all of it actually works every day. Because nobody buys AI for the demo. They buy it for production. And production has always been about reliability. The companies that get that right won’t necessarily have the biggest GPU clusters or the newest models. They’ll have infrastructure that’s resilient enough that nobody notices it’s there. Honestly, that’s always been the goal.
AI didn’t change that.
It just made it more important.
##
ABOUT THE AUTHOR

Don Boxley Jr is a DH2i Co-founder and CEO. He has more than 20 years in management positions for leading technology companies. Boxley earned his MBA from the Johnson School of Management, Cornell University.






