Opens in a new tab
vmblog logo 2024 wht (updated)

The Build-vs-Buy Trap Lurking in Your VMware Infrastructure Automation

Share: 

provisioning isn't lifecycle

By Rob Hirschfeld, CEO and co-founder, RackN

Capable teams rarely stall because they can’t build automation. They stall because they can’t maintain the automation they’ve already built.

I talk with many infrastructure leaders who are evaluating VMware alternatives and buying AI capacity.  Surprisingly, they don’t ask “which platform should we choose?” Their real struggle is more revealing: “We already know Ansible and Terraform, so why can’t we just build this ourselves?”

My honest answer tends to surprise people: “Yes, you could build it.”  And I always follow up with a question, “But do you want to own and maintain something that complex?  Are your teams so underworked that re-inventing bare metal automation is creating business value?”

Sadly, the challenge of managing bare metal drove many companies to the cloud. They bought into the idea that infrastructure was too complex and difficult.  Today, leading companies are returning to managing parts of their own IT needs to improve cost, security and governance with fantastic results.

Let’s examine why assumptions from the 2010s may be holding your organization back.

How DIY Provisioning adds risk

The build-vs-buy decision in infrastructure automation is often framed as a question of your team’s ability. It almost never is. Most ops teams have someone eager to build a one-off “bespoke” provisioning system cobbled together from open source tools.  We’ve all seen these utilities because they litter the operational landscape at every enterprise.  They are fragile, hardcoded, require specialized knowledge to maintain, and (worst of all) are mission-critical to getting anything done on your infrastructure.

Provisioning is not lifecycle

The biggest misconception with DIY automation is that provisioning is a one-time operation instead of a process.  Managing bare metal is an extended lifecycle over years: firmware updates that behave differently across vendors and hardware generations, operating-system reimaging, configuration drift, security remediation, hardware that fails in ways your scripts never anticipated, and eventually decommissioning.  These are predictable ways that your fleet keeps growing even when the person who wrote the original automation has moved on.

That distinction, one-time provisioning versus lifecycle, is where build-vs-buy decisions go wrong. Most do-it-yourself efforts only scope a small (easy) part of the real work.  The hard life-cycle stuff is rarely factored into the plans.

Where homegrown automation breaks

Let’s get a clear scope for the real work for bare metal automation.  While infrastructure is complex, we have delivered standardized processes covering hundreds of scenarios.  That’s important if you want to improve delivery time, reduce lock-in, and free your operators to focus on value-added work 

  • Hardware heterogeneity. The script that worked on one vendor’s server hits a firmware exception on others, multiplied across generations of hardware bought over many years.
  • Day-two remediation. The interesting failures, a node that half-provisions, a firmware version that regressed, a config that drifted, are exactly the ones that never make it into a clean module. They live in an engineer’s head.
  • Institutional memory. Your automation becomes a pile of scripts held together by the one or two people who remember why each workaround exists. That is a bus-factor problem disguised as a platform.
  • The maintenance tax. Every hardware refresh, every new server model, every firmware advisory becomes a small engineering project. The build was cheap; the upkeep compounds forever.

None of these show up in the DIY tool budget but all of them show up at even modest scale.  By then, the cost of unwinding a homegrown system is far higher than the cost of having chosen differently at the start.  That’s why many companies end up with multiple DIY provisioners!

You should leverage standard processes

When teams price out “build,” you are betting they have the expertise and knowledge to deliver in a highly complex environment; however, very few teams have the scale to maintain this level of experience.  Even if you have it, there is an opportunity cost for senior engineers spending time firefighting infrastructure exceptions instead of working on whatever actually differentiates your business when you could buy a solution that works out of the box.

A sharper set of questions helps:

  • Is infrastructure automation a source of competitive advantage?
  • Will my automation survive hardware refreshes, staff turnover, and five years of drift?
  • Are we committed to maintaining a platform, or is this accumulating technical debt?

For enterprises running VMware, Kubernetes, or AI fleets, it isn’t differentiated.  In fact, the overhead and delays caused by gaps in your bare metal lifecycle layer are adding significant cost and delays to more critical differentiated projects.  

This is not hypothetical: We’ve helped turn completely stalled VMware migration projects around in just two weeks by bringing in a robust bare metal platform with Digital Rebar.

Why this matters right now

This used to be an abstract debate. It isn’t anymore. Two forces are pushing teams into it at once. The VMware migration wave has thousands of organizations standing up a “second stack” — and discovering that standing it up was the easy part; operating it across real hardware at scale is where projects stall. And AI infrastructure is forcing those same teams to manage physical GPU fleets at a velocity and cost they have never dealt with before.

In both cases, the platform decision gets the attention, but the infrastructure operating model determines success. The teams that move fastest aren’t the ones trying to figure out bare metal along with a new platform. They’re the ones who decided early to standardize the hardware lifecycle.

Build-vs-buy was never really about whether your team is good enough. It’s about where you want your best people pointed. And if you’re already down the DIY road, don’t despair! It’s much (much) easier to onboard a bare metal platform than fix to a custom one.

##

ABOUT THE AUTHOR

rob hirschfeld photo

Rob Hirschfeld is the CEO and co-founder of RackN. He has worked in cloud and physical infrastructure for two decades, including four terms on the OpenStack Foundation board and a leadership role at Dell, where the founding team created Project Crowbar. He writes and speaks about infrastructure automation and the operational realities of running hardware at scale.