Broadcom kicked off its VMware Explore 2026 discussions this week, and there was a lot of meat on the bone. The headline news: a new thing called VMware AI Factory, which the company is positioning as the software-defined foundation underneath its broader VMware Private AI Cloud push. Translation for anyone skimming past the marketing language — Broadcom wants to take the painfully manual process of standing up AI infrastructure and turn it into something closer to a vending machine. You pick the hardware, pick the models, and the platform does the rest.
Before getting into what’s new, it’s worth pausing on the numbers Broadcom shared about where VMware Cloud Foundation already stands, because they set up why this announcement matters.
VCF9’s Report Card, One Year In
Paul Turner, chief product officer for the VMware Cloud Foundation Division, opened with an update on VCF 9, which shipped last year. The adoption numbers are honestly bigger than a lot of people probably expected: over 3,000 customer deployments and more than 19 million allocated cores running in production, spanning financial services, healthcare, public sector, manufacturing, and education across the Americas, EMEA, and APJ.
That’s not pilot-project scale. That’s real workloads, real budgets, real production systems.
A big driver behind the upgrade cycle, according to Broadcom, has been the ongoing hardware supply crunch — especially DRAM pricing, which has made server refreshes an expensive proposition for a lot of IT shops. VCF 9’s NVMe memory tiering, introduced last year, lets customers squeeze more workloads onto existing or new hardware, which softens that blow considerably. Pair that with an AI assistant built into VCF for troubleshooting things like ESXi availability issues, CPU or memory contention, or storage latency, and you start to see why customers keep pointing to VCF 9 as the reason they upgraded in the first place. The assistant is meant to flatten the learning curve for admins who don’t have deep storage or networking backgrounds, essentially giving every ops team a bit of borrowed expertise.
So What Exactly Is VMware AI Factory?
Here’s the thing about building AI infrastructure from scratch: it’s a slog. You need the right GPU form factor, networking, storage, Kubernetes for containerized workloads, an AI software stack, and then you need to pick from SLMs, LLMs, open source models, and commercial models — and somehow keep all of that patched and current. Most enterprises don’t have a team dedicated to babysitting that whole stack.
VMware AI Factory is Broadcom’s attempt to package that entire journey — what the company calls “metal to model” — into something pre-tested and validated. It’s not a rebrand of an existing product; it’s genuinely new, combining several pieces that didn’t exist together before:
- AMD partnership: VMware AI Factory pairs VCF with AMD Instinct MI350 Series GPUs and the open ROCm software ecosystem, with zero-touch provisioning handling the full stack from vSphere and vSAN through Kubernetes and the AMD GPU operator.
- Certified AI ReadyNodes: Broadcom has validated server configurations with Cisco, Dell Technologies, Lenovo, and Supermicro, so customers aren’t stuck guessing whether a given hardware combination will actually behave in production.
- MetalSoft integration: perhaps the most practical addition here, this brings heterogeneous bare-metal automation directly into the VCF operations console.
That MetalSoft piece deserves a beat of its own. Bare-metal provisioning across mixed-vendor server fleets has traditionally meant juggling separate tools for every hardware vendor’s firmware and lifecycle quirks. Lucas Roh, MetalSoft’s founder and CEO, put it plainly: “Managing physical servers has always been a separate operational domain from software, creating silos that slow down AI infrastructure deployment.” With this integration, Broadcom claims bare-metal provisioning that used to eat up weeks can now happen in minutes, all through a single console rather than a pile of vendor-specific tools.
Ram Velaga, president of Broadcom’s Infrastructure Software Group, framed the bigger picture behind all this: “VMware Private AI Cloud is the inflection point where enterprise private cloud and private AI infrastructure stop operating as separate disciplines and become one — enabling production inference workloads and agentic AI with the data sovereignty, compliance posture, and cost predictability their business demands.”
The Model Buffet: 150-Plus Options and Counting
One thing Broadcom kept circling back to during the briefing: not every AI use case needs a frontier model. Sometimes a small model handles the job fine and costs a fraction as much to run. So rather than pushing customers toward one flavor of AI, VCF now supports a governed catalog of more than 150 open source and commercial models, with a handful of newly validated names getting top billing:
- Nemotron 3 from NVIDIA — a hybrid Mamba-Transformer MoE model built for long-running agentic workflows with a 1-million-token context window.
- Gemma 4 from Google DeepMind — an open-weight, multimodal model aimed at developers building autonomous agents.
- cotomi from NEC — tuned specifically for Japanese language and business context, with a claimed 40% improvement in token efficiency.
- Qwen3.8-27B from Alibaba — a proprietary multimodal model with a one-million-token context window for global enterprises.
- GLM 5.2 from Z.ai (formerly Zhipu AI) — an open source model geared toward coding and reasoning agents running locally.
Chris Wolf, global head of AI and advanced services for the VMware Cloud Foundation Division, summed up the intent: “Working with the world’s leading AI model providers, we’re giving organizations a clear path to data sovereignty and cost-effective AI at scale, with leading models available securely and delivered as a service to their user community through VMware Cloud Foundation’s built-in services.”
There’s also a benchmark data point worth noting for the skeptics in the room — independent testing under MLPerf Inference v5.1 standards found VCF performing on par with bare metal. That matters for the argument that virtualization overhead is a bygone concern for AI workloads, which has been a lingering objection for years.
And the private cloud shift isn’t just a Broadcom talking point. According to the company’s own Private Cloud Outlook 2026 survey, 56% of enterprises are already running, or planning to run, production AI inferencing on private cloud. Cost, data privacy, and infrastructure control kept coming up as reasons, whether the workload is fine-tuning, RAG, or plain inferencing.
Under the Hood: Private AI Services That Handle the Boring Parts
Underneath all the hardware talk, Broadcom also rolled out a handful of software-level capabilities baked directly into VCF that quietly do a lot of the operational heavy lifting:
- Multi-tenant model sharing lets a model get deployed once and shared safely across lines of business through isolated namespaces, instead of each team spinning up redundant copies and eating GPU capacity they don’t need.
- AI Gateway enhancements add intelligent prompt routing, token and usage rate-limiting, and application authorization — basically a traffic cop for who gets to talk to which model and how much they’re allowed to spend doing it.
- Secure AI Sandboxes isolate dynamic, agent-generated code execution, so an agent that decides to write and run its own script doesn’t get free rein over the rest of the environment.
None of these are flashy on their own, but they’re the kind of guardrails that determine whether an AI deployment actually survives contact with a real budget and a real security team.
The Takeaway
Broadcom’s pitch here isn’t subtle: bring the AI to your data, not the other way around, and stop treating infrastructure procurement as a separate project from model deployment. Whether VMware AI Factory lives up to the “weeks to hours” claim in the wild remains to be seen — that’s always the gap between a briefing deck and a Tuesday afternoon in a customer’s data center. But given how much frustration enterprises have voiced about GPU costs, hardware fragmentation, and token economics over the past couple of years, it’s easy to see why Broadcom decided this was the moment to bundle it all together and give it a name.
Expect a lot more detail on the AI Factory partner ecosystem — including where NVIDIA fits in beyond model validation — as sessions roll out this week at Explore in Las Vegas.
##






