During the 56th edition of the IT Press Tour in Silicon Valley, VMblog had the opportunity to meet with Michael Buchel, CTO of Arc Compute, an innovative startup harnessing low-level GPU optimizations to condense AI/ML tasks and improve data center efficiency. Arc Compute’s suite of products – Arc HPC Nexus, Arc HPC Oracle, and Arc HPC Mercury – aim to solve pressing problems around GPU utilization and energy consumption in an era of chip scarcity and soaring operational costs.
“We have a situation where individuals are simply unable to get their own power chips or GPUs,” explained Buchel. “There’s a lot of reasons for it – approaches like ours have been around for quite a while, but when they were originally coming out, no one cared because GPUs grew on trees. Now they’re becoming a scarce commodity.”
Arc Compute’s software suite takes a unique approach to optimizing GPU utilization. Arc HPC Nexus provides a management solution that builds the environment needed to solve what Buchel terms the “old stirring problem” – matching tasks to GPUs and determining the optimal arrangement to maximize throughput. By deeply understanding workloads at a granular level and applying QoS at the individual VM level, Nexus can deliver sizable performance gains.
In testing with Sandia National Labs’ LAMMPS molecular dynamics simulator, a single A100 GPU running Arc HPC Nexus delivered over 9,000 atom timesteps per day. Splitting the A100 in half and running LAMMPS on both the upper and lower partitions concurrently nearly doubled throughput to around 17,000 timesteps per day – approaching the level of two full A100 GPUs for the price of one. Buchel believes further optimizations could push this to 19,000-20,000 timesteps, though CPU scheduling remains an area for improvement.
Arc HPC Oracle takes the jigsaw approach a step further, utilizing mathematical models to predict pipeline saturation, heat output, energy draw and wear patterns. By forecasting heat generation, data center operators can proactively reallocate workloads to avoid hotspots and equipment damage. Power consumption insights allow jobs to be scheduled to take advantage of lower energy rates without exceeding circuit capacity. This predictive modeling enables data centers to slash their power usage and cooling requirements.
“If you look at it from a mathematical standpoint, we absolutely can determine the heat generation,” noted Buchel. “We basically have to prevent thermals from spiking or walking away when a GPU gets too hot. By lowering the heat, you lower the amperage and you avoid this terminal walk away. That lowers the heating cost as well as lowers the amount of electricity you consume.”
Looking further out, Arc HPC Mercury aims to build optimal data center configurations based on anticipated workloads over a multi-year span. By designing an environment tailor-made for a customer’s needs using Oracle’s profiling capabilities, Mercury will help organizations deploy infrastructure that delivers the maximum achievable performance from Day 1.
While Arc Compute’s software currently focuses on NVIDIA GPUs, the underlying technique of intercepting code before it reaches the GPU is applicable across architectures. Support for AMD GPUs is on the near-term roadmap, and Buchel sees opportunities to bring Arc’s optimizations to FPGAs, APUs and other accelerators down the line. The company is also exploring unique integrations, such as the ability for virtual machines to share data with container-based applications by cloning VM disks into persistent volumes on the fly.
Under the hood, Arc’s software operates well below the CUDA level, optimizing the GPU-specific code generated by compilers before it reaches the hardware. This low-level approach makes Arc’s optimizations largely independent of the APIs and frameworks used for development. It also allows Arc to avoid the double-digit performance hit typically seen with virtualization, as the software can optimize across the hypervisor-guest boundary. This “bare metal” level of integration is key to Arc’s value prop.
On the competitive front, Buchel sees no direct rivals, with only NVIDIA’s own Multi-Process Service (MPS) coming close. However, MPS is limited to certain GPU architectures and operating systems, with limited programmability. Job schedulers and application-specific optimizations represent an adjacent but distinct market. Arc’s ability to holistically optimize across jobs and VMs fills a critical gap.
With 17 employees and steady revenue growth, mainly from direct sales to strategic accounts, Arc is already operating in the black. The company is currently working on a Series A round of funding to accelerate hiring low-level engineering talent and build out its channel partnerships. An NVIDIA acquisition could be in the cards, but Arc seems determined to remain independent and work across the ecosystem.
Sustainability is another area where Arc is poised to make a major impact. By maximizing utilization of existing GPUs and minimizing energy waste, Arc’s software could allow data centers to reduce their carbon footprint and e-waste output. Buchel estimates that once Oracle is fully productized, customers can expect to use 30% less energy for a given set of GPU-accelerated workloads. In a world facing chip shortages and climate change, that’s a powerful differentiator.
IT pros looking to keep GPUs lit up and costs down in an era of scarcity and soaring demand should keep a close eye on Arc Compute. With a strong technical foundation, early traction, and an ambitious roadmap, the company is well positioned to become a major player in the AI acceleration space. As more and more workloads shift to specialized silicon, Arc’s ability to wring out maximum performance from every watt and every dollar will only become more essential.
##






