At the IT Press Tour in Palo Alto, the Stanford-born startup, Clockwork, showed what happens when you pull the network cable and the power cord out of a live AI training job. The job kept going.
Suresh Vasudevan has been here before. Years ago, as CEO of Nimble Storage, he launched his company at an IT Press Tour event. This week he came back to Palo Alto with a different company, a different decade, and a very different kind of hardware problem: GPU clusters that cost more than most data centers and break far more often than anyone wants to admit.
The company is Clockwork.io. Along with Suresh, we heard from Gavin Cohen, VP of Products, who handled the deep technical walkthrough, and Dan Zheng, Chief Business Officer, who closed with customer stories and fielded the business questions. A couple of Clockwork engineers also joined us, one of them from a lab in the South Bay, to run a demo that I’d call brave. More on that in a minute.
Only days before the IT Press Tour meeting, the company laid out a couple of new announcements on the news wire: $31 million in new funding, production deployments at LinkedIn and Together AI, a bigger footprint at WhiteFiber, and a pair of new TorchPass capabilities. But the announcements are only half of the story. The other half is how a team of clock-synchronization researchers ended up as one of the more interesting fault-tolerance companies in AI infrastructure.
A Clock Company That Isn’t Really About Clocks
Here’s the thing. Clockwork started as a project at Stanford, and the original goal was to synchronize the clocks of tens of thousands of machines without special hardware. Co-founder Balaji Prabhakar is a Stanford professor and co-inventor of DCQCN, the congestion control scheme behind modern RDMA networks. Yilong Geng, his student, is now CTO. Mendel Rosenblum, the VMware founder, was involved early and serves as Chief Scientist today.
The software approach gets to tens of nanoseconds of accuracy. PTP can do that too, but it needs proprietary hardware, and it tends to top out at hundreds or maybe a few thousand nodes. Clockwork’s software reaches tens of thousands. From roughly 2017 to 2022, the customers were mostly financial: NASDAQ wanted to timestamp trades in the cloud, Bloomberg deployed it globally, and Wells Fargo uses it to keep transaction timestamps honest.
So why is a financial-timestamp company now talking about GPUs? Because once your clocks agree to within nanoseconds, you can measure something almost nobody can: the true one-way delay of a packet. The usual shortcut is to measure the round trip and divide by two, which assumes the trip out and the trip back took the same time. They often don’t.
For ordinary cloud apps that tolerate milliseconds, the difference barely matters. For a GPU cluster where messages hop between accelerators in single-digit microseconds, it matters a great deal. Around 2022, the team realized it was sitting on the right tool for the wrong market. They pivoted. The FleetIQ platform reached general availability in September 2025.
Why AI Clusters Break So Much
Suresh spent a good part of the session on a point that sounds obvious until you sit with it: AI workloads are the same state transformation repeated endlessly. A training iteration takes a few seconds, and a run can repeat it tens of thousands of times over weeks or months. Inference works the same way, one token at a time.
What makes this fragile is the lock-step nature of it all. Thousands of GPUs exchange data on every step, so the network is effectively part of the computer. One overheated GPU slows everyone down. One flapping link stalls a collective operation, and every other GPU sits and waits. Having spent his last stint at Sysdig around Kubernetes microservices, Suresh was blunt about the contrast: microservices degrade gracefully, while a synchronized GPU job does not.
The numbers he cited, all from published research, tell the story:
- A Meta study found a mean time to failure of about 7.9 hours for 1,024-GPU jobs. Meta also reported an unexpected interruption roughly every three hours during 54 days of Llama 3 training on 16,384 GPUs.
- A ByteDance study found that 42.5% of jobs ran at least 10% slower because of stragglers, which are the “fails slow” cases that never trigger an alarm.
- Manual detection and recovery can eat one to three hours when someone has to figure out what broke.
And when something does break, the bill arrives in four installments. You pay for detection time. You pay for a cold restart, which can run 20 to 30 minutes across thousands of GPUs. You pay to reload the last checkpoint. Then you pay again to recompute everything done since that checkpoint, which could be hours old. Suresh also pointed out that checkpointing is left up to each application, which would have sounded odd to the storage veterans in the room. “We stopped doing application-level recovery in storage about 20 years ago,” was the gist of his remark.
The Champagne Glass Problem
Gavin Cohen picked up from there with a slide of a champagne tower. His math: 256 nodes with eight GPUs each is 1,024 GPUs, and that’s somewhere around $150 million of equipment. Each GPU is a champagne glass. Touch any single one and the whole tower comes down.

It’s a good image because it frames Clockwork’s whole philosophy. Most of the industry has tried to make recovery faster. Clockwork tried to make the failure invisible to the job. That’s a different goal, and it requires three layers of defense.
Layer one: LinkPass, for the link flaps
LinkPass is an NCCL plugin that catches a failed network link at a very low level and moves the in-flight traffic, including queue pairs, onto the surviving NICs on the same host. Why does that need special software? Because NCCL doesn’t behave like TCP. When a link dies, it waits roughly a minute, times out, and PyTorch’s own timeout is 10 minutes by default. Clockwork detects and reroutes in hundreds of milliseconds, and fails back when the link recovers.
Link problems are a bigger share of the pain than you might expect. Gavin said 15 to 20 percent of failures come from links, and it was their first fault-tolerance product, in production for well over a year. It’s how they landed Oracle and LinkedIn.
LinkedIn’s story shows why. Before LinkPass, one InfiniBand NIC flap pulled an eight-GPU server out of service, and a switch port flap could drain a second server and double the damage to 16 GPUs. With LinkPass, the job keeps running and repairs happen later. LinkedIn says it prevents tens of thousands of GPU-hours of downtime every month. There’s a nice side effect, too. Switch maintenance used to mean tracking down job owners and making sure they’d checkpointed. Now the network team can just do the work.
“Failures are inevitable at AI scale. Losing hours of useful work to them should not be.”
— Suresh Vasudevan, CEO, Clockwork.io
Layer two: TorchPass, or vMotion for GPUs
If you’ve spent your career around virtualization, here’s the analogy Suresh used himself: this is like vMotion, except you’re moving a training workload from one GPU to another instead of a VM from one host to another.
When a GPU or node fails (or is about to), TorchPass quiesces the job at a consistent point, requests a replacement through the native scheduler, captures the state, moves it, and rewires the job so everyone else believes they’re still talking to the same peer. The new hardware simply takes over the identity of the old rank. A central orchestrator tracks job state, and scheduler plugins cover Kubernetes with Kubeflow, plus Slurm, Slinky and Soperator.
Gavin was candid that the quiescing step is the hard part, “the really, really hard piece that no one has solved before,” as he put it. The proactive cases are the easier ones: a GPU running hot, ECC errors piling up, a security patch to apply, a straggler to pull out. For sudden hard failures, TorchPass rebuilds state from a data-parallel replica, which nearly every large job has.
Layer three: snapshots, the safety net
What if there’s no spare, or no replica, or the cloud provider reclaims your spot instance? That’s where the new announcement lands. Clockwork now offers what it calls the industry’s first multi-node platform snapshots for training. Using the same state-capture technology (and, for the CPU side, Linux’s checkpoint/restore in userspace, married to GPU memory state that CRIU can’t see), TorchPass saves an entire running job without touching the training code.
Snapshots run 100 to 200 gigabytes per job and can go directly to storage or to DRAM with an asynchronous drain. According to Clockwork, platform snapshots take under 20 seconds, with about 1% overhead at a 25-minute interval. Teams willing to add a few hooks for TorchTitan, DeepSpeed or Megatron-LM can use model checkpoints instead, at under 5 seconds and about 1% overhead at a 5-minute interval.
Dan Zheng’s customer example made the value clear. A global tech company’s on-prem platform team runs hundreds of jobs across many research teams. Some checkpoint every 30 minutes, some every six hours, and some not at all. The platform team can’t force anyone to change their code, but it can take snapshots of everything. One policy, no code changes, and lost work is limited to the time since the last snapshot.
Does It Hold Up? SemiAnalysis Ran the Numbers
Clockwork’s own test used a 64-GPU mixture-of-experts model over 3,000 steps with a fault injected about once an hour. The baseline with no faults took roughly 400 minutes. With traditional checkpoint restart, the run took about 800 minutes, and the progress chart looked like a sawtooth, each fault rolling the job back. With TorchPass, the run tracked the baseline almost exactly.
Independent numbers backed that up. In SemiAnalysis’s ClusterMAX TCO calculator, using default values, TorchPass cut goodput loss on a 4,096-GPU gold-tier cluster from 13.68% to 2.76%. On a silver-tier cluster, it fell from 28.51% to 10.31%. SemiAnalysis founder Dylan Patel put it this way:
“Cluster fault tolerance used to be a training problem. It is now an inference problem too.”
— Dylan Patel, Founder, CEO and Chief Analyst, SemiAnalysis
That quote points to where this is going. Large models served across multiple servers lose an entire replica when one link dies. Reinforcement learning ties training and inference together, with rollout replicas that need fresh weights quickly. LinkPass keeps those replicas serving, and TorchPass’s fast application checkpoints get updated weights to them sooner.
Then They Pulled the Plug. Literally.
Suresh had teased this early on: “We’re taking a risk.” The risk was a live demo where engineers in Clockwork’s South Bay lab, connected by video call, would break a running training job on purpose. The job ran on two nodes, each with four GPUs and four RDMA NICs, with live charts for training steps and per-NIC traffic.
First, an engineer physically pulled a network cable. The step counter barely flinched. The traffic on the unplugged NIC dropped to almost nothing, the other three picked up the load, and when the cable went back in, the traffic evened out again.
Round two was riskier. The engineer powered down a node, and then pulled its power cables. (A glove was worn, safety was observed, and no engineer was harmed in the making of the demonstration.) The training stopped for about a minute. Then TorchPass brought in a standby node, moved the state across, and the job resumed at the exact same step. No checkpoint reload, no lost steps.
A reasonable question from the room: who has spare nodes lying around? Per Suresh, neoclouds typically carry 5 to 8 percent spare capacity, and it isn’t idle. It runs low-priority work that can be preempted when a spare is needed.
Storyline Two: Finding the Fault in a Fabric With 50,000 Transceivers
Fault tolerance is the dramatic half of the story. The other half is what Clockwork does to find the fault in the first place, and it’s where those nanosecond clocks pay off.
Gavin was careful to say that Clockwork isn’t trying to replace Datadog, Splunk or New Relic. It’s chasing the blind spot underneath them: deep inside the fabric, where NVIDIA’s DCGM and UFM and the switch vendors’ tools fall short. The scale alone is daunting. A DGX GB300 SuperPOD with 9,216 GPUs has about 25,000 network links and roughly 50,000 transceivers. Switch counters average over 30 to 60 seconds, which hides microsecond events. And so you end up with gray failures, where every dashboard is green and the job is still slow.
The probemesh and the “is it the network?” question
FleetLens works from a few foundations. Out-of-band probes go between NIC pairs every 10 milliseconds. Synchronized clocks turn each probe into a directional one-way delay measurement. Path decomposition tags every probe with the rail, leaf, spine, port and link it crossed. Triangulation then finds the narrowest shared element behind the degraded paths.
The goal is a verdict, not a pile of graphs. Is there a problem? Who is affected? Where should I start looking? And, maybe most useful for any operator who’s been blamed for everything: is the network actually to blame, or should the team go look at its own code?
We saw recorded demos, since the live cluster was, shall we say, still recovering. A downed NIC showed up as a detection the moment it failed. A lossy leaf-to-spine link was pinpointed down to the port and direction within seconds. A congested link showed one-way delays in the tens or hundreds of milliseconds, where microseconds should be. Gavin’s honest answer to how people find this today: “If you look at enough switch counters and you know where to look, you can track it down. It takes hours, and it takes expertise.” Clockwork wants to skip to the answer, and a future AI SRE layer that also explains the fix is on the roadmap.
FleetAudit, the pre-flight check
FleetAudit uses the same plumbing to validate a cluster before anyone runs a job on it, covering front-end and back-end networks, NCCL, NVSwitch, GPU health and PCIe topology. WhiteFiber, the GPU cloud, runs it before handing over clusters. The things it turned up are the kinds of faults that pass a basic test and then ruin a customer’s week: ConnectX-7 and BlueField NIC problems, marginal transceivers, MPO cable faults and ECN storms. Dan said work that used to take days of testing now takes minutes, and WhiteFiber is expanding across new clusters.
“Clockwork.io’s automated fleet audit validates every link and node at once, localizes faults in minutes, and lets us correct them before acceptance.”
— Tom Sanfilippo, CTO, WhiteFiber
A pilot at a large GPU cloud operator tells the FleetLens side. Its solutions architect called the network “a big black box,” and engineers had been taking leaf groups out one at a time to find a bad one. FleetLens isolated a single leaf-to-spine link and showed that exactly one job was affected.
The Business Side: Who Buys It and How
Dan was direct about the go-to-market. The two audiences are neoclouds and enterprises, with enterprises increasingly reached through partners. Together AI, which is bringing TorchPass and LinkPass to market as a service on its GPU Clusters, is close to a channel partner. Oracle Cloud Infrastructure embeds LinkPass. The customer list also includes Nebius, NScale, DCAI and Wells Fargo.
Pricing is per GPU per year for fault tolerance, with capacity-based deals for some large clouds. Clockwork also announced what I’d call a confident guarantee: the “you only compute once” (YOCO) promise, under which, the company commits that at least 90% of training failures on supported TorchPass workloads will be resolved through live GPU migration, with no lost training progress, no checkpoint rollback, and no recompute. If Clockwork.io falls short of that commitment in any contract year, customers receive a 25% credit against their next TorchPass renewal or expansion.
Competition is mostly the status quo. For observability, it’s whatever tools a customer already has. For fault tolerance, it’s checkpoint-restore and the question of whether the lost time is bad enough to act on. There’s no open source edition yet. Suresh, who ran an open source company before, said it takes real energy to make one succeed and Clockwork is focused on growing the business. The software runs on NVIDIA and AMD GPUs, with InfiniBand and RoCE, and builds on open-source pieces such as NCCL. TPUs and Trainium aren’t in the plan yet.
The funding supports that growth. The $31 million round was co-led by Premji Invest, Wing Venture Capital and Seligman Ventures, with existing investors NEA and e& Capital participating, and it brings total funding to $73 million. NEA’s Greg Papadopoulos, who first backed the company in 2021, summed up the thesis: “The one thing that scales perfectly is unreliability: put enough GPUs in one machine and something is always failing.”
Why You Should Pay Attention
If you run GPUs, or sell access to them, the arithmetic is hard to ignore. Every hour a $150 million cluster spends idle, restarting or recomputing is money gone. Hardware is scarce, power is scarce, and nobody can just buy their way out. As Dan put it near the end of the session, “efficiency is the new capacity.” The pitch from Clockwork, which cites 30 to 50 percent more GPU usage and 90 percent fewer disruptive checkpoint restarts, is that you can get that capacity from software.
What struck me is how unglamorous and how deep this problem is. Nobody puts “link flap” on a keynote slide. But it’s what actually stops a training run at 3 a.m., and Clockwork has paying customers who say it’s solving it, plus an independent analyst’s numbers to back them up. Add a team with VMware, NetApp, Nimble and Stanford networking in its DNA, and a live demo where they unplugged their own servers, and you have a company worth watching. If your platform team is still telling researchers to checkpoint more often and hope for the best, it may be time to take a look.
##






