Opens in a new tab
vmblog logo 2024 wht (updated)

When AI Projects Go Wrong: How Helikai's Approach Breaks the Failure Pattern

Share: 

David Marshall | Published: January 28, 2026

vmblog-helikai-approach 

There’s a familiar story playing out across Fortune 500 companies right now. Someone in the C-suite issues an AI mandate. Smart engineers get assigned. Budgets get allocated. Months pass. Then nothing happens. Or worse-something happens, but nobody can explain what it does or whether it’s actually working.

Jamie Lerner and Ross Fujii lived that story. Both veterans of massive tech companies, they watched their organizations throw resources at artificial intelligence only to watch those resources disappear into endless consulting engagements and perpetually incomplete projects. That experience led them to start Helikai, a company with a deceptively simple philosophy: maybe the problem isn’t that enterprises are thinking too small. Maybe they’re thinking too big.

At the 66th IT Press Tour in Palo Alto, the co-founders explained how their approach to AI, one built on narrow, purpose-built agents instead of all-encompassing systems, is actually delivering results with some of the world’s largest companies.

The “Boil the Ocean” Problem

Let’s talk about what usually happens when a company decides to deploy AI. The thinking goes something like this: We need to transform our entire operation with artificial intelligence. So they pick the most complex, most ambitious thing they can imagine. Maybe it’s automating the entire quote-to-cash pipeline. Or reinventing the customer data platform. Something big enough to justify the investment and the organizational disruption.

“I said, just give me something. Make something simple work,” Lerner recalled from his own experience at a previous company. “Just use AI to validate an address. Calculate shipping. Something simple.”

That shift-from “let’s solve everything” to “let’s solve one specific thing well”-became the foundation for how Helikai thinks about AI. Instead of asking “How do we automate our entire business?” Helikai asks “What’s one discrete task we do every day that’s boring, error-prone, or time-consuming?”

The difference matters. A lot.

Why Big AI Plans Usually Fail

There are a few reasons why enterprise AI projects drag on forever. The first is obvious: ambition collides with reality. When you’re trying to build a system that understands everything about your business, you need to teach it everything. That takes time. Years, usually.

The second reason is more subtle. Large language models are creative. They hallucinate. That’s actually great if you’re writing marketing copy or brainstorming ideas. It’s catastrophic if you’re calculating taxes. You can’t have your system tell a customer their shipping address is in Ohio one day and Arkansas the next. You can’t generate a quote and have it say the price is $50,000 today and $75,000 tomorrow when nothing has changed.

“Enterprise systems need true 99-plus percent accuracy,” Lerner explained. “You don’t get into an ERP system and say, ‘Yeah, this quote is probably fine. 89 percent chance this quote is accurate.’ You can’t run a business that way.”

Most AI approaches struggle with this requirement. When you train a model on vast amounts of data and tell it to be creative, you get hallucination. The bigger the model and the broader the training data, the worse the problem becomes.

The Micro AI Alternative

Helikai’s solution is to go in the opposite direction. Instead of building systems that know everything, they build systems that know almost nothing-except for the one thing they’re supposed to do.

They have an agent that reads purchase orders. It doesn’t know the entire history of your company, every book ever written, or the intricacies of your supply chain. It reads a purchase order and extracts the relevant information with near-perfect accuracy. Another agent validates shipping addresses. It knows whether an address is real, whether a truck can access it, whether it has a loading dock. Nothing more.

These agents are small enough, scoped enough, and focused enough that Helikai can push them to enterprise-grade reliability. They’re also small enough to run on actual hardware-not massive GPU clusters, not distributed computing infrastructure. A typical enterprise can run a full suite of these agents on a $22,000 server.

“We’re not trying to reconstruct the human mind,” Fujii said. “We’re not trying to digest every piece of data ever printed by mankind and house it in superintelligence. We’re doing the opposite.”

This distinction separates Helikai from the consultancy-driven, platform-heavy AI vendors that dominate the market. They’re not selling you a foundation to build anything. They’re selling you solutions for specific things.

The Workshop: Where Theory Meets Reality

Helikai doesn’t do traditional sales calls. If you want to work with them, you have to commit to spending a day or two actually doing work, understanding what you’re trying to achieve, where you actually are, and whether your ambitions make sense given your current maturity.

They use the MITRE AI Maturity Model as a framework. Do you have an AI platform? Do you have governance? Do you think about bias? Have you considered the geopolitical dimensions of the models you’re using? (This matters – a Chinese model and a U.S. model are trained differently, and those differences have real implications.)

“Very quickly, you realize most companies are not that mature,” Lerner said. “So we question why they want to start with something so complex.”

What emerges from this workshop is a roadmap. Not a five-year transformation plan or a “we’ll migrate everything to AI” kind of strategy. Instead: here’s what you should do first, given where you actually are. Start with simple, pre-built agents. Get a win. Then expand.

This approach sounds almost boring compared to the transformation narratives that usually dominate AI conversations. But boring, it turns out, is actually what works.

The Catalog: Solving Real Problems at Scale

Helikai has built a catalog of around 200 agents, with new ones shipping multiple times a week. Some of these are foundational-things like document processing, where an agent reads an invoice or purchase order and tears it apart into structured data: customer name, amounts, products, signatures.

Others are vertical-specific. In healthcare, they have agents that can look at medical images-pathology slides, X-rays, CT scans-and provide analysis. In media and entertainment, they have agents that generate subtitles in multiple languages, dub content, remove dust and scratches from old film, even generate background scenes for LED walls. In legal, they analyze contracts, generate motions based on precedent, and process discovery documents.

What’s interesting about this catalog isn’t the breadth-though it’s broad. It’s the depth. Take the example of a large furniture retailer. This company was spending millions of dollars and thousands of hours every year to generate product photography for millions of SKUs. They’d create 3D renderings, have artists manually paint shadows and textures onto those renderings, place them in virtual rooms, and hope the end result looked good.

Helikai built an agent that does all of that. It takes a 3D model, generates shadows based on light angle and intensity, applies textures, places the result in a background. The agent does what used to take artists hours in minutes. And it does it with perfect scientific accuracy-because shadows aren’t an art form, they’re physics.

But then the CEO of that company came back and asked for something else. “I want to talk to my data,” they said. “I want to ask my data warehouse questions.”

So Helikai built another agent-a semantic interface that lets executives ask natural language questions (“Why don’t these rugs sell in Europe?”) and get answers without waiting a week for a report. The company started with one problem solved, saw the value, and started buying more agents.

“Success breeds success,” Lerner noted. “And the opposite… if you have a big AI failure, everyone’s reticent to do anymore.”

Architecture That Actually Works in the Real World

The engine behind all this is called SPRAG-Secure Private Retrieval Augmented Generation. It’s worth understanding because it shows why Helikai’s approach produces enterprise-grade results.

A traditional RAG system ingests your documents, vectors them, and then uses large language models to generate answers. The problem is that all that data and all those vectors need to live somewhere, and for most companies, the somewhere is a third-party vendor’s servers. That creates data security and privacy concerns-especially for regulated industries.

Helikai’s SPRAG runs on your hardware. All your data stays behind your firewall. The system supports dozens of different models simultaneously-OpenAI, Google Gemini, Anthropic, open-source alternatives-so you’re not locked into any single vendor. When a new model comes out, Helikai drops it onto your server. When models improve in a few months, you get the benefit without rebuilding anything.

The clever part is the architecture itself. SPRAG isn’t just a vector database. It combines a vector database, a relational database, and traditional programming tools. Here’s why that matters: when you need to look up a tax rate in Texas, you don’t want AI doing that. You want a simple database query. Fast, accurate, deterministic. Same thing with checking inventory or validating customer credit.

“If I have to log into an Oracle database and say, ‘What is the tax rate in Texas?’, I don’t want AI to do that,” Lerner explained. “I don’t need artificial intelligence. That’s just classic, old-fashioned query a piece of data.”

So Helikai’s workflows weave between AI and non-AI components. For the ambiguous, context-dependent parts of a process, AI handles it. For the deterministic parts, traditional tools do. This hybrid approach is how they consistently hit 99-plus percent accuracy on business-critical processes.

The Human-in-the-Loop Layer

There’s another component called KaiFlow, and it’s where the system becomes genuinely interesting for regulated industries and high-stakes operations.

Every decision the AI makes gets logged. Every human interaction gets recorded. You can see the entire chain of reasoning-what data the agent considered, what it concluded, where it made decisions, and where humans intervened.

Importantly, humans can intervene anywhere. If the AI flags something unusual-a discount that’s bigger than policy allows, a customer it doesn’t recognize, an invoice that looks fraudulent-it can stop and ask for human approval. Or it can just flag it as suspicious and let a human review it later.

“We can look at everything that’s occurring in these complex pipelines and anywhere insert human interaction,” Lerner said. “Ask a human, stop. Go tell the sales manager and ask the sales manager what to do.”

This matters in practice. Helikai works with pathologists who look at microscope images all day. The system learns how pathologists analyze cells, then builds an agent to do that analysis. But the pathologist still reviews the results. They haven’t been replaced-they’ve been given a tool that handles the routine parts and flags the hard cases.

Similarly, if you’re processing invoices and the AI encounters a duplicate vendor number or conflicting data, it doesn’t just guess. It stops and asks a human to resolve it. This prevents garbage data from feeding into your systems and confusing the AI further down the line.

Cost, Speed, and Sanity

Here’s the part that matters to IT budgets: Helikai typically prices its services at 85 percent of what equivalent human work would cost. Not 20 percent cheaper, not 50 percent. Fifteen percent cheaper than hiring people to do the same thing.

That might sound modest until you consider the other variables. An AI agent processes work at machine speed. It doesn’t take vacations. It doesn’t make mistakes from fatigue. And it produces audit trails-something you don’t get with human hourly work.

“We had a customer with 10,000 old contracts. Human reading them was going to take two humans working full time for 600 days. That’s two years,” Lerner said. “We did it in two days and did it 15 percent cheaper. They just said, you have to pay us and get the work done.”

From a project management perspective, this changes everything. Most enterprise AI initiatives have unclear timelines and variable costs. You’re paying consultants by the hour, trying to figure out what’s actually happening, wondering when it will be done. With Helikai, you know what you’re getting, when you’re getting it, and what it costs.

“We know exactly what our agent costs, how long it takes to deliver it, because they’re doing such a scoped and defined item,” Lerner explained. “We’re not going to a company and saying we’re going to automate all things in your company. We pick very discrete tasks and we get rapid results and reliable business outcomes.”

Why This Matters for IT Leaders

If you’re an IT administrator or architect, the implications are worth thinking through.

First, budget certainty. You know what you’re paying for. You know when it will be done. You know what you’ll get.

Second, risk reduction. You’re not betting the company on a grand AI transformation that might fail. You’re solving one problem. If it works, you solve another. If it doesn’t work, you’ve only wasted one month, not two years.

Third, infrastructure simplicity. You don’t need massive GPU clusters or exotic hardware. A mid-sized SPRAG server handles most workloads. It can run on-premises, in a private cloud VPC, or hybrid. Your data doesn’t have to leave your facility.

Fourth, and this is subtle but important, Helikai’s approach insulates your organization from the constant churn in the AI market. New models ship every few weeks. New vendors emerge. New hype cycles begin. Helikai drops new models onto your server and you get the benefit automatically. You’re not locked in, and you don’t have to make infrastructure decisions every time OpenAI releases a new version of their model.

The Catalog Approach: A Blueprint for Enterprise AI

What Helikai has essentially done is solve the “custom project” problem that has plagued enterprise software for decades. They spent time understanding what problems come up again and again-document processing, data extraction, report generation, knowledge retrieval-and built best-in-class solutions for those problems.

A customer doesn’t have to reinvent the wheel. They start with agents that already exist, already work, already have the rough edges sanded off. The agents get trained on their specific data, deployed into their environment, and they’re done.

“We have a catalog of over pretty close to 200 agents now,” Lerner said. “So we pick agents right off the shelf that are built, tested, ready to go. We might train them off their data, the customer accepts them and then we help them operate and run those agents.”

As organizations mature and their ambitions grow, they can move to custom agents. But they don’t start there.

A Different Kind of AI Company

What makes Helikai different from the larger platforms and consultancies is how explicit they are about what they’re not trying to do. They’re not trying to build a “single source of truth” for your data. They’re not trying to replace your data warehouse or your CRM. They’re not pretending that one system will solve all problems.

They’re building agents-small, focused, effective tools that solve specific problems. And they’re willing to walk away from customers who aren’t ready for that approach or who want to start with something unrealistic.

“We question why they want to start with something so complex,” Lerner said. “What we do is let’s set you up for success. Let’s pick some simple use cases where we have a lot of known technology and then you start building a staircase to more and more complex use cases as the maturity goes up.”

This philosophy-focus on success, add complexity gradually, deliver measurable value early-might not make for exciting boardroom presentations. But it’s how you actually get AI working in enterprises.

And honestly? After years of watching AI initiatives stall, disappear, or deliver disappointing results, that pragmatism looks less like boring and more like refreshing.

##