Opens in a new tab
vmblog logo 2024 wht (updated)

CTERA Says the AI Value Gap Is a Data Problem, But Can Send Engineers to Fix It

Share: 

David Marshall | Published: October 6, 2026
ai value gap data problem

At the 70th Edition of The IT Press Tour in Palo Alto, CTERA’s CEO and CTO explained why enterprise AI keeps stalling at the pilot stage, and what the company is doing about it with a new Forward Deployed Engineering service.

Everybody is buying AI. Almost nobody can prove it paid off.

That was the opening punch from Oded Nagel, CEO of CTERA, when the company met with the IT Press Tour in Palo Alto this week. It’s a blunt way to start a briefing, but he backed it up. GPUs are on order, licenses are signed, and employees are happily asking chatbots to write their emails. Yet the budget holders can’t tie any of it to a business outcome, and the token bill keeps arriving anyway.

CTERA’s answer, announced today, is a bit of a surprise coming from a software company: it’s sending its own engineers to sit inside customer environments until AI actually runs in production. More on that shortly. First, the reasoning behind it, because it makes the whole story click.

The Gap Nobody Wants to Talk About at Budget Time

Nagel framed the problem with what CTERA calls the AI Value Gap. Adoption has shot up. Value hasn’t kept pace. Citing McKinsey’s State of AI research, the company’s slides noted that nearly 90% of enterprises now use AI, while most never get past pilots.

Nagel was candid about who is stuck. Banks, insurers, government agencies, healthcare providers: the biggest and most regulated organizations are using AI, but they aren’t in production with it. He told us about a conversation two weeks ago with a storage manager at one of the largest banks in the US, who said plainly, “We are not there yet.” Another year or two, he figured.

So where does it go wrong? CTERA sees three culprits.

  1. Data that isn’t ready. Most organizations don’t know what’s sitting on their file servers. CTERA’s own research suggests only about 4% to 10% of enterprise data is actively accessed, and when the company classified customer estates, it found things like MP3 libraries and stray personal data alongside the valuable material.
  2. Security that can’t keep up. The slides cited research showing 67% of executives believe their company has already suffered a breach from unapproved AI tools, and that GenAI data movement rose 80% year over year, slipping past legacy DLP.
  3. Not enough people who know how. Your storage admins are good at storage and cloud. Building agents that are secure, efficient, and pointed at the right data is a different trade. The slides noted the skills gap as an organizational challenge more than doubled in six months, from 4.8% to 10.4%.

Here’s the thing, though. None of this is the models’ fault. As Nagel put it, the AI didn’t fail. What’s missing is a trusted layer of context sitting between the models and the data, and few people have the skills to build it.

Move the AI to the Data, Not the Other Way Around

This is where CTERA’s long history matters. The company has spent roughly 18 years building a global file system that ties headquarters, branches, cloud VPCs, and endpoints into one namespace on top of object storage. Aron Brand, CTO, still lives and breathes it.

From that foundation, Nagel laid out a philosophy that ran through the entire session:

“Instead of moving the data to the AI, move the AI into your corporate data.”
— Oded Nagel, CEO, CTERA

Why does that matter? Because the usual approach, copying a petabyte into a cloud service so Copilot or a chatbot can read it, creates a second, ungoverned copy. Permissions don’t follow it. It doesn’t stay in sync. Brand compared it to the early Wi-Fi days when people plugged in their own access points, and said today’s version is shadow AI: files tossed into ChatGPT, creating repositories nobody knows about.

CTERA’s approach keeps the data where it lives and keeps the existing file permissions in force. If someone in marketing asks about a finance file they can’t open, the AI doesn’t hand it over. Simple to say. Hard to build.

Two Steps: Understand First, Then Activate

Brand walked us through the company’s product pieces using three real customers. I’ll be honest, I expected a parade of logos. What I got instead was a fairly logical, two-step story.

Step One: Know What You Have (CTERA InsightAI)

InsightAI is an administrator tool that reads only metadata. It never opens the files, so there are no egress fees and no months-long classification project. It looks at who touches files, what’s growing, file types, and names, then infers a surprising amount from that.

Brand’s example stuck with me. A traditional analytics tool sees two folders of audio files, both huge. InsightAI sees one full of files named after famous pop artists (a music library that probably shouldn’t be on a corporate file system) and another with a numbering scheme that looks like a call center, which is high-value, sensitive, and worth protecting.

The first customer story was a manufacturer with more than 1 PB across four regions and 33 CTERA filers. They had no folder-level view of growth, staleness, or access. InsightAI now gives them per-share growth analytics, automatic recommendations for archiving redundant, obsolete, and trivial data, and faster audit log searches for investigations.

The demo was a chat interface. Brand asked for the top active users and then for good archiving candidates, and the system built queries against its metadata catalog (OpenSearch underneath), reasoned about the results, and answered in plain language. It can also run scheduled reports, such as a monthly storage health check with recommendations that aren’t hard-coded.

Two details worth flagging:

  • It’s read-only by design. You can download an archive candidate list, but the AI doesn’t move anything on its own. Brand said a human-approval step is coming in the interface.
  • It helps with forensics. Because InsightAI keeps audit logs in a data lake, after a ransomware alert you can ask what that user did on that date and reconstruct which files were read or deleted. Brand argued that’s what lets you tell a regulator what was actually exposed, instead of assuming everything was.

One customer told them that half their support tickets are people who can’t find a file. Honestly, I laughed, because I’ve lived that.

Step Two: Put the Data to Work (CTERA Content Services)

Once you know which data matters, Content Services goes deeper. Unlike InsightAI, it reads the content, and it’s aimed at end users and AI agents. It runs OCR and transcription, converts everything to markdown, and handles full-text and semantic search, classification, summarization, and structured field extraction. It stores results in Postgres and the Qdrant vector database, and it is model-agnostic: Claude, OpenAI, DeepSeek, or a private LLM for sensitive material.

Brand’s second customer was a mortgage company with about 27 TB of loan files spread across Minneapolis, Washington D.C., and Azure. CTERA Portal runs on Azure Blob storage with edge filers in the two cities. The firm tags personal information automatically, runs access review cycles, and exposes agentic “experts” through MCP in Microsoft Copilot Studio. The goal is to turn dead files into curated datasets that can help predict loan performance.

Then came the demo that made the room lean in. Brand used a call-center scenario. Audio files dropped into a folder, and CTERA created transcripts and a JSON file of structured fields next to each recording. Then Claude Code, connected over SMB and MCP, built a churn-risk dashboard in about ten minutes, cross-checking Salesforce for upcoming renewals. The agent even removed read access for others on files flagged as containing personal information.

The smart part is the economics. The audio is processed once, the derived text is tiny (roughly 1% of the original data on average), and it can sit in the cloud without repeated egress. Cheap models handle the curation. You save the expensive models for the hard questions. As Brand put it, you shouldn’t burn top-tier tokens transcribing audio.

Wait, If the Demo Took Ten Minutes, Why the Gap?

Good question, and Nagel raised it himself. That dashboard looked easy. So why are enterprises stuck?

Because the ten minutes are the visible part. Building the workflow, picking the right schema, deciding what counts as reliable data, and wiring in permissions is where organizations flounder. Brand gave a sharp example. If you force a model to classify every document as one of three types and one doesn’t fit, it will hallucinate, because you left it no other option. Add an “other” category and the problem disappears. In a legal setting, he noted, an AI with no sense of what is reliable might treat a criminal’s testimony the same as a police officer’s evidence.

“If you give it low quality data, you’ll get hallucinations… very convincing nonsense.”
— Aron Brand, CTO, CTERA

That’s the essence of the third customer story, a small medical-legal firm that needed to turn scanned, handwritten medical files into case narratives suitable for legal review. CTERA built a Classify pipeline on real case files, with custom prompts and JSON schemas, models selected for quality, speed, and cost, and regression tests on every model change. And it was built with an engineer sitting alongside them.

Enter the Forward Deployed Engineer

That brings us to today’s news. CTERA announced CTERA Forward Deployed Engineering (FDE), a service that places CTERA engineers inside customer environments to put specific AI use cases into production on the customer’s own unstructured data.

The model borrows from Palantir, and Brand was upfront that CTERA didn’t invent it. But it’s a notable turn for a software vendor. Professional services used to be seen as a failure of product design, something to minimize. Brand said the situation has flipped: customers are “thirsty for knowledge” and want somebody to hold their hand.

Here’s how the engagement is structured:

  1. Discovery. Engineers read the file estate, metadata only, to validate use cases with real data.
  2. First use case. They prepare the data, build the workflow, validate it, and go live.
  3. Scale. The work extends to adjacent workflows and sites.
  4. Handover. The customer’s team runs and extends it.

Small teams pair a Forward Deployed Engineer with a Deployment Strategist, and the principles are straightforward. Start with the data. Classify, organize, and govern in place, so AI only touches content each user is already authorized to see. Integrate with tools the organization already uses, via open standards such as MCP and automation platforms like n8n. Then transfer ownership. Customer engineers work alongside CTERA from day one and document the configuration as it’s built, and an engagement is done when the customer can add the next use case without CTERA.

Futurum’s Brad Shimmin, quoted in the announcement, backs the premise. Futurum’s research found talent shortages, MLOps complexity, and integration problems are the top three data-related factors in AI project failures, and that more than 90% of enterprises face architectural bottlenecks building AI agents.

Nagel summed up the idea:

“Enterprises aren’t short of AI tools. What they’re short of is the time and specialized skills to connect those tools to their own data and workflows safely.”
— Oded Nagel, CEO, CTERA

The Practical Details

During Q&A we got some specifics. CTERA already has a small team ready to go, and they are prepared to scale up as needed. A typical engagement starts with at least a week of preparation and might run about a month, although larger customers may want a permanent FDE. Per the announcement, the service is available now to organizations using the CTERA Intelligent Data Platform, structured either on a time basis or as fixed-scope projects with milestone-based acceptance criteria.

Is CTERA turning into a consulting shop? Both executives said no. Nagel stressed that the strength still comes from the technology, and he doesn’t want “an army” of engineers. The pitch is that CTERA controls the whole stack, from file system and edge to search, embedding, and metadata, so the engineer isn’t stitching together five vendors. Brand added that the FDEs stay tightly connected to the core engineers. When they hit a gap in the infrastructure, they can call him and ask for a fix, so field knowledge flows back into the product.

So What’s the Takeaway?

Like most of the IT Press Tour briefings, this one had a clear thread. CTERA isn’t trying to build a better model, and Brand was blunt about that: the frontier labs are impressive but interchangeable. The company’s bet is that the lasting value is in the enterprise’s own data, plus the know-how to prepare it.

For IT teams sitting on decades of file shares, the sequence is easy to follow. Find out what you have with InsightAI. Curate the high-value sets and give agents context with Content Services. And if your team doesn’t yet have the skills, borrow some from engineers who’ve built these pipelines before, then take over.

Will it close the AI Value Gap? That’s a big claim, and CTERA’s own people admit they can’t predict where AI will be in a year. But as Brand joked, we’re living in a singularity. In the meantime, a vendor that offers to sit next to your admins and leave you able to run it yourself is making a promise worth watching.

##