Opens in a new tab
vmblog logo 2024 wht (updated)

The Data Management Platform Solving What Cloud Couldn't: Inside Arcitecta's Mediaflux

Share: 

David Marshall | Published: October 7, 2025

 

When Princeton University realized they needed a 100-year data management plan for their research data, they faced a sobering question: How many technology refreshes will happen over the next century? And more importantly, how will a Nobel Prize-winning researcher find their 2024 data in 2040?

This isn’t a theoretical exercise. It’s the kind of problem that Australian-based Arcitecta has been quietly solving while the rest of the industry chases the latest buzzwords.

At the 64th edition of The IT Press Tour in New York, Arcitecta CEO Jason Lohrey and Director of Product Marketing Eric Polet walked us through something refreshing: actual deployments, real customer problems, and solutions that work without requiring organizations to rip and replace their entire infrastructure.

Built Different: A Database at the Core

Here’s what makes Arcitecta genuinely different. “I think we’re the only ones that properly have a database at the core of a file system, and that gives some amazing benefits,” Lohrey explained.

That database is XODB – an XML-encoded object database created in 2010 as an early form of NoSQL. It’s not just a file system index. It’s a metadata powerhouse that sits at the heart of everything Mediaflux does.

The numbers tell the story: Arcitecta has a customer managing over a trillion objects in a single namespace. Not theoretical – actually deployed and running. The file system structure averages just 75 bytes per file, making it two orders of magnitude more dense than traditional file systems while remaining fully traversable.

When Cloud Promises Meet Reality

Dana-Farber Cancer Institute’s story should sound familiar to anyone who’s dealt with cloud repatriation. They went all-in on cloud early in their journey. “And boy did they get burned,” Polet said bluntly during the presentation.

The result? They brought 95% of their data back from the cloud. The culprit wasn’t cloud itself – it was the hidden costs, API charges, and egress fees that made AWS unsustainable for their research workloads.

Now they’re running Mediaflux with Wasabi cloud storage (for transparent pricing), Spectra Logic tape libraries, and colo facilities at Markley data centers in Boston. The kicker? Mediaflux manages all of it under a single namespace, giving them vendor flexibility without the migration headaches.

“After being burned by cloud and going into cloud and feeling the pain, they said, we’re going to take a much more measured and cautious approach to AI,” Polet explained. That measured approach led them to work with Arcitecta on vector database integration – making their data AI-ready without the hype.

Princeton’s TigerData: Free Storage as a Data Management Strategy

Princeton University took a different approach to their data management challenges. For 250 years, their library staff could manage physical artifacts in their heads. Then data started growing exponentially, and human memory couldn’t keep up.

Enter TigerData – a research data management platform built on Mediaflux that manages 200 petabytes of research data. The platform currently tracks nearly 497 million assets.

The clever part? Princeton offers free storage to researchers as the “carrot” to encourage proper data management and metadata tagging. Eventually they’ll transition to a paid model (the “stick”), but first they needed to change researcher behavior.

Chuck Betler, who runs TigerData at Princeton, explained their architecture: “Really, the Mediaflux application, and this is where we get to interact with all of that heterogeneous storage, begin to expand it to new technologies, new producing things. So we want to bring on cloud providers behind TigerData or a new technology that emerges in the field. We can do that.”

Here’s what’s particularly interesting: Princeton was locked into IBM with support contract increases they didn’t expect. They used Mediaflux to break free, adding Dell PowerScale, Dell ECS, and six IBM Diamondback modular tape libraries – all managed under that single Mediaflux namespace.

And here’s a technical nugget that matters: IBM’s Diamondback didn’t have S3 integration. Arcitecta built it for them, creating what IBM now calls Spectrum Deep Archive. IBM is literally using Arcitecta’s S3 implementation for their own hardware at Princeton.

The Modular Tape Renaissance

Something unexpected is happening in the data center: tape is making a comeback. Not the massive 40-frame libraries that require 96 feet of floor space, but modular units you can roll in, plug in, and expand as needed.

“Just in this past year, we see DFCI, Princeton, and MIT all add tape libraries,” said Graham Beasley, who runs Arcitecta’s US operations. “Whereas before we saw a lot of people replacing tape libraries and moving to Glacier and cloud. Now we’re seeing them come back.”

The reasons are practical: data sovereignty, privacy requirements, transparent costs, and the ability to air-gap data by simply unplugging a module. For research institutions dealing with sensitive data and hundred-year retention requirements, owning the physical media matters.

Mediaflux’s special sauce here? Take seven Diamondback libraries, stack them next to each other, and Mediaflux makes them look like one library while letting you buy and manage them as individual units.

AI Without the Nonsense

While everyone else rushed to slap “AI-powered” on their marketing materials, Arcitecta spent two years actually building vector database capabilities into XODB.

The implementation is practical: Mediaflux doesn’t try to generate vector embeddings itself. Instead, it orchestrates the process – sending data to specialized services like Wasabi AIR (formerly Granulate) for facial recognition, object detection, and optical character recognition, then storing those vector embeddings alongside metadata in XODB.

National Film and Sound Archive in Australia beta-tested this integration. They have petabytes of video and audio content that needs to be searchable not just by filename, but by who appears in the video, what objects are shown, and what’s being said.

The result? A unified search across files, metadata, and vectors. No jumping between different interfaces. No separate vector database to manage. Just data that’s ready for whatever AI tools make sense for the specific use case.

“We don’t like to talk about the things that we’re going to do. We like to talk about the things that we’re doing or have done,” Polet emphasized. “Vector database, AI – that’s something that had been bubbling in the background for a few years.”

Real-Time Replication: The Game Changer for Media

At Dell’s request, Arcitecta built something the broadcast industry didn’t know was possible: true real-time file system replication.

Mediaflux Real-Time lets someone stream data at location A while someone at location B reads that same growing file with latency measured in tens to hundreds of milliseconds. You can write via SMB at the source and read via NFS at the destination, with the data appearing on remote storage as if it had always been there.

For VFX studios, this means reviewing renders as they happen – catching color or saturation issues immediately instead of waiting for the entire project to finish. For broadcast, it means remote editors can work on live material without being co-located with the point of capture.

When Arcitecta showed this to a broadcaster, their response was succinct: “This is a game changer.”

The TU Dresden Story: When New Instruments Break Your Workflow

Technical University of Dresden’s challenge illustrates a problem many research institutions face. They acquired new instruments – better microscopes, improved sensors – that generated exponentially more data. Great for science, terrible for data management.

Each new instrument came with its own storage system and format, fragmenting what had been a unified workflow. Archive retrieval became slow. Finding data became difficult. Collaboration became nearly impossible.

“They said, you know, this is completely inefficient. This is not working. We can’t upgrade our infrastructure because it’s breaking it, and our scientists need new data. They need new instruments. How can we do that?” Polet recounted.

Arcitecta partnered with GRAU DATA’s XtreemStore (a tape gateway) to unify everything. Metadata-driven indexing made data discoverable again. Automated archiving policies moved cold data to tape. And researchers could collaborate across teams because everything lived in a centralized repository.

The benefits were tangible: optimized workflows through automation, seamless access even to tape-stored data, better collaboration, lower costs, and infrastructure that could actually scale with their research needs.

Datakamer: Building a Community

Perhaps the most interesting development has nothing to do with technology. Dana-Farber Cancer Institute approached Arcitecta about hosting an event where Mediaflux users could share experiences and learn from each other.

The result was Datakamer – a birds-of-feather style event in Boston with over 50 attendees from Princeton, MIT, Rutgers, Whitehead Institute, and other research organizations. No death by PowerPoint. No vendor pitches. Just practitioners discussing real challenges: 100-year archive plans, sharing data across organizational boundaries, compliance requirements, managing hundreds of petabytes.

The feedback? “Best event we’ve ever had. I actually learned something at an event.”

Now Princeton wants to host the next one. MIT volunteered to host another. TU Dresden in Europe wants to do one. National Film and Sound Archive in Australia is interested. What started as a one-time gathering is becoming a self-sustaining community.

“We want to create our own version of that for the northeast or for anywhere,” Polet said, referring to the Rocky Mountain Advanced Computing Consortium conference that’s been running for a decade in Colorado.

The Python Module and What’s Next

Arcitecta’s roadmap for the next 6-12 months focuses on making deployment easier and expanding capabilities:

Product enhancements include a Python module (debuting at Supercomputing 2025’s user group meeting) to make integration simpler. Python is ubiquitous in research environments, so this matters.

Mediaflux DAMS (Digital Asset Management System) is being genericized beyond its initial National Film and Sound Archive deployment, with more robust role-based access control and features suitable for galleries, libraries, archives, and museums – a market Arcitecta believes has been neglected by technology vendors.

Vector database expansion will add more third-party integrations and potentially enable conversational queries – asking your data questions in natural language and getting real answers.

Deployment streamlining addresses one of Arcitecta’s acknowledged pain points: getting from initial deployment to production-ready can take weeks or months. They’re working to compress that timeline.

The Partnership Ecosystem Expands

Arcitecta’s partner network has grown substantially. Dell signed a reseller agreement, giving them the ability to sell Mediaflux directly. IBM wasn’t even on the radar 18 months ago – now they’re using Arcitecta’s S3 implementation for Diamondback libraries.

Wasabi has been a consistent partner, from supporting the Datakamer event to co-developing the vector embedding integration. Spectra Logic continues their long relationship around tape integration.

And there’s an interesting hint: with Rob Mullard joining from HPE (where he worked with Cray, SGI, and others), HPE partnerships seem likely. “A lot of our employees are ex-SGI, ex-Cray,” Beasley noted. “We were very tight with the HPC team.”

Why This Matters

Here’s what makes Arcitecta’s story worth paying attention to: they’re solving real problems for organizations managing petabytes to hundreds of petabytes of unstructured data. Not with vaporware or slideware, but with production deployments managing trillion-object namespaces.

They’re not trying to be everything to everyone. “It’s better to be good at what we’re good at, rather than trying to stretch ourselves,” Beasley said when asked about Fortune 500 financial services companies.

Their sweet spot is unstructured data at research institutions, media and entertainment companies, and GLAM organizations – anywhere large instruments (microscopes, telescopes, satellites, broadcast cameras) generate explosive data growth.

The company remains around 50 employees globally, growing conservatively. 

The Vendor Lock-In Solution

Perhaps the most compelling aspect is vendor agnosticism. Mediaflux doesn’t care what storage you use – NetApp, Dell ECS, IBM, cloud blob stores, tape, whatever. It works with any protocol: NFS, SMB, S3, SFTP.

And here’s the part that should matter to IT decision-makers: you can exit Mediaflux anytime. Export your metadata. Access your data directly at the source. No lock-in. No hostage situations.

As one customer told Polet early in his tenure: “When we first started working with you guys, we always told you what we wanted you to do. What we learned over the first year is we needed to tell you the problem we had, not how we wanted you to solve it. You guys came through with answers that were better than what we thought.”

That’s the kind of partner relationship that keeps customers around – and gets them talking to peer institutions about the weird Australian company with the database-driven file system that actually works.

For organizations drowning in data growth, locked into expensive vendor relationships, or planning for decade-spanning data retention, Arcitecta’s Mediaflux deserves a serious look. Just don’t expect them to overpromise or chase buzzwords. They’ll be the quiet ones in the back of the room, working on the right answer.

##