Opens in a new tab
vmblog logo 2024 wht (updated)

StarTree 2025 Predictions: From Streams to Insights – 2025 Marks the Real-Time Analytics Revolution

Share: 

David Marshall | Published: January 3, 2025

 vmblog-predictions-2025 

Industry executives and experts share their predictions for 2025.  Read them in this 17th annual VMblog.com series exclusive.

By Peter Corless, Director of Product
Marketing at StarTree

While all the buzz is on AI, the actual
workhorse technology for 2025 will be real-time analytics. All the pieces have
fallen into place to make it mission critical for leading enterprises. This
year, you’ll see companies employ real-time analytics systems as the “last
mile” in an end-to-end data architecture, capable of ingesting and providing
sub-second insights against streaming data at petabyte scale.

Real-time analytics are already employed at
upper-quartile companies like Stripe, Uber, DoorDash, Cisco, LinkedIn. The
prevalence of open source solutions and hosted services will further
democratize and expand adoption across industries, as companies look to shift
increasingly from after-the-fact batch reporting delayed by hours or days to
real-time insights.

While you can argue real-time analytics are
nothing new – they’ve been a capability since Apache Druid (2011), ClickHouse
(2012), and Apache Pinot (2014) were first released – 2025 will hit different.
2024 saw a sea change in enterprise end-to-end architectures with all the
pieces falling into place.

First, data streaming. Over the past few
years, businesses have focused heavily on building out event streaming systems
like Apache Kafka, ensuring that data flows smoothly in real-time.  Confluent’s 2024 Data Streaming Report noted
69% were already using data streaming technologies for various critical
systems, and another 23% for non-critical systems. What the industry saw in
2024 was that Apache Kafka adoption was not the end, but a new beginning.
There’s now a growing field of Kafka-compatible systems and vendors, each
tailored and optimized towards various use cases, whether for highest raw
performance, or for volume/scale or for lowest cost/affordability. Kafka is the
ubiquitous protocol, but the implementation will be tailored to the
enterprise’s needs.

Next, stream processing. Data streaming
vendors like Confluent and Redpanda used 2024 to integrate stream processing as
a critical, core part of their services. And they are far from alone offering
stream processing solutions. Data from source systems needs to be combined,
enriched, filtered and de-duplicated before consumption. It’s not a
nice-to-have. It’s a need-to-have.

This leaves the final piece: real-time
analytics. Many organizations know that traditional analytic endpoints, such as
data warehouses and batch-based solutions, are unable to fully harness the
potential of these streams. And while they might employ search engines or
OLTP-based systems for a speed layer, those only work at smaller scales. They
were never designed for analytics in the first place, and certainly not at
terabyte-to-petabyte scale.

These legacy systems simply can’t deliver the
instant insights needed in today’s fast-paced environment. Or at least, not
affordably. In 2025, organizations will prioritize real-time analytics
platforms that can process and act on data insights instantly, closing the loop
and unlocking the true value of their streaming architectures. This shift will
enable innovative use cases such as hyper-personalized customer experiences,
real-time external-facing data products, and adaptive risk management systems-far
beyond the capabilities of traditional solutions.

That’s why companies like Uber are already
deploying a streaming + stream processing + real-time analytics real-time data
pipeline. In Uber‘s case, that’s Apache Kafka, Apache
Flink, and Apache Pinot back-to-back-to-back. This “KFP” stack, or equivalents
to it, are going to be seen at more enterprises in 2025 and in the years to
come.

As well, you’re going to see the accelerated
adoption of real-time analytics across observability use cases, displacing the
traditional mix of time series databases, search engines, and fast
transactional databases, both SQL and NoSQL. The reason being is simple:
real-time analytics provide faster insights at a lower cost-per-query. With the
disaggregation of observability stacks due to the embracing of Open Telemetry
(OTel), you’ll see more opportunities to employ real-time analytics for the
data storage and query layer.

If you want another industry leading example of real-time analytics in action
here at the tail end of 2024, check out Stripe’s Black Friday / Cyber Monday site. It
shows how they processed over 300 million transactions over this most recent
shopping holiday, accounting for $25 billion in cumulative payment volume.
Stripe has been generating real-time Black Friday / Cyber Monday statistics
for years, using Stripe
Radar
to watch for fraud detection on behalf of their vendors. All
powered by Apache Pinot.

In 2025, other companies’ CEOs will
increasingly ask, “Why can’t we answer our own questions in real time?”

##

ABOUT
THE AUTHOR

peter corless 

Peter Corless is the Director of Product
Marketing at StarTree.
Before StarTree he was the Director of Technical Advocacy at ScyllaDB. In his
long career in Silicon Valley he has also held various roles spanning from
Aerospike to Cisco Systems.