Sleep.ai CEO Colin Lawlor has a blunt message for enterprises racing to deploy AI: the model isn’t your problem anymore. In this candid Q&A, Lawlor argues that the real bottleneck in production AI is the data infrastructure underneath it — specifically, the ability to ingest, process, and act on continuous, real-world data streams at scale. “Garbage in equals garbage out has never been more true,” he warns, “but when you combine that risk with the confidence of an LLM, you have real problems — particularly in health.”
Lawlor speaks from direct experience. Sleep.ai works with one of the most complex continuous data types in human health: sleep. Unlike episodic or batch datasets, sleep data flows across days, weeks, and months, demanding pipelines built for persistence, variability, and real-time reliability — not just volume. It’s a challenge that exposes exactly where most enterprise AI deployments quietly fall apart.
In the conversation below, Lawlor breaks down the three biggest pipeline hurdles organizations face, explains how the company’s newly launched Sleep Sense API helps developers skip the hard infrastructure work, and makes a compelling case for why sleep — backed by more than 250 billion data points and validated against hospital-grade gold standard data — sits at the center of the next generation of predictive and preventative enterprise AI.
++
VMblog: There’s increasing discussion about the gap between AI model performance and the data infrastructure needed for real-world deployment. How are you observing this gap manifest today?
Colin Lawlor: Most AI models look impressive in controlled environments, but they break down when exposed to messy, real-world data. The gap shows up in latency, inconsistency, and the inability to handle continuous inputs. At Sleep.ai, we see this clearly—models trained on clean, episodic datasets struggle when faced with continuous behavioral data like sleep. The issue isn’t model capability anymore; it’s a combination of very high quality and well characterized data along with the infrastructure required to ingest, process, and act on these real-time, longitudinal data streams. That’s where most deployments fail. You will remember the age-old adage ‘Garbage in = Garbage out’ – it has never been more true, but when you combine that risk with the ‘confidence’ of an LLM, then you have real problems, particularly in areas like health.
VMblog: Sleep.ai specializes in working with continuous, time-series data. Why is this type of data so crucial for AI systems operating in production environments?
Lawlor: Because real life isn’t static—it’s continuous. Time-series data captures patterns, trends, and deviations over time. In sleep, a single night tells you very little (unless it is a night at a Hospital Sleep Lab). The real signal comes from consistency, variability, and change across days, weeks, and months. But you also need context day by day – did my intervention yesterday have an effect last night? For AI to deliver meaningful, actionable outcomes—especially in health—it needs context. Continuous data provides that context. Without it, you’re making decisions on snapshots, not reality.
VMblog: What are the biggest hurdles organizations face when designing data pipelines that can support always-on, real-world AI use cases?
Lawlor: The question assumes that the data going through those pipelines is of sufficient quality to serve the purpose – that is not always the case! (For example, in the case of sleep data, which wearable can I trust?). Assuming the data is well-characterized and understood, then three things:
- Data fragmentation – data comes from multiple sources, formats, and devices
- Reliability at scale – pipelines break when moving from batch to continuous ingestion
- Signal vs noise – most real-world data is messy, incomplete, and inconsistent
The biggest mistake is underestimating how hard it is to maintain quality over time. It’s not just about collecting data—it’s about making it usable, trustworthy, and consistent enough for AI to act on.
VMblog: Sleep.ai recently launched its Sleep Sense API. How does this enable developers and enterprises to better operationalize AI using real-world behavioral data?
Lawlor: Sleep Sense abstracts away the hardest part—turning raw, continuous sleep data into structured, actionable signals. Developers don’t need to build complex pipelines or models from scratch. They can plug into a normalized, validated layer of sleep intelligence that’s already designed for real-world variability.For enterprises, it accelerates deployment. Instead of experimenting with fragmented data, they can immediately integrate sleep as a behavioral input into their AI systems—whether that’s for health, performance, or engagement.
VMblog: How do you ensure data quality and reliability when dealing with long-duration, real-world data streams like sleep?
Lawlor: It starts with an enormous amount of high quality data which itself is god standard (in our case, Hospital Sleep Labs – but you can’t successfully train your models with small datasets because there are huge differences between people, their environment and what issues or illnesses they are suffering from). Then, it is about acknowledging that real-world data is inherently imperfect (and with hundreds of wearables and thousands of phones, you can’t find the ‘ground truth’ without the former).
We focus on:
- Normalization across devices and sources because we can assess versus gold standard
- Continuous validation and anomaly detection because we know the signals we are looking at
Sleep is a very complex long-duration signal, so reliability comes from enormous datasets with gold standard data (in our case more than 250 billion data points for example). Without a ‘ground truth’, it would be guesswork and that is not good enough when we are talking about peoples health.
VMblog: From an infrastructure standpoint, what changes are necessary for AI systems to transition from experimentation to truly scalable, production-ready deployments?
Lawlor: You need to move from batch thinking to continuous systems.
That means:
- Real-time or near-real-time data ingestion
- Scalable pipelines that handle variability, not just volume
- Feedback loops that allow models to be developed and adapted over time
- Strong observability and fully deterministic models—knowing when and why systems fail
Most importantly, infrastructure has to be built for persistence. Production AI isn’t a one-time model—it’s a living system that evolves with the data, and it can’t evolve successfully alone, it must have the guardrails built and monitored by human experts in the field.
VMblog: Looking ahead, how do you see continuous health and behavioral data influencing the next generation of enterprise AI applications?
Lawlor: Firstly, you cannot fix health without fixing sleep – so it starts with health and wellness companies recognizing that. It will fundamentally shift AI from reactive to predictive—and ultimately to preventative.
Continuous behavioral data like sleep becomes a leading indicator. It tells you what’s about to happen, not just what has happened. For example, the team at Stanford recently published a major study in Nature which showed that they could predict more than 130 separate chronic diseases (think neurodegenerative, cardiovascular and metabolic diseases) – all the key costs to the health system and the key drivers of poor quality of life. We believe this because of earlier work we did successfully predicting hospitalizations from data collected during sleep.
For enterprises, this opens up entirely new categories:
- Personalized health and performance optimization
- Early risk detection
- Adaptive user experiences based on real-world behavior
Sleep sits at the center of this. It’s the most underutilized, high-signal dataset in human health—and when integrated properly, it becomes a powerful input into next-generation AI systems.
##





