In an increasingly data-driven world, companies face significant hurdles in making their data usable and trustworthy. Eric Best, CEO of SoundCommerce, sheds light on these challenges and the critical role of data quality in the age of AI. In this exclusive VMblog interview, Best discusses the impact of siloed data, the importance of data meaning and understanding, and how SoundCommerce’s innovative Reactor pipeline is transforming the landscape of enterprise data management. As businesses rush to adopt AI technologies, Best’s insights offer valuable perspective on the evolving role of data engineers and the strategic shifts necessary for companies to build trusted AI systems.
VMblog: What are some of the biggest hurdles companies face to becoming more data-driven?
Eric Best: There are really two major hurdles companies face when it comes to making their data usable and trustworthy. The first issue is that data tends to be siloed. Your marketing team has access to marketing data, and your operations team has access to operations data, but these two aren’t unified in one place. Data might come from site tracking tags or browser cookies, or it might come from legacy system flat files or modern SaaS APIs. You can’t understand how profitable a marketing campaign offering free shipping actually was, until you have the variable shipping costs, data on returns, and customer acquisition costs all in one place.
Another issue is data meaning and understanding. Enterprise data is often incomplete, outdated, or inaccurate. Data transferred from one platform to another, for example, might not be formatted correctly and end up causing a problem. If you’re looking to make real-time decisions about your business but you only have access to data up to last week, it’s not timely enough. If you have customers who visit you in stores and also shop online, but you don’t have their profiles merged across all platforms, you have an incomplete picture of your customer. Performing analytics on low quality data means you can’t truly trust the output.
VMblog: With the emergence of AI, has data quality become even more critical?
Best: Absolutely. If the data being fed to an AI model isn’t high quality, how can the model come to accurate conclusions or provide up-to-date information?
One area where enterprises are not focusing enough as they adopt AI is the semantic layer of data. For models to understand and interpret data accurately, that data needs to be labeled and defined. Those definitions and labels need to be easily understood, and commonly defined by everyone who uses the data. Without these labels and definitions, AI models cannot interpret the intent of business users who are asking it questions conversationally. These semantic descriptions must reflect conversational language and company or industry jargon that is commonly used.
Snowflake, a partner of ours, recently launched its Cortex Analyst, an agentic AI tool that offers a conversational interface where businesses can talk to their data. Cortex translates text prompts into SQL, queries data, and provides answers. But Snowflake requires customers to provide semantic descriptions of their data during setup to enable this. Unfortunately, many companies are in no position to do so.
Most companies use data catalogs to define, label and map data to semantic concepts and entities, but these governance tools are often applied closer to the end of the data flow. Applying meaning to data after the fact, unfortunately, means that it’s most likely garbage in, and AI outputs cannot be trusted.
What we need is an automated tool that collects, unifies, defines, labels and maps AI-ready data from the start. And that’s exactly what we’re doing with our new low code intelligent pipeline, Reactor.
VMblog: Tell me more about Reactor, your intelligent data pipeline. What makes Reactor different, and what challenges does it solve for data professionals?
Best: Reactor offers a fast, efficient path to useful data models that are ready for generative AI, analytics and activation. It’s an intelligent “ETLT” (extract/transform/load/transform) pipeline that provides fully prepped and modeled data for faster time-to-value, while creating cost efficiencies to cut down the ever-growing costs of enterprise data management.
Reactor ingests, maps and models data to be hosted in modern data warehouses like Snowflake and Google Cloud BigQuery. Its data quality is natively accessible for LLM platforms like Snowflake Cortex and Google Gemini, offering a fast option to enable generative AI for businesses. Its pre-built data collectors can ingest data from nearly 100 SaaS and on-prem solutions and applications.
One great thing about Reactor is that it features a drag-and-drop user interface that allows you to connect data from many different sources, and create a semantic layer to define the data, with just a few mouse clicks. We’re inching closer to a world where everyone is going to need to be able to access and use data, regardless of technical ability, so we wanted to make it as user-friendly as possible.
VMblog: You recently announced a partnership with data activation and reverse ETL platform Census Embedded. How does this benefit users?
Best: Census Embedded provides a complete reverse ETL and data activation for SoundCommerce customers. This enables our customers to more efficiently transform, govern and activate data using Census embedded in SoundCommerce, while Reactor provides more and better data to Census for activation. This reduces data engineering time and cost, while providing faster time to engagement with new and loyal customers thanks to a greater variety of signals and attributes for campaign automation.
VMblog: How is this impacting the role of data engineers?
Best: This is taking a great deal of manual, repetitive work off the hands of data engineers. Data engineers no longer have to spend as much time manually processing data – they can rely on AI to handle it, while they focus on more strategic work. I believe that in the coming years, we’ll see many data engineers upskill and take on new roles thanks to AI. This means better outcomes for organizations who today are dedicating vast resources to data engineering.
VMblog: What are some other ways companies are pivoting their data and AI strategies to make these systems more trusted?
Best: There’s been a lot of hype around LLMs, but for many organizations, little ROI from these investments. That’s because they are extremely expensive. An AI model that includes the menu of every restaurant with a website, everything Shakespeare ever wrote, and every article ever written by a newspaper is extremely bloated for most use cases. That’s why we’re seeing many companies shift to SLMs, or small language models. SLMs are often more specific, and trained on a company’s proprietary data. This improves model accuracy and relevance, while minimizing computing and cloud storage costs. However, at the end of the day, data quality, meaning and understanding still reign supreme to ensure the output can be trusted.
##






