Opens in a new tab
vmblog logo 2024 wht (updated)

Unreal data can solve real problems

Share: 

David Marshall | Published: September 30, 2022

The metaphor “data is the new oil” originally coined by Clive Humby, British mathematician and data scientist, in 2006 has grown to become more apt as the decades turn. Data is indeed the fuel that powers today’s AI engines, from recommendation systems in e-commerce, to language processing. Without data the entire AI/ML ecosystem would be crippled. Although we are generating vast caches of data through our myriad interactions with technology, not all of it is usable or relevant. Collecting data and then labelling, annotating and preparing it so it can be used to train AI is an arduous and costly affair. Organizations are now facing severe scarcity of the right kind of data needed to train AI. The cost of acquiring quality data maybe too high in many circumstances, or they may not exist or be recorded at all for a variety of outlier events. In many scenarios, there may be inherent bias in the data, as it may not portray an adequately exhaustive array of possibilities, or it may not reflect the richness of all the diverse features needed for training the AI.

How Good is Synthetic Data?

Data generated artificially by algorithms to represent and augment real-world data, is called synthetic data. While synthetic datasets, themselves, do not represent real world objects, events, or people, it can be of high quality and mimic real attributes of the entities they represent to near perfection. These techniques and algorithms are so designed that they can replicate all the aspects of real time data and even simulate data of outlier scenarios, which are hard to find in the real world. This artificial data generation capability can help AI detect and predict patterns and events that previously were deemed impossible. Synthetically generated data contains all the statistical and scientific properties of real events and can also simulate real world relationships between data.

An abundance of computing power, and newer deep learning models have made it possible to generate massive quantities of synthetic data modelled on real records with less costs and time and this can be a game-changer for many industries. This practice is growing and according to a report by Gartner Inc, 60% of the data used for the development of AI and analytics projects will be synthetically generated. Research from institutions like MIT show that in most cases, using synthetic data, data scientists can get almost the same level of accuracy that they would from real-world data.

Circumstances where synthetic data can bring value

There are instances where synthetic data can play a crucial role. For example, in cases like credit-card fraud detection, AI models look for specific patterns of suspicious behavior, extracted from troves of real- world data. But collecting adequate data on rare types of fraud or various mechanisms of fraud is an arduous task considering that rarity of their occurrence. In these situations, using synthetic data to generate a collection of fraud instances can be immensely beneficial to bolster and augment the training data.

In scenarios where the data requirement is high, for example when training complex computer vision models for autonomous driving, collecting, and storing these huge volumes of image data can be prohibitively time- and cost-intensive. In addition, replicating scenarios of crashes and accidents is difficult, expensive, and hazardous too. In these cases, synthesizing photo-realistic images to train the computer vision models makes so much more sense. In situations, where the data landscape is constantly changing, say in a factory where materials and defects to be tracked are constantly changing, capturing images frequently may not be possible and using synthetic images simulating the anomalies can make the training process faster.

Gaining access to the necessary data might be difficult due to privacy laws concerning PII and PHI. For example, in the highly regulated medical industry, collecting certain types of sensitive patient data is not permitted. Collecting and storing data may also expose organizations to potential data breaches and consequent penalties and lawsuits. Generating synthetic data based on real patient records like generating synthetic images of rare afflictions can make AI-predictions more accurate. This is so much more effective than privacy enhancing technologies like data masking and anonymization.

There might even be scenarios where historical data does not capture all the patterns and trends. In risk assessment, for example there may not be adequate data on several adverse scenarios, including black swan events that probably has not even occurred in human history. For predicting such scenarios, we would need synthetically generated data based on intuition and knowledge that can mimic real-world data.

In product testing and development, for a new product there might not be significant customer interaction and buying behavior data. Generating artificial datasets for training AI models based on similar products in a related category can prove useful. Synthetic data can unlock new pathways for innovation.

AI models can develop biases with regards to sex, race, personal affiliations due to biases inherent in the humans designing the models. Generating and testing these against synthetic data sets can also uncover the biases of existing AI models and balance out the biased datasets. What a fine way to help organizations develop Responsible AI! Potential applications can be in human resources, sales and marketing and corporate governance.

For consumer-facing industries like retail, where storing customer data is not possible, or customer survey data proves to be inadequate, synthetic data sets can make a difference. AI-powered marketing and recommendation systems work on synthetic datasets and generate marketing communication that can be tailored and targeted to customers.

Impact brought forth by synthetic data

Small organizations like start-ups, unlike the incumbent giants, do not have access to a rich trove of data. Synthetically generated datasets can enable them to develop AI solutions faster and with greater accuracy. By democratizing this aspect of data accessibility, boundless innovations will be unlocked as more and more players will foray into the AI space and solve previously unimagined problems. It will give rise to a new ecosystem of AI-driven organizations who will now have the capability to compete with deeply entrenched players thus creating a more level playing field.

This will also transform data infrastructure, not requiring the landscape to be driven by fast and efficient data pipelines, thus reducing costs. This can also speed up model development as time for designing and calibrating real world experiments to collect data is no longer needed.

Developing superior capabilities to generate synthetic data responsibly is one of the fundamental requirements to develop related technologies like digital twins which are digital models of real-world physical factories, buildings and systems. It will also form the foundation of the metaverse.

Due to rising demands, the technology to generate synthetic data will also become more mature and sophisticated, with a rich ecosystem of tools and platforms to generate synthetic data. This can potentially herald a new age of AI, creating the impetus to scale and embed AI into the fabric of business, culture, and life itself!

##

ABOUT THE AUTHOR

Bali D.R., Head of AI and Automation 

Bali D.R. 

Balakrishna, popularly known as Bali D.R., is the Head of AI and Automation at Infosys where he drives both internal automation for Infosys and provides independent automation services leveraging products for clients. Bali has been with Infosys for more than 25 years and has played sales, program management and delivery roles across different geographies and industry verticals.