By John Durcan, Chief Technologist, IDA Ireland
Artificial Intelligence has dominated the news this year, making many companies take notice of emerging trends in this rapidly advancing field. An important trend that will happen in 2024 and beyond is how the world will start paying more attention to data governance in AI, particularly for the many U.S. companies doing business in the large European market. Top of mind is the EU’s AI Act, which goes into effect in 2024. Driving a goal of assuring AI is ethical and trustworthy, this first-of-its-kind legislation will ensure that consumers trust in AI products and that companies have clear direction. Therefore, U.S. companies selling AI-related products in Europe will seek more knowledge and data governance skills next year to assist in user acceptance, technology development, market expansion and protection of profits.
What is data governance in AI?
It’s a framework that covers how AI-related data is obtained, managed, used and secured by an organization. Adopting such a strategy lets entities trust the integrity of their AI and machine learning models by ensuring that their data originates from reliable sources. For U.S.-based AI companies, a key part is to ensure a safe and usable product for consumers. Potential risks for AI products include:
1. Hallucinations
A byproduct of how models source information is that untrue data can be presented as otherwise, sometimes called “hallucinations” in technology. Passing on false truths or perpetuating incorrect data that might contain harmful lies is an issue with any data set including Large Language Models (LLMs). Therefore, methods need to be developed for assuring the accuracy of LLM information including visible source references, with the latter currently the focus of extensive research.
2. Hidden algorithms or bias
Some AI methodologies put the procured data into a black box, from which answers emerge. However, this approach has no transparency of the underlying algorithm, thus complicating understanding and amendment. This is yet another area that companies should address to improve their products and lessen their vulnerabilities. Given that some AI models pull their data from the internet, companies need to be cautious of a built-in bias toward any one group.
3. Copyrights
LLMs and other AI tools that pull information from the internet could run afoul of copywrite sources, be it books, art or anything else that can be legally protected. Several related issues are now under rigorous study such as for how long such protected data can be passed on and under what circumstances. If companies train a model, then delete the suspect sources, will they have legal issues in the future, such as under the AI Act?
4. Privacy
Companies doing business in Europe will already be aware of legislated privacy issues, most notably in the General Data Protection Regulation (GDPR), which governs how personal data can be used, processed and stored. Experience with GDPR is helpful in understanding proper data protection methods but LLMs and advanced AI represent a larger realm, with the strong likelihood of legal challenges and many questions regarding what data has been amassed, its sources, where it’s stored, for how long and related issues.
5. Collaborating to better address data governance
Current AI regulations and the doubtlessly future ones that will be developed demand new approaches to product development and business practices while also straining the resources of all the companies that need to add data governance to their to-do lists. According to Empower Executive Director Denise Manton, (based in Maynooth University), “One of the big challenges within companies around data governance is insufficient resources assigned to these areas, producing a skills gap.”
Empower is a collaborative research program whose premise “is to co-create best practices in data governance,” she says. Co-funded by Science Foundation Ireland (SFI) and a number of industry partners, it will likely be joined by similar programs around the world to address the rising need for data governance in AI. This is another probable future trend since drawing on shared resources between government, industry, research institutions and academia is a more effective way to achieve progress in making a powerful technology like AI more trustworthy and ethical. Among the many companies involved in Empower are Meta, Siemens, Huawei, Genesys and Analog Devices.
Manton describes the coming era of AI data governance as best managed via a multi-disciplinary process such as that of Empower. “It starts with data sharing but moves on to creating new data marketplaces and ecosystems, then to privacy and preserving technologies, to creating sandboxes where companies can test out their data governance policies. It includes the legal side and regulatory side and focuses on what’s becoming even more important: the ethical side and the people-centered design.”
6. Teaming to improve data governance
Programs like Empower are a model for the future but it is already on its way to creating helpful breakthroughs that will increase data protection internationally and promote ethical data usage. As part of Empower, SFI’s research center for chronic and rare neurological diseases is collaborating with the Irish operations of pharmaceutical leader Novartis and IQVIA, a U.S.-based life sciences data analytics leader, on two programs. One is developing a prototype learning health system and the other explores how to develop a culture of trust that promotes the safe use of health data.
According to Manton, AI data governance isn’t just a rulebook or source of limitations but also presents a revenue opportunity. Vast amounts of data have been amassed over the years “but it’s still locked in,” she says. “Having really good data governance practices allows companies to share their data safely without giving away all their secrets, but they can also start using the data to make new revenue streams.”
These opportunities can include new products but she sees an even bigger possible profit source: “Service innovation,” notes Manton. “The service sector is really growing.”
##
ABOUT THE AUTHOR
John Durcan is Chief Technologist at IDA Ireland, the national investment development agency for Ireland. IDA Ireland partners with companies worldwide to provide financial assistance, on-the-ground support and advice to help them establish and transform their operations in Ireland. Durcan’s current key focus areas are artificial intelligence (AI), quantum computing, cyber security and the semiconductor sector. www.idaireland.com.





