Opens in a new tab
vmblog logo 2024 wht (updated)

ODPi 2018 Predictions: Data Governance, Data Science & the Hadoop Landscape

Share: 

David Marshall | Published: December 21, 2017

Industry executives and experts share their predictions for 2018.  Read them in this 10th annual VMblog.com series exclusive.

Contributed by John Mertic, director of program management for ODPi and Open Mainframe Project at the Linux Foundation

2018 Predictions for Data Governance, Data Science & the Hadoop Landscape

ODPi is a nonprofit organization committed to simplification & standardization of the big data ecosystem. As the Director of Program Management, John Mertic has a unique perspective on the state of the landscape along with the below observations for where the ecosystem is heading in years to come.

Rise of Data Governance → As industry demand for the integrity, availability and security of enterprise data increases, the governance of this valuable insight has become crucial. Over time, as the landscape has worked to remove the manual nature of governance within AI & machine learning practices, it’s become hard to scale governance efforts when blocks to its efficacy exist. In the coming year, I think the trends around governance will center less around how can we invest in this framework and more around how best to automate these efforts at a reasonable cost and high scale.

Hadoop Ecosystem Evolution → In March of 2016, Merv Adrian of Gartner declared that Hadoop “stack expansion has ground to a halt” – but I predict that 2018 will mark a huge uptick of maturation within the Hadoop stack. With generic commodity Hadoop platforms like Apache BigTop being widely embraced by a number of cloud and emerging big data vendors, there is less competition at this level – leaving room for the development of and innovation around boutique services. This trend would add value at the top of the stack, as all offerings can be based on the same three or four lineages, and allow for fewer individual Hadoop stack roll outs – leading to the building of differentiated, vertically-targeted apps for uni-purpose, full-stack solutions.

Data Science Boom → While efforts around Data Science are ever-growing, there’s an enormous amount of maturity still needed in this space – and open source will be at the center of it. Here, I expect to see huge growth in usage, as well as ecosystem around both R and Python, more enterprise investment in and maturity around open source tools and packages that support building user-facing applications, and better integration of this booming field into the infrastructure layer around Spark and Hadoop. Along the lines of open source solutions, I think we can also look to see them making it easier to build and deploy shiny new apps for internal organization use.

##

About the Author

John Mertic 

John Mertic is director of program management for ODPi and Open Mainframe Project at the Linux Foundation. John comes from a PHP and open source background. Previously, he was director of business development software alliances at Bitnami, a developer, evangelist, and partnership leader at SugarCRM, board member at OW2, president of OpenSocial, and a frequent conference speaker around the world. As an avid writer, John has published articles on IBM Developerworks, Apple Developer Connection, and PHP Architect and authored The Definitive Guide to SugarCRM: Better Business Applications and Building on SugarCRM.