Industry executives and experts share their predictions for 2024. Read them in this 16th annual VMblog.com series exclusive.
Meeting the Data Performance Demands Associated with AI
By Ido Bukspan, CEO, Pliops
There are many obstacles for data centers to conquer when it comes to AI, including performance and scalability challenges. As we enter 2024, I anticipate several trends to emerge related to data and AI – from Generative AI to Large Language Models (LLMs) to AI-generated media content. Read on for more.
- As inference workloads and costs grow significantly, there will be a rise in startups developing lower-cost performance silicon chips while hyperscalers pursue their custom silicon for inference and startups. With the proliferation of AI applications across industries such as healthcare, autonomous vehicles, finance, and more, the volume of inference workloads has surged. These workloads involve tasks like image recognition, natural language processing (NLP), and recommendation systems, all of which require real-time processing. Meeting these demands efficiently and economically is crucial. As inference workloads continue to witness a substantial increase in demand and associated costs, there is a growing imperative for the development and utilization of dedicated custom silicon solutions.
- The race to build larger, smarter, comprehensive LLM models will continue, resulting in large organizations spending a fortune in the range of $100 million to $0.5B. Staying ahead of the curve in AI-powered language capabilities can lead to market dominance. The race to build larger language models has prompted big enterprises to commit substantial financial resources, approaching or exceeding billion-dollar budgets. These investments are motivated by the potential for competitive advantage, market expansion, improved customer experiences, data monetization, research and development, infrastructure, and the anticipation of significant ROI.
- There will be increased adoption of key-value stores from traditional applications to Generative AI applications for caching in the Prompt and Token phase of LLMs. The increased adoption of key-value stores represents a pivotal shift in the technological landscape, extending from traditional applications to the emerging domain of Generative AI applications, particularly in the context of caching during the Prompt and Token phases of LLMs. Often with hundreds of billions of parameters, LLMs demand substantial computing resources for both training and inference. To mitigate latency and reduce computational overhead, organizations are turning to key-value stores as a means to efficiently cache frequently used prompts, tokens, and intermediate model states.
- AI-generated media content will become the norm, from basic website designs to generating songs and making Hollywood movies. AI-generated media content isn’t limited to a single medium; it spans a wide range of creative outputs. These include not only static elements like website designs, logos, and graphics but also dynamic and interactive content such as videos, animations, virtual reality experiences, and interactive storytelling. AI-powered tools can significantly expedite the creative process.
- There will be a rapid adoption of Retrieval Augmented Generation (RAG) in enterprises using Vector Databases to complement the LLM foundation model. To derive higher business value from large language foundation models, we will witness a large number of organizations leveraging vector databases for building Retrieval augmented Generation systems with enterprise internal data sources. RAG is an increasingly significant area in the field of NLP and GenAI to provide enriched responses/answers to customer queries with domain-specific information in chatbots & conversational systems. Microsoft, Google AlloyDB, Oracle, and Amazon RDS services like Aurora, and Pinecone, weaviate provide vector database functionality to serve as a platform for organizations to build RAG systems.
To keep up with the AI demands on data, 2024 will be as important as ever to properly shape data infrastructure optimization and workload acceleration so that data centers have all of the tools needed to succeed in the AI era.
##
ABOUT THE AUTHOR
Ido Bukspan is the CEO of Pliops. Prior to joining Pliops, Ido was the senior vice president of the Chip Design Group at NVIDIA and one of the leaders at Mellanox before it was acquired by NVIDIA for nearly $7 billion. He holds a BSEE and an MBA from Tel Aviv University.





