Industry executives and experts share their predictions for 2026. Read them in this 18th annual VMblog.com series exclusive.
By Val Bercovici, Chief AI Officer and Shimon Ben-David, Chief Technology Officer at WEKA
As we approach 2026, the AI industry is undergoing a fundamental economic transformation from venture-subsidized growth to sustainable, profit-driven business models. The market is maturing beyond breakthrough announcements toward usage-based pricing, infrastructure sovereignty, and measurable enterprise ROI.
The biggest shifts of 2026 won’t be model capabilities or benchmarks. Instead, the focus will be the end of free AI as providers transition to outcome-based pricing, sovereign nations building independent infrastructure, and enterprises demanding cost predictability over impressive demos.
We’re approaching critical inflection points in AI economics: how workloads are priced and consumed, how reinforcement learning enables sustainable business models, and where competitive advantages emerge between customer-facing applications versus infrastructure efficiency.
Read more below to learn how the year 2026 has the potential to reshape the AI industry through the end of subsidized pricing, the rise of token economics, sovereign infrastructure development, and the shift toward measurable enterprise ROI.
Val Bercovici, Chief AI Officer
AI surge pricing arrives as the subsidized agent era ends
Real market rates for AI inference will emerge by the end of 2026 as the industry’s subsidy era collapses under trillion-dollar CapEx realities, finite energy constraints, and upside-down AI token unit economics. Leading AI model providers can no longer afford to underwrite usage as customer adoption of AI agent applications balloons. Multiple token classes with dynamic pricing tiers will become standard, and users will experience market segmentation and maturation as “free tier” offerings evaporate, and cost-plus models reveal true inference economics.
In the latter half of 2026, AI agent pricing may shift from token-based to outcome-based
The cost of inference at scale will create pricing ripple effects throughout next year, catalyzing the rise of new pricing models. In the first half of the year, AI model providers will follow Anthropic and Z.ai’s lead, starting to charge for memory persistence (aka long-term KV cache storage). In the second half of the year, pricing may shift from token-based to outcome-based economics, with businesses measuring AI value by results, such as completed software projects or investment returns, rather than units consumed. The marginal utility of individual tokens matters far less than the holistic value of AI-driven outcomes. This transition will reshape business models across the AI industry, though it remains risky for providers in the near term.
By 2026, a trailblazing cohort of Sovereign AI nations will start to track Gross Token Production (GTP) alongside GDP as a key metric that defines global economic output
Gross Token Production (GTP) will emerge as a critical macroeconomic indicator alongside GDP, potentially even tracked through a GTP Index that measures a nation’s AI sovereignty and competitive standing. As AI inference becomes embedded in every sector of the economy, a nation’s ability to generate high-quality tokens at scale will directly correlate with productivity, innovation velocity, and economic output, making GTP a defining metric of leadership in the AI era.
Cache haves and have-nots will be the 2026 AI infrastructure efficiency divide. AI providers will split into two camps-those achieving 95%+ KV cache hit rates will serve 10x more tokens, with inference at 10x lower latency and cost than competitors, fundamentally restructuring the competitive landscape
Having the smartest model means nothing if you can’t serve it profitably. Infrastructure efficiency, not model capabilities, determines which AI providers build sustainable inference businesses. Competitors with cutting-edge models but wasteful infrastructure either raise prices (losing customers) or subsidize usage (burning cash). Efficiency becomes the moat that can’t be bought, only built.
Reinforcement and Continual Learning will become the new AI tracks in the race to AGI
The convergence of training and inference workflows through reinforcement and continual learning (RL & CL) will become the dominant path toward more capable AI systems. This unified approach blends the best practices of both disciplines into iterative loops that advance model capabilities. Leading labs including OpenAI, Anthropic, and DeepSeek are pouring resources into RL as the most promising scaling law toward Artificial General Intelligence. Labs still focused on traditional training approaches fall behind, as continual RL-trained models demonstrate superior reasoning and problem-solving capabilities. The infrastructure, talent, and investment will follow as CL + RL becomes the defining characteristic of frontier AI development.
Shimon Ben-David, Chief Technology Officer
By the end of 2026, 40% of enterprise AI projects initiated in 2024-2025 will be defunded for failing to demonstrate ROI, forcing a brutal reckoning in AI budgets
Organizations that experimented with AI in the last two years will begin to demand measurable returns in 2026. The “let’s try AI” era ends as investors and CFOs require clear ROI timelines and sharpen their focus on margin impact. Labs pitching vague productivity gains will lose to competitors demonstrating specific cost reductions or revenue increases. AI budgets shift from innovation teams to line-of-business owners who need to justify every dollar. Projects that can’t prove ROI within a defined timeframe will be at risk, forcing organizations to build economic models into their deployment plans.
Customer-facing AI becomes the primary deployment model
Internal AI experimentation gives way to production AI deployments that generate revenue. By 2026, the majority of enterprise AI projects will be customer facing: personal AI assistants, agent swarms managing support workflows, and AI-powered product features that drive measurable returns. ROI pressure forces organizations to prioritize consumer-facing AI over internal efficiency projects. For organizations that deploy production AI applications successfully, the business shift will be dramatic: AI will move from back-office cost centers to front-line profit centers.
Sovereign AI clouds will unlock a new era of public-sector innovation at scale
The rise of sovereign AI clouds established over the last year will reach operational maturity by 2026, transforming public sector innovation capabilities. Countries building sovereign data centers will start deploying AI for population-scale services across healthcare, transportation, and economic applications. Data sovereignty stops being just about security and is now laying the foundation for the development of domestic AI capabilities and technology ecosystems with real-world impact. Governments that invest in AI infrastructure now will own the economic advantages that define the next era of global leadership.
##
ABOUT THE AUTHORS
Val Bercovici, Chief AI Officer

Valentin (Val) Bercovici is WEKA’s Chief AI Officer. He has extensive experience in the data infrastructure industry, having previously been the Chief Technology Officer at NetApp/SolidFire, where he drove innovation in cloud storage and data management solutions. Val co-authored the Windows Shadowcopy snapshots and has made significant contributions to the storage standards community.
As Co-chair of the Storage Networking Industry Association’s (SNIA) Solid State Storage Initiative, Val helped to establish the first NAND Flash SSD storage standards. Additionally, Val served as the Chair of the SNIA Cloud Storage Initiative (CSI), where he led the development of the international S3 standard CDMI (ISO 17826). He was also a founding member of the Kubernetes Cloud Native Computing Foundation’s Governing Board, helping to shape the global direction of container orchestration.
Val holds patents in AI agent smart contracts, streaming data integrity, and augmented reality (AR) for data center maintenance. His work continues to push the boundaries of what’s possible at the intersection of AI, cloud, and emerging technologies.
Shimon Ben-David, Chief Technology Officer
Shimon is WEKA’s Chief Technology Officer (CTO), responsible for engaging with the company’s customers and partners to track emerging trends and bring actionable feedback and insights to its Engineering and Product Management teams. He previously held leadership positions in the company’s customer success and sales engineering functions.
Shimon brings deep enterprise IT domain expertise to WEKA, having previously run IT and support services at Primary Data, XtremIO (acquired by EMC), IBM, and XIV Storage, where he met WEKA’s founding team. He studied Computer Science and Philosophy at Ramat Gan University.





