The single most revealing number in enterprise AI today is not a model benchmark or a cloud revenue figure. It is this: 83% of organizations that operate GPUs run them at 50% utilization or less, while nearly half plan to evaluate AI-specialized clouds — a category almost none of them currently use — within the next twelve months. That is the compute gap, and it is widening. A new wave of VentureBeat Pulse Research, surveying 107 qualified enterprise respondents (organizations with more than 100 employees) in Q2 2026, documents a market in which infrastructure investment is accelerating well ahead of the instrumentation needed to govern it. Enterprises are buying more compute faster than they can account for what they already own, and the decisions they are about to make will reshape the competitive landscape of AI infrastructure.
How the Research Was Conducted
This survey was fielded as part of the ongoing VentureBeat Pulse Research series, focused specifically on enterprise AI infrastructure, compute, and inference economics. Responses were filtered to organizations with more than 100 employees (n=107), drawn from a single Q2 2026 (June) wave. Because this is one wave rather than a pooled multi-month sample, the results read cross-sectionally and do not infer month-over-month trends. Several questions allowed multiple selections, so shares can sum to more than 100%. The sample skews toward the mid-market (36% at 101–250 employees, 27% at 251–1,000) and toward earlier-stage adopters, making it best read as the view from organizations actively building out AI infrastructure rather than from the largest hyperscale operators. Respondents span managers (38%), individual contributors (28%), VPs and directors (19%), and the C-suite (13%), with 45% holding final decision-making authority and another 30% serving as recommenders or influencers for AI solutions. The sample is large enough to read directionally but should be treated as a directional signal rather than a precise measurement.
Finding 1: Ambition Far Outpaces Production Maturity
Only one in five enterprises runs AI in production at scale. The deployment journey remains heavily front-loaded: 38% are still experimenting with proofs of concept, 37% have some workloads in production but not across the organization, and just 21% describe themselves as operating at scale. Another 4% are not yet running AI workloads at all. This foundational reality shapes every subsequent finding: the infrastructure decisions documented in this report are being made largely by organizations still early in their deployment journey, whose compute footprint — and whose costs — are about to grow significantly. The evaluation and switching intentions discussed below represent the leading edge of that build-out, not the settled preferences of mature operators.
Finding 2: The Current Stack Runs on Hyperscalers and Model APIs
When asked which providers and platforms they currently use to run AI, enterprises point to a familiar set of incumbents. Google Cloud leads at 48%, followed by Microsoft Azure at 29%, AWS at 22%, and Oracle Cloud at 22%. On the model side, Google’s Gemini models are used by 41%, OpenAI by 40%, and Anthropic by 12%. The specialized “neocloud” GPU providers that dominate AI-infrastructure headlines — CoreWeave, Lambda, Crusoe, Nebius, Together, Fireworks, and their peers — each register at or below 2% among these enterprises. Only 6% run their own on-premises or co-located GPU clusters, and 4% operate a custom open-source self-managed stack. For now, enterprises are running AI on the providers they already buy from, with an average of 2.1 providers per respondent. This makes the evaluation intentions in the next finding all the more striking.
Finding 3: The Next Dollar Is Aimed at Infrastructure They Do Not Yet Run
Asked where they plan to evaluate AI infrastructure over the next twelve months, enterprises point decisively away from their current stack. The single most-cited planned evaluation area is AI-specialized clouds (CoreWeave, Lambda, Crusoe, Nebius) at 45% — the very category that barely registers in current usage. Nearly a third (32%) plan to evaluate non-NVIDIA accelerators such as AWS Trainium, Google TPU, AMD Instinct, Intel Gaudi, or in-house ASICs, while 28% plan to evaluate next-generation NVIDIA Blackwell (GB300) GPUs. Even decentralized or distributed compute networks (16%) and sovereign or region-specific compute (11%) draw meaningful interest. The direction-of-travel question confirms this pattern: every infrastructure approach is net-expanding, but specialized AI clouds carry the highest net momentum at +24, edging out even the hyperscalers at +22. This continues a trend observed in an earlier survey wave, where the most-cited planned change was moving workloads to specialized AI clouds. Two waves, two differently worded questions, one consistent picture: the type of cloud enterprises are most eager to assess is the type they have barely begun to use.
Finding 4: A Switching Wave Is Building — Most Will Move Within the Year
For a category as foundational as compute infrastructure, the intended movement is remarkable. A clear majority of enterprises (64%) plan to switch or add an infrastructure provider within twelve months, and 38% intend to do so within the next quarter alone — this was tied for the most common answer. Only 36% have no plans to change. However, the providers drawing the most switching consideration are again the incumbents: Microsoft Azure and Google Cloud at 33% each, OpenAI at 30%, and Gemini at 22%. This suggests that much of the near-term movement is reshuffling among the majors and consolidating spend rather than defecting to new entrants. The neocloud interest documented in Finding 3 is a twelve-month evaluation thesis; the switching expected in the next quarter is primarily incumbents trading share among themselves.
Finding 5: Nobody Buys on Token Price — Integration and TCO Decide
When selecting an AI infrastructure provider, enterprises do not optimize for the metric vendors compete on most aggressively. The top factor cited is integration with the existing cloud and data stack at 41%, followed by total cost of ownership at 35%. Performance — latency and throughput — matters to 24%, while security and compliance, autoscaling for spiky workloads, and GPU access and availability each draw 19%. The headline metric — cost per million tokens — is the deciding factor for just 8%, placing it dead last. This pattern is coherent but exposes a tension: buyers say TCO matters most, yet, as the next finding shows, most cannot yet measure it rigorously. The stated priority and the measured capability are out of step.
Finding 6: Expensive GPUs Sit Idle Most of the Time
The compute already in place runs cold. Among enterprises that operate GPUs, 83% report utilization at 50% or less. Specifically, 37% run at 26–50% utilization, 34% at 10–25%, and 15% at under 10%. Only 12% exceed 50% utilization, and a further 8% do not measure utilization at all. Another 7% consume AI via API and operate no GPUs of their own. Idle accelerators are expensive accelerators, and this is the clearest single measure of the compute gap: enterprises are planning to buy more GPUs and specialized compute while the capacity they already own sits substantially unused. The efficiency headroom in the current fleet is large and largely unmeasured.
Finding 7: Spending Fast, Measuring Slowly — Cost Visibility Lags
Fewer than half of enterprises (44%) rigorously track the cost and return of their AI compute. A further 39% track it only partially, 20% cannot quantify it yet, and 6% say it is not a priority. This measurement gap is consequential given that total cost of ownership was the second-ranked buying criterion: enterprises are choosing providers on an economic basis they mostly cannot yet measure. Satisfaction with current infrastructure averages 4.0 on a five-point scale, but the softest scores land on value for money (3.9) and ease of implementation (3.8) — the dimensions hardest to judge without rigorous measurement. Enterprises are spending quickly and accounting slowly.
Finding 8: The Next Bottleneck — Memory Bandwidth — Is Barely on the Radar
As inference scales, the binding constraint shifts from GPU compute to memory bandwidth, specifically KV-cache capacity. Yet this frontier is not yet a priority for most enterprises. Asked which approach they would rely on to address this emerging constraint, enterprises scatter: Dell (PowerScale / Project Lightning) leads at 31%, NVIDIA (Dynamo / ICMSP) follows at 16%, and the rest fragments across storage vendors such as Hammerspace (10%) and DDN (9%), open-source KV-cache tooling, model-level efficiency techniques, and a long tail of other options. Most tellingly, roughly one in five enterprises (18%) either do not recognize the constraint (9%) or have not begun to address it (8%). For a shift that will fundamentally reshape inference cost and architecture, this is an early and unsettled market — and consistent with the measurement gap, one where many enterprises simply do not yet have a view. It is the next chapter of the compute gap, arriving before most have closed the current one.
What This Means for Enterprise Decision-Makers
The compute gap is not a capacity problem that more hardware will solve on its own. It is, first, a problem of seeing what the hardware already costs. Enterprises that treat infrastructure investment primarily as a procurement exercise — selecting providers on headline token price or brand familiarity — risk compounding an already expensive idleness problem. The organizations that will navigate this transition best are those that invest first in measurement: rigorous tracking of utilization, unit economics, and total cost of ownership across the full stack. The re-platforming wave is coming — specialized clouds, alternative accelerators, and new inference architectures are all drawing serious evaluation intent — but switching without visibility is not optimization; it is just spending in a different direction. The open question for the next twelve months is whether enterprises build that visibility before the re-platforming arrives, or buy the next layer of infrastructure as blind to its economics as the last.
Based on survey responses from 107 qualified enterprise respondents (100+ employees), drawn from a single Q2 2026 (June) wave. The sample is self-selected, skews mid-market, and leans toward earlier-stage adopters. Respondents include managers, individual contributors, VPs/directors, and the C-suite, with buyer-credible purchasing authority, across Technology/Software, Healthcare/Life Sciences, Financial Services, Retail/E-commerce, and other industries.