TL;DR
- A Gartner survey of 1,303 senior executives published September 1, 2026 found only 22% of organizations have successfully scaled AI across multiple business units. - Despite this, 85% of functional leaders plan to increase AI spending in 2026, after spending an average of 12% of their functional budgets on AI in 2025. - MIT's Project NANDA found 95% of enterprise generative AI pilots deliver zero measurable profit-and-loss impact. McKinsey found only 39% of organizations report any enterprise-wide EBIT impact from AI. - The failure mode is almost never the model. MIT, McKinsey, BCG, and IBM all point to the same root causes: no defined success metrics before launch, data foundations that were never built for AI, and workflow integration that never happened. - Organizations classified as high performers by Gartner — those that track AI ROI continuously and treat initiatives as a portfolio — report positive returns on 81% of their AI projects. Low performers cannot state the return for 29% of theirs. - The 5% of organizations that actually capture AI value at scale share a specific pattern: one specific pain point, workflow redesign, and data infrastructure fixed before deployment.
Every week there is a new announcement about an AI deployment somewhere. A bank cuts processing time by 40%. A retailer personalizes recommendations at scale. A healthcare system reduces administrative burden. The announcements are real. What is less visible is everything else: the 78% of pilots that never reached production, the 95% of generative AI deployments that produced no measurable financial return, and the organizations spending 12% of their budgets on technology they cannot evaluate.
On September 1, 2026, Gartner published a survey of 1,303 senior executives across organizations with at least $50 million in annual revenue. The finding that should have made more headlines: only 22% of organizations have successfully scaled AI across multiple business units or adopted an AI-first approach. Nearly four in five organizations with real AI budgets have not managed to move from pilot to meaningful enterprise deployment.
The same survey found 85% of functional leaders plan to increase AI spending in 2026. Eleven percent of organizations do not even know how much their function spent on AI in 2025.
This is not a story about technology failing. Every reputable study on enterprise AI failure says the same thing: the models mostly work. The organizations often do not.
What the Numbers Actually Show
Multiple research institutions ran large-scale surveys on enterprise AI performance in 2025 and 2026. They used different methodologies and different definitions of success. They arrived at numbers that are not identical but point in the same direction.
MIT's Project NANDA reviewed over 300 publicly disclosed AI initiatives, conducted interviews with 52 organizations, and surveyed 153 senior leaders. Their finding on generative AI specifically: 95% of pilots delivered no measurable profit-and-loss impact. Only 5% of integrated systems created significant value.
McKinsey surveyed nearly 2,000 respondents across 105 countries. Only 39% of organizations reported any enterprise-wide EBIT impact from AI. The share qualifying as genuine high performers — organizations attributing more than 5% of EBIT to AI — was around 5% to 6%.
BCG analyzed 1,250-plus global firms and found that 60% were not achieving material AI value at all. The distribution, they noted, is not a bell curve. It is winner-take-most, with a small cohort pulling away from the field.
IBM's CEO study, covering 2,000 CEOs across 33 countries, found that only 25% of AI initiatives delivered expected ROI, and only 16% had scaled enterprise-wide.
A separate study from Prosigns covering 1,200 enterprise leaders found that 82% of enterprise AI initiatives never reach production, up from 73% in their 2024 baseline. Teradata's Wakefield Research survey of 1,000 senior technology leaders found that 63% report no more than a small or emerging positive return on agentic AI investments to date, despite 90% planning to increase investment over the next 12 months.
The numbers differ because the definitions differ. But the consistent finding across all of them is that fewer than 10% of organizations are capturing most of the AI value.
Why the Gap Exists
The reflexive explanation — that models are not yet good enough — is wrong, and the research is explicit about it. MIT concluded that the divide between the few organizations capturing value and the many that are not is determined by approach, not model quality. BCG's high-performing cohort was not using better models. It was using models differently.
The failure modes the data actually points to are organizational, not technical.
No defined success metrics before launch. In 73% of failed AI projects, there was no agreed definition of what success looked like before the project began. Sixty-one percent were approved with a projected ROI figure that was never measured again after go-live. You cannot retroactively evaluate a return you did not set up a way to measure.
Data foundations not built for AI. Gartner has repeatedly identified data infrastructure as the single biggest recurring blocker to AI at scale. Teradata's research found that 77% of executives report that 20% or less of their enterprise data is sufficiently described and contextualized for AI agents to use. The models are capable. The data pipelines they depend on are not.
Workflow integration that never happened. MIT found the highest-performing 5% of pilots were co-designed with the frontline teams who would use them daily. The failing 95% were largely IT-led or innovation lab-led, with limited end-user input until launch. A tool that sits beside an existing workflow rather than inside it does not change how work gets done. It generates a demo. Then it gets shelved.
FOMO-driven investment. IBM's CEO study names a failure mode that does not appear in the other research but is instantly recognizable: organizations launching AI initiatives because competitors are launching AI initiatives, not because they have identified a specific problem AI is the right tool to solve. The result is a portfolio of proofs-of-concept that demonstrate capability without establishing business value.
The Measurement Gap Is Getting Worse, Not Better
Gartner's survey includes a finding that should concentrate executive attention: 11% of organizations do not know how much their function spent on AI in 2025. This is not a small organization problem. The survey covered companies with at least $50 million in annual revenue.
A separate KPMG survey of organizations with at least $1 billion in annual revenue found that only 26% had full, real-time visibility into AI operating costs. That means 74% of large enterprises are spending significant money on AI without being able to see clearly what it costs to run.
Gartner's description of high performers is instructive precisely because it implies what everyone else is not doing. High performers continuously track the ROI of AI initiatives, manage them as a portfolio rather than as individual projects, and regularly reallocate resources or discontinue initiatives that do not perform. They treat AI spend the way a sophisticated investor treats a portfolio: actively managed, with clear criteria for when to hold, when to increase, and when to cut.
The result: high performers report positive returns from 81% of their AI initiatives. Low performers cannot state the return on 29% of their initiatives.
The difference is not which AI model they use. It is whether they built the measurement and governance infrastructure before they started spending.
The 5% Pattern
BCG identified specific structural traits shared by the organizations actually achieving AI value at scale. Three of them appear consistently across the research.
They identify one specific business pain point rather than launching broad pilot portfolios. MIT's data suggests that the multi-pilot, multi-function approach produces 95% failure rates. The focus that actually works is a level of specificity that most organizations find politically difficult — because committing to one use case means not pursuing five others simultaneously, and that requires executive conviction rather than optionality.
They redesign workflows around AI rather than layering AI onto existing processes. The Teradata research identifies what they call the "action bridge" problem: AI output currently lives outside the systems where consequential work actually happens. When intelligence is surfaced inside the tool where someone is already working, action follows. When it lives in a separate dashboard, it usually does not.
They fix data infrastructure before deployment, not after. MIT found that off-the-shelf generic tools plateaued fast because they could not retain context or adapt to a specific team's workflow over time. Narrower tools built around one workflow, with governed, well-described data underneath them, kept improving. The data foundation is not a cleanup task. It is a precondition.
Prosigns' research found that organizations with systems in production for 12 months or more report median annualized ROI of 287%. Organizations still in pilot-only mode report median ROI of negative 34%.
The gap between those two groups is not talent, compute budget, or model selection. It is whether the organization reached production.
What This Means for AI Investment Decisions in 2026
The pattern in the data suggests a specific reframing for any organization currently evaluating or scaling AI.
The question is not whether to invest in AI. The 85% planning to increase spending in 2026 will be right that AI creates value. The question is whether the preconditions for capturing that value are in place before the spending accelerates.
Those preconditions are not exciting. They do not generate press releases. Defining what success looks like before a pilot starts, identifying the specific workflow the AI needs to change and redesigning it, fixing the data quality and governance issues that will block production deployment — these are boring operational tasks, not capability announcements.
But the data is consistent across four major research institutions, multiple methodologies, and thousands of organizations: the 5% that are pulling away from the field did those boring things first.
Gartner's Trough of Disillusionment framing for generative AI in 2026 fits the historical pattern of the Hype Cycle. After peak hype comes a period where the gap between expectation and delivered value becomes impossible to ignore, investment gets more selective, and the organizations that built real infrastructure during the hype phase begin pulling away from those that did not.
The 22% figure suggests that trough is already here for enterprise AI. Most organizations are spending more, not less, on technology they cannot evaluate, in pursuit of returns they cannot measure, in workflows they have not changed.
The organizations on the right side of that gap did not get there by waiting for a better model.



