
The MIT study behind the 95% AI failure stat is real, but it measured whether custom enterprise AI tools produced marked P&L impact within six months, not whether AI works. For Shopify merchants, the useful finding is where the returns actually appeared.
The most quoted piece of AI skepticism in business today was produced by a group that builds agent infrastructure, and its recommended fix is agent infrastructure.
In the last twelve months, I have had the same stat quoted at me from two directions by people who wanted opposite things. An agency founder used it to argue that AI is mostly theater and merchants should ignore it. Three weeks later, a vendor used it in a pitch deck to argue that merchants need a done-for-you agent build because 95% of the ones who try it themselves fail.
Both were citing the same document. Neither had read it.
The study exists, the number is quoted accurately in most cases, and the underlying finding is genuinely interesting. What almost nobody carries across is what was being counted. That gap matters because, at $500K to $2M in revenue, the difference between “AI does not work” and “custom AI tools bought from vendors did not produce measurable profit inside six months” is the difference between two completely different budget decisions.
The study measured one narrow thing: whether custom or vendor-sold generative AI tools reached production with a marked and sustained productivity or profit impact inside roughly a six month window. That is the definition of success the authors used, stated in their own research note, and it is the number that became “95% of AI projects fail.”
The document is called The GenAI Divide: State of AI in Business 2025, published in July 2025 by Project NANDA at MIT. The methodology is stated plainly in the full report and its stated methodology: 52 structured interviews, 153 survey responses gathered at four industry conferences, and a review of more than 300 publicly disclosed AI initiatives, all collected between January and June 2025. The cover page describes the work as preliminary findings.
The 95% comes from a funnel. For embedded or task-specific tools, 60% of organizations investigated them, 20% reached pilot, and 5% reached successful implementation. Subtract and you get the headline. The report attaches its own caution to that exhibit, noting the figures are directionally accurate based on individual interviews rather than official company reporting, and that success definitions varied between organizations.
Here is the part that gets lost completely. That funnel applies to custom and vendor-sold tools. The general-purpose LLM funnel in the same exhibit runs 80% investigated, 50% piloted, 40% implemented. People using ChatGPT and Copilot were not in the 95%. And the report’s own list of myths opens by rejecting the claim that AI will replace most jobs in the next few years, stating that researchers found limited layoffs concentrated in already-affected industries. If you have seen this study cited as evidence that AI failed to replace workers, that framing is the opposite of what the document says.
The report’s recommended fix is agentic AI built on interoperability protocols, and the group that wrote it builds interoperability protocols. NANDA stands for Networked Agents And Decentralized Architecture. The acknowledgments section states that the project builds on Anthropic’s Model Context Protocol and the Google and Linux Foundation A2A standard to create infrastructure for distributed agent intelligence at scale.
The conclusion then argues that organizations should stop investing in static tools, start partnering with vendors offering custom systems, and prepare for an Agentic Web built on NANDA, MCP, and A2A. That is not a hidden agenda. It is printed in the document. But it does change what the 95% is doing rhetorically. A finding that current tools fail, published by people whose stated work is the successor to current tools, is a sales argument wearing a lab coat.
This is the same filter I applied to the AI search panic, and it is worth reusing verbatim. When someone predicts the future of a technology and the prediction terminates in their own product, go read what the free documentation says instead. That was the whole method behind what Google actually published about AI search, and it holds here.
Others reached the same place through methodology rather than motive. Paul Roetzer of the Marketing AI Institute published a detailed breakdown of where the methodology falls apart, arguing the success definition ignored efficiency gains, cost reductions, churn improvements, and pipeline velocity, and that a zero-return finding resting on 52 interviews the authors themselves called directional does not carry the weight the headline gives it. His advice was blunt: do not put weight on this study.
None of that makes the report worthless. It makes it one input, produced by an interested party, from a small self-selected sample, over a window too short to catch anything but fast wins. Treat it accordingly. If you want the neutral version of where agent protocols are genuinely landing in commerce, Shopify’s UCP and Catalog infrastructure is the documentation to read, and it is free.
Three months later, a larger survey found that roughly three in four enterprises were already seeing positive returns from generative AI, which is close to the inverse of the MIT figure. Wharton Human-AI Research and GBK Collective published the third annual enterprise AI adoption report in October 2025, reporting 74% positive ROI, fewer than 5% negative, and 82% of leaders using generative AI at least weekly.
The sample is the interesting part. Knowledge at Wharton documented how that survey was constructed: a cross-sectional survey of more than 800 senior leaders at companies with over 1,000 employees and at least $50 million in revenue, now in its third consecutive year. That is roughly five times the respondent count of the MIT survey, with a repeated year-over-year design instead of a single snapshot.
Neither study is about you. Both surveyed enterprises far larger than any Shopify brand reading this. The reason to hold them side by side is that two credible groups asked adjacent questions six months apart and got answers 70 points apart, which tells you the measurement is doing most of the work. That is not a reason to dismiss both. It is a reason to define your own measurement before you spend anything.
The most useful number in the MIT report is not 95%, it is the budget allocation: roughly half of generative AI spend went to sales and marketing, while the documented cost savings clustered in back office functions nobody was funding. The report is candid about why. Sales and marketing metrics map cleanly to board-level KPIs, so the spend goes where the attribution is easy rather than where the return is largest.
I have watched the identical pattern play out in Shopify stores for years, and it is the mechanism underneath the premature complexity problem that stalls brands between $500K and $2M. The AI budget goes to the storefront, because the storefront is where the founder looks. An AI product description generator, an AI merchandising widget, an AI ad creative tool, an AI popup that personalizes the offer. Six subscriptions, all pointed at the most visible surface, none of them with a number attached.
Meanwhile the actual time sink sits in the back office and gets nothing. Support macros that have not been rewritten since 2023. Returns processed by hand in a spreadsheet before anyone touches Shopify admin. Supplier emails that a human retypes every Monday. Product data hygiene that quietly determines whether an agent can recommend your product at all. Those are unglamorous, they do not demo well, and per the report’s own findings they are where the payback periods were fastest.
The stage-aware version is simple. Under $50K per month, the highest-return AI work is product data and support templates, not a merchandising tool. From $50K to $500K per month, the first paid AI tool should attach to whichever workflow currently consumes the most staff hours, which for most brands is customer support, and Gorgias plus a well-configured macro set will usually beat a net-new AI platform. Above $500K per month, you can afford parallel bets, and the discipline shifts from picking one to making sure each has an owner and a number.
Two conclusions were available from the same finding: AI does not work, or AI budgets are pointed at the wrong half of the business. The second one is supported by the data and almost nobody quotes it.
The report found that employees at more than 90% of surveyed companies used personal AI tools for work while only 40% of those companies had bought an official subscription, which means the working deployments were the cheap ones nobody procured. The authors call this the shadow AI economy, and it sits awkwardly next to the 95% headline, because it describes AI succeeding at scale through the exact channel the study was not measuring.
One example in the report is worth the whole document. A corporate lawyer’s firm spent $50,000 on a specialized contract analysis tool, and she kept defaulting to ChatGPT because the purchased tool produced rigid output she could not iterate on. A $20 per month general-purpose tool outperformed a $50,000 bespoke system on the work that actually got done.
Translate that to a seven figure Shopify brand and it is uncomfortably familiar. The $29 per month subscription your ops lead uses to draft supplier emails, clean up CSV exports, and rewrite 200 product descriptions in an afternoon is almost certainly returning more than the $299 per month AI app sitting in your admin that nobody has opened since the trial. One of those has a login someone uses daily. The other has a renewal date.
This is also why the platform-level direction matters more than the app-level one. When Shopify shipped Shopify’s own agent tooling built on MCP, it moved store management into the general-purpose clients merchants already use rather than asking them to adopt another dashboard. That is the shadow AI pattern being formalized, and it is a better bet than most standalone AI apps precisely because it does not require anyone to change where they already work.
You are in the 95% if you cannot state, right now, the specific number each AI tool in your stack was bought to move. That is the whole test, and it takes about 45 minutes to run honestly.
Open your Shopify app billing page and your card statement. List every AI tool with its monthly cost. Next to each one, write the metric it was supposed to change and the value of that metric before you installed it. Most merchants doing this exercise for the first time can complete the metric column for two or three tools out of eight, which is the finding, not a failure. The tools with no metric are not producing zero value necessarily, but you have no way to know, and neither did 95% of the organizations in the study.
Then cut in one pass rather than gradually. A tool with no metric and no daily user gets cancelled this week, not reviewed next quarter. The 30 day gradual audit never happens, because there is always a launch. For the tools that survive, attach a number and a review date, and keep the number small. The most disciplined version I have seen is the three metrics that tell you a deployment is working, which for a support-side agent are first response time, resolution rate without human intervention, and satisfaction on AI-handled tickets against your human baseline.
Apply the 18 month filter to whatever you keep. Will this tool matter in 18 months, or is it a wrapper around a model capability that the platform will absorb? Product data quality will matter in 18 months. Support workflow automation will matter. A tool whose only function is calling an API you could call yourself for $20 will not. That filter would have saved most of the organizations in the MIT study a great deal of money, and it costs nothing to apply.
No, that is a compression of a narrower finding. The MIT NANDA report found that 95% of organizations in its sample saw no marked or sustained profit impact from custom or vendor-sold generative AI tools within roughly six months. General-purpose tools like ChatGPT and Copilot showed a 40% successful implementation rate in the same exhibit, and the report separately found that employees at more than 90% of surveyed companies used personal AI tools for work. The study also explicitly rejected the idea that AI is replacing most jobs. “95% of AI projects fail” is not what the document says.
It was published in July 2025 by Project NANDA at MIT, authored by Aditya Challapally, Chris Pease, Ramesh Raskar, and Pradyumna Chari. Treat it as one input rather than settled fact. The report labels itself preliminary findings, rests on 52 interviews and 153 conference-gathered survey responses, and states that its central figures are directionally accurate rather than drawn from company reporting. It also recommends agentic infrastructure built on protocols that Project NANDA itself develops, which is a relevant interest to know about before you weight the conclusion.
The report found that roughly half of generative AI budgets went to sales and marketing while the clearest documented cost savings came from back office functions that were underfunded. The authors attribute the imbalance to measurement convenience rather than actual value, since front office metrics map more easily to board-level reporting. For a Shopify merchant, the practical read is that AI spend aimed at the storefront is usually the visible choice, while support workflows, returns processing, and product data hygiene are where the faster payback tends to sit.
Less than most operators at that stage are currently spending, and concentrated in one workflow rather than spread across six. A store doing roughly $1M annually is usually better served by one general-purpose subscription that the team actually uses daily, plus one workflow-specific tool attached to whichever function consumes the most staff hours, typically customer support. That is often $50 to $400 per month total. The number matters less than the rule: every paid AI tool needs a named owner and a metric with a baseline recorded before installation.
Record the baseline before you install it, then check three numbers at 30 and 90 days. For a support tool, track first response time, the share of tickets resolved without a human, and satisfaction on AI-handled tickets compared to your human baseline. For a content or merchandising tool, track hours saved on the specific task and whether conversion or organic visibility moved on the pages it touched. If you cannot produce a before figure, you cannot evaluate the tool, and the honest move is to set the baseline now and re-evaluate in 90 days rather than renewing on instinct.