Skip to content

Research

The 95% Problem: What MIT’s GenAI Divide Study Actually Means for Your P&L

Eric Avery·Founder & CEOJune 8, 20264 min read

This past summer I sat in on a renewal call I wasn’t technically part of. A friend who runs finance at a mid-sized logistics company asked me to listen in while her team tried to get a straight answer out of an AI vendor about what the tool had actually delivered over the previous year. Forty minutes into the call, the vendor’s rep was still talking about seats activated and prompts submitted. Nobody in that room, on either side of the table, could tell you what the company had gotten back for the money it spent. I left the call thinking about how often that scene must be repeating itself, in conference rooms all over the country, week after week.

A few weeks later, a group of researchers at MIT’s NANDA initiative put a number on it. Their report, published in July and picked up everywhere from Fortune to the trade press, looked at hundreds of enterprise generative AI deployments and reached a finding that should have dominated the news cycle for a month: despite an estimated thirty five to forty billion dollars in enterprise spending on generative AI, roughly ninety five percent of pilot programs showed no measurable impact on the bottom line. A small handful of organizations were pulling real, provable value out of their AI investments. The overwhelming majority were not.

What stopped me wasn’t the failure rate on its own. New technology cycles are full of failed pilots, and anyone who has spent time around venture-backed companies has watched worse odds play out in other categories. What stopped me was the explanation underneath it. The researchers didn’t conclude that the models themselves were incapable of producing value. Their read was closer to the opposite: the tools were live, usage was real in a lot of these organizations, and the value may well have been there in some form. What was missing was any real mechanism for knowing whether it had shown up, where it had shown up, or how much of it there was.

That distinction matters enormously for how a board or an investor should read this moment. A ninety five percent failure rate sounds like a verdict on the technology. Read closely, it’s a verdict on measurement. Companies spent freely, adopted quickly, and then found themselves unable to answer the one question every finance team eventually asks about every investment: what did we get for this. Without an answer, a pilot that quietly created value and a pilot that quietly wasted it look identical on a spreadsheet. Both get canceled, or both get renewed, for reasons that have nothing to do with what actually happened.

This is the conversation we have with finance and procurement teams every week, and it’s rarely about whether AI works. Most of the people we talk to have already seen enough to believe it can. What they don’t have is a defensible way to say how much it’s working, for which vendor, in which part of the business, measured against what it actually cost to get there. That gap between belief and proof is the entire market Nymbral was built to close, and MIT just handed the rest of the industry the data to prove it exists.

The organizations that separate themselves over the next two years won’t be the ones that spent the most on AI, or even the ones with the best models. They’ll be the ones that can walk into a board meeting, or a due diligence room, and answer the question with a number they can defend under questioning. That’s a solvable problem. It just requires treating AI spend the way every other category of capital allocation gets treated: measured, attributed, and held to the same standard of evidence as everything else on the balance sheet.