Skip to content

Research

Self-Reported Versus Independently Verified: Two Studies, One Year, Two Very Different Realities

Eric Avery·Founder & CEOJuly 22, 20264 min read

Two credible research organizations asked roughly the same question in 2025 and came back with answers that don’t just differ, they contradict each other. Deloitte surveyed thousands of business and technology leaders for its State of AI in the Enterprise research and found that the majority of companies reported their advanced AI initiatives meeting or exceeding ROI expectations, with a meaningful share reporting returns above thirty percent. Around the same time, MIT’s NANDA initiative published research finding that ninety five percent of enterprise generative AI pilots showed no measurable impact on the bottom line. Same year, same broad market, opposite conclusions.

The instinct is to assume one of these studies is simply wrong, and I understand the instinct because it would be a more comfortable explanation. But both are methodologically serious pieces of research, conducted by organizations with real reputational stakes in getting it right. The more useful explanation isn’t that one study is bad. It’s that they measured two fundamentally different things while using the same word to describe them both.

Deloitte’s numbers are self-reported. A leader inside an organization was asked whether their AI initiative met expectations, and the answer that came back reflects that leader’s own read of their own program, filtered through whatever expectations they’d set going in and whatever incentive they might have to describe their own initiative favorably. MIT’s numbers came from something closer to an outside audit: researchers examining public deployments and structured interviews, looking for actual evidence of measurable financial impact rather than taking anyone’s word for it. One study asked people how they felt about their AI investment. The other went looking for proof.

That distinction, self-reported against independently verified, is not a small methodological footnote. It’s the entire reason a category like AI ROI measurement needs to exist in the first place. If self-reported sentiment were reliable enough to run a board conversation on, nobody would need an independent layer sitting between a company’s belief about its AI program and the evidence for that belief. The gap between what Deloitte’s respondents said and what MIT’s researchers found is, in miniature, the exact gap Nymbral was built to close: the space between how a program feels from the inside and what an outside party, using a defensible methodology, can actually prove.

If you’re a finance leader trying to decide which of these studies applies to your own AI portfolio, the honest answer is that you don’t know yet, and neither does anyone reporting up to you, until somebody runs the independent version of the analysis. That’s an uncomfortable thing to sit with heading into a budget cycle. It’s also, if you get the measurement right, a genuinely solvable problem.