Skip to content

Finance

The Vendor Grading Its Own Homework

Eric Avery·Founder & CEOJuly 8, 20263 min read

Late last year, Microsoft published the results of a study showing its 365 Copilot product delivering, in the best case scenario modeled, up to three hundred and fifty three percent return on investment for small and mid-sized businesses. The study was conducted by Forrester Consulting, based on interviews with real customers, and by the standards of vendor-commissioned research it was a serious piece of work. I want to be clear about that up front, because the point I’m about to make isn’t that the study was dishonest. It’s that the study, however careful, was never going to be independent, and the difference between those two things is the entire reason a category like ours needs to exist.

Think about the structure for a second. A vendor pays a research firm to study its own product’s impact. The research firm interviews customers the vendor selected or approved for participation. The resulting number gets published under the vendor’s name, on the vendor’s website, as part of the vendor’s sales motion. None of that means the number is wrong. It means the number was produced by the party with the most to gain from a favorable result, using a sample the same party had a hand in choosing. In finance, we have a word for that kind of evidence. We call it interested, not because anyone involved is lying, but because the incentive structure makes it structurally difficult to trust as a neutral read.

Microsoft isn’t unusual here. Nearly every major AI vendor now publishes some version of an ROI study about its own product, and the numbers are consistently, almost suspiciously, flattering. That’s not a knock on any single company. It’s what you’d expect from a market where the loudest voice on the question of value is also the party selling it. If you only ever heard the seller grade the sale, you’d think every purchase in history had been a good one.

This is the gap independent measurement is built to fill, and it’s why we’ve never built Nymbral to measure a single vendor’s product in isolation, let alone to take a check from the vendor whose numbers we’re producing. Our methodology doesn’t care whether a tool is Microsoft’s, Anthropic’s, Salesforce’s, or a startup nobody’s heard of yet. It applies the same research-anchored, conservative approach to all of them, funded by the customer who wants the real answer, not the vendor who wants a specific one. That’s a small structural choice with a large consequence: a board can actually trust the number, because nobody in the room built it to sell something.