Aleksei Ogarkov
← // Writing
// WRITING

Measuring AI ROI: did it move the number, or the dashboard?

Two figures about AI’s return are circulating, and they cannot both be true. In Google Cloud’s 2025 ROI of AI report, 74% of executives said they were already seeing a return within the first year of generative AI. In MIT’s Project NANDA study, 95% of enterprise gen-AI pilots showed no measurable impact on the P&L, against an estimated $30-40 billion invested. The same leaders, in the same quarter, booking a triumph and a rounding error.

That gap is not survey noise. I have run programs on both sides of it, and the difference was never the model. One camp measures the dashboard; the other measures the number. Measuring AI ROI honestly starts with being able to tell which one you did.

95%
of enterprise gen-AI pilots showed no measurable P&L impact, against an estimated $30-40 billion invested — MIT Project NANDA ↗

Attribution is not incrementality, and ROI is an incrementality question

A dashboard tells you what the AI touched. ROI asks what it caused. Only the second belongs on a P&L.

Attribution assigns credit across the touchpoints it can observe on the path to an outcome. Incrementality asks the counterfactual: what would have happened anyway, without the intervention? Marketing fought this war a decade ago. Haus, an incrementality firm, reports that when multi-touch attribution credits retargeting with, say, 3.8x ROAS, a geo-holdout on the same campaign routinely shows 40-60% of those conversions would have happened regardless. As their Ike Armstrong puts it, “correlation-based tools will always systematically overfund the measurable and underfund the unmeasurable.”

Now transpose that to AI. An assistant is in the room when the deal closes, so the deal is credited to the assistant. A copilot is open in the editor when the feature ships, so the feature is credited to the copilot. Baseline demand, the salesperson’s own competence, and a dozen concurrent initiatives all get swept into the AI’s ledger, because nothing was ever netted out. Presence is attribution. It is not lift. This is why IDC’s Andrea Siviero argues that agentic AI is “breaking your ROI model”: when both the benefit and the cost of a system float, dividing one fixed number by another describes nothing.

The number that moved might be your own perception

Most “AI ROI” evidence is self-report, and self-report is biased upward with confidence.

The sharpest result of the last two years is not an ROI survey. It is METR’s 2025 randomized trial of experienced open-source developers working in their own repositories. Before the tasks, they expected AI to cut their completion time by 24%. Afterward, they believed it had sped them up by 20%. Measured against a randomized control, AI made them 19% slower. That is a roughly 40-point gap between what these experts felt and what the clock recorded, and it points the wrong way.

That is the engine under the 74% dashboard: surveyed satisfaction, estimated hours saved, adoption counts. Ask people and you are measuring enthusiasm; the output never enters the survey. And once a proxy becomes the target a program is funded on, Goodhart’s Law takes over: adoption gets manufactured, deflection gets booked on tickets that would have self-resolved, usage gets mandated. Ron Kohavi, after years running controlled experiments, put the bar plainly: getting numbers is easy; getting numbers you can trust is hard.

When you can’t randomize, build a control anyway

When you can’t cleanly randomize the individual, you manufacture the counterfactual from structure.

The honest number needs a design, not a dashboard. I have measured this way: a Next-Best-Action program I led at a top-5 pharma in the CIS returned +7% incremental Rx, read visited-versus-not-visited, indexed 100 → 107.

RECEIPT
Scope
Next-Best-Action program, top-5 CIS pharma
Baseline
indexed 100, visited-vs-not-visited
Role
AI lead & builder
Result
+7% incremental Rx (100 → 107)

You almost never get to randomize the individual, so you build the untreated arm out of structure instead:

DesignUse whenWatch for
Geo holdoutyou can split regionsspillover across borders
Ghost / PSA holdoutyou can withhold at the unit levelstatistical power and cost
Synthetic controlone treated unit, clean pre-periodpoor pre-period fit biases it
Stepped-wedgewithholding forever is unacceptableconfounding by time

Each one builds the comparison the dashboard skips. Google’s Meridian, built for exactly this, is described in its own docs as “designed for causal inference, not prediction.” The payoff is a number no attribution model would produce. eBay ran a geo holdout across 68 media markets and found its branded paid-search ads moved sales by an amount statistically indistinguishable from zero.

When your three numbers disagree, the holdout is the referee

Attribution, marketing-mix modeling, and a holdout will give you three different answers for the same program. That spread maps where your uncalibrated models are lying; averaging it away only hides them.

Mature teams triangulate, but not by splitting the difference. The order of trust is fixed:

  1. The experiment is ground truth.
  2. Its result becomes the prior that calibrates the mix model.
  3. Attribution stays for relative, day-to-day diagnostics and never carries the ROI line by itself.

When the three disagree, the holdout is the referee and the other two are witnesses you re-weight against it. Shortcuts also inflate in a predictable direction: across 663 advertising experiments, Gordon and colleagues found that methods skipping the randomized control turned real lifts of 29 / 18 / 5% into 83 / 58 / 24%, inflated every time and always upward.

// The objection

A holdout costs money and might still tell you nothing

Both objections are true. You answer them by sizing for power, pricing the alternative, and calibrating once instead of re-testing forever.

The sophisticated pushback is never “do we need a control.” It is “withholding costs real revenue, and the read may still be too noisy to use.” Lewis and Rao, across 25 field experiments, found the median confidence interval on advertising ROI is over 100 percentage points wide.

The way through is design. Ghost-ad and geo methods recover most of the precision at an order of magnitude less cost than naive holdouts. So “too expensive” usually means you chose the expensive design. If withholding forever is politically impossible, randomize the timing instead of the treatment and read the effect from a staged rollout. And price the alternative honestly: Uber turned off about $100 million of a $150 million ad budget and installs did not move; P&G cut $200 million of digital spend and sales held while reach rose. That money burned for years precisely because no one held anything out. The suppression cost of a test is the tuition. The waste it exposes is the refund. Once one clean experiment anchors the mix model, you stop re-testing every quarter and calibrate against that reference.

The one question that separates real AI ROI from a dashboard

AI can pay. But a number is only an ROI if a counterfactual stood behind it. My +7% earns its credibility from the untreated arm, the same reason eBay’s “zero” is more trustworthy than the 3.7x a Microsoft-sponsored IDC survey reported with no control at all.

So before you believe the multiple, ask one thing: what didn’t get the treatment? If the answer is “nothing,” you are reading the dashboard, not the number.

Wiring that standard of evidence into how a commercial team plans and spends is the operating-model work I do, and the discipline behind the Next-Best-Action program I led at a top-5 CIS pharma. If your AI portfolio is long on dashboards and short on defended numbers, start with the holdout.