How Do You Measure the ROI of AI Agents?

Baselines, deltas, and the three tests of a defensible number

James Proctor
James Proctor
Subscribe

Updated:

Published:

You measure AI agent ROI by comparing instrumented process outcomes against a pre-deployment baseline, converting the delta into financial terms through unit economics, and attributing it honestly through phased rollout.

That is the entire method, and its rarity is the scandal. Most published agent ROI is constructed differently: vendor-reported activity multiplied by assumed rates, with no baseline, no attribution, and no way for anyone, including its authors, to defend it.

What Are the Three Tests a Defensible ROI Must Pass?

A defensible ROI number must pass three tests before it leaves the building:

  • The baseline test: is the improvement measured against a documented pre-agent state? A claim without a before state is not valid, cannot be validated, and is therefore unfalsifiable.
  • The attribution test: can you separate the agent’s contribution from everything else that changed? Perfection is not required; a phased rollout, deploying to part of the operation while the remainder briefly serves as comparison, produces evidence that is directionally reliable and honest about its limits.
  • The audit test: would the number survive your own internal auditors applying the scrutiny they apply to any other capitalized claim? Not a hostile regulator. Your own people, on an ordinary Tuesday.

Most agent ROI circulating today fails all three, which leads to the position I hold and most programs resist: the most honest AI ROI figure many enterprises could publish right now is unknown, and unknown is a respectable starting point. It is fake precision that should embarrass people.

A number that fails the baseline test is not a result. It is a wish.

What Belongs in the ROI Construction?

Only realized value belongs in the ROI construction: deltas that reach the P&L or a stated strategic outcome.

Cycle time should be converted through its cost and revenue consequences. Error and rework reduction should be converted through its cost of quality. Capacity freed counts only when a reallocation decision exists, a discipline covered in the companion post on why hours saved misleads.

On the cost side, count everything: platform, inference, integration, and the human oversight the agents require.

Two further disciplines keep the construction honest. Measure after stabilization, not launch week, because early readings capture novelty and tuning, not steady state. And keep a two-sided ledger: count the regressions, the new error classes, and the added oversight alongside the gains, because a ledger with no debit column is an advertisement.

An ROI built from full costs and realized value only will be smaller than the vendor deck’s number, and it will be yours, which matters more.

What Does Defensible ROI Look Like in Practice?

Consider a national specialty pharmacy that deployed agents in benefits verification and prior authorization — processes where turnaround time is both cost and revenue.

Before deployment, it baselined ninety days of history: verification turnaround, first-pass yield, rework rates, and cost per completed verification. It rolled agents into half its verification teams first, leaving the rest as comparison for one quarter.

The measured result — a substantial turnaround reduction and first-pass improvement in agented teams against flat comparison teams — converted through unit economics into a defensible annual figure that finance ratified.

The deliberately unglamorous headline was the point: the number was roughly a third of what the platform vendor’s calculator had promised, and unlike the vendor’s number, it survived every question the CFO asked. Funding for expansion followed the smaller, truer number.

Building baseline-and-attribution measurement into agent programs is part of Inteq’s Agentic AI consulting services. This post accompanies the full white paper on measuring agentic AI.

Related White Paper

Read the foundational white paper on why process outcomes, decision quality, and unit economics replace deployment counts as the measures that make agentic AI success defensible.