Measure hours returned, cost per completed outcome, and conversion rate on delegated work against the same work done by a human. Record all three for two weeks before deploying anything, because without a baseline you will have an opinion rather than a result. Ignore output metrics — emails sent, posts published, leads touched — which rise automatically and tell you nothing about whether the agent worked.
The most common way to get AI agent ROI wrong is to measure the thing that is easiest to count. Emails sent, posts published, leads touched — these go up the moment you deploy anything, which makes them feel like proof and makes them worthless as evidence.
Record the baseline before you deploy
This is the step almost everyone skips, and skipping it is unrecoverable: once agents are running you can no longer measure what the previous month looked like.
Twenty minutes, written down, before anything is connected:
- How many leads got a second touch last month?
- How many hours went into that, honestly?
- What was the reply rate on those touches?
- How many reached a booked call?
- What did the whole motion cost — tools plus a fair estimate of the time?
The three numbers worth tracking
1. Hours returned per week
The most immediately felt return, and the easiest to verify. Track your own time for a fortnight before and after. If the figure is under three hours a week, something is wrong with the setup rather than the concept — usually the agent was given a description of your business instead of live access to it.
2. Cost per completed outcome
A qualified lead, a booked meeting, a resolved ticket, a published piece. This is the only number comparable across different pricing models *and* against the cost of a person doing the same work.
| Founder doing it | Agents doing it | |
|---|---|---|
| Leads given a second touch | ~60 | 200 |
| Hours per month | ~18 | ~2 (review) |
| Reply rate | 12% | 7% |
| Replies | 7 | 14 |
| Cost per reply | ~$180 of founder time | ~$5 in actions |
Note the reply rate *fell* and the result improved. That is the shape of most honest agent deployments, and it is why rate-based metrics alone mislead.
3. Conversion on delegated work vs your own
Compare the agent's output against the same work when you did it. If it converts at half your rate but happens four times as often, that is a win. If it converts at a tenth, the agent lacks context and no amount of volume fixes it.
The metrics to ignore
- Emails sent, posts published, tasks completed. These measure activity, and activity always rises.
- Time saved, as estimated by the vendor. Estimate your own or do not count it.
- Anything the platform reports without showing its working. If you cannot reconstruct the number, do not put it in a board deck.
Making the cost side countable
Cost per outcome only works if the cost side is legible. This is one argument for action-based pricing over seat-based: when one credit is one action and every action appears in a log, dividing spend by outcomes is arithmetic rather than estimation. Operater prices this way — 150 actions free each month, then $39 for 400, $199 for 2,000, $599 for 6,000, and $0.14 per action beyond the allowance — so the denominator is something you can count rather than something you have to model.
Under a per-seat model the same calculation requires you to guess how much of a subscription belongs to which outcome. More on the trade-offs in how AI agent pricing works.
When to conclude it is not working
Give it two weeks on a daily workflow, then compare against the baseline. Three outcomes:
- Clearly better — expand scope on this workflow before adding another agent.
- Roughly neutral — nearly always a context problem, not a capability one. The agent does not know enough about who you sell to.
- Worse — check that the workflow genuinely runs daily and that the agent has live tool access. It is almost always one of those two.
The honest version: if you cannot compute cost per completed outcome after a month of running, that is itself the finding. A platform that will not show you what it did is not one you can evaluate.
Key takeaways
- Record the baseline first. Twenty minutes before deployment is worth more than any dashboard after it.
- Cost per completed outcome is the only figure comparable across pricing models and against a human.
- Output metrics always improve. That is exactly why they are useless as evidence.
- A lower conversion rate at four times the volume is usually a win. Judge the product, not the rate.
- If you cannot compute cost per outcome after a month, the platform is not giving you enough visibility.