Zentor (formerly Moclaw) · Client sample, shown with permission
How to measure AI agent ROI through verified work, value capture, full costs, and payback. A framework for recurring work, not demos.
By 8:00 on Monday, a market brief is waiting in the shared folder. It used to take an analyst most of the morning. Now it arrives before the team logs in, neatly formatted and ready for review.
The speed is easy to see. The harder question is what the business actually gains from it.
Some weeks, review takes ten minutes. A changed website, weak source, or missed detail can turn the next one into an hour.
That swing in review time is not unusual. A 2025 survey of 1,854 executives across Europe and the Middle East found that only 10% of organizations using agentic AI were already seeing significant ROI. Half expected returns within three years, and another third expected them within three to five. Although the survey is regional, the tension is familiar.
At this point, I stop looking at speed and check three things: what passed review, what the workflow cost, and where the returned time went.
Key Takeaways
Count work that passes review, not prompts, drafts, or attempted runs.
Treat returned capacity as potential value until the business puts it to use.
Include setup, tools, review, failed-run recovery, and upkeep.
Track cost per verified completion alongside ROI and payback.
Scale only if the return holds when use, quality, or cost changes.
What AI Agent ROI Really Measures
AI agent ROI compares the financial value a workflow creates with the full cost of building and running it.
ROI % = (Realized benefits − Total costs) ÷ Total costs × 100
Realized benefits are gains the business can trace and price. They may come from lower costs, more accepted work, higher gross profit, fewer errors, faster service, or smaller losses.
The full cost goes far beyond the model bill and includes setup, software, paid data, human review, repairs, monitoring, and upkeep.
The formula also works for agentic AI. But as an agent takes on more of the task, retries, review, and recovery can rise too. Those costs belong in the business case.
Follow the Value Through the Workflow
The work an agent could handle is rarely the value the business can claim.
I trace the value through six stages:
Eligible work is everything the agent could reasonably handle. Adoption reduces that pool to the work people actually send through it.
Agent-covered work reaches the point where the agent’s job ends. Verified work also passes review without a major correction.
Even verified work is not financial value on its own. The benefit appears only when the returned time or added output changes cost, capacity, service, risk, or profit. Count the part the business can price and keep the rest as supporting evidence.
Subtract the full workflow cost to find net benefit. Divide that net benefit by total cost to calculate ROI.
Take a workflow with 1,000 eligible tasks. Staff send 80% of them to the agent, which completes 75% from start to finish. Of those results, 85% pass review.
1,000 × 80% × 75% × 85% = 510 verified completions
The workflow produced 510 verified results. It did not create value from all 1,000 possible tasks. The financial case must still ask how much value the business captures from those 510 completions.
Set the Baseline and the Finish Line
Start with the old process.
Pick one unit you can count, such as an approved research brief, a routed support case, or a checked invoice. Then observe a normal period of work.
During that period, record task volume, hands-on time, review and correction time, loaded labor cost, and normal rework. Add cycle time when speed affects service or revenue.
Use the same finish line for both the manual and agent paths.
A research brief, for example, may need current facts, approved sources, a set format, and a reviewer’s approval. If one required source is missing, a polished draft has not reached the finish line.
Using the same finish line keeps partial agent output from receiving credit for the whole job.
Set the baseline before rollout. Use a comparison group where practical. Also track where reclaimed time goes, because time creates value only when it moves into useful work.
Keep the evidence behind the numbers. Record what ran, what finished, what passed review, and what needed repair. Without that record, too much of the ROI estimate depends on memory.
Count the Costs That Appear After the Run
The model bill appears on an invoice. The rest of the cost builds up in smaller pieces.
Initial work may include process mapping, data cleanup, integrations, tests, access rules, and staff training. Monthly costs can include the platform, model use, APIs, paid sources, browser sessions, storage, and support.
Together, these items form the AI agent’s total cost of ownership.
Then comes the human tail.
People still review outputs, handle exceptions, repair failed work, and update the workflow when a source, rule, or tool changes. These small tasks add up across the month and often go uncounted.
Discarding a weak draft may take only a few minutes. Recovery costs rise after the agent changes a customer record, sends a message, or updates another system. Then someone may need to find the mistake, reverse it, check the repair, and explain what happened.
One production analysis calls the verification and rework around agentic systems an “agency tax.” Recovery costs rise once an error reaches a live system.
Run cost can vary too. A 2026 study of coding agents found that repeated runs on the same task could differ by as much as 30 times in token use. Higher use did not reliably improve accuracy.
The study covered coding tasks, so the 30x figure is not a benchmark for every business agent. It still shows why one average run can hide a wide cost range.
For day-to-day control, track:
Cost per verified completion = Total cost of the agent-handled workflow for the period ÷ Verified agent completions in that period
This measures the cost of usable output rather than the cost of each prompt or attempted run.
That metric covers only the agent-handled path. To price the entire operation, use a separate measure:
Cost per accepted outcome = Total cost of the full workflow ÷ All accepted outcomes
Use cost per verified completion for the agent-handled path and cost per accepted outcome for the whole operation, including manual work. Keeping them separate prevents unrelated costs from entering the same calculation.
A Complete AI Agent ROI Calculation
Michael leads a small research team. The example below is illustrative, not a reported customer result.
His team prepares 20 market briefs each month. One brief takes four hours, and the fully loaded labor rate is $60 per hour. The current process therefore carries $4,800 in monthly labor cost.
The team sets up an agent for recurring research and first drafts. During the pilot, staff send 80% of briefs through the agent. The agent covers the full routed task, and 80% of its briefs pass review.
Each attempt still needs 30 minutes of checking. When a brief fails review, the team spends one extra hour diagnosing the problem and returning the brief to the normal manual process.
This calculation counts only that extra hour because the normal manual completion time is already part of the baseline. Platform and usage costs total $300 per month, and setup costs $9,000.
Work That Reaches the Business
Staff route 16 of the 20 briefs through the agent:
20 × 80% = 16 briefs
At an 80% review-pass rate, the workflow produces an average of 12.8 verified briefs per month:
16 × 80% = 12.8 verified briefs
The decimal is a monthly planning average. It does not mean the team completes a fraction of one brief.
At four hours and $60 per hour, those verified briefs represent $3,072 in gross monthly capacity value. This is before the value-capture rate and workflow costs:
12.8 × 4 hours × $60 = $3,072
Realized benefit = Gross capacity value × Value-capture rate
Review, Recovery, and Operating Cost
Reviewing all 16 attempts costs $480 per month:
16 × 0.5 hours × $60 = $480
An average month has 3.2 failed briefs. Recovering them adds $192:
3.2 × 1 hour × $60 = $192
After the $300 platform and usage cost, the recurring workflow costs $972 per month:
$480 + $192 + $300 = $972
That puts the recurring cost per verified completion at about $76:
$972 ÷ 12.8 = about $76
This figure includes monthly review, recovery, platform, and usage costs. It excludes the one-time setup cost.
First-Year Return
First-year measure
Calculation and result
Annual benefit at 100% value capture
$3,072 × 12 = $36,864
Total first-year cost
$9,000 + ($972 × 12) = $20,664
Net first-year benefit
$36,864 − $20,664 = $16,200
First-year ROI
$16,200 ÷ $20,664 × 100 = about 78%
Monthly recurring net benefit
$3,072 − $972 = $2,100
Setup payback after stable performance
$9,000 ÷ $2,100 = about 4.3 months
The ROI and payback figures assume 100% value capture. Michael’s team must find useful work for all the capacity returned. The payback period starts only after the workflow reaches the stable monthly performance shown above.
If the team captures 75% of the returned capacity, annual benefit falls to $27,648 and first-year ROI to about 34%. Capture only half, and the annual benefit no longer covers first-year cost; ROI drops to about −11%.
Nothing changed in agent quality or operating cost. Michael’s team simply used less of the capacity it received.
I do not count an hour as savings until I can see where it went.
What Real Results Can Safely Show
Metrovacesaoffers a useful company-reported example because its figures move beyond agent activity.
Its customer system processed 4,577 portal leads and automated 84.5% of engagements. The company reported around 3,000 confirmed visit requests, a 38% cut in customer-service time, and 56% of conversations handled outside normal business hours.
Together, the figures show substantial use, customer action, and less service work.
The case also links the confirmed visits to €70 million in potential business volume. That number should remain labeled as potential. A full ROI calculation would still need setup and running costs, human oversight, sales conversion, linked revenue, and gross margin.
FletcherTech’s three-month trial tells a different story. Its internal AI system delivered 31,778 answers to 222 employees and was credited with returning more than 2,500 hours.
That supports a strong claim about adoption and returned capacity. The public case does not show how much of that time became financial value.
When I read a customer story, I look for the outcome behind the activity. The missing numbers tell me what the case cannot yet prove.
Read the Monthly Pattern Before You Scale
A single ROI figure can hide where the workflow gains or loses value.
Keep a short scorecard for recurring work:
Adoption: Whether people send eligible work through the agent
Verified completion: How often results reach the finish line
Review time: How much human effort remains
Recovery cost: How expensive failed work becomes
Cost per verified completion: The price of usable agent output
Value-capture rate: How much returned capacity reaches the business
Look for a pattern across several runs. One poor brief may mean little. Repeated failures from the same source or handoff show where the workflow needs work.
Before adding volume, test a weaker month. Lower the pass rate, increase review time, or raise tool costs. The assumption that damages ROI most should guide the next round of testing.
Pair early signs, such as adoption and review-pass rate, with later proof, such as avoided cost and ROI. Give one person ownership of the review and hold it on a fixed schedule. A dashboard has little value when nobody acts on it.
Scale, Revise, Narrow, or Stop
By the end of a pilot, the team should know what it will do next.
Scale Carefully
Add volume or authority only after use is steady, quality holds, and the numbers still work under conservative assumptions. Change one part at a time so the team can see what shifts the economics.
Revise the Weak Stage
A poor source, slow handoff, or long review step may weaken an otherwise sound workflow. Name the fix and the result it should produce before funding the change.
Narrow the Scope
When one part works well, and another consumes most of the benefit, reduce the agent’s role. It may gather and organize evidence while a person keeps the final judgment. A narrower role may still produce a better return.
Stop When the Value Disappears
Stop when people avoid the workflow, review time keeps growing, or a simpler method produces the same accepted result for less. Ending the pilot is a valid decision when the agent no longer adds enough value to justify its cost and risk.
How MoClaw Helps Make Agent Work Measurable
Recurring work is easier to measure when its inputs, outputs, and run history stay together.
MoClaw gives the agent a private cloud computer with a real filesystem, browser, shell, and persistent state. Files, installed tools, and browser sessions can remain available between sessions.
Its scheduling tools can run recurring jobs in the cloud. The Schedules dashboard lets users inspect jobs and review run history.
In Michael’s case, the same workspace could preserve the approved source list, older briefs, working files, and finished drafts. A scheduled run can then gather the next inputs and save a draft for review. Its run history shows whether the task finished, failed, or needed repair.
His team could connect that record to the monthly scorecard: attempted briefs, verified completions, review time, recovery work, and cost per verified completion.
MoClaw helps preserve the operational evidence behind recurring work: files, run history, and outputs. Michael’s team still defines what counts as an accepted brief, records review and recovery work, assigns value to the result, and decides whether to scale.
Frequently Asked Questions
How Long Should an AI Agent Pilot Run Before ROI Is Measured?
Use enough normal work to observe common inputs, exceptions, review time, and cost changes. A weekly process may need several weeks. A high-volume queue may reveal a stable pattern sooner.
Can Negative First-Year ROI Still Support an Investment?
Sometimes. A workflow may have a longer payback period or create benefits that grow with volume. State those benefits, measure them, and compare the return with other uses of the budget.
When Is Fixed Automation Cheaper Than an AI Agent?
Fixed automation often fits stable rules, known inputs, and predictable outputs. An agent earns its added cost when the work needs judgment, changing context, or flexible use of tools.
Use the Numbers to Make the Next Decision
After several ordinary weeks, the pattern matters more than one impressive run. The team can see whether people use the workflow, whether its outputs hold up, and whether the value survives review and recovery costs.
From there, the choice becomes clearer: scale what works, repair the weak stage, narrow the agent’s role, or stop. The goal is not the largest possible workflow. It is one that produces reliable value at a cost the business can justify.