Qonnex
Evidencias5 min de lectura

How to measure AI automation returns honestly

Three tests separate real automation gains from wishful arithmetic, with a worked example: full costs, measured hours, and break-even in months.

Austin Vu · Managing Partner, Qonnex

Every automation proposal contains a savings number. Sixty hours a week. A three-month payback. The numbers are round, confident, and rarely survive contact with the finance team a year later.

The gap is usually not dishonesty. It is that returns get calculated once, before implementation, by the people most motivated to believe them, and then never measured again. Honest measurement differs in three specific ways, and each one has a test attached.

Three tests most claims fail

The counterfactual test: would this have improved anyway? Processes improve for reasons that have nothing to do with new software. A new hire finishes ramping. A supplier finally standardizes its invoice format. Volume dips after peak season. If your team was already planning a cleanup that would have saved eight hours a week, the automation does not get credit for those eight hours. The only defense is a dated, written baseline plus a list of every other change in flight during the measurement window.

The full-cost test: did you count everything you spent? The software subscription is usually the smallest line. Implementation and integration are real money. Training is paid time for every person in the room. And the largest hidden line is your own team's hours: the operations lead sitting in every configuration call, the accounting manager rebuilding approval rules, the two staff who cleaned vendor records over three weekends. If those hours went into the project, they belong in the cost column at loaded rates.

The durability test: does the gain survive month three? Month one has novelty and executive attention behind it. Champions push adoption, the vendor's success team answers fast, and edge cases get handled by enthusiasm. By month three the champions have moved to the next project and the exception queue has found its natural size. Whatever is still measurable then is the real gain. Measure at month one and you are recording the honeymoon.

A return that survives all three tests is smaller than the brochure number, arrives later, and is real.

One worked example, arithmetic included

Take a 200-person distributor processing about 4,000 supplier invoices and 2,500 delivery documents a month, with five accounts-payable staff doing the keying, matching, and chasing. It deploys a document-processing automation that reads incoming documents, matches them to purchase orders, and posts clean transactions to the accounting system.

The cost side, all of it:

  • Implementation and integration: $38,000, one time
  • Internal team hours during rollout: 300 hours at a $55 loaded rate, $16,500
  • Training: 120 staff-hours across thirty people, $6,600
  • Upfront total: $61,100
  • Ongoing: $2,400 a month for software plus $600 a month for support, $3,000 in total

The gain side, measured rather than claimed. The vendor's case study promised 60 hours a week. The counterfactual test removed eight of those hours, because a planned invoice-template standardization would have delivered them regardless. The automation also created new work: about six hours a week reviewing the exception queue. And by month three, once the early attention faded, the sustained, timestamp-verified figure was 34 hours a week.

Claimed hours saved vs hours that survive the three tests
Vendor case-study claim
60 hrs/week
After the counterfactual test
52 hrs/week
After counting new work created
46 hrs/week
Still measurable in month three
34 hrs/week

At a $40 loaded rate, 34 hours a week is roughly $5,900 a month. Add $1,600 a month in duplicate payments caught and early-payment discounts no longer missed, both visible in the ledger, and the measured gain is about $7,500 a month. Net of the $3,000 in ongoing costs, the automation produces $4,500 a month.

Break-even: $61,100 divided by $4,500 lands close to month fourteen. Stated honestly, eleven to eighteen months depending on adoption and wage assumptions.

The vendor's arithmetic on the same product counted 60 hours against the subscription price alone: payback in under four months. Same software, same company. Only one of the two numbers deserves a place in a board deck.

Why a range beats a point estimate

"Fourteen-month break-even" sounds precise, but the precision is borrowed. The real drivers, adoption rate, exception volume, and wage assumptions, are each uncertain, so the output is uncertain too. A range forces you to name your assumptions; a point estimate hides them.

The decision rule is simple. Build the range from a conservative case and an expected case. If the conservative end still clears your hurdle, proceed. If only the optimistic end clears it, you are not approving a project, you are placing a bet.

Why "time saved" surveys overstate

Asking staff how much time an automation saves them is the most common measurement method and the least reliable. People report the task time they remember, and memory keeps the worst version: the invoice that took forty minutes, not the median one that took six. Reclaimed minutes also scatter through the day in fragments too small to redeploy, which is why survey totals never reappear as output anywhere.

In our experience the gap between self-reported and timestamp-measured savings commonly runs a third or more. The rule that fixes it: an hour counts as saved only when a timestamp, a throughput number, or a payroll line moves. Everything else is a hypothesis.

A worksheet you can run

One page, six columns per process:

  1. Baseline, measured and dated: hours, cost, and error rate before anything changes.
  2. Claimed gain: the vendor's or the team's number, recorded so it can be checked later.
  3. Counterfactual deduction: improvement attributable to anything else already in motion.
  4. New work created: exception handling, review queues, maintenance.
  5. Verified gain, measured at month three and again at month six, from system data rather than surveys.
  6. Full cost: one-time spend, internal hours at loaded rates, and recurring charges.

Break-even is upfront cost divided by net monthly verified gain, written as a range. Any automation that cannot fill in column five within a quarter of going live should be treated as unproven, whatever the deck said.

This worksheet is the same discipline we apply in our own engagements, because we have to: the analysis and strategic review are free; beyond that, you don't pay unless verified savings exist. How that verification works is described on the how it works page, and the analysis applies this arithmetic to your numbers before you commit to anything.

Lo que preguntan los compradores

Usted no paga si no le hacemos ahorrar.

Ver cómo funciona el proyecto

Usamos cookies de analítica para entender cómo se usa el sitio y cómo nos encuentran los visitantes. El almacenamiento estrictamente necesario permanece siempre activo; nada más se carga hasta que usted elija. Política de cookies