All posts

Research / Economics and governance

Build a Baseline Before Claiming AI Saved Money

Learn how to measure AI cost savings with a fair pre-launch baseline for time, volume, errors, quality, and labour, plus a worked comparison.

By Waypoint ExponentialPublished Revised
Two parallel rows of teal cubes on a brass measuring grid, with terracotta cubes marking exceptions in a before-and-after AI workflow comparison

“AI saved us 30%” sounds precise. It often means someone timed a few easy tasks, ignored the work needed to check the output, and multiplied the difference across a year's volume. A credible claim starts before launch: define a completed case, measure how the current process handles it, and choose a comparison that can survive a change in demand or case mix.

Choose the unit of work

Pick a unit that ends with an outcome the business accepts: a customer query resolved, an order entered correctly, or a report delivered with its review complete. A draft answer, model call, or generated document isn't the unit if a person still has to finish the job. Give the unit a start and end event in your systems, and decide how to count reopened or abandoned cases.

Write down which cases belong in the measure. A customer-support team might include billing questions received through email and web chat, then report cancellations and fraud separately. If a team adds phone calls halfway through the test, its average handling time may change even when the AI has no effect. Count both the number of eligible cases and the share the new system actually touches.

Record five baseline measures

Take a pre-launch sample from normal work, including exceptions and busy periods. Keep the underlying case IDs so someone can check the calculation. A simple record for each case should capture the following:

  • Volume and mix: cases arriving and completed, by type and difficulty. Record backlog and reopenings so a team can't improve its average by leaving hard cases unfinished.
  • Active labour time: minutes people spend reading, deciding, writing, checking, and correcting. Measure waiting time separately. A shorter queue is valuable, but it doesn't directly reduce paid handling hours.
  • Errors and rework: cases corrected before completion, reopened later, or sent to the wrong person. State the error definition and check a consistent sample of completed work.
  • Quality: an outcome measure such as correct resolution, customer follow-up, or policy compliance. Report serious mistakes on their own, even if the average improves.
  • Labour and other cost: paid hours, loaded hourly cost, overtime, contractors, and the cost of existing tools. Note which costs vary with volume and which stay fixed.

Use the same definitions after launch. If reviewers change what counts as an error, you need to rescore both periods or disclose the break. Microsoft’s guidance on measuring agent value recommends setting a baseline, combining system data with reported time, and tracking where reclaimed time goes.

Make the comparison fair

A before-and-after average is a starting point, not proof that AI caused the change. Demand may fall; experienced staff may replace new hires; a policy change may remove the hardest cases. The UK government's Magenta Book guidance on impact evaluation explains why a baseline alone doesn't show what would have happened without the intervention.

Where the workflow permits it, route comparable cases to an AI-assisted group and a group using the old process during the same period. Assign cases by a rule agreed in advance, then compare outcomes within case types. If that isn't possible, phase the rollout by team or location and keep a contemporaneous group. Document differences in staffing, demand, policy, and seasonal patterns. A simple historical comparison can still inform a decision, but label its limits.

Keep every eligible case in the denominator, including cases the tool declines, escalates, or fails to process. Measure the completed result after human review. If the AI drafts a reply in 20 seconds but an employee takes six minutes to verify it, those six minutes belong in the assisted workflow's time. Compare error severity and quality alongside speed; a faster wrong answer is a worse result.

Calculate a support-workflow result

Consider an illustrative support team handling 1,200 eligible billing queries each month. Before AI, staff spend 12 active minutes per completed query. That is 240 hours a month. At an assumed loaded labour cost of £35 an hour, labour costs £8,400, or £7 per completed query. The team records 30 queries a month that need correction after completion: 2.5% of the 1,200 cases. These are example inputs, not a benchmark.

After rollout, the same case mix and volume take eight active minutes per completed query, including prompt use, checking, edits, and escalation. Staff spend 160 hours, or £5,600 at the same hourly cost. The AI service and ongoing support cost £600 a month, bringing the assisted workflow to £6,200, or about £5.17 per completed query. The difference is £2,200 a month of potential capacity value and 80 staff hours. Suppose 18 of the 1,200 completed queries need correction, or 1.5%, with no increase in serious errors.

That result is credible only if both groups received similar billing questions, the team measured all review work, and the correction checks used the same rules. If assisted staff handled mostly simple address changes while the old process kept disputes, the averages don't answer the cost question. Compare like cases first; then weight each case type by the mix the business expects at full use.

Separate capacity from cash savings

In the example, 80 hours became available. They don't automatically appear as £2,800 in the bank. If the same employees use those hours to clear a backlog, you gained capacity and perhaps faster service. If the company avoids planned overtime, reduces a contractor bill, or handles growth without a budgeted hire, it can identify a cash effect against that budget. Show the actual line item and the date it changed.

Also account for implementation work and costs outside the monthly £600: integration, data cleanup, training, monitoring, rework, and incident response. Separate one-time spend from recurring cost, then show results at low, expected, and high volume. A fixed support cost of £600 becomes £1 per case at 600 cases; the unit economics change. NIST's AI Risk Management Framework calls for measuring benefits and costs, including effects that aren't monetary, and documenting uncertainty.

Keep a decision-ready measurement record

Before the first launch, save the case definition, sample period, data source, timing method, quality rubric, labour rate, comparison design, and decision owner. During the test, keep version changes and any shifts in staffing or routing beside the results. Report counts and rates together: “18 corrections out of 1,200 completed queries” tells a reader more than “quality improved by 40%.”

Present three numbers separately: the measured change in completed work and quality, the value of capacity released, and the cash effect confirmed in a budget. Add the assumptions that would make each number change. That gives an investor, board, or operating team a claim they can test and a clear reason to keep, revise, or stop the rollout.