All posts

Research / Economics and governance

What Should an AI Pilot Prove Before It Scales?

Set a clear scale decision for an AI pilot with baseline results, quality and safety gates, human review time, cost per completed task, and stop rules.

By Waypoint ExponentialPublished Revised
A small terracotta cube passes three teal and cream checkpoints and a brass gate toward a larger stack of teal cubes, representing an AI pilot's scale decision

A pilot can produce good demos and enthusiastic feedback while making the real workflow slower or less safe. Before you extend it to more teams, decide what result would justify that move. The answer needs a baseline, a quality threshold, the time people spend reviewing output, a full cost per completed task, and a person who can stop the rollout.

Write the scale decision before the test

Define the task, the first group of users, and the decision date. State the volume you expect at full use and the action the system may take. A tool that drafts invoice fields for an accounts-payable clerk has a different risk from a tool that approves a payment. Name the process owner, the person responsible for technical operation, and the person who can approve wider use.

Write three outcomes in advance: expand within a defined boundary, revise and run another test, or stop. Set a maximum spend and a clear end date for the pilot. Without those terms, a team can keep changing the success measure until the test appears to pass. The NIST AI Risk Management Framework calls for repeatable measurement and an explicit decision about whether a system should proceed, alongside plans for monitoring after deployment.

Compare completed work with a baseline

Measure the old process before changing it: cases completed, active time per case, waiting time, rework, serious errors, and cost. Split routine cases from exceptions. If you test only clean examples, the pilot won't represent the work that staff must handle after rollout.

Compare the pilot with similar cases processed under the old method during the same period when you can. Demand, staffing, and policy changes can otherwise look like an AI effect. Count the final outcome after a person reviews the output, because a fast first draft has no value if the team spends longer correcting it. Microsoft’s guidance on measuring agent value recommends a baseline, a comparison group where possible, and measures of both use and business outcomes.

Set quality and safety gates

Choose an acceptable result for each case type before reviewers see the AI answer. Measure whether the completed case is correct, whether the first draft needed correction, and whether staff caught errors before they reached a customer or payment system. Report severe errors separately. A 98% overall pass rate can hide a small number of mistakes that the business cannot accept.

Write hard stops for events such as an unauthorized action, a cross-customer data leak, or a wrong payment instruction. Keep the system in draft mode if the pilot cannot show that it respects approval and data boundaries. A small sample also cannot prove that a rare, serious error has disappeared. Expand in stages, monitor real cases, and keep a route back to the old process.

Count human review and full unit cost

Track minutes spent checking every output, handling exceptions, and fixing mistakes. Ask reviewers to record why they changed a draft. If staff feel they must reread every source document, the model may have shifted work instead of removing it. Measure how often they use the tool voluntarily once the pilot team stops prompting them.

For the unit cost, include staff time, model and infrastructure charges, software fees, monitoring, and ongoing support. Divide the total by completed, acceptable cases, not by model calls or drafts. Show the calculation at expected production volume, because fixed support costs spread differently when use grows or falls. Time returned to staff is capacity; call it a cash saving only when a budget or paid workload actually changes.

Work through an invoice pilot

Consider an illustrative company that processes 1,000 invoices a month. Staff currently spend an average of 10 active minutes on each completed invoice. At an assumed loaded rate of £30 an hour, that is £5 of staff time per invoice. These numbers are examples for the calculation, not a target for another company.

The AI system drafts invoice fields but cannot approve or send a payment. During the pilot, staff spend an average of six minutes reviewing and handling exceptions per completed invoice, or £3 of staff time. Model and infrastructure cost £0.20 per invoice, monitoring costs £0.50, and £500 of monthly support costs £0.50 per invoice at the forecast volume. The full cost is £4.20 per completed invoice, leaving £0.80 of potential capacity value per invoice against the £5 baseline. At 1,000 invoices, that is £800 a month before any rollout cost.

That arithmetic doesn't pass the pilot on its own. The owner must also check correctness after review, supplier mix, exceptions, and near misses against the old process. If review rises to eight minutes in another team, staff time alone becomes £4 per invoice and erases the expected margin. At half the forecast volume, the fixed support cost doubles per invoice. A small pilot that selects only easy suppliers gives a misleading scale estimate.

Expand, revise, or stop

Expand only if the completed work meets the pre-agreed quality bar, serious failures stay within the approved risk boundary, review time falls, and full unit cost improves at realistic volume. Give the next group a limited share of cases and an owner who checks quality and incidents each week. Keep the same evaluation cases and add new exceptions as the workflow changes.

Revise if the problem is specific and fixable, such as missing supplier records or a review screen that makes staff repeat work. Set a date and a new test for that change. Stop if the system crosses a hard safety boundary, cannot beat the old process after review, or has no credible route to a better full cost. A stopped pilot still gives the company an answer: the evidence does not support broader use of that design.