All posts

Research / Economics and governance

Measure AI Quality in Business Terms

Turn AI evaluation scores into correct orders, resolved tickets, and fewer costly errors. Define outcome metrics, denominators, review windows, and a business scorecard.

By Waypoint ExponentialPublished Revised
A teal cubic assembly fits a cream recess beneath a brass square frame while a terracotta cube sits beside an empty recess, representing a finished result checked against a standard

Measure AI quality by the work your business accepts: an order with the correct items and terms, a ticket that stays resolved, or a close completed without a material correction. Define what counts, check the evidence, and keep unfinished work in view. Model scores help diagnose problems, while an outcome scorecard tells an operating owner whether the whole process delivers acceptable work.

Define the business outcome before choosing a score

An AI assistant can extract every field in an order email correctly and still create the wrong order. A product code may match the message but refer to a discontinued item. The application may use an outdated price or promise stock that another customer has already reserved. Each component needs testing, and the business needs a check on the resulting commitment.

Start with the process owner's acceptance rule. For order entry, specify the customer, authorised items, quantity, agreed terms, and valid fulfilment commitment. Include the required approvals. The final record must agree with the request and current business rules; an assistant's statement that it completed the task doesn't establish this.

Anthropic's January 2026 guide to agent evaluations distinguishes a trial's transcript from the resulting state in the environment. Its booking example checks whether a reservation exists. Apply that principle to your business system: verify the actual order, ticket, or reconciliation against its acceptance rule.

A recent Reddit discussion about business results and agent scores asks why better component performance can leave the overall process unchanged. The poster discloses a commercial interest, so treat the discussion as a question readers have, rather than evidence of a measured industry trend. Your answer needs records from the process you operate.

Keep diagnostic measures beside the outcome measure. Field accuracy, retrieval quality, and tool-call failures help engineers locate the cause of a bad order. They don't replace the acceptance check. A supplier lookup can improve while an approval delay still prevents the customer receiving a valid commitment on time.

Different jobs need different definitions. For support, a closed ticket can count as resolved only after you check the agreed issue outcome and an observation period for repeat contact. For a finance close, specify the deadline and what makes a later correction material. Write these rules with the people who own the work.

Write a metric contract that two people can apply

A metric contract is a short definition of how the team counts a result. It names the business unit, acceptance criteria, evidence source, reporting period, and owner. It also states what belongs in the population and how you treat missing evidence. Two reviewers should reach the same verdict on the same case.

Use a stable case ID that survives retries and transfers. An order can involve several model calls, a coordinator's edit, and a delivery-system check. Those events describe a single business request. Counting each successful call as completed work inflates throughput and hides repeat effort.

For an illustrative order assistant, define the unit as a distinct customer order request entering the agreed intake queue. Define eligibility before the assistant runs, using supported request types and account rules. Link each request to its final order, approval record, and any correction. Preserve the original request so a reviewer can compare it with the commitment.

Define first-time acceptance as meeting the business rule at the initial completion, without a repair during the agreed observation window. An authorised review that confirms a correct proposal can fit that definition. An edit that fixes an incorrect quantity is a repair, even when staff catch it before sending the order.

Count successful escalation separately. An assistant that correctly refers an unsupported request to a specialist has followed its routing rule. That request hasn't yet reached the final order outcome. The scorecard can recognise correct routing and still track the specialist's eventual result and handling time.

The NIST AI Risk Management Framework's Measure function calls for documented metrics, performance checks under conditions similar to deployment, and monitoring in production. It also includes feedback from domain experts. The contract described here gives an operating team a concrete way to apply those principles; it doesn't establish certification.

Keep the denominator and unfinished work visible

Every percentage needs its count and population. “94% accepted” means little until you know whether the team counted all eligible requests, only completed requests, or only cases the assistant chose to attempt. Publish the numerator and denominator together, including the exclusions.

Report intake, eligibility, attempts, and completion as separate counts. Eligibility limits the supported scope; attempts show actual coverage inside that scope. Keep a record of eligible cases the system skipped or sent to a person. A system that handles only the simplest requests can have a high pass rate with little effect on the queue.

Track the unserved population outside eligibility too. Those requests still consume staff effort. A team can reduce the assistant's supported scope to improve its score while increasing the manual team's workload. The operating owner needs to see both the supported process and the wider intake.

Distinguish conditional quality from delivered coverage. Accepted cases divided by completed cases describes quality among recorded completions. Accepted cases divided by eligible intake describes how much acceptable work that intake received. Show both, with the pending count, so the scorecard can't hide a growing backlog.

Give unfinished cases a separate state. At a reporting cutoff, some requests may still wait for a customer, an approval, or a system repair. They contribute to waiting time and the intake count. Don't silently remove them or classify them as confirmed incorrect without inspecting the reason.

Missing outcome evidence also needs its own count. If the team can't link a request to its final order, it hasn't verified acceptance. Exclude it from the verified numerator and show the evidence gap. Keep that gap distinct from a proven business error so operations knows whether to repair the work or the measurement.

Read a worked order-processing scorecard

Consider an illustrative cohort of 1,000 eligible order requests. The team observes each request through a defined service deadline and a further 14-day repair window. The reporting cutoff comes after all those windows have elapsed. These numbers illustrate counting rules; they aren't a customer result or a recommended target.

Of the 1,000 requests, 960 have a recorded completion and 40 still lack a completed outcome. Full case review finds 880 that meet acceptance at first completion and need no repair during the observation window, 60 that reach acceptance after a repair, and 20 that still fail the acceptance rule despite their completed status. The team therefore verifies 940 accepted outcomes.

Illustrative quality measures for the same 1,000 requests
MeasureCountMeaning
Recorded completion.960 / 1,000 = 96%.The system marks the case complete.
Verified acceptance.940 / 1,000 = 94%.The business accepts the result.
First-time acceptance.880 / 1,000 = 88%.The case needs no repair.
Unfinished work.40 / 1,000 = 4%.The outcome still needs attention.

Among the 960 recorded completions, first-time acceptance is about 91.7%. Across eligible intake, it's 88%. Neither calculation is an error; each answers a different question. Label them clearly. Management needs the intake view when deciding whether the process serves demand.

Suppose 850 of the 940 accepted outcomes also meet the service deadline. Accepted on-time coverage is 850 / 1,000, or 85%. A correct order after the customer's required service window doesn't pass that timeliness measure. Report the quality result and deadline result together instead of averaging them into an unexplained score.

Now assume the team incurs A$4,400 in processing costs across the entire cohort, including review, repair, and the unsuccessful cases. Dividing by 940 verified accepted outcomes gives about A$4.68 per accepted request. Dividing by all 1,000 requests gives A$4.40 per incoming eligible request. Both measures can help planning, but they describe different units.

Those assumed costs don't prove a saving. Compare them with the old method at similar volume and quality, including the costs your calculation covers. See our guide to building a baseline for AI cost claims for that comparison. Keep any omitted implementation cost visible.

Separate severe failures from routine corrections

A wrong punctuation mark and an unauthorised price commitment shouldn't carry the same operating consequence. Define severity using the effect on the customer, business obligations, and recovery effort. Record the actual event and consequence, rather than assigning seriousness from how confident the generated text sounds.

Keep severe incidents as a separate count with the relevant exposure. In the illustrative cohort, three cases could involve an unauthorised commitment, including cases staff later repair. Report three incidents among 1,000 eligible requests and describe the affected action. Those incidents overlap other outcomes; don't add them as another mutually exclusive row in the completion breakdown.

Use the exposure that fits the failure. A wrong-price commitment rate needs the number of price commitments as its denominator. A customer-record disclosure rate needs the relevant access or disclosure opportunities. Intake counts still provide context, but they can't substitute for a clearly defined event population.

Report near misses too. If review catches a wrong item before fulfilment, the customer outcome may stay correct while the preparation step needs repair. Track both the prevented error and the effort that caught it. The human-in-the-loop workflow guide explains how to connect review to the final action.

A weighted error score can help prioritise improvements when its weights reflect a stated business rule. Keep the underlying counts visible. A large number of easy successes mustn't cancel an event that the owner has defined as a reason to stop automated execution.

Set incident responses before reviewing a dashboard. Specify who can narrow the permitted scope, return work to manual processing, and authorise a restart after evidence of a fix. An aggregate metric won't make those decisions for the team.

Check the evidence behind each reported result

Build the outcome record from the source request, authoritative system state, and any subsequent correction. Include timestamps and the version of the workflow that handled the case. A model-generated success flag provides a diagnostic event; the acceptance check needs independent evidence.

For orders, compare required fields and terms against approved records and inspect the resulting commitment. For support, connect repeat contact to the same underlying issue where possible, including contact through another channel. A customer who opens a new ticket can expose a failed resolution without reopening the original ticket ID.

Use automated checks for conditions that your system can verify reliably. Domain reviewers can assess ambiguous outcomes using a written rubric. If a model helps grade free text, compare its verdicts with human decisions and inspect disagreements. Anthropic's evaluation guide recommends calibrating model graders against expert judgement; the same principle applies to a live quality audit.

State whether a result covers all cases or an audited sample. The worked scorecard above assumes full case review. If reviewers instead inspect 100 cases and find eight failures, report eight failures in that 100-case sample, its selection method, and the population it represents. Don't present the sample's accepted count as a census of the whole queue.

Sample across the real case mix, and inspect severe incidents and known failure patterns separately. If you deliberately review more difficult cases, account for that selection when estimating a population rate. Keep a targeted defect review distinct from a representative quality estimate.

A small sample with no observed serious errors doesn't establish that serious-error risk is zero. Show the sample size and observation period, and gather more evidence as exposure grows. Keep the operational stop rules active while uncertainty persists.

Give reviewers a way to mark the outcome unknown when the evidence doesn't settle it. Assign someone to reconcile those cases. Forcing a pass or fail from incomplete records introduces invented certainty into the dashboard.

Compare like work and allow time for failures to appear

Group cases by intake period and follow them through a consistent observation window. A dashboard of completions this week can mix requests from several weeks while omitting this week's unfinished intake. An intake cohort makes that delay visible and preserves the correct denominator.

Choose the window according to how failures emerge. A support team may need time to see repeat contact. A finance team may need the next reconciliation or review to discover a correction. Fourteen days in the example is a stated assumption, not a universal quality standard.

Label recent cohorts provisional until their windows mature. At the weekly meeting, report their pending work and early signals, while using mature cohorts for the settled quality comparison. Record late-discovered failures and revise the relevant cohort with a visible measurement history.

Compare similar request types, channels, and operating conditions. A shift toward simple repeat orders can raise overall acceptance even when neither simple nor complex case performance improves. Publish quality by important case groups and the count in each group.

If you compare the AI workflow with the old method, keep the acceptance rule and observation window consistent. Use concurrent comparable work where possible, and record staffing, policy, and demand changes. An improving business result doesn't establish that AI caused the improvement if several other inputs changed.

Include review and repair effort from receiving teams. Faster order entry can increase the warehouse's correction work. Look across the handoff before reporting a benefit, and check elapsed customer time alongside active staff time.

Use the measures to make a specific operating decision

Bring the owner a compact report: eligible volume, verified first-time acceptance, accepted on-time coverage, unfinished work, serious events, and cost per accepted outcome. Include the definitions and evidence gaps. Keep technical diagnostics available for investigating a change.

Pair each adverse result with a responsible team and a next check. Incorrect item matches may need a catalog fix. Long approval waits may need staffing or a narrower intake boundary. Review the evidence before deciding that a different model will solve the problem.

Set thresholds from the operating requirements and consequences of failure. A draft-only assistant can have a different acceptance boundary from an executor that commits orders. A startup, portfolio company, and corporate division need to agree their own limits; a headline percentage from another deployment doesn't supply them.

For VC diligence, ask the startup to reconcile its evaluation score with customer outcomes, coverage, and repair effort. Inspect a sample of cases behind the claim. For a PE operating team, make sure portfolio reports use comparable definitions before ranking projects. Corporate leaders need a receiving owner who can explain unresolved work and act on it.

Keep the technical test suite connected to the outcome failures. When an incorrect order reveals a repeatable defect, add it to the agent evaluation set and verify the fix before the next release. Continue measuring the live outcome, since a passing regression test doesn't establish quality across changing customer work.

Begin with a recent, fully traceable intake cohort. Write the acceptance rule with its owner, reconcile the completions, and inspect the repairs. If you can't reconstruct the claimed outcome from those records, fix the evidence trail before using the percentage to approve wider automation.

Put the work into practice

AI implementation and delivery

We help SMEs and scale-ups put AI into a specific business workflow. We define the problem, prepare the data, build the software, and help your team operate it in production.