All posts

Research / Venture investing

Ask an AI Startup for Its Evaluation Set

Assess an AI startup's evaluation set, failure severity, human intervention, version history, and production results before accepting its demo claims.

By Waypoint ExponentialPublished Revised
A brass inspection frame surrounds a grid of teal test cubes, with terracotta exceptions beside a separate approved stack, representing AI evaluation diligence

Ask an AI startup to show how it tests the job its customers pay it to do. A convincing evaluation set contains representative inputs, explicit acceptance criteria, and results you can trace to a specific product version. For venture investors, that evidence helps test whether a promising demo can support repeatable customer delivery and the operating economics in the investment case.

Turn the product claim into a testable job

“Our agent handles customer support” leaves the buyer's actual job undefined. Ask whether the product drafts a reply, resolves an eligible ticket after approval, or independently changes an account. Each promise needs different evidence. Write down the input, permitted action, expected business outcome, and the person who accepts that outcome.

Consider a startup that processes order-change requests. A successful case needs the right customer and order, a change that policy permits, and a confirmed update in the order system. If the product only prepares a draft, judge the draft and the review effort. Don't give a draft-only product credit for unattended completion or penalise a correctly designed escalation as though it promised full autonomy.

Andreessen Horowitz's enterprise AI builder analysis describes the gap between a polished demo and a product handling messy customer data and varied user behaviour. For diligence, translate that observation into a concrete request: show the ordinary cases and the exceptions for the workflow that customers actually buy.

A recent Reddit discussion asks how teams test agents before exposing them to users. Replies raise concerns about actions that look successful while the underlying state stays wrong. That discussion identifies a reader question; it doesn't establish a measured failure rate. Your diligence needs the startup's own records.

Ask the founder to define the scope before selecting examples. Include which customer configurations the product supports and which it excludes. An investment case based on expanding into larger accounts needs evidence about the additional workflows and constraints those accounts bring. Success in a narrow pilot supports a narrower conclusion.

Request a compact evidence package

You don't need unrestricted access to customer records to begin. Ask for a redacted package that preserves the logic of each test and its result. If removing data changes the task, arrange controlled review with the customer or startup rather than treating the simplified case as equivalent. Keep private evaluation data in the agreed diligence environment.

An initial AI evaluation diligence request
EvidenceWhat it establishesFollow-up question
The case register and scoring rules.The tested scope and definition of success.Which requests never enter the score?
Results linked to runs and versions.What the team actually ran.Can you reproduce a selected failure?
Intervention and delivery records.The work people perform around the product.Who repairs an incomplete task?
Production samples and incident history.How the test relates to customer operation.What changed after a real incident?

For each case, request a stable identifier, the relevant input, the expected outcome, and the reason the team included it. For each run, request the attempted outcome, score, error classification, and any human involvement. Link those records to the model configuration and application release. The evidence can live in a spreadsheet if the team can reconcile it with the actual executions.

Ask for a recent complete run before asking for a handpicked success. Select a few cases yourself, including a failure and an escalation. Have the team walk through what happened without silently changing the configuration. A transparent failure with a specific repair plan often tells you more about engineering discipline than a perfect example chosen for a pitch.

Match the request to company stage. A pre-product startup may have a small manually reviewed set and no production history. A company selling unattended processing at scale needs stronger operating evidence. Record missing evidence as a limit on the claim you can support, with a named next step and owner.

Inspect who and what the cases represent

Start with the population behind the score. Does the set represent paid customers, a pilot customer, synthetic requests, or examples written by the founders? Ask how the team sampled cases and whether it discarded any after seeing results. Keep the selection rule visible so a later score doesn't improve simply because the team removed difficult inputs.

OpenAI's evaluation guidance recommends tests tied to the task, coverage of typical and difficult cases, and calibration of automated scoring against human judgement. Apply that principle to the customer's operating conditions. A general model benchmark can't tell you whether this product correctly handles this customer's order changes.

For the order-change example, ask for cases with missing order numbers, conflicting instructions, and requests after dispatch. Include an account that can view an order but can't edit it. Check whether the set includes the languages and document formats the startup sells support for. A rare case can deserve explicit testing even when it contributes little to average volume.

Separate a traffic sample used to estimate routine performance from a deliberately difficult challenge set. Report both. Filling half the sample with unusual cases changes the meaning of the overall score; excluding them entirely hides the limits of the product. Ask the team to explain any weighting using observed traffic, and retain the results for each group.

Protect a reserved set that the team doesn't use while tuning the product. Split related cases together where appropriate: near-duplicate requests from the same account can leak the answer across development and test groups. A later time period or previously unseen customer may reveal a different limitation. State which type of generalisation the split actually tests.

Ask who approved the reference answers and what they do when reviewers disagree. Business policy can leave several valid outcomes. Preserve those alternatives instead of forcing a single wording. If the product needs context available only after the decision, the test may accidentally give it information that real users don't have.

Separate completion from serious failures

Agree on failure classes before reviewing the score. A draft with an awkward sentence and an update to the wrong customer's order create different consequences. Record material failures separately, including the affected action and the control that stopped or allowed it. A high average score can't cancel a serious permission failure.

Anthropic's agent evaluation guide distinguishes an execution record from the final state in the environment. It also describes combining code checks, model grading, and human review. Ask to see evidence of the resulting order state alongside the conversation, using the appropriate check for each claim.

For a permitted address change, compare the authorised destination with the saved destination. Confirm that the integration changed only the intended order. For a refusal, check that no write happened. For a timed-out update, inspect whether the operation committed before a retry. Those checks test application behaviour around the model as well as the generated response.

Use an explicit stop condition for sensitive actions. The responsible customer and product team should decide what prevents deployment and what requires review. Don't let an investor's generic threshold override a customer's access rules. Test the enforcement in the application: asking a model to respect permissions doesn't prove that its tools enforce them.

Small samples limit the claims you can make about rare failures. Zero observed serious incidents in a short run isn't evidence of zero future risk. Request counts and denominators, repeat important cases where output varies, and record what the test didn't exercise. If there isn't enough evidence for unattended action, narrow the initial permission scope and state the additional evidence needed.

Count retries and human intervention

Separate a first attempt from the best result after several attempts. If the product retries automatically, inspect the retry policy and its effect on latency and cost. If an engineer reruns failed cases until a result passes, record that intervention. The customer's experience includes the waiting and repair work.

Here is an illustrative diligence calculation, not an industry benchmark. A company receives 240 requests. It excludes 40 as outside the supported scope and reports 190 successes among 200 eligible cases, a 95% pass rate. Twenty of those 190 successes need a staff member to repair context or retry the task. The count of completions without that intervention is 170.

On the eligible population, that gives 85% completion without intervention. Across all arriving requests, it gives about 71%. The product may still meet a valuable narrow use case. The diligence issue is whether the pitch and staffing forecast use the same scope and definition as the evidence. A planned mandatory approval step also needs its own review-time measure; don't hide it inside a label such as “automated”.

Suppose the 20 interventions take six minutes each. They add two hours of work in this illustrative batch. You still need the time for expected approvals, the excluded requests, and the ten eligible failures. Measure those separately before estimating labour savings. Use the same workload for the previous process so that a changed case mix doesn't manufacture an improvement.

Ask who performs this work and who pays for it. Founder effort, a customer's administrator, and a paid delivery engineer all affect scalability differently. Our guide to an AI startup's inference bill covers another part of delivery economics. For evaluation diligence, calculate the full cost per accepted outcome, including repair, rather than dividing model spend by generated answers.

Compare product versions on comparable evidence

Request a release history that connects changes to evaluation results. A model upgrade can coincide with a new prompt, changed retrieval data, and a different tool implementation. If several things changed together, record that uncertainty. Ask which evidence supports the team's explanation of the improvement.

Use a fixed set for direct version comparisons and a separately reported refreshed set for newly discovered cases. Keep the old scoring rules or explain their changes. A rising chart with a moving denominator doesn't show how the product improved on the same task. It may show that the team changed what it measures.

Ask to inspect the current release, the previous release, and a release that caused a regression. Look for a case that improved and a case that got worse. Compare the business outcome, intervention effort, and elapsed time. A gain in fluency can coexist with weaker action accuracy or more expensive processing.

Keep a record of the input sources available during each run. If the current customer policy changed after an older run, don't score the old decision as though the agent had the new policy. Conversely, a test against a frozen policy snapshot doesn't establish that the deployed system retrieves current rules correctly. Request a separate check for that dependency.

A reproducible package needs enough context for an authorised reviewer to understand the run without exposing every private record. Prefer stable case IDs, versioned fixtures, and controlled access to source material. If a third-party model version can't be rerun exactly, preserve the original results and state that limitation instead of claiming exact reproduction.

Connect test results to deployed performance

Ask how the startup reconciles evaluation results with customer operation. Request a recent reporting period with arriving volume, supported volume, accepted outcomes, and manual handling. Include failures discovered later, such as a ticket reopened after an apparently successful resolution. A metric measured at the instant of response can miss the work created the following day.

Inspect results by customer and important workflow segment. An average can conceal a large account with poor completion or a configuration that requires constant engineer help. Ask whether the production product uses the same tools and access controls as the test environment. Mock tools establish different evidence from supported integrations operating under customer permissions.

Follow a deployed incident through discovery, containment, repair, and regression testing. Check whether the team added a case that reproduces the original problem. Ask who monitors recurrence and who can pause the affected action. Evidence that the team learns from failures strengthens a delivery assessment, but it doesn't erase the original impact.

Compare the evaluation suite with the backlog of customer complaints. A test that never covers the most common complaint needs explanation. It may measure the wrong outcome or omit an integration the customer depends on. Speak with an authorised customer contact about review workload and exceptions, and reconcile their account with the startup's reporting.

Keep experimental comparisons within approved boundaries. You can compare drafts in a controlled replay without letting both versions change live orders. An experiment involving real actions needs the customer's approved process. A shadow test also has limits: it can inspect proposed behaviour while leaving actual completion and operational load unproven.

Write a conclusion with explicit limits

Summarise the supported workflow, evidence period, tested population, and product version. State what the startup demonstrated and which claim still lacks evidence. Identify the failure classes that matter to the customer's operation, and connect manual effort to the delivery assumptions in the investment case.

A strong conclusion can be specific: the product prepares eligible order changes for staff approval, with measured review time and traceable failures across named customer configurations. That supports an assisted workflow. Expanding into unattended writes or new configurations requires further evidence. Avoid compressing those different claims into a single judgement that the technology “works”.

If a gap matters to the proposed growth plan, ask for a bounded follow-up: a reserved test period, a broader customer sample, or an operating report that includes intervention hours. Agree the acceptance criteria before the team runs it. Document the responsible owner and the date you'll review the evidence.

The same evidence package helps a PE operating team assess a portfolio deployment and helps a company buying AI compare suppliers. For a venture investor, its value is a clearer account of delivery capability and its limits. Combine that account with customer demand, retention, and the commercial terms; a well-run evaluation alone doesn't establish an attractive investment.

Put the work into practice

AI technical due diligence

We help VC, PE, and family-office investment teams examine the technical claims behind an AI business. We connect product evidence, architecture, delivery effort, and operating economics to the questions you need answered.