All posts

Research / Venture investing

Does Proprietary Data Give an AI Startup an Advantage?

Test an AI startup's proprietary-data claim with evidence of authorised use, measured product improvement, collection costs, and an advantage that persists.

By Waypoint ExponentialPublished Revised
A tall teal block holds a terracotta cube beside a line of cream cubes ending in a brass cube, representing proprietary information and measured product improvement

Ask an AI startup to show the product's performance with and without its claimed proprietary data. Keep the task, model, and acceptance criteria comparable, then test against a credible public or licensed alternative. Check the rights to use the data, the cost of maintaining it, and what happens as competitors and models improve. Those checks turn a data claim into an investment question you can investigate.

Specify the data and the claimed advantage

Replace the pitch phrase with a description of the asset. Ask which records the startup has, where they came from, when they were created, and what the product does with them. Request a source inventory and a small authorised sample through an appropriate diligence process. A document count alone doesn't establish coverage, quality, or commercial value.

Separate customer-specific context from data that improves the shared product. A customer's current contract can help the system answer that customer's question. That doesn't establish a permission to train a shared model or an improvement for another customer. Ask the founder which of those benefits they claim and which the architecture actually delivers.

Describe the comparison too. The startup may have an exclusive source, a better cleaned version of an obtainable source, or outcome labels that another supplier would take time to collect. Those claims need different evidence. Cleaning can create product value even when rivals can obtain the raw records; its cost and reproducibility belong in the assessment.

A recent discussion about AI startup moats asks what protects a company when rivals have access to the same models. It's a reader question, not a market study. For a data claim, answer it with the specific source, the measured benefit, and the difficulty of obtaining an equivalent result.

Write the claimed outcome in business terms. For an industrial maintenance assistant, that might mean identifying the right fault procedure for a specific machine and revision. For a distributor, it might mean mapping an incoming order to the correct item. Measure the completed task and its error consequences rather than the amount of text the system can retrieve.

Check the permitted use and continuity of access

Have the appropriate legal and data owners review the actual intended uses. Ask for evidence covering collection, storage, model-provider processing, retrieval, training, and any sharing across customers. Keep unresolved permissions explicit in the diligence findings. The fact that the startup can open a record doesn't prove that every proposed use has approval.

Build a rights register tied to the source inventory. Record the counterparty, agreement or other approved basis, permitted purposes, restrictions, term, and review owner. Include the scope of any claimed exclusivity. Limit inspection to what the diligence team is authorised to see; public blog examples don't require access to confidential customer records.

Ask who controls access when a customer leaves or a supplier changes terms. Identify export, retention, deletion, and derived-data obligations for specialist review. If the business case depends on keeping labels or model improvements after access ends, ask the reviewer to address those items specifically rather than assuming the original access agreement covers them.

Test the technical implementation against the reviewed scope. A cross-customer retrieval leak can invalidate a deployment even when the source is valuable. Check tenant boundaries and field-level access where needed, including evaluation datasets, logs, and support tools. The shared product's learning process must follow the approved boundaries too.

Put access dependencies beside the revenue forecast. If a single supplier provides the critical feed, request the renewal terms and a fallback plan. Evaluate an interruption using a restricted or stale copy in a safe test environment. A commercially valuable dataset can still create an operating dependency that the investor needs to understand.

Trace how the data changes a product outcome

Ask the team to trace a case from source record to accepted result. Identify whether the data supports retrieval, structured lookup, model adaptation, evaluation, or a combination. Each step needs a different check. A correct record that the retriever never returns can't improve the final answer.

Inspect missing context and versioning. A maintenance note without its machine identifier, service date, and applicable revision can mislead the system. A product that reconstructs those relationships may add value through its preparation process. Test that process separately from any claim about exclusive access.

Anthropic's Contextual Retrieval experiments show how adding document context to chunks changed retrieval performance on its evaluated corpora. That is evidence that preparation can affect access to relevant information. It doesn't establish a startup's exclusive-data advantage or its customer outcomes; the startup must measure those on its own task.

Separate retrieval success from decision success. A system can find the right manual and still select an obsolete instruction. Grade the relevant evidence, answer or action, and final business state. Require the product to handle absent evidence and conflicting sources rather than reward a confident completion on every case.

Sun and colleagues' May 2026 EnterpriseRAG-Bench preprint includes synthetic enterprise information with near-duplicates and conflicts, and questions that test missing information. Its corpus is synthetic, so it doesn't prove performance on a target company's private records. It provides concrete failure types to include alongside authorised customer cases.

Run a fair comparison against alternatives

Define the decision before running the experiment. Ask whether the claimed dataset improves an accepted task enough to justify its costs and whether a rival can obtain comparable performance. Fix the model version, tools, task inputs, scoring rules, and allowed review effort for the first comparison. Then test a practical alternative as a separate experiment.

Use a representative held-out set that the team hasn't used to tune the system. Group related cases, documents, or accounts to avoid placing near-duplicates on both sides of the split. When the claim concerns new customers, hold out customer groups. When it concerns future outcomes, make the split by time and restrict each input to information available at the decision date.

Record the knowledge snapshot as well as the question and expected result. Cohen and colleagues' WixQA benchmark releases a matching knowledge-base snapshot with its question sets and distinguishes expert-written, simulated, and synthetic cases. Apply that discipline to the target's evaluation: label case origins and preserve the evidence the system had at the time.

Compare a general baseline, a credible obtainable-data alternative, and the claimed proprietary-data version. Removing all account information can make a product fail by design; that weak baseline doesn't show that a competitor lacks access to equivalent context. Give each alternative the information its intended product would legitimately obtain.

Run both a controlled comparison and a tuned alternative. The controlled run isolates a data change within one system. The tuned run asks what a capable competing implementation can achieve with its own preparation and retrieval. Report their engineering effort and cost separately. Keeping a broken retrieval configuration fixed isn't a fair commercial comparison.

Use repeated trials where outputs vary and preserve the traces. Calibrate subjective grading with experienced reviewers, and inspect high-consequence errors individually. Anthropic's January 2026 agent evaluation guidance distinguishes the transcript from the resulting environment state and explains multiple trials and grader types. For diligence, verify the outcome the buyer needs, not only the assistant's account of it.

Read an illustrative diligence result

Consider an illustrative assistant that matches distributor orders to catalogue items. The buyer accepts a case only when every required item identifier and quantity passes review and the proposed order has no unauthorised substitution. The held-out set contains 200 orders. The figures below are invented to demonstrate interpretation, not client results or sector benchmarks.

Illustrative comparison on the same 200 held-out orders
VersionAccepted ordersWhat the run examines
Baseline with the current catalogue140 / 200 (70%)Performance with the account information a competing product also obtains.
Obtainable supplier cross-reference data162 / 200 (81%)The same system with a credible additional source available to rivals.
Reviewed proprietary matching history170 / 200 (85%)The same system with the claimed asset in place of the additional source.
Tuned obtainable-data alternative168 / 200 (84%)A practical rival design with separately recorded preparation effort.

The proprietary version accepts 30 more orders than the catalogue baseline, a 15 percentage-point difference. Against the obtainable-data version, it accepts eight more, a four-point difference. Against the tuned alternative, the difference is two orders, or one point. Each comparison answers a different question.

These totals don't establish a reliable future advantage. Inspect which individual cases changed from failure to success and which regressed, with reviewer agreement and uncertainty around the differences. A repeated run can change the totals. A large average improvement can also conceal a damaging substitution error.

Next, inspect time and cost per accepted order. The proprietary history may need costly review while the obtainable source needs substantial initial preparation. Record both. If the small difference comes from rare but expensive customer exceptions, price that business consequence explicitly; don't discard or exaggerate it from the headline acceptance rate.

The example supports further investigation into the proprietary version's remaining advantage, preparation costs, and performance on new accounts. It doesn't establish a data moat. A positive product result and a durable competitive position require their own evidence.

Inspect the collection and correction process

Follow a new record through collection, verification, and use. Ask what happens when a customer corrects an item match and whether the correction captures the reason and eventual outcome. An accepted draft doesn't necessarily establish a correct order; the final confirmation or downstream reconciliation may provide different evidence.

Request conversion counts for that process. How many events become approved records, how many receive reliable labels, and how many enter a tested release? Show missing outcomes and rejected examples. A growing event log can contain little information that changes product quality.

Inspect whose feedback the process captures. Active customers and easy cases can dominate the dataset while abandoned workflows and manual overrides disappear. Include the relevant failures in analysis, and check whether label quality differs by account or operator. Hold the operating boundary stable when claiming that more collected data caused an improvement.

Calculate the recurring acquisition and preparation cost, including staff corrections, licensing, validation, and refresh work. Connect those costs to the cost per accepted outcome using a complete workflow budget. Keep a customer's unpaid review burden visible even if it doesn't appear on the startup's supplier bill.

Ask for successive dataset and product releases on an unchanged held-out suite plus fresh cases. Compare the added records with the change in quality, error severity, and review effort. This tests the claimed feedback loop. If the team also changed the model and workflow, use separate comparisons to determine which change produced the gain.

Test whether the advantage persists

Run the evaluation with a newer general model and the obtainable-data alternative. Preserve the earlier configuration so the team can explain what changed. A provider improvement may close the gap on common tasks while leaving a local, current-information advantage intact. Inspect that remaining task coverage rather than assuming either outcome.

Evaluate fresh accounts and new time periods. A matching history can perform well for recurring orders from existing customers while adding little during onboarding. Ask whether the startup can transfer approved learning between accounts and whether the test measures that transfer. Keep customer-specific value visible without describing it as a universal product improvement.

Ask a domain expert what obtaining an equivalent asset requires. List plausible sources, collection channels, review skills, and time to verify outcomes. Compare purchase and partnership options as well as building a dataset. A rival doesn't need identical records if another source supports the same accepted result.

Examine freshness and retention. Some assets need constant updates to preserve their value; others contain historical outcomes that take time to reproduce. Model an interruption or loss of an important source and rerun the relevant cases. Include the cost and operating work required to maintain access.

Test customer alternatives too. If the buyer supplies the same records to any chosen supplier, the startup's advantage may sit in processing, integration, or service quality. Inspect those strengths on their own terms. Our broader defensibility guide covers distribution and workflow dependence alongside data.

Turn the evidence into a diligence decision

Request a compact evidence package: source and rights registers, dataset release history, a controlled comparison, the tuned alternative, and the collection-cost record. Add authorised customer evidence for the claimed outcome. Use a secure diligence process and summaries where direct access would breach the approved scope.

State the finding at the level the evidence supports. The dataset may improve a defined task for existing accounts while offering uncertain value on new customers. Its collection process may be promising while rights or costs need resolution. Don't collapse these separate findings into a single proprietary-data label.

Name what changes the conclusion. Examples include a repeatable lift on fresh customer groups, an approved source agreement, lower verified collection cost, or a credible alternative closing the gap. Assign each open question to a person and a dated next step so diligence produces a decision.

A PE investor can apply the same checks to an acquired company's historical records before budgeting an AI programme. An SME or corporate team can test its own operational data against a bought product. Access to internal records creates an opportunity to investigate; it doesn't establish the project's return.

Support the claim of a durable data advantage when reviewed access, measured outcomes, and ongoing collection evidence survive the relevant alternative tests. Where evidence supports only a preparation or integration advantage, assess that strength directly. The investment case needs a reproducible business result and a credible account of how the company keeps delivering it.

Put the work into practice

AI technical due diligence

We help VC, PE, and family-office investment teams examine the technical claims behind an AI business. We connect product evidence, architecture, delivery effort, and operating economics to the questions you need answered.