Research / Venture investing
10 Technical Questions to Ask Before Investing in an AI Startup
Ten technical due diligence questions for AI startup investors, with the evidence to request on product quality, data rights, security, model risk, costs, and adoption.

An AI demo tells you what a product can do once, with a chosen input. Technical diligence asks what happens across real customer work: awkward inputs, changed models, restricted data, outages, and a bill that grows with use. These ten questions help an investor ask for evidence without treating every early-stage startup like a mature software company.
1. Can you show a completed customer job?
Pick a real task from intake to final action. For a support product, follow a customer request through policy lookup, draft response, approval, and the record saved in the support system. Ask the team to show which step the model handles, which step a person handles, and how the application knows the job is complete. A screen that produces a fluent answer may leave most of the work with staff.
Request a redacted trace from a normal customer case, including the source records, tool calls, edits, and final outcome. If a founder can show only a prepared prompt, ask to watch an operator complete an ordinary task. The gap between the two tells you what the product still needs to build.
2. How do you measure the result?
Ask for the evaluation set, the scoring rule, and the last few results by product version. The set should contain routine cases and the cases that make staff stop and think. For a support assistant, separate correct policy answers from plausible but unsupported replies, and record whether a human had to rewrite the answer. Compare the same cases with the customer's current process or a simple baseline.
Check who chose the cases and who judged the outputs. If the team tuned prompts on every case in the test set, request fresh held-out cases or a sample reviewed with a customer. OpenAI's evaluation guidance recommends task-specific tests, evaluation after changes, and adding new failure cases from production. A single aggregate score can hide a serious failure in a small but important group of requests.
3. Which failures matter most, and how does the product respond?
Ask the team to rank failures by their effect on the customer. A wrong suggestion that an agent catches is different from an unauthorised refund or a reply that exposes another customer's record. Look for a log of actual failures, their causes, the time it took to detect them, and the changes made afterward. The response should match the severity: show uncertainty, ask for review, stop an action, or hand the job back to a person.
Request one recent incident or near miss and trace it through detection, customer impact, correction, and a regression test. An early product may have a small sample. It can still show that the team records and learns from mistakes. The NIST generative AI risk profile treats measurement and ongoing risk management as part of the system's life cycle.
4. What happens when the model or provider changes?
Map each model call to a named task. Ask which provider and model version the product uses, what it costs, and which outputs fail when that dependency changes. Then request a run of the same evaluation set on an alternative model or a newer version. Compare answer quality, latency, and cost together. A claim that the team can switch providers in a day needs a working demonstration.
Also ask who approves a model upgrade and how the team rolls it back. Some applications depend on the exact format of a model's output or on later steps reading earlier responses. If one changed response can break a downstream action, the investor should see the test and release process that catches it.
5. Where does the data come from, and what may the company do with it?
Trace one document or record from collection through retrieval, logging, evaluation, and deletion. Ask who owns it, which agreement permits each use, and whether the startup may use it to improve a shared product. A right to process a customer's records for that customer's service doesn't by itself establish a right to train on them or show them to another customer. Review the relevant contract language with counsel rather than relying on a data-room slide.
Request a data inventory that names source, owner, purpose, access rules, retention, and deletion path. For a claimed proprietary data advantage, compare quality with and without the data on a named task. NIST's generative AI profile includes actions for tracking content provenance and its interaction with privacy and security.
6. How do you protect one customer's data from another?
Ask the engineer to follow a request through authentication, retrieval, the model provider, application logs, and any human review queue. Who can read the prompt and retrieved documents at each point? How long are they kept? Show a test in which a user from customer A tries to retrieve customer B's records. Permissions must come from the application and source system, not from an instruction telling the model to respect boundaries.
Request the access-control design, a redacted cross-tenant test result, and the provider's data-handling terms. Check whether logs contain raw customer content and who can export them. OWASP's LLM risk guidance lists sensitive-information disclosure among the risks teams need to address.
7. What can the AI do without a person's approval?
List every tool the model can call and the credentials behind it. A read-only lookup, a draft reply, and a payment action need different controls. Ask for the exact point where a person approves a consequential change, how the system checks the action against business rules, and what audit record survives afterward. A prompt that says "ask first" isn't an authorisation boundary.
Have the team run a malicious instruction hidden in a retrieved document or customer message. The test should show whether that text can trigger a tool call, bypass an approval, or leak data. OWASP describes excessive agency as a risk when an application gives a model more functions, permissions, or autonomy than its task needs.
8. What happens when production breaks?
Ask what happens after a provider timeout, a failed tool call, duplicate delivery, or an unavailable source system. Can the customer finish the job through a person or an existing workflow? Will a retry send the same refund twice? Request a redacted trace of a failed request, the alert it raised, and the steps an operator took to recover. Look at slow requests as well as average speed; a customer waiting at the end of a long queue experiences the slow case.
The team should know which component owns each completed action and which system holds the final record. If a model response and a database update disagree, there needs to be a way to reconcile them. This is ordinary production engineering, but a fluent demo can make it easy to overlook.
9. What does a successful customer job cost?
Start with a completed task, not a model-call price. Add model use, retrieval and hosting, retries, human review, support, and rework. Separate the easy cases from the difficult ones and show cost by customer or workload tier. Ask how many jobs the product completes at acceptable quality without staff intervention. A low token cost means little if the team must review every answer or fix many actions later.
Request a recent usage cohort, its fully loaded variable cost, and the revenue tied to those jobs. Model a higher-volume customer using that cohort's actual mix of requests, then test a provider price or usage change. OWASP includes unbounded consumption in its LLM risk list because resource use can become an operating and cost problem.
10. Do customers keep using the product after the pilot?
Ask for usage by paying customer and by the people who do the work, with pilots marked separately. Which workflow runs regularly, which customers expanded it, and which teams reverted to their old process? A technical product can pass its test set and still create more review work than customers will accept. Talk to an operator as well as the buyer and ask what they use when the product fails.
Match product events to invoices, renewals, and customer references. At seed stage, a few observed jobs and candid customer calls may be the strongest available evidence; at growth stage, expect repeat-use and retention cohorts. If the company can't yet show either, name adoption as an open investment risk and agree on the proof needed before more capital goes in.
Use these questions to write a short diligence memo with the evidence seen, the gaps, and the consequence of each gap. A startup doesn't need a perfect answer to every question. You do need to know which claims already hold up under real work and which ones still depend on the next experiment.
