All posts

Research / Technical implementation

Is Your Data Ready for AI? A Practical Guide

Assess data readiness for one AI workflow. Check missing records, conflicting definitions, stale information, permissions, and unwritten business rules before launch.

By Waypoint ExponentialPublished Revised
Five groups of cream, teal, and terracotta cubes connect to a central teal cube, with a missing piece and a brass access gate showing data readiness gaps

“Is our data ready for AI?” has no company-wide yes or no answer. Your data may support an assistant that drafts an internal note but fail a tool that sends a binding customer quote. Start with the decision the AI will help make, then check the records, meanings, update times, access rights, and business rules that decision needs.

Start with one decision

Consider a B2B distributor that wants an assistant to prepare sales quotes. A representative enters a customer, product, quantity, and requested delivery date. A sound quote needs the customer's contract price, current stock, freight terms, and the rule for approving a discount. The task ends when an authorised person checks and sends a quote, not when a model fills a text box.

List the facts required for that task and name the system that owns each fact. The customer relationship system may hold the account and contact. The enterprise resource planning system may hold stock and orders. Contract files may hold negotiated price and delivery terms. If two systems disagree, name the person who decides which record governs the quote. This is a narrower and more revealing test than counting how many terabytes the company owns.

AWS's generative AI workload assessment asks about completeness, accuracy, consistency, timeliness, and relevance for a specific use case. NIST's AI Risk Management Framework likewise calls for checking whether the data selected for an AI system is available and suitable for its intended context. Treat readiness as evidence about this quote process, not a badge for the whole company.

Find missing records

Missing data has two forms. A required field may be blank, such as a contract end date. Or an entire event may never enter a system: a sales manager agrees to special freight terms by email, and the contract store has no record. A dashboard that counts blank cells will miss the second problem.

Take a sample of recent quotes and trace each fact back to its source. Count how many lack a customer ID, current price, stock record, or approved delivery term. Then ask staff where they found the answer when the system did not have it. Do not fill gaps with plausible text. A missing price should stop an automated quote and route the case to the contract owner.

Fix the capture point where a recurring gap begins. If the sales team records special terms only in email, add an approved place to save them and a person who checks them. If a source system has the data but the assistant cannot reach it, solve the connection instead. These are different repairs with different owners.

Settle conflicting definitions

Two systems can have complete records and still disagree on what a term means. “Active customer” might mean a recent order in sales and an open account in finance. “Available stock” might include reserved units in one report and exclude them in another. A model that sees both labels cannot decide which definition the company intended for a quote.

Write a short definition for each term that changes the decision. Include the owner, source, unit, and calculation. For stock, state whether the quote uses on-hand units, unreserved units, or an allocation the warehouse has confirmed. For price, state whether the figure includes tax, currency conversion, and contract discounts. Use stable customer and product identifiers so the same entity joins across systems without a guess based on a name.

Microsoft's data-quality guidance separates completeness, consistency, timeliness, and uniqueness. A high score on one measure cannot settle a business definition. The process owner and data owner need to agree on the meaning that governs the quote, then make the rule visible in the application and its tests.

Check how old the facts are

A record can be accurate when stored and unsafe when used. Yesterday's product catalogue may be adequate for an internal draft; yesterday's inventory may be wrong for a delivery promise. Record when the source changed, when the assistant last received that change, and how old the information may be for this task.

Test a normal update: sell the last unit, change a contract price, and revoke an expired offer. Measure when the assistant sees each change. If stock updates every night, do not let it promise today's availability from a morning search index. It can draft the rest of the quote and ask the representative to confirm stock in the live system. Show the timestamp and source beside the recommendation so the reviewer can see what they are approving.

The required update speed depends on the decision. A public product description can have a longer review interval than stock or customer credit status. Set that interval for each source and stop or fall back to a person when the data is older than the agreed limit.

Carry permissions into retrieval

A salesperson may see their own accounts but not every customer's negotiated price. The assistant needs the same boundary. Copying all contracts into a single searchable store without their access rules can expose information that the original application kept private. Passing a user's name in a prompt does not enforce an access check.

Test the assistant as two real roles. A representative for Account A should see only the records needed for Account A. A manager with wider authority may see more, but the system should still record which identity made the request. Test a former employee's revoked access and a document shared with the wrong group. AWS's secure-retrieval guidance recommends access controls and source metadata at ingestion and retrieval. Authorisation belongs in the application and data layer before material reaches the model.

Also decide which data may leave the company for a model service and how long the service or application keeps it. If the answer is unknown, use masked test cases while the owner settles the policy. A pilot does not need a copy of the full contract archive to prove that quoting can improve.

Write down the rules people know

Experienced staff carry rules that no database column describes. A regional manager may need to approve a discount above a threshold. A product may ship only from a particular warehouse for one customer. A promised date may change when an order includes installation. If the assistant learns only from past quotes, it may copy an old exception as if it were current policy.

Ask a representative and approver to walk through recent straightforward quotes and exceptions. Record the trigger, rule, authorised owner, and date the rule took effect. Put hard limits in ordinary application checks where possible, and show the current policy text to the person reviewing an exception. Do not ask a model to infer approval authority from examples or invent a rule when the written policy is silent.

Run a small readiness test

Choose twenty recent quotes that cover common orders, contract prices, changed stock, special terms, and rejected discounts. For each case, record the correct final quote and which sources and approvals produced it. Run the same questions through a read-only version of the proposed assistant. Do not use live customer data without the agreed access controls.

  1. Coverage: Count the cases with every required record. Separate a blank field from a missing event or document.
  2. Meaning: Mark every case where sources disagree or a term has no agreed definition.
  3. Age: Compare the source and retrieval timestamps with the limit set for that decision.
  4. Access: Repeat selected cases as different roles and confirm that each sees only authorised records.
  5. Rules: Check whether the assistant stops at each approval boundary and cites the rule the reviewer should apply.

Publish the counts and examples with a named owner for each gap. A pilot can proceed with a narrow draft-only task when staff can verify its sources and catch its errors. Do not let it send quotes or promise stock while required facts, permissions, or approval rules remain unsettled. Fix those gaps at their source, rerun the cases, and expand the assistant only when the evidence supports the added authority.