AI Technical Diligence Checklist
An AI technical diligence checklist: evidence to request on product claims, data rights, evaluation, vendors, security, agent authority, and IP, plus red flags.
Direct answer
Direct answer: AI technical diligence checklist
AI technical diligence should test the claims the investment depends on, not survey the whole technology stack. Start from the three or four statements in the deck that matter most, request the evidence behind each, and record what that evidence does and does not show. When a claim can only be settled by running an experiment, say so and scope a separate validation study rather than stretching a document review to cover it.
The question
“What evidence should we request to test a company’s technical claims before deciding whether deeper work is needed?”
Who it is for
Investors, investment committees, corporate development teams, and venture builders assessing an AI company or product.
What this gives you
A technical evidence request list, red-flag questions, and conflict-disclosure prompts.
- 01 An evidence-request list grouped by area, ready to send to the company
- 02 Red-flag questions for the management session
- 03 Conflict-disclosure prompts to settle before anyone reviews the company
- 04 A clear basis for deciding whether a deeper review or experimental validation is needed
Before you start: the claims that matter
- □List the three to five technical statements the investment case depends on, in the company’s own words.
- □For each, note whether it is a capability claim, a performance figure, a cost or margin claim, or a defensibility claim.
- □Agree which claims a document review can settle and which would need an experiment.
Product and performance claims
- □Evaluation reports behind every headline figure: dataset, baseline, sample size, date, and who ran it.
- □Examples of failures and how often they occur, not only successful demonstrations.
- □Which results come from production use and which from controlled tests or pilots.
- □Customer usage data showing that the AI feature is used and retained, where the claim depends on it.
- □How much human review or correction sits behind the product today.
Data rights and provenance
- □Sources of training, fine-tuning, and retrieval data, with the rights or licences for each.
- □Customer data terms: what the company may use for training or improvement, and what it may not.
- □Personal data processed, where it is processed, and under which agreements.
- □How data quality and labelling are checked, and by whom.
- □Any data the product depends on that a third party could withdraw.
Evaluation practice
- □The evaluation suite the team runs before each release, and recent results.
- □How the team detects regressions when a model, prompt, or vendor changes.
- □Whether test data is kept separate from training and tuning data.
- □How quality is monitored once a feature is live.
Model and vendor dependencies
- □Every model and AI service the product depends on, with contract terms and pricing exposure.
- □What happens to cost and quality if the main model provider changes price, terms, or behaviour.
- □Whether the product could move to another provider, and what that would take.
- □Gross-margin sensitivity to inference costs at current and projected volumes.
Security and agent authority
- □What the product’s AI components can read, write, or execute in customer environments.
- □How high-impact actions are approved, logged, and stopped.
- □How prompt injection, data leakage, and misuse have been tested.
- □Incident history and how incidents were handled.
Team and intellectual property
- □Who built the core system, and whether they are still with the company.
- □How much of the technical advantage is in code, data, workflow, or relationships.
- □Patent position, distinguishing granted patents from pending applications and provisional filings. A provisional filing is not a granted patent.
- □Open-source dependencies and their licences.
- □Key-person risk and documentation.
Red-flag questions for the management session
- □Which of your headline figures has someone outside the company reproduced?
- □Show us a case where the system was wrong. How did you find out, and what changed?
- □If your main model provider doubled prices tomorrow, what would happen to margin?
- □What does a customer have to do manually today that the deck implies is automated?
- □Which part of the product would a well-funded competitor find hardest to copy, and why?
- □What would make you stop or change the current AI approach?
Conflict-disclosure prompts
Settle these before any reviewer sees confidential material.
- □Has the reviewer built, advised, invested in, or competed with the company or its close competitors?
- □Does the reviewer have an existing or expected commercial relationship with any party to the deal?
- □Would the reviewer be offered implementation or advisory work if the investment proceeds?
- □Record each disclosure, and decide whether to proceed, narrow the scope, or use a different reviewer. A different brand or domain does not make a conflicted reviewer independent.
Evidence
What this draws on
The Correctness Illusion in LLM-Generated GPU Kernels
Operator-aware testing caught 10 of 10 seeded defects while 16 of 16 correct controls stayed clean across five GPU classes. A controlled corpus result, not a deployed-model defect rate.
Inspect the source ↗Boom investor commitments
Boom recorded more than $2.5M in investor commitments and commercial letters of intent. That is not money raised, recognised revenue, or production adoption.
Read more →Six companies founded, 2008–2024
Kwippy, Jaja.tv, Octo.ai, ExpressMOJO, Boom, and Neul Labs, spanning social software, interactive media, ML and analytics, logistics, blockchain infrastructure, and applied AI. Advisory and investment records are not counted as founded companies.
Read more →Provisional patent filings
90+ provisional filings across Hike and Orangewood Labs in AI, computer vision, robotics, messaging, and consumer systems. Provisional filings, not granted patents.
Scope and limits
What this is not
- —This checklist supports technical judgement. It is not legal, financial, valuation, or investment advice, and it does not replace legal, financial, or commercial diligence.
- —A document review can show whether evidence exists and what it covers; settling a performance claim may need a separate experimental study.
- —Completing the checklist does not certify the company, its product, or its claims.
If the need is different
Common questions
Answers before you commission
What should an AI technical-diligence review cover?+
The claims the investment depends on: product and performance evidence, data rights and provenance, evaluation practice, model and vendor dependencies, security and agent authority, and team and IP. When a claim can only be settled by experiment, scope a separate validation study.
What should be in an AI data room?+
Evaluation reports behind each headline figure, data sources and rights, model and vendor contracts, the release evaluation suite and recent results, security testing and incident history, and documentation of the core system and who built it.
How do we handle conflicts of interest in AI diligence?+
Ask the reviewer to disclose any building, advisory, investment, competitive, or expected commercial relationship before they see confidential material, then decide whether to proceed, narrow the scope, or use someone else.
When is a document review not enough?+
When the investment depends on a performance claim that has not been reproduced outside the company, such as accuracy on a particular task or the correctness of generated code. That needs a controlled test, not a reading of the company’s own report.
Related