Sphere Partners
Technical Due Diligence for Startups: A Checklist

AI Due Diligence: How to Evaluate an AI Company or AI Product Before You Invest

Investors are being asked to pay AI multiples for products that may be thin layers over someone else's model. Here is how to test whether a target's AI is real, owned, economically sound and compliant before you price it.

10 min read
In this article

AI due diligence is the part of a technology assessment that tests whether a target's AI capability is real, owned, economically viable and compliant. It checks whether the product is proprietary or a thin layer over a third-party model, whether the company holds the rights to its data and models, what evidence supports its quality claims, how inference costs scale, and which regulations apply. The answers decide how much of the valuation the AI actually supports.

This guide is for investors and acquirers evaluating an AI-first company or a product with AI at its core. It covers diligence on AI, not the use of AI tools to speed up a deal review. AI diligence is one of six practice areas in Sphere's technical due diligence practice, and the questions below are the ones we see decide valuations.

Why Does AI Need Its Own Due Diligence?

A standard technology assessment asks whether the software works, scales and is secure. An AI product adds questions that a code review cannot answer on its own. Its behavior depends on data the company may not fully control and on models that may belong to someone else. Its quality is probabilistic rather than pass-fail. Its gross margin depends on inference costs that move with usage and vendor pricing. And it falls under regulation, such as the EU AI Act, that classifies systems by use case rather than by technology.

Those differences are why the AI in a pitch deck and the AI in production can be very different things. The general scope of an assessment is covered in our overview of technology due diligence; this article focuses on what changes when the asset is AI.

Is the AI Real, or a Wrapper Around Someone Else's Model?

Building on a foundation model is not a red flag by itself. Most AI products today call models from providers such as OpenAI, Anthropic or Google, and that is often the right engineering choice. The diligence question is what the company has built around the model that a competitor could not copy quickly. Evidence of real, defensible capability usually looks like this:

  • Proprietary or licensed data that improves results and that competitors cannot easily obtain.
  • Fine-tuned or custom models, with training code, data lineage and version history in the repository.
  • Retrieval, orchestration and guardrail layers that encode domain knowledge, kept under version control and tested.
  • An evaluation suite that runs on every change and shows how quality has improved over time.
  • Workflow integration deep enough that customers would face real switching costs.

If most of the value sits in a prompt and a user interface, price the company as a software business with an AI feature, not as an AI company.

Who Owns the Data and the Models?

In an AI deal, data rights often decide whether the asset can be transferred at all. Diligence should trace authorization for every significant data source: training data, fine-tuning sets, customer data used to improve the product, and any scraped or third-party content. The questions are whether the company had the right to use that data for training, whether customer contracts allow it, and whether those rights survive a change of control.

Model rights need the same scrutiny. Check the licence terms for any open-weight models in use, since some restrict commercial use or attach conditions, and the provider terms for any API-based model, including how customer data is handled. A target that cannot show where its training data came from is carrying a liability, not an asset.

Data readiness belongs in this lens too. When Sphere assessed Inverite, a regulated open-banking platform, for Marble Financial, a dedicated machine learning specialist evaluated whether the target's data infrastructure could support the buyer's analytics and data product roadmap. That ran alongside a review of more than 15 security control domains, and the report was delivered in three weeks. The Marble Financial and Inverite case study describes the full scope.

What Evidence Supports the Quality Claims?

AI products fail in ways traditional software does not. They can be confidently wrong, degrade when inputs drift, and behave differently after a model version changes. A target should be able to show how it measures quality, not just demonstrate good outputs.

Ask for the evaluation datasets and how they were built, the metrics tracked and the thresholds that block a release, results over time across model and prompt changes, and how production failures are monitored and fed back into development. Then run your own tests on cases the target did not choose. A curated demo proves very little; a regression suite that has caught real problems proves a lot. AI-written code inside the product raises its own maintainability questions, which we cover in our piece on AI and technical debt. For how to review a codebase written largely with AI coding tools, see our guide to due diligence on AI-generated code.

Do the Inference Economics Hold Up at Scale?

Model calls are a cost of goods sold, not an R&D line. Every query, document or agent step has a marginal cost, and that cost scales with usage in a way traditional SaaS hosting does not. A product with healthy unit economics at pilot volume can lose money at enterprise volume if each customer action triggers long prompts, large context windows or chains of agent calls.

Diligence should rebuild cost to serve per customer, per transaction or per seat, and test it against the growth plan. Check whether the target routes simple requests to cheaper models, caches repeated work and monitors spend by feature. Then run the sensitivity: what happens to gross margin if a provider changes its pricing, or if usage per customer doubles?

How Dependent Is the Company on One Model Vendor?

Model providers deprecate versions, change pricing, adjust rate limits and update usage policies. A target built tightly around one provider's specific model inherits all of that risk. Diligence should ask what breaks if the main model is retired or repriced, how long a switch to an alternative would take, and whether evaluation suites exist to prove the alternative works. An abstraction layer and a tested fallback are signs of a team that has thought about this. Prompts hard-tuned to a single model version are signs it has not.

What Is the Governance and Regulatory Exposure?

Regulation now applies to AI by use case. The EU AI Act entered into force on 1 August 2024. Prohibited practices have applied since 2 February 2025, obligations for general-purpose AI models since 2 August 2025, and most of the Act became applicable on 2 August 2026. According to the European Commission (opens in new tab), obligations for high-risk systems in sensitive areas such as employment, education and critical infrastructure now apply from 2 December 2027, and for high-risk AI built into regulated products from 2 August 2028. Under Article 99 (opens in new tab), fines for prohibited practices can reach €35 million or 7% of worldwide annual turnover, whichever is higher.

For a buyer, the questions are practical. Does the target sell into the EU or serve users there? Has it classified its systems by risk tier? Does it keep the technical documentation, logging and human oversight its tier requires? Sector rules add to this in financial services, healthcare and hiring, and several US states have passed their own AI laws. Our EU AI Act risk classification guide explains how the tiers work. Governance gaps should be costed like any other finding, because closing them takes engineering time.

Does the Team Have Real AI Depth?

AI capability lives in people as much as in code. Find out who designed the data pipelines, who runs evaluations, and who can explain why the system behaves the way it does. Look for an evidence culture: decisions backed by experiments, documented failures, and a habit of measuring before shipping. Then test concentration risk. If one or two people hold all the model knowledge, retention belongs in the deal terms and knowledge transfer belongs in the first 100 days.

It helps to benchmark the target against how strong AI engineering firms operate. The criteria in how to choose an AI software development company work just as well for judging a target's own team. For earlier-stage targets, pair this lens with our startup technical due diligence checklist.

AI Due Diligence Checklist for Investors

Use this table to scope the AI workstream or to pressure-test findings from a diligence provider.

AreaQuestion to askRed flag
CapabilityWhat would a competitor need to rebuild this?The value sits in prompts and the interface
Data rightsWhere did training data come from, and do the rights survive a change of control?No data lineage; customer contracts silent on training use
Model rightsWhich models are used, under what licence or API terms?Licence terms restrict the intended commercial use
EvaluationHow is quality measured, and what blocks a release?Curated demos only, no regression suite
EconomicsWhat is cost to serve per customer at plan volume?Gross margin falls as usage grows
Vendor dependencyWhat breaks if the main model is retired or repriced?No fallback; prompts tuned to one model version
GovernanceWhich regulations apply, and where is the compliance file?No risk classification or technical documentation
TeamWho can improve the system, and what if they leave?One or two people hold all model knowledge

The AI workstream sits inside a wider review of architecture, security, IP and infrastructure, covered in our full software due diligence checklist. For the methodology behind all of it, see the technical due diligence and code audit guide.

Frequently asked questions

It is an assessment of a target's AI systems before an investment or acquisition. It tests whether the capability is proprietary, whether the company owns or has rights to its data and models, how quality is measured, what inference costs at scale, how dependent the product is on one model vendor, and what regulation applies.

AI tools can speed up document review, contract analysis and data room search in any deal. That is different from AI due diligence, which evaluates a target's own AI. Many deal teams now do both, but tools that accelerate review do not replace hands-on testing of a target's models, data and costs.

Ask what a competitor would need to rebuild the product. If the answer is an API key, prompts and an interface, it is a wrapper. Real capability shows up as proprietary data, custom or fine-tuned models, tested retrieval and guardrail layers, and an evaluation suite with history.

Capability and defensibility, data rights, model licences, evaluation evidence, inference economics, vendor dependency, governance and regulatory exposure, and team depth, each with a specific question and a red flag that would change the price or terms.

It can. The Act applies to providers that place AI systems on the EU market and can apply where a system's output is used in the EU, regardless of where the company is based. A US target with EU customers should be able to show how it has classified its systems.

As a workstream, it usually fits inside a standard two-to-four-week technical due diligence. A narrow AI-only review can be faster, while targets with many models or regulated use cases take longer.

Know what the AI is worth before you sign

Sphere's technical due diligence practice assesses AI products and AI-first companies alongside architecture, security and team, with costed findings in four weeks or less.

Related posts

We'd love to hear from you!

Please provide your contact details, and our team will get back to you promptly.