
AI Due Diligence: How to Evaluate an AI Company or AI Product Before You Invest
Investors are being asked to pay AI multiples for products that may be thin layers over someone else's model. Here is how to test whether a target's AI is real, owned, economically sound and compliant before you price it.
- Luke SunejaClient Partner
In this article
- Why Does AI Need Its Own Due Diligence?
- Is the AI Real, or a Wrapper Around Someone Else's Model?
- Who Owns the Data and the Models?
- What Evidence Supports the Quality Claims?
- Do the Inference Economics Hold Up at Scale?
- How Dependent Is the Company on One Model Vendor?
- What Is the Governance and Regulatory Exposure?
- Does the Team Have Real AI Depth?
- AI Due Diligence Checklist for Investors
AI due diligence is the part of a technology assessment that tests whether a target's AI capability is real, owned, economically viable and compliant. It checks whether the product is proprietary or a thin layer over a third-party model, whether the company holds the rights to its data and models, what evidence supports its quality claims, how inference costs scale, and which regulations apply. The answers decide how much of the valuation the AI actually supports.
This guide is for investors and acquirers evaluating an AI-first company or a product with AI at its core. It covers diligence on AI, not the use of AI tools to speed up a deal review. AI diligence is one of six practice areas in Sphere's technical due diligence practice, and the questions below are the ones we see decide valuations.
Why Does AI Need Its Own Due Diligence?
A standard technology assessment asks whether the software works, scales and is secure. An AI product adds questions that a code review cannot answer on its own. Its behavior depends on data the company may not fully control and on models that may belong to someone else. Its quality is probabilistic rather than pass-fail. Its gross margin depends on inference costs that move with usage and vendor pricing. And it falls under regulation, such as the EU AI Act, that classifies systems by use case rather than by technology.
Those differences are why the AI in a pitch deck and the AI in production can be very different things. The general scope of an assessment is covered in our overview of technology due diligence; this article focuses on what changes when the asset is AI.
Is the AI Real, or a Wrapper Around Someone Else's Model?
Building on a foundation model is not a red flag by itself. Most AI products today call models from providers such as OpenAI, Anthropic or Google, and that is often the right engineering choice. The diligence question is what the company has built around the model that a competitor could not copy quickly. Evidence of real, defensible capability usually looks like this:
- Proprietary or licensed data that improves results and that competitors cannot easily obtain.
- Fine-tuned or custom models, with training code, data lineage and version history in the repository.
- Retrieval, orchestration and guardrail layers that encode domain knowledge, kept under version control and tested.
- An evaluation suite that runs on every change and shows how quality has improved over time.
- Workflow integration deep enough that customers would face real switching costs.
If most of the value sits in a prompt and a user interface, price the company as a software business with an AI feature, not as an AI company.
Who Owns the Data and the Models?
In an AI deal, data rights often decide whether the asset can be transferred at all. Diligence should trace authorization for every significant data source: training data, fine-tuning sets, customer data used to improve the product, and any scraped or third-party content. The questions are whether the company had the right to use that data for training, whether customer contracts allow it, and whether those rights survive a change of control.
Model rights need the same scrutiny. Check the licence terms for any open-weight models in use, since some restrict commercial use or attach conditions, and the provider terms for any API-based model, including how customer data is handled. A target that cannot show where its training data came from is carrying a liability, not an asset.
Data readiness belongs in this lens too. When Sphere assessed Inverite, a regulated open-banking platform, for Marble Financial, a dedicated machine learning specialist evaluated whether the target's data infrastructure could support the buyer's analytics and data product roadmap. That ran alongside a review of more than 15 security control domains, and the report was delivered in three weeks. The Marble Financial and Inverite case study describes the full scope.
What Evidence Supports the Quality Claims?
AI products fail in ways traditional software does not. They can be confidently wrong, degrade when inputs drift, and behave differently after a model version changes. A target should be able to show how it measures quality, not just demonstrate good outputs.
Ask for the evaluation datasets and how they were built, the metrics tracked and the thresholds that block a release, results over time across model and prompt changes, and how production failures are monitored and fed back into development. Then run your own tests on cases the target did not choose. A curated demo proves very little; a regression suite that has caught real problems proves a lot. AI-written code inside the product raises its own maintainability questions, which we cover in our piece on AI and technical debt. For how to review a codebase written largely with AI coding tools, see our guide to due diligence on AI-generated code.
Do the Inference Economics Hold Up at Scale?
Model calls are a cost of goods sold, not an R&D line. Every query, document or agent step has a marginal cost, and that cost scales with usage in a way traditional SaaS hosting does not. A product with healthy unit economics at pilot volume can lose money at enterprise volume if each customer action triggers long prompts, large context windows or chains of agent calls.
Diligence should rebuild cost to serve per customer, per transaction or per seat, and test it against the growth plan. Check whether the target routes simple requests to cheaper models, caches repeated work and monitors spend by feature. Then run the sensitivity: what happens to gross margin if a provider changes its pricing, or if usage per customer doubles?
How Dependent Is the Company on One Model Vendor?
Model providers deprecate versions, change pricing, adjust rate limits and update usage policies. A target built tightly around one provider's specific model inherits all of that risk. Diligence should ask what breaks if the main model is retired or repriced, how long a switch to an alternative would take, and whether evaluation suites exist to prove the alternative works. An abstraction layer and a tested fallback are signs of a team that has thought about this. Prompts hard-tuned to a single model version are signs it has not.
What Is the Governance and Regulatory Exposure?
Regulation now applies to AI by use case. The EU AI Act entered into force on 1 August 2024. Prohibited practices have applied since 2 February 2025, obligations for general-purpose AI models since 2 August 2025, and most of the Act became applicable on 2 August 2026. According to the European Commission (opens in new tab), obligations for high-risk systems in sensitive areas such as employment, education and critical infrastructure now apply from 2 December 2027, and for high-risk AI built into regulated products from 2 August 2028. Under Article 99 (opens in new tab), fines for prohibited practices can reach €35 million or 7% of worldwide annual turnover, whichever is higher.
For a buyer, the questions are practical. Does the target sell into the EU or serve users there? Has it classified its systems by risk tier? Does it keep the technical documentation, logging and human oversight its tier requires? Sector rules add to this in financial services, healthcare and hiring, and several US states have passed their own AI laws. Our EU AI Act risk classification guide explains how the tiers work. Governance gaps should be costed like any other finding, because closing them takes engineering time.
Does the Team Have Real AI Depth?
AI capability lives in people as much as in code. Find out who designed the data pipelines, who runs evaluations, and who can explain why the system behaves the way it does. Look for an evidence culture: decisions backed by experiments, documented failures, and a habit of measuring before shipping. Then test concentration risk. If one or two people hold all the model knowledge, retention belongs in the deal terms and knowledge transfer belongs in the first 100 days.
It helps to benchmark the target against how strong AI engineering firms operate. The criteria in how to choose an AI software development company work just as well for judging a target's own team. For earlier-stage targets, pair this lens with our startup technical due diligence checklist.
AI Due Diligence Checklist for Investors
Use this table to scope the AI workstream or to pressure-test findings from a diligence provider.
| Area | Question to ask | Red flag |
|---|---|---|
| Capability | What would a competitor need to rebuild this? | The value sits in prompts and the interface |
| Data rights | Where did training data come from, and do the rights survive a change of control? | No data lineage; customer contracts silent on training use |
| Model rights | Which models are used, under what licence or API terms? | Licence terms restrict the intended commercial use |
| Evaluation | How is quality measured, and what blocks a release? | Curated demos only, no regression suite |
| Economics | What is cost to serve per customer at plan volume? | Gross margin falls as usage grows |
| Vendor dependency | What breaks if the main model is retired or repriced? | No fallback; prompts tuned to one model version |
| Governance | Which regulations apply, and where is the compliance file? | No risk classification or technical documentation |
| Team | Who can improve the system, and what if they leave? | One or two people hold all model knowledge |
The AI workstream sits inside a wider review of architecture, security, IP and infrastructure, covered in our full software due diligence checklist. For the methodology behind all of it, see the technical due diligence and code audit guide.
Frequently asked questions
Part of
Related posts

Choosing an AI development company? Learn the 5 signals of a genuine AI-native partner, red flags to avoid, and the questions to ask before you sign.

A practical tech due diligence checklist covering 250+ items across 11 areas: code quality, security, AI/ML, scalability, to prep for your next funding round.
