Sphere Partners
Senior engineer taking notes beside an open laptop while reviewing a codebase in an engineering hub

Diligence on AI-Generated Code: What Buyers Need to Check Before Acquiring an AI-Built Codebase

AI-built codebases can look clean and still hide licence gaps, unreviewed code, shallow tests and key-person risk. Here is what buyers should check, and how an AI code audit verifies it before close.

10 min read
In this article

Diligence on AI-generated code checks everything a normal technical review checks, plus how the code was produced. Before acquiring an AI-built codebase, buyers should verify licence and IP provenance, how much generated code was merged without real review, whether the tests actually prove behavior, common AI security patterns, hidden dependencies and whether anyone on the team can maintain the system. An AI code audit answers those questions from repository evidence, not from management's description.

Most engineering teams now use AI coding assistants, and that alone is not a problem. The problem is that a codebase written quickly by agents can look clean and still carry risk that only shows up after close, when the buyer's team tries to change it. This article covers the six risks to check, how an audit verifies each one, what separates a red flag from normal AI-assisted development, and what to ask the seller before the work starts. If the AI is the product rather than the tool that built it, our guide to AI due diligence covers model provenance, data rights and AI economics.

Why Does an AI-Built Codebase Need Different Diligence?

Traditional code review assumes a human wrote each change and another human checked it. AI-assisted development breaks the first assumption and often weakens the second. Code can now be produced faster than any team can read it, so review coverage becomes uneven: careful in some modules, a rubber stamp in others.

The output also hides its own weaknesses. Generated code is usually well formatted, consistently named and plausibly commented. It reads like senior work, which is exactly why a quick skim during diligence tells you very little.

Independent research points the same way. Veracode's 2025 GenAI Code Security Report (opens in new tab) tested more than 100 large language models on 80 coding tasks and found that the generated code introduced security flaws in 45% of cases. GitClear's analysis of 211 million lines of code (opens in new tab) found that 2024 was the first year in which copy-pasted lines outnumbered moved (refactored) lines, and that code blocks with five or more duplicated lines rose roughly eightfold that year.

None of this makes an AI-built product a bad acquisition. It means the code audit has to examine the build process as well as the result.

What Should Buyers Check in an AI-Generated Codebase?

Six risks account for most of what goes wrong. The table summarizes each risk, how an audit verifies it and what to request from the seller.

RiskHow an audit verifies itEvidence to request
Licence and IP provenanceLicence scan of dependencies, snippet matching against open-source code, review of AI tool termsAI tool list and plans, usage policy, SBOM or dependency manifests
Duplicated or unreviewed codeDuplication analysis, pull request review metrics, churn on recently generated modulesRepository access with full history, PR and review records
Missing or shallow testsCoverage plus test-quality checks such as mutation testing on critical pathsCI configuration, coverage reports, recent failed builds
Security patternsStatic and dependency scanning, manual review of auth, input handling and secretsScanner output, pen test reports, incident log
Hidden dependenciesTransitive dependency review, abandoned or unknown packages, embedded third-party and model APIsPackage lock files, vendor and API contracts, cloud and API bills
Knowledge concentrationCommit authorship distribution, module walkthroughs with the engineers who own themOrg chart, on-call rota, architecture and runbook docs

Licence and IP provenance

Two questions matter. First, does the generated code reproduce open-source code under a licence the buyer cannot accept, such as a copyleft licence in a proprietary core? An audit runs licence scanning on dependencies and snippet matching on the source itself, because generated code rarely arrives with an attribution header.

Second, how much of the code is protectable? In January 2025 the US Copyright Office concluded (opens in new tab) that AI output is copyrightable only where a human determined sufficient expressive elements, and that prompts alone are not enough. Human selection, arrangement and modification still count. For a buyer, that makes the review and editing history part of the IP story. Counsel should own the legal conclusion; the audit supplies the facts about how the code was produced and reviewed.

Duplicated or unreviewed code

Assistants are good at producing a new function and poor at noticing that a similar one already exists. The result is the same business rule implemented three slightly different ways, which works until someone changes one copy. An audit measures duplication directly and checks review records: how many merged pull requests had substantive comments, how large they were, and how quickly they were approved. A 2,000-line pull request approved in four minutes is a review record in name only.

Missing or shallow tests

Generated tests often assert what the code currently does rather than what it should do, so they pass and prove little. Coverage percentages can look healthy while critical paths are barely exercised. Our piece on why "it looks right" is not a test strategy goes deeper; in diligence, the audit samples tests on revenue-critical flows and checks whether they would fail if the logic broke.

Security patterns

AI-generated code tends to repeat a recognizable set of weaknesses: missing authorization checks on internal endpoints, unvalidated input, permissive CORS settings, hard-coded secrets and outdated cryptography. Automated scanning catches some of these. Context-dependent flaws, such as whether a user should be able to reach a given record at all, need a person who understands the application.

Hidden dependencies

Assistants add packages freely, sometimes outdated, sometimes barely maintained, and occasionally ones that do not exist until someone registers the name. Products built with AI also tend to call model APIs directly, which creates cost, rate-limit and data-handling exposure that may not appear in the architecture deck. The audit maps every external dependency and every outbound API call, then ties them to contracts and bills.

Knowledge concentration

When one or two people prompted most of a system into existence, they may be the only ones who understand why it works. That is a classic key-person risk with a new cause. The audit checks commit authorship, asks owners to walk through specific modules and looks for documentation that would let a new team take over.

How Does an AI Code Audit Verify These Risks?

An AI code audit follows the same discipline as any code audit, with extra attention to the build process. The sequence below is how a buyer-side audit usually runs.

How a buyer-side AI code audit runs

  1. 1Scope and accessAgree the repositories, environments and decision at stake. Get read-only access to full history, not a snapshot.
  2. 2Build-process reviewWhich AI tools were used, under what terms and policy, and where human review gates sit in the workflow.
  3. 3Automated analysisLicence and dependency scans, static security analysis, duplication and churn metrics across the codebase.
  4. 4Targeted manual reviewSenior engineers read the high-risk modules: auth, payments, data access, integrations and anything with weak review records.
  5. 5Change testMake small, realistic changes to critical modules and see what breaks. Fragility under change is the risk unique to fast-generated code.
  6. 6Costed findings and roadmapEvery finding tied to a file or dependency, with severity, remediation effort and a recommended fix.

Timelines depend on scope. Sphere's code audit service runs from a one-week executive scorecard to a four-week deep dive with prioritized findings and dollar and time estimates for remediation, which fits inside most deal timelines.

Where the target handles regulated data, security deserves its own workstream. In Sphere's diligence for Marble Financial's acquisition of Inverite, the team reviewed more than 15 security control domains in three weeks, giving the acquirer a quantified list of security and compliance gaps to take into closing negotiations.

Is AI-Generated Code a Red Flag in an Acquisition?

Not by itself. Heavy AI use with clear review gates, real tests and scanning in CI is simply modern engineering, and often a sign of a productive team. Our conversation on how AI is changing technical debt covers why the upside is real.

The red flags are about missing controls, not the tools:

  • No record of who reviewed generated code, or review that is clearly perfunctory on large changes.
  • No AI usage policy, or developers using personal accounts on tools whose terms allow training on submitted code.
  • Copyleft-licensed code or unexplained snippets inside the proprietary core.
  • Secrets, customer data or credentials found in prompts, logs or the repository.
  • Core modules that nobody on the current team can explain without re-reading them.

Each of these should appear in the technical due diligence report with a severity and a remediation cost, so the deal team can decide whether it belongs in the price, the purchase agreement or the 100-day plan.

What Should You Ask the Seller Before the Audit?

A short questionnaire up front saves days of the audit itself:

  1. Which AI coding tools do engineers use, on which plans, and under what data-retention terms?
  2. Is there a written AI usage policy, and how is it enforced?
  3. What review is required before generated code is merged, and is that enforced in the repository settings?
  4. Which parts of the product call external model APIs, and what do those calls cost each month?
  5. Who owns each core module today, and who else could maintain it?
  6. Has any licence, security or data incident involved generated code?

Sellers can run the same list on themselves before going to market; our guide to sell-side technical due diligence covers the rest of that preparation.

If you are on the building side rather than the buying side, the same questions are a useful standard for your own team. Our guide to AI app development services and how to choose a partner covers how to hold a delivery partner to it.

For the wider picture of scope, process and partner selection, see the technical due diligence and code audit guide.

Frequently asked questions

An AI code audit is an independent review of a codebase that was partly or largely written with AI coding tools. It covers the usual areas, such as security, architecture, dependencies and tests, and adds checks on how the code was produced: review coverage, provenance and licensing, duplication and maintainability.

In the US, the Copyright Office's 2025 guidance says AI output is protectable only where a human determined sufficient expressive elements; prompts alone are not enough, while human selection, arrangement and modification can be. Buyers should have counsel assess the IP position using the audit's evidence.

Common issues include missing authorization checks, unvalidated input, hard-coded secrets, outdated cryptography and risky or unnecessary dependencies. Veracode's 2025 research found security flaws in 45% of AI-generated code samples tested, so scanning plus manual review of sensitive modules is essential.

No. AI code review tools are useful inside a team's workflow, but an acquisition needs an independent view that covers process, licensing, key-person risk and business impact, with findings costed for the deal team. Tool output is one input to the audit, not a substitute for it.

A high-level scorecard can take about a week. A deep-dive audit with prioritized findings and remediation estimates typically takes around four weeks, and multi-repository or regulated environments can take longer.

Know what's in an AI-built codebase before you buy it

Sphere's code audit traces every finding to a real file, function or dependency, with a fix and a cost. Results in 1–4 weeks.

Related posts

How AI Is Transforming Tech Debt, Data Modernization, and the Future of Engineering — Insights from Alex Ter-Zakhariants

In this episode of SphereCast, Field CTO Alex Ter-Zakhariants breaks down what engineering teams actually face when bringing AI into real systems: tech debt, disorganized data, and infrastructure that wasn’t built to scale. From data modernization to AIOps to AI copilots, Alex shares a practical roadmap for building systems that can adapt, not just react. No hype, no shortcuts—just clear thinking about what makes engineering work in an AI-driven world.

We'd love to hear from you!

Please provide your contact details, and our team will get back to you promptly.