AI IP due diligence: training data, model ownership and licences

AI IP due diligence is the review an investor or buyer carries out before backing or acquiring an AI company: where the training data came from and on what terms, who owns the model, code and weights, which licences the business depends on, whether its EU AI Act documentation is in order and which claims are pending. In AI deals, the value often sits in assets that ordinary IP checklists miss, such as datasets and model weights. This guide is for investors, corporate buyers and founders preparing for a funding round or sale.

Key takeaways

  • Start with training data provenance: sources, licences, scraping practices and respect for opt-outs.
  • Providers of general-purpose AI models must keep technical documentation, have a copyright policy and publish a training content summary under Article 53 of the AI Act; ask to see all three.
  • The European Commission can fine providers of general-purpose AI models up to 3% of worldwide turnover or EUR 15 million, whichever is higher, and its enforcement powers apply from August 2026.
  • Check that code, weights and datasets were actually assigned to the company by employees, contractors and founders.
  • Pending AI copyright cases in Europe show the litigation risk is real; price it into warranties, indemnities or escrow.

Why is AI IP due diligence different from a standard IP review?

A traditional review checks registered rights: patents, trade marks, domain names. In an AI company, the core assets are often unregistered: a curated dataset, model weights, training code and know-how. Their value depends on how they were built. A model trained on data the company had no right to use can carry claims that follow it into the buyer’s hands.

The law is also moving. In Germany, the Munich Regional Court I held on 11 November 2025 (GEMA v OpenAI) and 31 July 2026 (GEMA v Suno) that works memorised in a model are reproduced without authorisation; both are first-instance rulings. And from 2 August 2025, the EU AI Act imposed specific obligations on providers of general-purpose AI models, with a transition until 2 August 2027 for models placed on the market before that date (Article 111(3)).

Training data: what to ask for

  • A dataset inventory: each source, its volume, when it was collected and on what legal basis (licence, own data, public domain, text and data mining exception).
  • Copies of data licences, checking that they expressly cover training, fine-tuning and any retrieval use, for the right models, territory and term.
  • Scraping practices: crawler names, whether robots.txt and other machine-readable reservations were respected, and whether any paywalls or technical protections were bypassed.
  • Removal and complaint records: requests from rights holders and how they were handled.
  • Personal data in training sets, which raises separate GDPR questions.

The legal backdrop is Article 4 of the DSM Directive: commercial text and data mining is permitted only where rights holders have not reserved their works. The copyright chapter of the General-Purpose AI Code of Practice (10 July 2025) commits signatories to follow robots.txt, not to circumvent effective technological measures and to exclude websites recognised as infringing on a commercial scale. Whether the target signed the Code, or how else it shows compliance, is a useful first question.

The AI Act Article 53 file

Under Article 53(1) of the AI Act, providers of general-purpose AI models must keep up-to-date technical documentation (a), provide information to downstream providers that integrate the model (b), put in place a policy to comply with EU copyright law, including identifying and respecting reservations of rights (c), and publish a sufficiently detailed summary of training content using the AI Office template (d). Models released under a free and open-source licence with public weights are exempt from (a) and (b) unless they present systemic risk; (c) and (d) apply to all.

The Commission published the template for the training content summary on 24 July 2025. According to the Commission’s follow-up to the Parliament’s resolution on copyright and AI, the summary covers public and private datasets, large public datasets and scraped data, including crawler names, collection period and the top 10% of domains scraped (top 5% or 1,000 domains for SMEs). The published summary is the easiest way to test whether what the target tells you matches what it told the market.

Area Documents to request Red flags
Training data Dataset inventory, data licences, crawler logs, opt-out policy No inventory; licences silent on AI; paywalled or pirated sources
AI Act compliance Article 53 technical documentation, copyright policy, public training summary No summary for a model on the EU market; policy exists only on paper
Model and code ownership Employment and contractor agreements, founder assignments, open-source register Contractors without written assignment; licence terms restricting commercial use of a base model
Inbound and outbound licences Upstream model or API terms, customer contracts, indemnities granted Uncapped IP indemnities to customers; upstream terms prohibiting competing models
Disputes Claims, warning letters, takedown and complaint logs Unreported complaints; outputs reproducing third-party works
Trade secrets and registered IP Confidentiality measures over weights and data; patents and trade marks Weights shared without protection; AI named as inventor

Who owns the model, the code and the weights?

Ownership follows the people who built the assets. In Spain, the Intellectual Property Act gives the employer the exploitation rights in software created by an employee in the course of their duties, unless agreed otherwise (Article 97(4)). That rule does not cover freelancers or founders working before incorporation: their contributions need a written assignment (Articles 43 and 45), limited to the uses it expressly names. Latin American laws have their own rules, so each contributor’s country matters.

Model weights and curated datasets are usually protected as trade secrets rather than by registration. Under Spain’s Trade Secrets Act (Law 1/2019, Article 1), information only qualifies if its holder took reasonable measures to keep it secret, so access controls and confidentiality agreements are part of the diligence. A substantial investment in a database may also attract the sui generis database right (Article 133 of the Intellectual Property Act). On patents, the EPO Legal Board of Appeal held in J 8/20 (21 December 2021) that a machine cannot be an inventor, so filings must name the human inventors.

What this means for your business

  1. Investors and buyers: send an AI-specific request list early and give the target time to build a dataset inventory if it has none.
  2. Translate findings into the deal: specific warranties on data provenance and AI Act compliance, indemnities for identified claims, escrow or price adjustments.
  3. Founders: prepare the data room before the round, with licences, assignments and the Article 53 file ready.
  4. Cross-border targets: a Spanish start-up with Latin American developers or Brazilian datasets needs each chain of title checked under local law.

Our team for AI and digital asset IP runs the AI-specific workstream, and our cross-border IP due diligence and valuation service integrates it with the wider transaction.

Where AI deals go wrong on IP

  • Relying on the founders’ word about the data. Without an inventory and licences, provenance cannot be verified, and warranties only shift the risk.
  • Missing contractor assignments. Code and weights built by freelancers may not belong to the company at all.
  • Ignoring upstream terms. A model built on another provider’s base model or API inherits its licence restrictions.
  • Treating the AI Act as a later problem. Enforcement powers over general-purpose AI providers apply from August 2026, and fines reach 3% of worldwide turnover or EUR 15 million (Article 101).
  • Pricing litigation risk at zero. Claims such as those brought by GEMA show that rights holders do sue, and that courts may scrutinise what a model memorises.

Frequently asked questions

What should AI IP due diligence cover?

At a minimum: the provenance and licensing of training data, ownership of code, weights and datasets, inbound licences such as base models and APIs, outbound commitments to customers, AI Act documentation for general-purpose models, trade secret protection and pending claims. Findings should feed directly into warranties, indemnities and price.

Does the AI Act apply to a start-up we are buying outside the EU?

Under Article 2(1)(a), the AI Act applies to providers placing general-purpose AI models on the EU market, whether established in the EU or in a third country. A Latin American or US target serving EU customers with its own model may therefore need the technical documentation, copyright policy and training summary, and a buyer inherits any gap.

Who owns code written by contractors for an AI company?

Generally the contractor, unless there is a valid written assignment. In Spain, the employer only owns software created by employees in the course of their duties automatically; freelancers and pre-incorporation founders must assign their rights in writing, specifying the uses transferred.

Can IP Global Guard run IP due diligence on an AI company across several countries?

Yes. We build the request list, review data licences, assignments, AI Act documentation and disputes, and report findings in deal terms. Where the target or its data sit in Latin America or Africa, we coordinate local correspondents, and the META Channel group covers AI Act and GDPR compliance.

How IP Global Guard supports your AI transaction

In an AI deal, the questions that decide value are often where the data came from and who really owns the model. IP Global Guard, the IP services line of META Channel Corporation Limited, carries out due diligence, valuation and portfolio structuring across more than 25 jurisdictions in Europe, Latin America and Africa, from a single point of contact.

Share the deal timeline, the target’s main models and products and the countries where its team and data sit. We will propose a focused review plan and the documents to request first. Contact our transactions team.

This article is general information, not legal advice, and does not replace a review of the specific transaction.

Sources