StandpointInternational Business Law
Data and AI in contracts: what may your counterparty do with your data?
Data license, machinery supply contract, SaaS: why control over your data hinges on the line between adaptation and transformation, and how the contract captures AI training and competing products.

In 1981 Lynn Goldsmith photographs the musician Prince. Three years later she licenses the photo to a magazine for 400 dollars, as an “artist reference” for an illustration, expressly for one-time use. The illustrator the magazine hires is Andy Warhol. He delivers more than the commissioned illustration; the photo becomes 16 works. When the Warhol Foundation licenses one of them to Condé Nast for 10,000 dollars in 2016, Goldsmith learns of the series for the first time. In 2023 the US Supreme Court rules in her favor: the one-time reference license did not cover all that, and no statutory limitation replaced it.1
This is not art history. It is the base case of every data contract: one side hands over material, the other makes more of it than was agreed, and the conflict surfaces only years later. With AI the case accelerates – what used to take an artist’s lifetime, a model does overnight: enrich, rebuild, pour into new products.
The setting: your data, their model
The pure case is the data license. A provider licenses market data, maps, or industry figures; the licensee enriches them, builds products, and licenses on to end customers. The real subject of negotiation is what the licensee, its affiliates, and its sublicensees may do: only use, or also modify, blend, analyze with AI, and pass on?
The same questions now sit in contracts nobody calls a license. In the machinery supply contract, the equipment generates operating and sensor data, and the manufacturer wants them for maintenance, product improvement, and the training of its models. In the AI or SaaS service, the provider wants to keep inputs and outputs to improve its systems. In the development contract, both sides feed in data and later dispute who owns the enriched version.
And the roles flip. In the data license, the data user is the customer. In the machinery purchase, it is the supplier. With SaaS, the provider. Whether the contract calls someone “buyer” or “seller” says nothing about who is using whose data; all that counts is the direction of the data flow. My first question in every negotiation is therefore the same: which data are you handing over, and what may the other side do with them?
Since the Data Act, this is no longer a matter of negotiation alone
For connected products the starting position has shifted. The Data Act has applied since 12 September 2025 and inverts the default: connected products must be designed so that the product data is accessible to the user easily, securely, free of charge, and in a machine-readable format by default (Article 3(1)); where the user cannot access the data directly, the data holder must make it available without undue delay (Article 4(1)). Even before the sale, rental, or lease, the supplier must disclose what data the machine generates, whether continuously and in real time, where it is stored, and how the user gets to it (Article 3(2)).
More important for drafting is Chapter IV. It subjects clauses on data access and data use that one company imposes unilaterally on another to a fairness review of their own: if they are unfair, they do not bind (Article 13(1)). Three points hit the data owner’s toolkit squarely. A clause is unfair where it gives the imposing party access to and use of the other side’s data in a way that seriously harms that side’s legitimate interests, in particular where trade secrets or protected rights are involved (Article 13(5)(b)). A clause is equally unfair where it prevents the other side from making reasonable use of the data it provided or generated (point (c)), and where it denies that side a copy of that data during the term or within a reasonable period after it (point (e)).
The lever is unilaterality: a clause is unilaterally imposed where one party supplies it and the other cannot influence its content despite trying to negotiate – and the burden of proving otherwise falls on the party that supplied it (Article 13(6)).2 Present your data clauses as a non-negotiable block, then, and you lose twice in a dispute: on the review and on the burden of proof. Why the counter-clause, that everything was individually negotiated, does not save it, I set out in the article “No individual agreement”.
The line on which control hangs
Where that line runs, I show by the example of US copyright, the venue of the current AI litigation. As long as the counterparty adapts your data, you as the data owner hold the longer lever. The adaptation (derivative work) is, in US law, the author’s exclusive right (17 U.S.C. § 106(2)): whatever the contract does not permit stays reserved, and the result of an unlawful adaptation does not even enjoy protection of its own (17 U.S.C. § 103(a)).3 The danger lies in transformation: if the use tips into transformative use, the fair-use limitation applies and copyright control ends. It is exactly this line that the Supreme Court shifted in favor of rights holders in Warhol v. Goldsmith: the transformation needed for transformative use must go beyond what already makes a work derivative, and where both works serve the same commercial purpose, a new message does not help. I set out the doctrine in detail in the reference article “IP clauses in contracts: who owns what, and who may use it”.
For the contract this means two things. First: “modifications” is the most contested word of the data contract. Define it, and allocate the modifications expressly, through reassignment, license, or prohibition. Be careful with the word ownership: German law knows no ownership of data, and the debate about it has expressly run the other way.4 What gets allocated are powers of use, duties to hand over, and duties to delete – and only the contract carries them. Second, the honest limit: the contract binds only the parties. Against your own counterparty, a modification bar works beyond the copyright line as a contractual duty; the chain of affiliates, end customers, and sublicensees is reached only if the contract binds it expressly. In practice that means flow-down: the data user must impose restrictions on its customers that are no less restrictive than its own, up to and including bans on compiling competing databases from the data or publishing representations that enable third parties to scrape them.
How much these clauses are needed shows in how far the statute alone carries: not far. The sui generis database right bites on distributed access only where the users act in conscious and deliberate concert, and repeated retrieval of insubstantial parts does not infringe it as long as it is not aimed at reconstituting the database. And in the scraping case, even a contractual bar on automated retrieval was not enough for the Federal Court of Justice to treat the querying of freely accessible data as unfair obstruction.5
Training and competing products: the new core
The sharpest AI question is not enrichment but training. A model, once trained, preserves the value of your data permanently, even after termination and deletion of the raw data. And the US courts showed in 2025 that copyright is no reliable brake here:6
| Decision | Outcome | Key point |
|---|---|---|
| Thomson Reuters v. Ross (D. Del., 11 Feb 2025) | fair use denied | AI training for a directly competing product, same market purpose |
| Bartz v. Anthropic (N.D. Cal., 23 June 2025) | training transformative, pirated library infringes | fair use does not cure unlawful sourcing; later settled |
| Kadrey v. Meta (N.D. Cal., 25 June 2025) | fair use only on this record | express warning about market dilution, the fourth factor |
This line is first-instance and fact-dependent, ended in one case by settlement, so it is not settled law. But one thing Thomson Reuters v. Ross shows clearly: fair use fails most readily where the data are used to train a competing product, for the same market purpose. That is also where the contractual line has settled in large data license agreements: whether training is permitted is negotiable. What should not be negotiable is the bar that trained systems must not produce products that compete with the data owner’s data.
In Europe the lever sits elsewhere, and it is a formality with teeth: text and data mining is permitted as long as the rightholder does not expressly reserve it (Section 44b(2) and (3) UrhG). How much that reservation carries was indicated by the Regional Court of Hamburg in 2024, without having to decide it: the claim failed on the research exception, yet the court thought it likely that a licensee too can declare the reservation and that a reservation in natural language is machine-readable enough.7 The judgment is not final, and those passages do not carry the decision. For the contract something solid follows all the same: declare the reservation, oblige your counterparty to carry it forward on every onward transfer, and set it technically as well – precisely because the law is unsettled.
What does carry are two other routes, and the contract has to support both. Machine-generated data can be protected as a trade secret, but only where it is secured by reasonable steps to keep it secret – access restrictions, marking, confidentiality duties down the chain. And a data pool can attract a database right of its own where the investment in obtaining, verifying, and presenting it is substantial.8 Neither arises by itself: the protection hangs on measures that have to be in the contract.
The clause map
Both roles need the same catalog, with the signs reversed.
As the data owner:
- Define the permitted uses exhaustively; whatever is not expressly permitted stays reserved.
- Define “modifications” and allocate them: reassignment, license, or prohibition. Not through ownership of data, but through powers of use, handover, and deletion.
- Address AI training expressly: prohibit it, or permit it as a separate, separately priced license track, always with a bar on competing products. Declare the reservation under Section 44b(3) UrhG and oblige your counterparty to pass it on.
- Bind the chain: affiliates, sublicensees, and end customers (flow-down no less restrictive than the main contract), plus audit and reporting rights.
- Settle the end of the contract: deletion duties that cover modifications, and a clear end to further use in products, at most with a defined tail period.
As the data user:
- Does the license cover the plan: use, modification, AI analysis, future use cases? Purpose-bound licenses end sooner than roadmaps do.
- Secure ownership of your own value added: US law protects only what you added yourself, and only where the adaptation was licensed (17 U.S.C. § 103(b)). With AI-generated enrichment comes the additional protection gap of missing human authorship, which only the contract closes; German law faces the same question for inventions, where according to the Federal Court of Justice a human inventor nonetheless remains the rule.9
- Provenance and indemnity: warranties on the chain of rights in the sourced data and an indemnity for infringement cases; how the indemnifying party confines it, I set out in the article “Indemnity: how the supplier can reduce its liability”.
Two model building blocks that carry the core. The highlighted terms are placeholders for the data-providing and the data-receiving party; replace them with the contract’s own designations.
Conclusion
Lynn Goldsmith received 400 dollars and fought her control back before the Supreme Court decades later. Do not rely on that route: it takes years, it is expensive, and on the near side of the transformation line it is open to you only if the contract is clean. Whoever hands over data without settling training, modification, and the chain is not licensing their product – they are licensing their successor.
Notes
-
Andy Warhol Foundation for the Visual Arts, Inc. v. Goldsmith, 598 U.S. 508 (2023); the details of the 1984 license per the syllabus of the decision. ↩
-
Regulation (EU) 2023/2854 (Data Act), OJ L of 22 December 2023; Articles 3(1) and (2), 4(1), 13(1), (5), and (6), and 50 (application from 12 September 2025). For context Hennemann/Steinrötter, Der Data Act, NJW 2024, 1. ↩
-
17 U.S.C. §§ 101, 103, 106, 107; foundational Campbell v. Acuff-Rose Music, Inc., 510 U.S. 569 (1994). ↩
-
On the rejection of data ownership Hoeren, Datenbesitz statt Dateneigentum, MMR 2019, 5; Kühling/Sackmann, Irrweg “Dateneigentum”, ZD 2020, 24. ↩
-
BGH, judgment of 22 June 2011 – I ZR 159/10 (Automobil-Onlinebörse), headnotes 1 and 2 on Section 87b(1) UrhG; BGH, judgment of 30 April 2014 – I ZR 224/12 (Flugvermittlung im Internet), on the fair-trading assessment of automated retrieval despite contrary terms of use; set aside and remanded there. ↩
-
Thomson Reuters Enterprise Centre GmbH v. Ross Intelligence Inc. (D. Del., 11 February 2025); Bartz v. Anthropic PBC (N.D. Cal., 23 June 2025), settlement announced in September 2025; Kadrey v. Meta Platforms, Inc. (N.D. Cal., 25 June 2025). ↩
-
LG Hamburg, judgment of 27 September 2024 – 310 O 227/23, not final; the claim failed on Section 60d UrhG, and the passages on the reservation under Section 44b(3) UrhG do not carry the decision. ↩
-
On the protection of machine-generated data under the German Trade Secrets Act Hessel/Leffer, MMR 2020, 647; on the data pool as a trade secret Krüger/Wiencke/Koch, GRUR 2020, 578; on the sui generis database right Sections 87a et seq. UrhG. ↩
-
Thaler v. Perlmutter, 130 F.4th 1039 (D.C. Cir. 2025); U.S. Copyright Office, Copyright and Artificial Intelligence, Part 2: Copyrightability (January 2025); on the human inventor where AI is used Gotthardt, Regelungen des Einsatzes von künstlicher Intelligenz in Forschungs- und Entwicklungsverträgen, GRUR 2024, 1870. ↩
Reference: Poleacov, P. (2026). Data and AI in contracts: what may your counterparty do with your data?. INN.LAW. https://inn.law/en/perspectives/data-ai-contracts/