By CrossBorder IP · Published August 4, 2026
AI and IP is the most searched, most misunderstood corner of intellectual property right now — and after a run of landmark decisions in 2025 and 2026, the legal picture is finally sharp enough to act on. The headlines make it sound chaotic. It is not. A clear principle has emerged from the case law, and once you understand it, the compliance steps for your business become obvious.
This guide translates the key rulings into plain terms and gives you a checklist for using AI without walking into an IP problem — whether you are training models, building AI features, or just generating content with off-the-shelf tools.
In the US, training an AI model on copyrighted works is likely to be treated as fair use IF the copies were obtained legally. Downloading the same works from pirate sources is not. How you source the data matters more than the act of training itself.
That single distinction — lawful sourcing versus piracy — explains most of what you have read in the news. The courts have been drawing a consistent line between authorized copies and pirated ones, and the outcomes turn heavily on which side of that line a company was standing on. The picture that has emerged since mid-2025 is more nuanced than either side originally claimed: it is not that AI training is automatically legal, nor that it is automatically theft. Provenance is the pivot.
In the closely watched Bartz v. Anthropic litigation, the court found that training on copyrighted books could qualify as fair use, while separately holding that downloading and storing pirated copies of those books was not protected. That second finding drove a settlement reported at roughly 1.5 billion dollars — on the order of 3,000 dollars per work across a very large set of titles — described as the largest copyright recovery of its kind. Notably, the settlement resolved past claims only; it did not license future training or cover model outputs.
(For transparency: Anthropic is the maker of the AI assistant used to help draft parts of this article. The facts stated here are matters of public record.)
A parallel case against Meta over training its Llama model reached a similar fair-use conclusion on the training question, while claims tied to the acquisition of pirated copies during the data-gathering process remained live. The pattern repeats: the training step tends to survive scrutiny; the sourcing step is where liability lives.
The training question is only half the story. A separate line of cases concerns whether AI outputs themselves infringe — for example, litigation between major news organizations and AI developers over both training and generated outputs, where courts have allowed discovery to proceed broadly. The unsettled question of output liability is why indemnification terms in your AI contracts matter so much (more on that below).
On the ownership side, the human-authorship requirement held firm. The US Supreme Court declined to take up Thaler v. Perlmutter in March 2026, leaving in place the refusal to register a work generated autonomously by an AI system with no human author. Separately, the Copyright Office has signaled that fair use in AI training is highly fact-specific, and that human involvement remains the touchstone for whether an AI-assisted work can be registered at all.
The practical takeaway: to claim copyright in something made with AI, a human being needs to have contributed enough creative choice to be considered the author. Purely machine-generated output, with no meaningful human authorship, may not be protectable — which affects what you can own, license, and enforce.
These are US outcomes. Other jurisdictions are drawing their own lines, and they do not all match. Most countries require substantial human involvement for an AI-assisted work to be eligible for protection, but the treatment of training data, text-and-data-mining exceptions, and output liability varies widely. If you operate internationally, do not assume a US fair-use analysis travels — your exposure can look very different in Europe or Asia. This is precisely the kind of cross-border gap that catches companies off guard.
Whether you are a SaaS company shipping AI features, a startup using AI internally, or a brand generating marketing assets with AI tools, three exposures deserve attention.
Run this against your own operations. Most companies can close the biggest gaps in a single focused pass.
Pro tip: Treat data provenance like a chain of title. If you cannot show where a dataset came from and that you had the right to use it, assume a counterparty in your next deal or dispute will ask — and plan accordingly. Buyers in M&A now probe exactly this.
The AI-IP landscape is less a minefield than a set of clear lines. Source your training data lawfully. Keep a record of the humans behind anything you want to own. Put the right terms — especially indemnities — in your contracts. And do not assume US outcomes govern your exposure abroad. Companies that do these things are in a strong position; the ones that get caught out are almost always the ones that could not answer a simple question about where their data or their authorship came from. If you want a structured review of your AI-related IP exposure, that is work we handle for clients regularly.
Ready to protect your IP?
Book a free 15-minute strategy call with Cameron Reid.
Book a Free Strategy CallAbout the Author
Cameron Reid is the cofounder of CrossBorder IP, where he advises SaaS companies, tech startups, e-commerce brands, and in-house legal teams on international IP strategy. With over 20 years of experience spanning Big Law, in-house counsel roles, and startup advisory, Cameron specialises in helping businesses protect and scale their IP globally — particularly across the US, Europe, and Asia-Pacific markets.
Disclaimer: This article provides general information about IP strategy and should not be relied upon as legal advice. IP laws vary significantly by jurisdiction and every business situation is unique. Consult qualified counsel about your specific circumstances.