AI Policy, Law & Safety · AI Copyright & Intellectual Property
Is It Legal to Train AI Models on Copyrighted Books and Articles?
This is genuinely unsettled: AI companies argue that training on copyrighted books and articles qualifies as fair use, while authors and publishers have filed lawsuits arguing it constitutes copyright infringement, and courts are actively working through the question with no single, final, universal answer yet.
Legal disclaimer
This page provides general information only and is not legal advice. Laws vary by jurisdiction and change over time. Consult a licensed attorney in your jurisdiction before making decisions based on this content.
Key takeaways
- Multiple authors, publishers, and news organizations have filed lawsuits against AI companies alleging unauthorized use of copyrighted works to train AI models.
- AI companies commonly defend the practice as fair use, arguing that training involves transformative use of the material rather than reproducing it for the same purpose as the original.
- Rights holders argue that training on their work without permission or compensation undermines the market for and value of their original content.
- Courts have been actively considering these arguments, and outcomes have varied by case and by the specific facts involved, without a single definitive resolution covering all AI training practices.
- Some AI companies have separately pursued licensing deals with publishers and content owners, suggesting the industry sees value in agreements even amid legal uncertainty.
An Open Legal Question, Not a Settled Answer
There isn’t currently a single, clean “yes” or “no” answer to whether training AI models on copyrighted books and articles is legal. Instead, this is one of the most actively contested legal questions in the entire AI industry, playing out through numerous lawsuits filed by authors, publishers, news organizations, and other rights holders against AI developers. The legal system is still working through how existing copyright doctrine — much of it developed long before large-scale machine learning existed — applies to the practice of using vast quantities of text to train a model.
Anyone looking for a definitive, universal answer right now is likely to be disappointed, because the honest state of affairs is ongoing litigation and genuine disagreement among courts, rights holders, and AI companies about where the law should land.
The Core Argument on Each Side
AI companies training on copyrighted material generally lean on the doctrine of fair use, a long-standing part of copyright law that permits certain uses of copyrighted work without permission under specific circumstances — commonly considering factors like whether the use is transformative, the nature of the copyrighted work, how much of it was used, and the effect on the market for the original. AI developers typically argue that training is transformative: the model isn’t republishing the book or article for readers to consume in place of the original, but rather analyzing statistical patterns across huge datasets to build a general-purpose system, which they argue is a fundamentally different purpose from the original work’s.
Authors, publishers, and news organizations pushing back generally argue the opposite on several of those same factors: that copyrighted works were copied in their entirety as part of the training process, that this was done without permission or compensation, and that AI systems trained this way can go on to generate output that competes with or substitutes for the very content used to train them — potentially harming the market for original works, which is one of the specific factors fair use analysis considers. Some rights holders have also raised concerns about how training data was obtained in the first place, including whether copyrighted works were used from sources without proper licensing.
Why the Outcome Is Genuinely Uncertain
This dispute is difficult to resolve cleanly because it doesn’t map neatly onto earlier fair use precedents, which were developed around more familiar situations like search engine indexing, parody, or academic quotation. Training a large language model is a distinct technical process, and courts are having to apply decades-old legal tests to a use case that didn’t exist when those tests were formulated. Different courts examining different sets of facts — different AI companies, different datasets, different alleged uses of the resulting models — have reached varying conclusions and are still working through appeals and related cases, meaning there is no single ruling that has settled the question industry-wide.
Adding another layer, some AI companies have separately chosen to pursue licensing agreements directly with publishers and content owners, paying for the right to use their material for training. This suggests that even amid legal uncertainty, parts of the industry see commercial and legal value in securing permission rather than relying solely on a fair use defense.
Bottom Line
Whether training AI models on copyrighted books and articles is legal remains a genuinely open, actively litigated question rather than a settled matter, with fair use arguments cutting both ways depending on the specific facts of each case. Anyone directly affected — as a rights holder or an AI developer — should follow current case developments closely and consult legal counsel rather than assume a fixed answer.
Go deeper
Important caveats
- This is one of the most actively litigated areas of AI law, and the legal landscape can shift significantly as courts issue new rulings.
- This is general information, not legal advice; rights holders or AI developers with a specific dispute should consult an attorney rather than rely on general industry trends.
Frequently asked questions
Have any courts ruled definitively on whether AI training is fair use?
Various courts have issued rulings and decisions touching on aspects of this question, but the legal landscape remains actively contested and evolving across multiple ongoing cases, rather than settled by one single definitive, universally applicable ruling.
Why do AI companies argue training on copyrighted material is fair use?
They generally argue that training a model involves analyzing patterns across vast amounts of text for a transformative purpose — building a statistical model — rather than republishing or competing directly with the original works themselves.
Why do authors and publishers push back against this argument?
They generally argue that their copyrighted works were used without permission or payment to build a commercial product, and that AI-generated output can end up competing with or diminishing demand for their original work.
Related questions
- What Is 'Fair Use' and How Does It Apply to AI Training Data?
- Can ai companies be compelled to disclose their training data sources?
- Can You Copyright Something an AI Helped You Write?
- Who Owns the Output of an AI Image Generator?
- Can an AI Be Listed as an Inventor on a Patent?
- What Legal Cases Have Artists Filed Against AI Companies?
Sources
- [1]US Copyright Office — United States Copyright Office
- [2]Congress.gov — Library of Congress
Written by Editorial Team
Last updated July 25, 2026
Get one well-sourced answer a week
No spam. Unsubscribe anytime.