AI & Technology

Is Training AI on Copyrighted Books Legal? The Issue Is Complicated

DROPIDEA By Admin
August 23, 2026 24 views
DROPIDEA | دروب ايديا - Is Training AI on Copyrighted Books Legal? The Issue Is Complicated

Many users today know that the language models powering tools like ChatGPT, Gemini, and Claude are trained on massive databases comprising hundreds of millions of books, articles, academic papers, and everything that can be found on the internet. What's worse is that most authors have contributed — without their knowledge or consent — to developing the very tools that may threaten their livelihoods. This may seem illegal at first glance, but the legal reality is far more complex.

Why Is the Case So Thorny?

Cathy Gellis, an attorney specializing in intellectual property, copyright, and technology, believes the overlap between this legal field and rapid technological development makes matters extremely complicated. Feelings are sharply divided between supporters and opponents, and the laws themselves were not designed to confront such challenges.

The fundamental problem is that U.S. copyright law has not been updated since 1976, meaning judges are forced to interpret texts nearly half a century old to rule on cases that may shape the future of the entire AI industry.

A Ruling That May Seem Like a Victory for Authors... But It Isn't

Last year, Judge William Alsup issued one of the first rulings of its kind, requiring Anthropic to pay a massive settlement of $1.5 billion to a group of writers whose works were used to train its models. However, a deeper reading of the ruling reveals an important paradox: the judge deemed training models on protected works legal, and that the penalty was due to piracy — namely, obtaining those books from illegal digital libraries.

The judge compared the way a language model absorbs trillions of words to a writer studying literature, writing that the models learn from works not to copy or replace them, but to produce something different. Gellis considers this ruling to favor AI companies more than writers — what is the value of a $1.5 billion fine to a company expecting annual revenues approaching $200 billion by 2028?

The Crux of the Dispute: The Concept of "Fair Use"

Most of these cases revolve around what is known as "Fair Use," an exception within copyright law that permits the use of protected material without explicit permission for purposes such as criticism, parody, and education. Judges rely on specific criteria when adjudicating it, including:

  • The purpose and nature of the use.
  • The amount of the original work used.
  • The extent of the use's impact on the commercial market for the work.

Attorney Jason Henderson, founder of the intellectual property and media practice at JWL International, explains that copyright always aims to protect and develop the market. He notes that courts tend to reject training when its purpose is direct competition with the work's owner, while they find justifications to allow it if there is no direct competition.

A Notable Legal Precedent

Henderson cites a case filed by Thomson Reuters against the research company Ross Intelligence, which copied its content to build a competing AI-based legal platform. Judge Stephanos Bibas ruled that Ross's use was not transformative, because it did not carry a purpose or character different from Reuters' work, and therefore fair use did not apply. Authors could theoretically argue that chatbots compete with them by generating new synthetic books, but this argument has not succeeded in court so far.

Another Facet of the Problem: Who Owns the Generated Content?

Gellis believes it is useful to distinguish between two different issues: training on protected works on one hand, and ownership rights over content produced by AI on the other. In the case of Thaler v. Perlmutter, the court ruled that a work generated 100% by AI is not subject to copyright protection, which opens thorny questions: How do we prove a work was produced by AI? And what proportion of it is machine-generated?

Gellis gives a simple example: when you write your novel in Word and use its spell checker, we don't feel the program owns your novel. But AI is now forcing us to reconsider decisions we long ignored.

An Uncertain Future Awaiting Resolution

Most AI companies remain mired in pending lawsuits, meaning a final resolution to these issues will not appear soon. Gellis affirms that the initial rulings exert a clear influence on the landscape, but this influence may be overturned if other courts rule otherwise, and it may require subsequent stages of litigation to determine which direction will ultimately prevail. Nevertheless, these decisions remain influential in everything that is happening, and it would be unwise for AI companies to ignore them.

✦ بقلم فريق دروب أيديا

DROPIDEA

We hope this article has added real value to you. At DROPIDEA, we always strive to deliver high-quality content that helps you grow and evolve in the digital space. Follow us for more useful articles and guides.

Share Article