AI Company Faces Massive Lawsuit Over Pirated Content

Aug 31, 2026 5:15 PM
Premium
Advertisement
AI Company Faces Massive Lawsuit Over Pirated Content
AP Photo/Eugene Hoshiko
VIP

Artificial intelligence company Anthropic is facing a major lawsuit alleging that the company stole books, music, and other forms of intellectual property to train its technology.

Sony Music Publishing and Warner Chappell Music on Friday filed a copyright infringement lawsuit against Anthropic PBC and co-founders Dario Amodei and Benjamin Mann in the U.S. District Court for the Northern District of California.

The lawsuit’s objective is to stop the company from using thousands of copyrighted musical compositions, sheet-music books, and song lyrics to train and develop its Claude chatbot models.

The plaintiffs seek a jury trial and statutory damages of up to $150,000 for each piece of pirated content.

This case is part of a wider debate over how artificial intelligence companies collect data to develop their technology. Developers have claimed fair use while companies and creators are demanding protection against the unauthorized exploitation of their intellectual property.

“In a blatant violation of copyright law, Defendants have unlawfully acquired troves of Music Publishers’ musical compositions, and then systematically copied those works multiple times, including as the inputs to train Anthropic’s Claude AI models and in the outputs those models generate,” the lawsuit reads.

Mann allegedly used BitTorrent to download at least five million pirated books from Library Genesis in June 2021. Staff downloaded two million more from Pirate Library Mirror in July 2022.

Users of these peer-to-peer networks upload files to others for them to download. The plaintiffs contend that each transfer constitutes an infringement on distribution rights.

"Dr. Amodei and Mr. Mann are personally liable for their respective roles in this illegal torrenting of pirated copies of Music Publishers' works from LibGen and PiLiMi,” the lawsuit argues.

The complaint further details Anthropic’s data harvesting operation that takes copyrighted lyrics from online and physical sources. The company scraped lyrics from licensed websites such as MusixMatch and LyricFind in violation of the sites’ terms and conditions. The company also allegedly extracted text from physical books by purchasing second-hand copies in bulk.

The complaint states, "Anthropic employs a sophisticated ‘destructive scanning’ operation, modeled after the Google Books project, which involves purchasing physical copies of millions of second-hand books, scanning ink and paper into digital text, then destroying the remains."

Then, the company “adds the newly scanned text to its central library for potential later use in AI training,” according to the complaint.

The plaintiffs further allege that Anthropic violated the Digital Millennium Copyright Act by removing Copyright Management Information (CMI). Under copyright law, CMI refers to identifying details that are attached to a piece of work. These include song titles, names of lyricists and composers, ownership data, and formal copyright notices.

Federal law prohibits removing this information when it conceals or enables copyright infringement. The publishers accuse Anthropic of stripping this metadata so Claude could output song lyrics without displaying credits that reveal infringement.

"After unlawfully harvesting Music Publishers' works, Anthropic processes the text from those works to remove repetitive ‘garbage’ portions that it does not want to use in training its models and with the knowledge and purpose of concealing its other infringements," the complaint alleges.

Companies like Anthropic, OpenAI, and others have faced a deluge of criticism ever since their technology started becoming more ubiquitous at the consumer level. When it was revealed that they fed vast collections of copyright-protected written text into networks during training, many cried foul, pointing out that the companies did not obtain permission to use these works to improve their algorithms.

The companies countered these arguments by arguing that their activities constitute transformative fair use. A federal court recently approved a $1.5 billion settlement to resolve class-action lawsuits from authors over Anthropic’s use of pirated books, according to The Associated Press.

An Anthropic spokesperson said, "We disagree with the publishers' claims and we intend to defend ourselves robustly in court." 

Federal lawmakers introduced bipartisan legislation last year that would allow Americans to sue AI companies that use protected creative works to train their AI models without prior consent.

Recommended

Trending on Townhall Videos