Sony Music Publishing and Warner Chappell Music have sued Anthropic, accusing the AI company and its co-founders Dario Amodei and Benjamin Mann of illegally acquiring and using thousands of copyrighted musical works to develop and operate Claude. The lawsuit seeks statutory damages of up to $150,000 per infringed composition and could expose Anthropic to potentially billions of dollars in damages.
The case, filed on August 28, 2026, in the U.S. District Court for the Northern District of California, represents a major escalation in the growing copyright battle between generative AI companies and the music industry. The publishers allege that Anthropic did not merely encounter copyrighted lyrics as part of ordinary web crawling. Instead, they claim the company engaged in a deliberate, large-scale campaign involving torrenting, scraping and downloading copyrighted material from pirate repositories and other sources.
Importantly, this is a lawsuit and the allegations have not been proven in court. Anthropic disputes the publishers’ claims and says it intends to defend itself robustly.
The dispute is particularly significant because it connects allegations involving copyrighted songs and lyrics to the same underlying training-data practices that previously resulted in Anthropic agreeing to a $1.5 billion settlement with authors and publishers over pirated books.
What Sony Music Publishing and Warner Chappell Are Accusing Anthropic Of
The complaint describes what the publishers characterize as a “brazen campaign of illegally torrenting, scraping, and downloading copyrighted works on a massive scale” to develop and operate Anthropic’s Claude AI models.
According to the lawsuit, Anthropic allegedly obtained copyrighted material through multiple channels rather than relying exclusively on properly licensed datasets.
One of the central allegations concerns Library Genesis, commonly known as LibGen, a massive shadow library containing unauthorized copies of books and other copyrighted material.
The complaint alleges that in June 2021, Benjamin Mann used BitTorrent to download approximately five million pirated books from LibGen. The publishers further allege that Anthropic employees subsequently downloaded at least another two million books from Pirate Library Mirror, or PiLiMi, in July 2022.
The significance for the music lawsuit is that many of those books allegedly contained copyrighted song lyrics and sheet music. The publishers argue that copyrighted musical compositions were therefore swept into Anthropic’s training-data pipeline alongside the books themselves.
The complaint identifies numerous well-known compositions that the publishers say were present in the allegedly infringing material, including songs associated with Bon Jovi, Marvin Gaye, Earth, Wind & Fire, Leonard Cohen, Mariah Carey and Taylor Swift. Examples cited in reporting on the lawsuit include Livin’ on a Prayer, Ain’t No Mountain High Enough, September, Hallelujah, All I Want for Christmas Is You and Paper Rings.
📬 Stay Ahead of Cyber Threats
Get the latest cybersecurity news, critical vulnerabilities, threat intelligence, tutorials, and exclusive giveaways delivered straight to your inbox. No spam. Unsubscribe anytime.
Subscribe to the Newsletter →The lawsuit also alleges that Anthropic obtained lyrics through scraping and other sources, including websites associated with licensed lyric services such as Musixmatch and LyricFind.

Why the LibGen and PiLiMi Allegations Matter
The LibGen and PiLiMi allegations are not appearing in isolation.
Anthropic has already faced extensive litigation over its acquisition of copyrighted books from these same repositories. In 2025, a federal court ruled in the broader Bartz v. Anthropic litigation that Anthropic’s use of legitimately acquired copyrighted books for certain training purposes could qualify as fair use, while the company’s acquisition of pirated copies presented a separate problem.
Anthropic subsequently agreed to a $1.5 billion settlement with authors and publishers over claims concerning millions of pirated books. The settlement documents specifically identify LibGen and PiLiMi as the datasets at issue and state that Anthropic denied the allegations and maintained its fair-use position.
The legal question is not simply whether an AI model can ever be trained using copyrighted material. One of the central questions in these cases is how the material was obtained, what copies were made, how the material was processed, how it was used during training, and what the resulting model can reproduce.
The new music lawsuit attempts to apply that broader factual record to copyrighted musical compositions.
Anthropic Founders Dario Amodei and Benjamin Mann Are Personally Named
The lawsuit does not name Anthropic alone.
Both Dario Amodei, Anthropic’s CEO and co-founder, and Benjamin Mann, another Anthropic co-founder, are individual defendants. The publishers assert both direct and contributory infringement theories against the founders in connection with the alleged torrenting activity.
The complaint alleges that senior Anthropic personnel were aware of the nature of the material being acquired. According to allegations described in the filing, Anthropic’s own Archive Team had characterized LibGen as a blatant copyright violation, while internal discussions allegedly recognized the questionable nature of the repository before the downloads occurred.
The publishers are therefore attempting to establish something more significant than an accidental inclusion of copyrighted material in a giant internet dataset. Their theory is that decision-makers knowingly participated in obtaining copyrighted material from sources they understood to be unauthorized.
That distinction could become central to the case because knowledge and intent can materially affect copyright remedies, particularly when plaintiffs seek enhanced statutory damages for willful infringement.
The Lawsuit Also Targets Claude’s Ability to Reproduce Copyrighted Lyrics
The case is not limited to how Anthropic allegedly obtained the material.
The publishers also allege that Claude can reproduce identical or near-identical portions of copyrighted works in response to users’ prompts.
That raises a technically and legally important distinction between training data acquisition and model output.
A large language model does not ordinarily store a conventional searchable database of documents from which it simply retrieves text. During training, text is converted into numerical representations and used to adjust the model’s parameters. The resulting model learns statistical relationships between tokens and can generate new sequences based on those learned relationships.
However, sufficiently repeated or distinctive material can sometimes be memorized by a model. In such circumstances, carefully constructed prompts may cause a model to reproduce text that closely resembles material appearing in its training corpus.
That is particularly relevant to song lyrics because lyrics can be highly structured, distinctive and relatively short compared with an entire book.
The publishers argue that Claude’s ability to reproduce protected lyrics demonstrates that copyrighted works were not merely incidental to the training process and that the resulting system can potentially substitute for licensed lyric sources.
The Copyright Management Information Allegation
Another technically significant component of the case concerns Copyright Management Information (CMI).
The publishers accuse Anthropic of removing or altering copyright-related information during the processing of training data. The lawsuit contains a separate claim concerning the alleged removal or alteration of CMI.
CMI can include information identifying the copyright owner, author, copyright notices and other information associated with the management and identification of copyrighted works.
This issue has appeared in earlier litigation involving Anthropic’s processing of lyrics. Music publishers have alleged that Anthropic’s data-cleaning systems removed copyright notices and related information while preparing material for AI training. Earlier litigation also alleged that tools used to clean scraped data were selected in part because they removed information considered unnecessary for model training.
The new lawsuit therefore potentially involves two distinct questions:
- Was copyrighted material unlawfully copied or acquired?
- Was copyright-management information knowingly removed or altered from that material?
The second theory is particularly important because the publishers are seeking additional statutory damages under the Digital Millennium Copyright Act framework.
Anthropic Could Face Billions in Potential Damages
The headline figure is $150,000 per infringed work.
Under U.S. copyright law, statutory damages can reach $150,000 for a work in cases of willful infringement. The publishers are alleging infringement involving tens of thousands of musical compositions, meaning the theoretical exposure could reach billions of dollars if the court ultimately finds liability and awards maximum statutory damages.
The publishers are also seeking up to $25,000 for each instance involving the alleged removal or alteration of copyright management information, according to reporting on the complaint.
These figures should not be interpreted as a prediction of what Anthropic will actually pay.
Statutory damages are determined through the legal process, and the final amount can depend on findings concerning infringement, willfulness, the number of qualifying works and other legal issues. A complaint’s maximum requested damages are therefore very different from an actual judgment.
The case nevertheless creates potentially enormous financial exposure for Anthropic if the publishers succeed on their core claims.
This Is Not Anthropic’s First Music Copyright Battle
The new Sony Music Publishing and Warner Chappell lawsuit arrives after several other music publishers have already sued Anthropic over Claude.
A separate case filed by Concord Music Group and other music publishers alleges that Anthropic copied copyrighted lyrics, used those lyrics in training Claude and allowed users to obtain unauthorized reproductions through the model. That litigation has also involved allegations concerning the removal of copyright management information.
The broader music industry litigation has consequently developed into a multi-front legal challenge involving questions about:
- copyrighted lyrics in training datasets;
- scraping and web crawling;
- pirate repositories;
- fair use;
- model memorization;
- copyrighted material reproduced in AI outputs;
- copyright-management information;
- contributory infringement; and
- whether AI companies must obtain licenses for copyrighted training material.
The latest lawsuit adds some of the world’s largest music-publishing catalogs to that fight.
Why the Case Matters Beyond Anthropic
The lawsuit could become important for the entire generative AI industry because it focuses on a fundamental problem in AI development: where training data comes from.
Modern foundation models require enormous quantities of text and other information. Developers can obtain that material from licensed datasets, public-domain sources, user-generated content, web crawling and other sources. The legal status of each source can differ substantially.
The music industry presents an especially complicated environment because songs contain multiple layers of rights. A musical composition and a sound recording are not necessarily the same copyrighted work, and ownership can be divided among publishers, songwriters and other rights holders.
That makes large-scale automated ingestion particularly sensitive.
A dataset containing millions of webpages may inadvertently contain copyrighted lyrics. But if a company knowingly downloads a repository whose primary purpose is distributing unauthorized copies, plaintiffs can argue that the acquisition itself is fundamentally different from ordinary web indexing.
That is one of the central themes running through the Anthropic litigation.
Anthropic’s Previous $1.5 Billion Settlement Does Not Automatically Decide This Case
It would be easy to assume that Anthropic’s previous $1.5 billion settlement means the music publishers have already won the new case.
That is not how the legal process works.
The earlier settlement resolved claims involving a particular group of authors and publishers and particular works. Anthropic did not admit wrongdoing as part of that settlement. The settlement materials explicitly state that Anthropic denied the allegations and maintained that its use of the datasets was protected by fair use.
The new lawsuit involves different plaintiffs, musical compositions and legal claims.
However, the previous litigation could still be highly relevant because the publishers are relying on factual allegations and evidence concerning the same alleged acquisition practices, including LibGen and PiLiMi.
That creates a potentially important legal and evidentiary bridge between the book and music disputes.
What Happens Next
The lawsuit was filed in the Northern District of California as case 5:26-cv-09217, with a jury trial requested.
The next stages will likely involve motions to dismiss, discovery and extensive disputes over Anthropic’s training data, data-processing systems and model behavior.
Discovery could be particularly consequential.
The publishers are seeking information about Claude’s training data, and reporting on the complaint indicates that they also want infringing copies destroyed and an accounting of the material used to train Claude.
That means the litigation could eventually force detailed examination of questions that have remained difficult to answer publicly, including exactly what datasets Anthropic used, how those datasets were processed, which copyrighted works were present, and how those works influenced the models.
Anthropic, meanwhile, has rejected the publishers’ allegations and said it intends to defend itself.
The Bigger AI Copyright Question
At its core, the Sony Music Publishing and Warner Chappell lawsuit is about more than whether Claude knows the lyrics to famous songs.
It asks where the legal boundary lies between learning from information and copying protected expression at scale.
AI companies argue that training models involves transformative computational processes and that broad access to information is essential for developing useful systems. Copyright owners argue that companies should not be able to build highly valuable commercial AI systems by copying enormous quantities of protected works without authorization, particularly when the material is obtained from clearly unauthorized sources or can later be reproduced by the model.
The Anthropic cases are helping courts separate those questions.
The distinction between lawfully acquired copyrighted material, pirated copies, training use, model memorization and infringing output is likely to become one of the defining legal issues of the generative AI era.
For Anthropic, the stakes are particularly high. The company has already agreed to a $1.5 billion settlement over alleged book piracy. Now, some of the world’s largest music publishers are asking a federal court to hold Anthropic and two of its founders responsible for allegedly using a similar pipeline to acquire copyrighted musical works.
And if the publishers’ allegations are ultimately established, the financial consequences could extend into the billions of dollars.
For the broader AI industry, the case could be even more consequential: a ruling on how copyrighted music can be acquired, processed, used to train an AI model and reproduced through that model could establish another important precedent for how future foundation models are built.
Important: The allegations described above come from the publishers’ complaint and reporting on the case. Anthropic disputes the claims, and no court has determined that the alleged conduct occurred as described.









