{"id":78500,"date":"2026-08-30T09:42:31","date_gmt":"2026-08-30T13:42:31","guid":{"rendered":"https:\/\/overcentral.com\/en\/?p=78500"},"modified":"2026-08-30T09:42:31","modified_gmt":"2026-08-30T13:42:31","slug":"sony-warner-sue-anthropic-copyright-78500","status":"publish","type":"post","link":"https:\/\/overcentral.com\/en\/sony-warner-sue-anthropic-copyright-78500\/","title":{"rendered":"Sony Music, Warner Sue Anthropic Over Copyright Theft"},"content":{"rendered":"<p><a href=\"https:\/\/www.sonymusicpub.com\" target=\"_blank\" rel=\"noopener noreferrer\" data-iacss-external=\"1\">Sony Music Publishing<\/a>, <a href=\"https:\/\/www.warnerchappell.com\" target=\"_blank\" rel=\"noopener noreferrer\" data-iacss-external=\"1\">Warner Chappell<\/a>, and numerous other music publishers have filed suit against <a href=\"https:\/\/www.anthropic.com\" target=\"_blank\" rel=\"noopener noreferrer\" data-iacss-external=\"1\">Anthropic<\/a> and its co-founders, Dario Amodei and Benjamin Mann, accusing the AI lab of a \u201cbrazen campaign\u201d of copyright theft centered on the <a href=\"https:\/\/overcentral.com\/en\/mit-study-reveals-ai-art-lacks-traceable-training-data\/\" title=\"MIT Study Reveals AI Art Lacks Traceable Training Data\" data-iacss-internal=\"1\">training data<\/a> behind Claude, its flagship large language model. The complaint, filed late Friday in the U.S. District Court for the Northern District of California, alleges that Anthropic illegally torrented, scraped, and downloaded thousands of copyrighted works, including books filled with lyrics and sheet music, without authorization. The case arrives after a series of earlier copyright challenges that already produced a $1.5 billion judgment against Anthropic.<\/p>\n<h2>Inside the Copyright Theft Lawsuit Against Anthropic<\/h2>\n<p>The lawsuit <a href=\"https:\/\/overcentral.com\/en\/ai-search-moves-cognitive-load-does-not-remove-it\/\" title=\"AI Search Moves Cognitive Load, Does Not Remove It\" data-iacss-internal=\"1\">does not<\/a> treat Claude\u2019s output as the problem. Instead, the publishers focus on the resource required to build the model: the training corpus. They say Anthropic engaged in systematic unauthorized copying that reached into the music industry\u2019s most carefully guarded assets \u2014 the words of songs and the printed notation of music. The complaint accuses Anthropic of \u201cblatant theft\u201d by using copyrighted works on a mass scale to train its AI model, and it frames the company\u2019s conduct as an ongoing campaign rather than a one-time mistake.<\/p>\n<p>That distinction matters. Earlier litigation over <a href=\"https:\/\/overcentral.com\/en\/cloudflare-googlebot-block\/\" title=\"Cloudflare Blocks Googlebot from Indexing and AI Training\" data-iacss-internal=\"1\">AI training<\/a> data has often centered on whether a model that is trained on copyrighted text effectively transforms that text into something new. The publishers\u2019 complaint shifts the focus to how Anthropic obtained the material in the first place. Whether the works came from peer-to-peer networks, automated scraping tools, or direct downloads, the plaintiffs argue that Anthropic built its training library on copying that no license was ever sought to authorize.<\/p>\n<p>For Anthropic, the legal risk is not simply that a court might find its use of the works to be unfair. The risk is that a court could find the underlying copying unlawful from the start. That would sever the connection between fair-use arguments and the facts of the case, because fair use does not typically protect a defendant who acquires an infringing copy through deliberate piracy.<\/p>\n<h2>Why Music Publishers Are Bringing the Case<\/h2>\n<p>Music publishing is a licensing business. Publishing companies own or administer the copyrights in lyrics and musical compositions, while recorded-music companies typically control the sound recordings rather than the written works. Every use of a song\u2019s words or printed music \u2014 in a streaming service, a film, a television show, a karaoke system, or a print publication \u2014 ordinarily requires a license and a royalty payment.<\/p>\n<p>Books that contain lyrics and sheet music are therefore more than ordinary literature for music publishers. They are repositories of protected commercial assets. If those books were copied into an AI training corpus without permission, the publishers argue, the copying invaded the same exclusive rights that govern traditional music licensing. The fact that the copies were made for a machine-learning model, rather than for direct distribution to listeners, does not make the copying innocent in their view.<\/p>\n<p>That is why this case feels different from a typical music-industry dispute. There is no allegation that Claude played a song on the radio or offered a lyric sheet to a streaming subscriber. The accusation is more fundamental: the works were copied, without permission, in bulk, and then absorbed into a proprietary system.<\/p>\n<h2>Before This Case, There Was a $1.5 Billion Judgment Against Anthropic<\/h2>\n<p>Anthropic has already faced serious legal consequences over the same class of claims. In the landmark Bartz v. Anthropic case, a group of authors accused the company of using copyrighted works to train Claude and related products. The court eventually ordered Anthropic to pay $1.5 billion. The ruling was not a total rejection of AI training on copyrighted material. The judge concluded that Anthropic could lawfully use copyrighted works to train its models, but that it could not lawfully acquire those works through piracy.<\/p>\n<p>The new lawsuit arrives against that backdrop. Some of the same lawyers behind this action also represent Concord Music Group and Universal Music Group in a separate music-related case filed in January, and those same lawyers led the Bartz case against Anthropic. The parallels are not accidental. The new complaint adopts and expands on arguments from the earlier litigation, but it also adds a layer that the authors\u2019 case did not emphasize: the use of illegal torrenting to obtain millions of books, including books that contain lyrics and sheet music.<\/p>\n<p>That level of specificity is significant. Instead of relying only on a general claim that Anthropic copied protected works, the publishers say they can identify both the method and the scale of the alleged infringement. They describe the conduct as \u201cflagrant piracy,\u201d and they argue that it amounts to one of the most direct cases of AI training data theft yet presented to a court.<\/p>\n<h2>The Allegations in Three Parts: Torrenting, Scraping, and Downloading<\/h2>\n<p>The complaint describes three related methods of obtaining copyrighted material. Each method is treated as part of a broader campaign, not as an isolated technical detail.<\/p>\n<ul>\n<li><strong>Illegal torrenting:<\/strong> The publishers say Anthropic used peer-to-peer file-sharing networks to obtain millions of copies of books, including books that contain lyrics and sheet music. Torrenting is efficient for moving very large files because it distributes the load across many computers, which makes it attractive for sharing massive training datasets.<\/li>\n<li><strong>Scraping:<\/strong> The complaint also says Anthropic used automated tools to pull copyrighted text from websites and other online sources. Web scraping can capture enormous amounts of material quickly, and when the material is protected by copyright, it is no less a reproduction than a direct download.<\/li>\n<li><strong>Direct downloading:<\/strong> The publishers further allege that Anthropic simply downloaded copyrighted works in bulk. This part of the claim is important because it broadens the case beyond torrenting and scraping to include any unauthorized act of copying that fed the training pipeline.<\/li>\n<\/ul>\n<p>Together, the three methods paint a picture of data procurement that was, in the publishers\u2019 telling, both aggressive and indifferent to copyright law. The complaint says the campaign was \u201cbrazen\u201d precisely because it was so systematic and so large.<\/p>\n<h3>What are the music publishers accusing Anthropic of?<\/h3>\n<p>The music publishers are accusing Anthropic of carrying out a \u201cbrazen campaign\u201d of copyright infringement by illegally torrenting, scraping, and downloading copyrighted works to train its Claude AI model. The complaint specifically identifies thousands of copyrighted works, including books with lyrics and sheet music, and calls the conduct \u201cblatant theft.\u201d The alleged theft focuses not on Claude\u2019s user-facing answers but on the unauthorized copying that made the model\u2019s training possible.<\/p>\n<h2>How the New Case Builds on the Bartz Ruling<\/h2>\n<p>The Bartz ruling created a split that the music publishers are now trying to exploit. A judge said Anthropic could train on copyrighted works and still be within the bounds of the law, but that it could not obtain those same works through pirated channels. The music publishers have taken that distinction and applied it to a different portion of the copyright economy: songs, lyrics, and sheet music.<\/p>\n<p>The latest lawsuit is particularly broad. It is not limited to one author, one book, or one system. It names Anthropic, its co-founders Dario Amodei and Benjamin Mann, and describes a training-data operation that allegedly touched millions of books. By adding the co-founders as individual defendants, the publishers are signaling that they intend to push beyond corporate liability and hold the people who oversee the company accountable for decisions about data acquisition.<\/p>\n<h3>How does this lawsuit differ from the Bartz case?<\/h3>\n<p>The Bartz case was brought by authors who said Anthropic used their copyrighted books to train Claude. The new lawsuit is brought by music publishers who say Anthropic used copyrighted books containing lyrics and sheet music, and it places far greater emphasis on illegal torrenting as the method of obtaining that material. The Bartz case established that unlawful acquisition could overcome a fair-use defense; the new case applies that principle to music copyrights on a much larger scale.<\/p>\n<h2>Why Torrenting Is a Critical Legal Detail<\/h2>\n<p>Torrenting is a peer-to-peer method of distributing files across a network. It is used for legitimate purposes, such as delivering open-source software and large public-domain archives, but it has also become one of the most common ways to share pirated movies, television shows, and books. For AI training, torrenting is a particularly efficient way to move enormous datasets between servers, because each participant in the network contributes bandwidth as well as storage.<\/p>\n<p>The publishers say Anthropic used that kind of file-sharing system to obtain millions of copies of books. If proven, that allegation would turn the case into something closer to a traditional piracy matter than a novel AI fair-use dispute. Fair-use doctrine can be complex when a model learns patterns from protected text, but there is far less complexity when the text itself was acquired through an illegal distribution channel.<\/p>\n<p>Scraping and downloading play supporting roles in the complaint. Scraping is the automated extraction of data from websites; downloading is simply transferring a file from one system to another. The publishers treat all three as evidence of one coordinated effort to amass a training library without paying or asking. The result is a legal theory that does not depend on proving exactly what Claude says to users. The infringement, according to the publishers, happened during the creation of the training set.<\/p>\n<h3>What is illegal torrenting in the context of AI training?<\/h3>\n<p>Torrenting is a peer-to-peer way of distributing large files over many connected computers. When the files being shared are copyrighted works that the distributor does not own or have permission to share, downloading and distributing those files can be illegal. The publishers say Anthropic used this method to obtain millions of books, including books containing lyrics and sheet music, without authorization.<\/p>\n<h2>Anthropic\u2019s Response: \u201cWe Intend to Defend Ourselves Robustly\u201d<\/h2>\n<p>\u201cWe disagree with the publishers\u2019 claims and we intend to defend ourselves robustly in court,\u201d an Anthropic spokesperson said in an emailed statement. The statement was short and did not address the specific allegations about torrenting, scraping, or downloading.<\/p>\n<p>A robust defense could take several forms. Anthropic may argue that the books in its training data were obtained from sources that did not make their origins obvious, or that the company reasonably believed it had permission to use the material. It could also argue that the publishers have not shown a direct connection between any particular file and Claude\u2019s training process. And Anthropic could challenge the naming of Dario Amodei and Benjamin Mann as individual defendants, especially if it argues that the founders had no direct role in assembling the specific datasets at issue.<\/p>\n<p>None of those arguments, however, answer the core factual allegation that copyrighted material was copied on a massive scale. The company\u2019s legal defense will therefore need to address the provenance of its data, not just the nature of AI training.<\/p>\n<h2>Fair Use, Piracy, and the Legal Line at the Center of the Case<\/h2>\n<p>Most music copyright disputes are about public performances or reproductions of finished works. A song plays in an unlicensed podcast, a clip is used without permission, or a master recording is uploaded to a platform without clearance. This case is different because the alleged infringement is entirely internal. The publishers\u2019 theory is that copying protected books into a training dataset is itself a violation, regardless of whether any lyrics ever appear in Claude\u2019s output.<\/p>\n<p>That theory places the case at the center of the AI industry\u2019s most important legal debate. Many AI developers argue that training on publicly available text is transformative, and therefore covered by fair use. Copyright holders respond that copying protected works into a training corpus is an act of infringement even if the model later transforms the information. The Bartz ruling offered a subtle middle path: training on copyrighted works can be legal, but the method of obtaining those works cannot be piracy.<\/p>\n<p>The music publishers have followed that path directly to the front of the courtroom. They have made the case about methods, not outcomes. They are asking the court to apply the same rule to Anthropic that was applied in Bartz, but this time to works that are associated with writers, composers, and the broader music ecosystem.<\/p>\n<h2>Potential Stakes for the AI Industry and the Music Business<\/h2>\n<p>The potential damages in this case could be substantial. The publishers say the alleged infringement involves thousands of copyrighted works, and the scale of the alleged copying reaches millions of books. Under U.S. copyright law, statutory damages can be awarded for each work infringed, which means the number of works and the number of copies matter enormously to the final calculation.<\/p>\n<p>Beyond damages, the case could reshape how AI companies assemble training data. If the court accepts the publishers\u2019 framework, Anthropic and other developers may be forced to audit their training corpora for pirated content, to create records showing how every file was obtained, and to seek licenses before using protected material. That would add cost and friction to a process that has so far operated with very little transparency.<\/p>\n<p>For the music business, the case is about control. Lyrics and sheet music are valuable assets, and the publishers depend on their ability to license those assets in clear, contracted ways. If an AI model can be trained on lyrics without payment, the publishers argue, the value of their catalogues is diminished and their control over licensed uses is destroyed.<\/p>\n<h2>A Licensing Market Could Emerge From the Dispute<\/h2>\n<p>If the publishers prevail, one likely outcome is a new licensing market for AI training data. Music publishing companies already license lyrics for streaming, print, and broadcast uses, and they have the infrastructure to license their catalogues to AI developers. A ruling that requires AI labs to obtain training licenses would create a direct commercial bridge between the technology sector and the music publishing business.<\/p>\n<p>That market would not only benefit the large publishers who filed this suit. It would also set a precedent for songwriters, composers, authors, and other copyright holders whose work is used to train models without payment. The deeper issue is whether training data has economic value that must be paid for, or whether it is simply raw material that AI companies are free to harvest.<\/p>\n<p>The music publishers have made their position clear: the answer by their lights is that the works have value, the value belongs to their owners, and the owners are entitled to be paid.<\/p>\n<h2>A Defining Test for AI Training Data Provenance<\/h2>\n<p>The new lawsuit will unfold in the same district where the Bartz case produced its landmark judgment, and it involves many of the same lawyers, the same defendant, and the same underlying tension between lawful training and unlawful acquisition. What makes this moment different is the plaintiffs\u2019 explicit focus on the pipe, not the model. Torrenting, scraping, and downloading are not questions about whether AI deserves fair-use protection; they are questions about whether a company can build a valuable product on top of data it illegally obtained.<\/p>\n<p>For Anthropic, the stakes are clear. A court that finds the company liable for the composition of its training set could impose damages far beyond the costs of settlement, and it could produce a ruling that forces the entire AI industry to change how it sources data. For music publishers, the case is a chance to protect the songs and compositions that underpin their business. And for every other developer building on large datasets, the lesson is already visible: the origin of the data may matter more than the quality of the model.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Sony Music Publishing, Warner Chappell, and numerous other music publishers have filed suit against Anthropic and its co-founders, Dario Amodei and Benjamin Mann, accusing the AI lab of a \u201cbrazen campaign\u201d of copyright theft centered on the training data behind Claude, its flagship large language model. The complaint, filed late Friday in the U.S. District [&hellip;]<\/p>\n","protected":false},"author":7,"featured_media":78505,"comment_status":"closed","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/pub-4d4fc17555de4152be07eaf2a416a31e.r2.dev\/en\/ocie_1788097374176.jpg","fifu_image_alt":"Sony Music, Warner Sue Anthropic Over Copyright Theft","footnotes":""},"categories":[31],"tags":[],"class_list":["post-78500","post","type-post","status-publish","format-standard","has-post-thumbnail","category-technology"],"fifu_image_url":"https:\/\/pub-4d4fc17555de4152be07eaf2a416a31e.r2.dev\/en\/ocie_1788097374176.jpg","fifu_image_alt":"Sony Music, Warner Sue Anthropic Over Copyright Theft","_links":{"self":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/78500","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/users\/7"}],"replies":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/comments?post=78500"}],"version-history":[{"count":0,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/78500\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media\/78505"}],"wp:attachment":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media?parent=78500"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/categories?post=78500"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/tags?post=78500"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}