Publishers and Authors Sue Google, Alleging Gemini Was Trained on Stripped-Down Books
Three of the world's largest publishers and one of America's best-known novelists have accused Google of building its Gemini artificial-intelligence models on a foundation of stolen books, and of deliberately scrubbing away the copyright markings that would have exposed the theft. In a proposed class-action complaint filed July 10, 2026, in the U.S. District Court for the Southern District of New York, Hachette Book Group, Cengage Learning, Elsevier, author Scott Turow and the writers' group S.C.R.I.B.E., Inc. call Google's conduct "one of the most prolific infringements of copyrighted materials in history."
The case, docketed as No. 1:26-cv-05870, lands Google in the same courthouse crush of AI copyright fights that has already ensnared OpenAI, Meta and Anthropic. But the publishers are pressing a claim that sets their suit apart from the fair-use battles dominating the field: they allege Google did not just copy their work, it altered it to hide what it had done.
The allegations
At the center of the complaint is Google's long, once-cooperative relationship with the publishing industry. For years, publishers and authors handed Google their books for narrow, agreed-upon purposes, most prominently Google Books, which lets users search text and view short snippets, along with Google Play Books and Google Scholar. The plaintiffs say those arrangements never authorized generative-AI training, and that Google knew it.
According to the filing, Google "illegally copied works from all these scope-limited programs for AI training, knowing it lacked authorization to do so," then copied those works "many times over" to train what the complaint describes as a multi-billion-dollar system. The publishers say Google supplemented that trove with unauthorized web scrapes of "virtually the entire internet," including material pulled from behind paywalls and from known pirate sources.
The most distinctive allegation concerns copyright management information, or CMI: the titles, author names and ownership details that identify a work and its rights holder. The plaintiffs contend Google stripped that information from the books to "conceal" that its Gemini models had been trained on what the complaint calls stolen materials. That is a claim under Section 1202 of the Digital Millennium Copyright Act, a distinct legal theory from the fair-use question of whether training on copyrighted text is lawful in the first place. In all, the suit brings four claims: direct infringement, contributory infringement, removal or alteration of CMI, and related DMCA violations.
The publishers also quote what they say are Google's own internal documents. One allegedly warned that training on "Publisher Provided" copyrighted books from Google Play Books would be "highly problematic for Google," flagging "$10Bs-$100Bs in potential fines." Others, the plaintiffs say, noted that publishers were "sensitive about training on their data" and that there was "heightened risk around fair use defenses." Google did not immediately respond to press requests for comment on the complaint, and no court has ruled on any of the claims, which remain allegations.
Why the CMI angle matters
The timing sharpens the stakes. Two closely watched California rulings in June 2025 sided with AI developers, finding that training on copyrighted books can qualify as fair use. Yet Anthropic still agreed to pay roughly $1.5 billion to settle claims that it pirated the works it trained on, a settlement approved in July 2026 and the largest in the history of U.S. copyright law. The lesson many lawyers drew was that how a company acquired its training data can matter as much as what it did with it.
That is precisely the ground the publishers are fighting on. By emphasizing the alleged removal of copyright information, they sidestep the hardest part of the fair-use debate and lean into a theory about deception and concealment. DMCA Section 1202 carries its own statutory damages, and a CMI-stripping claim is harder to wave away as a transformative, non-expressive use. It reframes the dispute from an abstract argument about how machines learn into a more familiar one about whether a company hid its tracks.
Filing in New York rather than California is also deliberate. It puts the questions before a different bench, one not bound by the California fair-use decisions, at a moment when no appellate court has settled the core issues. The publishers had initially planned to intervene in the consolidated In re Google Generative AI Copyright Litigation, but chose a standalone suit to preserve claims that fall outside that case.
What to watch
The immediate tests are procedural but consequential: whether the court certifies a class that could sweep in hundreds of thousands of authors and publishers, and how Google answers the CMI allegations and the internal documents quoted in the complaint. Watch, too, for Google's expected fair-use defense, and whether it echoes the company's public argument that training on public data is transformative. If the DMCA theory survives an early motion to dismiss, it could become a template for the next wave of AI copyright suits, shifting the fight from whether models can learn from books to whether their makers were honest about how they got them.
"Google illegally copied works from all these scope-limited programs for AI training, knowing it lacked authorization to do so."— Hachette v. Google complaint, Plaintiffs' filing, S.D.N.Y.