For nearly two decades, publishers handed Google their books on a promise: you may scan them, index them, and show readers a few lines at a time. Now three of the largest names in publishing say Google took that library and fed it to a machine built to compete with them.

Hachette Book Group, Cengage Learning, Elsevier, novelist Scott Turow and S.C.R.I.B.E., Inc. — the corporation holding Turow's copyrights — filed a putative class action against Google LLC on Friday, July 10, in the U.S. District Court for the Southern District of New York. The 57-page complaint, docketed as Case No. 1:26-cv-05870 and picked up by the press early this week, accuses Google of "willful infringement of millions of textual works" to build its Gemini large language models.

"Desperate to maintain its online dominance, Google abandoned its early motto of 'Don't be evil' and engaged in one of the most prolific infringements of copyrighted materials in history," the complaint alleges. Google has not publicly responded; it did not reply to requests for comment from TechCrunch, TheWrap or Al Jazeera. None of the claims has been tested in court.

The Google Books problem

Most training lawsuits allege scraping or piracy: the defendant went and got the books. This complaint alleges something more intimate — that Google already had them, legitimately, because publishers handed them over. "For years, publishers and authors provided Google with massive amounts of copyrighted works for the express, limited purpose of making books searchable via Google Books," the filing states. They "never authorized Google to copy the works they received for Google Books for the completely separate purpose of training its AI models and building a multi-billion dollar competing business."

The complaint extends that theory to Google Play Books and Google Scholar, each a "scope-limited program" in plaintiffs' framing — a trove, the filing says, that "none of Google's AI competitors can access."

The legal architecture is deliberately narrow. Plaintiffs plead four counts: three for direct infringement under 17 U.S.C. §§ 106(1) and 501 — copying from Google services, copying via web scrapes, copying during training — plus one under DMCA § 1202(b) for removal of copyright management information. Every infringement count is reproduction-only; no derivative-works or output-based claim. "This conduct constitutes infringement of the exclusive right of reproduction," the complaint repeats, "regardless of any subsequent use of those copies in model training." That pins liability to the copying itself, sidestepping whether any Gemini output resembles any particular book.

On the DMCA count, plaintiffs allege Google stripped copyright notices during pre-processing, deliberately: "Google removed this CMI to conceal from rightsholders, Gemini users, and the general public that its Gemini Models were trained on stolen materials."

The 2015 ruling Google won, and why plaintiffs say it doesn't help

Google has been here before, and it won. After a decade of litigation, the Second Circuit ruled in 2015 that Google's wholesale scanning for Google Books was fair use. The complaint concedes that ruling — then tries to cage it.

That court found copying lawful, plaintiffs write, "but only because" Google made its copies "for the purpose of providing the public with [Google Books] search and snippet view functions." The decision, they argue, "did not authorize, permit, or deem it fair use for Google to make copies of copyrighted works for the new and separate purpose of training or developing commercial AI models." Google Books' legality, the filing says, "hinges on a narrow, heavily litigated premise that Google would provide a free searchable books index to the public—and nothing else."

That is purpose-scoping, and it is the whole ballgame. If it lands, Google's most famous copyright victory becomes a liability — a finding that its corpus was lawful for one purpose, implying fresh reproduction whenever the purpose changes.

Plaintiffs also lean hard on alleged internal documents. Google flagged that using "Publisher Provided [] copyrighted books" from Google Play Books for AI was "highly problematic for Google," warning of "$10Bs-$100Bs in potential fines," the complaint says, and cited "[h]eightened risk around fair use defenses." Most quotably, it attributes this line to Gemini's lead engineer: "we don't do deals for data we already have or already possess."

Where it fits

The AI copyright docket has cut both ways. In Kadrey v. Meta, a California judge found for Meta on fair use in 2025 — a ruling this complaint cites in footnote 2, borrowing its market-dilution reasoning as an affirmative theory. Bartz v. Anthropic produced a record $1.5 billion settlement over pirated books, later rejected by the judge as "nowhere near complete." NYT v. OpenAI grinds on. Filing in S.D.N.Y. puts the question before a different bench than the California rulings — and in the circuit that decided Google Books.

Kirk Sigmon, founding partner at KellDann Law, told Al Jazeera the acquisition angle is what matters: "The idea, in short, is that any fair use argument that Gemini has would arguably be mooted by the fact that they allegedly acquired the books unlawfully."

Proof is the recurring problem. "Once the egg is baked into the cake, it is extremely difficult to identify it, quantify its contribution or prove precisely which copy of a book was used," Oli Huggins, CEO of ExpertEdge and VP of partnerships at Packt Publishing, told Al Jazeera.

Per the complaint, Google licenses training content from the Associated Press, Reddit and Shutterstock, and is in talks with at least 20 more news publishers — but, as Adweek reports, unlike Meta, Microsoft, Amazon, OpenAI and Anthropic, it has struck no licensing deals with publishers.

What to watch

Plaintiffs, represented by Oppenheim + Zebrak and Keller Rohrback, seek statutory damages "up to the maximum provided by law," an injunction, an accounting of Gemini's training data, and destruction of infringing copies. No dollar figure is pleaded.

Three things will decide this: whether the court certifies the proposed class of ISBN and DOI holders; whether Google's fair use defense survives plaintiffs' purpose-scoping read of its own 2015 win; and whether discovery surfaces those internal documents in full, because if the quoted lines hold up, willfulness gets much easier to argue. Watch, too, for consolidation with In re Google Generative AI Copyright Litigation — the AAP said publishers filed separately to preserve claims outside that case's putative class.

Google's answer is due in the coming weeks. So far, it has said nothing.

“The idea, in short, is that any fair use argument that Gemini has would arguably be mooted by the fact that they allegedly acquired the books unlawfully.”
— Kirk Sigmon, Founding partner, KellDann Law
1:26-cv-05870
Case number, S.D.N.Y.
4
Counts pleaded
18
Sample works listed
July 10, 2026
Complaint filed