NEW YORK / RankWire.AI / – Three leading U.S. publishing companies have initiated a lawsuit against Google, accusing the tech giant of copyright violations related to its Gemini artificial intelligence platform. Hachette Book Group, Cengage Learning, and Elsevier filed a proposed class action alongside author Scott Turow and his organization, S.C.R.I.B.E. They submitted their complaint on July 10 in a federal court in New York. The case alleges that Google duplicated millions of copyrighted books and journal articles without authorization during the development and training of Gemini models.

According to the plaintiffs, Google acquired material via Google Books, Google Play Books, and Google Scholar. The publishers and authors had provided these works to facilitate search, sales, and research functions, the complaint states. The filing claims that these arrangements did not grant Google permission to copy the content for commercial AI training purposes. It also accuses Google of utilizing web-scraped datasets that included works from pirate sites and subscription-based services protected by paywalls.
Google faces four allegations outlined in the 57-page complaint. Three of these claims concern illegal reproduction through Google services, web scraping activities, and the training or development of Gemini. The fourth claim invokes the Digital Millennium Copyright Act. The plaintiffs allege that Google removed or altered copyright management information, including author names, ownership details, and publication data. As of July 15, the court had yet to rule on these allegations or grant class-action status.
Four allegations focus on Gemini’s training data
The proposed class includes owners of registered U.S. copyrights in books and journal articles. Eligible books must have an International Standard Book Number, while eligible articles must possess a Digital Object Identifier or International Standard Serial Number. This definition pertains to works that Google allegedly copied from its services, downloaded via web scraping, or reproduced during the development of Gemini. Eligibility is also limited to works registered within the deadlines specified in the complaint.
The complaint cites works from Hachette, Cengage, and Elsevier as examples of the alleged copying. The categories affected include fiction, textbooks, and scholarly publications. The filing further references internal assessments by Google regarding legal risks tied to publisher-supplied books. One internal review reportedly warned of potential fines ranging from $10 billion to $100 billion, according to the plaintiffs. The court has not issued any rulings regarding these internal documents.
Plaintiffs seek monetary damages and transparency
The plaintiffs are seeking statutory or actual damages along with profits attributable to any proven infringement. They also request an injunction, legal fees, and a jury trial. Their proposed court order would compel Google to disclose the materials and methods used to train Gemini. Additionally, they ask the court to oversee the destruction of any unauthorized copies under Google’s control. The complaint does not specify an exact damages total.
This New York case follows an earlier effort by Hachette and Cengage to join separate copyright litigation involving Google’s AI in California. The Association of American Publishers noted that the new case preserves claims not included in that proceeding’s proposed class. The current lawsuit also involves Elsevier, Turow, and S.C.R.I.B.E., and seeks to have the New York court determine whether Google’s Gemini training practices and data collection violate federal copyright laws and the Digital Millennium Copyright Act.
