NEW YORK / RankWire.AI / – Hachette Book Group, Cengage Learning, and Elsevier have initiated a lawsuit against Google regarding its Gemini artificial intelligence platform. Author Scott Turow and his organization, S.C.R.I.B.E., have joined the class action complaint. The plaintiffs filed their lawsuit on July 10 in the U.S. District Court for the Southern District of New York. They accuse Google of copying millions of copyrighted books and scholarly articles without authorization during the development and training of Gemini. As of July 15, the court had yet to rule on the claims or certify the proposed class.

The complaint states that Google acquired material via Google Books, Google Play Books, and Google Scholar. Publishers and authors had contributed works for specific functionalities, including search, sales, and research purposes. The plaintiffs argue that these arrangements did not permit broader commercial AI training. They also claim Google downloaded extensive web-scraped datasets containing copyrighted content. The filing notes that some of this material originated from known piracy sources and paywalled services.
The 57-page complaint outlines four claims under federal law. Three relate to alleged reproduction through Google services, web scraping activities, and the development or training of Gemini. The fourth invokes the Digital Millennium Copyright Act. The plaintiffs allege that Google removed or altered copyright management information from training data. The filing also references internal discussions about utilizing publisher-provided books. One assessment estimates potential fines ranging from $10 billion to $100 billion. These allegations have not yet been tested in court.
Class members include owners of registered works
The proposed class encompasses owners of registered U.S. copyrights for qualifying books and scholarly articles. Eligible books must bear an International Standard Book Number (ISBN). Eligible articles require a Digital Object Identifier (DOI) or an International Standard Serial Number (ISSN). The class includes works allegedly copied from Google services or downloaded during web scraping. It also covers works purportedly reproduced during Gemini’s development or training phases.
Registration timing also limits who can join the class. One criterion requires registration within five years of publication and before Google’s alleged reproduction or distribution. Another mandates registration within three months of publication. The complaint excludes government entities, Google affiliates, certain court participants, and individuals who properly opt out. The court must approve the class designation before the case moves forward on behalf of the broader group.
Legal claims include damages and accounting requests
The plaintiffs seek either statutory damages or actual damages related to proven infringements. They also request that Google account for profits attributable to any confirmed copyright violations. Their proposed remedies include an injunction, legal expenses, and a jury trial. The complaint does not specify a total damages amount but asks Google to disclose Gemini training data, methods of data collection, and known capabilities through a court-ordered accounting.
Such an accounting would identify copyrighted works used in Gemini’s training process and detail how Google collected, copied, processed, and encoded those materials. The plaintiffs also request court-supervised destruction of unauthorized copies under Google’s control. Earlier, Hachette and Cengage sought to join separate legal actions involving Google’s generative AI efforts in California. The New York lawsuit expands the list of plaintiffs to include Elsevier, Turow, and S.C.R.I.B.E., focusing on claims related to Google services, web scraping, and Gemini’s training activities.
