Google Faces Major Copyright Lawsuit Over Gemini AI Training

Major publishers have filed a lawsuit against Google, claiming the company illegally used millions of copyrighted books to train its Gemini artificial intelligence system. The case, filed in federal court in New York, represents what the publishers describe as "one of the most prolific infringements of copyrighted materials in history."

The lawsuit was brought by Hachette Book Group, Cengage Learning, Elsevier, and bestselling author Scott Turow. The publishers argue that Google misused books originally provided for limited purposes through services like Google Books, Google Play Books, and Google Scholar. These platforms allowed Google to display searchable excerpts and sell ebooks, but the lawsuit claims Google went beyond these agreements by copying the works to train a commercial AI product.

According to the complaint, Google copied copyrighted books without permission or payment to build Gemini, despite internal discussions acknowledging potential legal risks. The filing states Google identified that it could face "10 billion to 100 billion dollars in potential fines" for using texts provided through Google Play Books. The publishers contend that Google's actions harm authors and the publishing industry, particularly as AI-generated content could replace sales of original works.

The suit includes a striking example of the concern. Gemini could generate "a 100-page murder mystery set in a quiet seaside town filled with secrets" in 20 minutes for 39 cents, potentially substituting for original copyrighted works on which the model trained. The publishers argue that "no publisher or author can compete with that."

The complaint names specific copyrighted books allegedly used without permission, including N.K. Jemisin's The Fifth Season and Lemony Snicket's Who Could That Be at This Hour?

This lawsuit adds to expanding legal battles over generative AI and copyright protections. Authors and publishers have filed multiple cases against Google, OpenAI, Anthropic, and Meta, alleging unauthorized use of copyrighted works for AI training. A judge previously ruled in Meta's favor in a copyright case brought by authors. However, Anthropic recently agreed to pay 1.5 billion dollars to authors who alleged their pirated books trained the AI chatbot Claude.

Earlier this year, thousands of authors including Kazuo Ishiguro, Philippa Gregory, and Richard Osman published an "empty" book protesting AI firms' use of their work without permission.

This new case follows an earlier attempt by Hachette and Cengage to join an existing copyright lawsuit against Google filed by authors and illustrators in 2023. Google opposed their participation, prompting the publishers to launch separate legal action.

The plaintiffs are seeking statutory damages, an injunction preventing continued alleged infringement, and a court order requiring Google to destroy unauthorized copies used in training its AI systems. Google did not respond to requests for comment.