Meta CEO Mark Zuckerberg at an F8 conference in 2018 (Photo: Anthony Quintano Wikimedia Commons,)
Cover Meta CEO Mark Zuckerberg at an F8 conference in 2018 (Photo: Anthony Quintano, Wikimedia Commons)
Meta CEO Mark Zuckerberg at an F8 conference in 2018 (Photo: Anthony Quintano Wikimedia Commons,)

Meta is under fire for allegedly scraping pirated books from LibGen to train its AI model. As authors push back and lawsuits unfold, here’s what you need to know about the Kadrey vs Meta case and its impact on generative AI

Large language models (LLMs) are among the most widely used forms of generative artificial intelligence. These systems can analyse, understand and generate complex human-like text. But to perform effectively, LLMs must be trained on massive amounts of data—most of which is scraped from the internet, including web pages, public text archives and digital libraries.

Examples of popular LLMs include the GPT series from OpenAI, Claude by Anthropic, Google Gemini, Meta’s Llama, and DeepSeek-R1. These AI systems underpin a wide range of applications—from chatbots and code generators to translation tools, search engines and virtual assistants. With tech companies racing to develop ever-more advanced models, questions are being raised about where all this training data comes from—and whether some of it was legally obtained.

Also read: What to know about Bluesky’s Jay Graber, the tech CEO competing with Elon Musk’s X

LibGen in the spotlight

In March, an investigative report published by The Atlantic revealed that Meta may have scraped millions of books and academic papers from Library Genesis, a notorious online repository of pirated material. Launched in Russia in 2008, LibGen has grown into one of the largest shadow libraries on the web, reportedly offering over 7.5 million books and 81 million research papers for free download.

This massive trove of copyrighted content was allegedly used to train Meta’s LLM, Llama. According to unsealed court documents, internal Meta emails suggest the company was aware of the questionable nature of this data source—but proceeded with training anyway.

Don’t miss: AI’s creative conundrum: From copyright infringement to artistic originality

Kadrey vs Meta court case: Authors are pushing back

The controversy surfaced during Kadrey vs. Meta, a class action lawsuit filed by authors including Richard Kadrey, Sarah Silverman, and Christopher Golden, who allege their copyrighted books were used without permission. In internal discussions cited in the case, Meta employees debated licensing the books, but reportedly opted against it, believing they could rely on the “fair use” clause in U.S. copyright law.

One filing revealed that Meta torrented at least 81.7 terabytes of data across multiple shadow libraries through the site Anna’s Archive, including at least 35.7 terabytes of data from Z-Library and LibGen”.

The revelations have sparked widespread outrage among authors and copyright advocates. The Society of Authors called Meta’s actions “appalling”, with chief executive Ana Ganley saying: “Rather than ask permission and pay for these copyright-protected materials, AI companies are knowingly choosing to steal them in the race to dominate the market. This is shocking behaviour by big tech that is currently being enabled by governments who are not intervening to strengthen and uphold current copyright protections. As part of the Creative Rights in AI Coalition, the SoA has been at the heart of the fight and is continuing to lobby against these unlawful and exploitative activities.”

Meta’s statement

In response, Meta filed a motion to dismiss the case, stating: “At the crux of this case is an issue of extraordinary importance to the future of generative AI development in the United States: whether Meta’s use of publicly available datasets to train its open-source large language models constitutes fair use under US copyright law.”

The company maintains that training AI with data freely available online falls under fair use—an argument that may have far-reaching consequences for the future of AI development.

Read more: Global companies are rolling back DEI programmes—here’s what you need to know

What happens next?

As Kadrey vs Meta unfolds in the United States, other copyright infringement lawsuits have also been filed against Meta. Meanwhile, some authors are exploring ways to remove their work from pirate libraries like LibGen and Z-Library.

With mounting legal pressure and global scrutiny, the outcome of this case could set a precedent for how AI companies source their training data—and whether creators will finally get a say in how their work is used.

Topics