AIs can generate near-verbatim copies of novels from training data

AIs can generate near-verbatim copies of novels from training data

Recent investigations reveal that leading AI models can produce nearly identical reproductions of popular novels, sparking new concerns regarding their claims of not retaining copyrighted content. Studies conducted by researchers have uncovered that large language models (LLMs) from prominent companies such as OpenAI, Google, Meta, Anthropic, and xAI have a greater capacity for memorization of their training data than previously acknowledged. Experts in AI and law have pointed out that this memorization capability could significantly impact the defense of AI organizations facing numerous copyright lawsuits worldwide. This challenges their fundamental assertion that LLMs learn from copyrighted materials without actually storing them. Yves-Alexandre de Montjoye, a professor of applied mathematics and computer science at Imperial College London, emphasized, "There’s growing evidence that memorization is a bigger issue than previously believed." Historically, AI companies have contended that they do not memorize content. For instance, in a 2023 correspondence with the US Copyright Office, Google insisted that "there is no copy of the training data—whether text, images, or other formats—present in the model itself." Furthermore, the industry maintains that utilizing copyrighted texts for training falls under “fair use,” asserting that the process transforms the original work into something new and distinct. However, a recent study by researchers from Stanford and Yale Universities demonstrated that by carefully prompting LLMs from OpenAI, Google, Anthropic, and xAI, they could elicit thousands of words from 13 notable books, including titles like A Game of Thrones, The Hunger Games, and The Hobbit. In a striking result, the model Gemini 2.5 reproduced 76.8 percent of Harry Potter and the Philosopher’s Stone with impressive accuracy, while Grok 3 achieved a 70.3 percent reproduction rate. This evidence raises critical questions about the ethical implications of AI technologies and their interaction with copyrighted literature.

Sources : Ars Technica

Published On : Feb 23, 2026, 15:40

AI
China's Kimi K3: A New Contender in the Global AI Race Sparks US Concerns

The tech world is buzzing once again as a new Chinese AI model, Kimi K3, developed by the startup Moonshot AI, challenge...

CNN | Jul 24, 2026, 02:05
China's Kimi K3: A New Contender in the Global AI Race Sparks US Concerns
AI
Global Leaders Unite to Promote Secure Open-Source AI at APEC Summit

In a significant development at the recent APEC summit in Chengdu, China, member nations, including the United States, e...

CNBC | Jul 24, 2026, 02:45
Global Leaders Unite to Promote Secure Open-Source AI at APEC Summit
Mobile
Unlocking Your Google Account Just Got Easier with Selfie Video Verification

In an innovative move, Google has introduced a new recovery feature called the 'Selfie Video' method, designed to simpli...

Business Today | Jul 24, 2026, 06:50
Unlocking Your Google Account Just Got Easier with Selfie Video Verification
Automotive
Mobileye's Visionary Leader Transitions as Company Shifts Focus to Robotics and Robotaxis

Amnon Shashua, the founder and CEO of Mobileye, has announced his intention to step down from his leadership role after ...

TechCrunch | Jul 23, 2026, 23:15
Mobileye's Visionary Leader Transitions as Company Shifts Focus to Robotics and Robotaxis
Startups
Flipkart Enters the Food Delivery Arena, Challenging Major Competitors

Walmart's e-commerce powerhouse, Flipkart, is gearing up to launch its own food delivery platform, igniting competition ...

Business Today | Jul 24, 2026, 07:15
Flipkart Enters the Food Delivery Arena, Challenging Major Competitors
View All News