Are bad incentives to blame for AI hallucinations?

Are bad incentives to blame for AI hallucinations?

Recent research from OpenAI delves into the persistent issue of hallucinations in large language models, including GPT-5 and chatbots like ChatGPT. These hallucinations are characterized as “plausible but incorrect statements generated by these models.” Despite advancements in technology, OpenAI acknowledges that hallucinations continue to pose a significant challenge for all large language models, a problem that is unlikely to be fully resolved. To highlight this issue, the researchers conducted a test using a popular chatbot, asking it about the title of Adam Tauman Kalai’s Ph.D. dissertation. The chatbot provided three different, incorrect answers. When they inquired about Kalai’s birthday, the model again generated three incorrect dates. This raises an important question: how can a chatbot present such confidently inaccurate information? The researchers propose that these hallucinations stem partly from the pretraining phase, which emphasizes the models’ ability to predict subsequent words without the benefit of true or false labels. The models are trained on fluent language examples, leading them to approximate general patterns. However, low-frequency facts, such as specific dates or uncommon knowledge, cannot be inferred from these patterns alone, resulting in hallucinations. Interestingly, the paper suggests that the root of the problem may lie not only in the training process but also in the evaluation methods used for large language models. The current evaluation frameworks do not directly cause hallucinations, but they create incentives that encourage guessing. The researchers likened these evaluations to multiple-choice tests, where random guessing can yield correct answers, but leaving a question blank results in a guaranteed zero. The solution proposed focuses on reforming the evaluation process. Instead of solely grading accuracy, which drives models to guess, the researchers advocate for a system that penalizes confident errors more heavily than uncertainty. This approach would be akin to standardized tests that incorporate negative marking for incorrect answers and offer partial credit for unanswered questions, discouraging blind speculation. The researchers emphasize that merely introducing a few uncertainty-aware tests is insufficient. The prevalent accuracy-based evaluations must be overhauled to deter guessing behavior. If the main evaluation metrics continue to reward random lucky guesses, the models will persist in learning to guess rather than express uncertainty when appropriate.

Sources : TechCrunch

Published On : Sep 08, 2025, 09:23

Automotive
Uber's Former CEO Makes Waves with New Ventures Amidst Tesla's Earnings Update

In the ever-evolving landscape of transportation, recent developments have come to the forefront, particularly surroundi...

TechCrunch | Jul 26, 2026, 16:25
Uber's Former CEO Makes Waves with New Ventures Amidst Tesla's Earnings Update
Startups
The Boring Company Eyes $4 Billion Funding Boost Amid Expanding Tunnel Ventures

Elon Musk's tunneling enterprise, The Boring Company, is reportedly negotiating a substantial funding round of $4 billio...

TechCrunch | Jul 25, 2026, 19:50
The Boring Company Eyes $4 Billion Funding Boost Amid Expanding Tunnel Ventures
Startups
The Future of Work: Executives Weigh In on AI's Impact on Gen Z Careers

As generative artificial intelligence continues to rise, uncertainty looms for the incoming Gen Z workforce. Leaders fro...

Business Insider | Jul 26, 2026, 10:15
The Future of Work: Executives Weigh In on AI's Impact on Gen Z Careers
AI
Hugging Face CEO Calls for Action Following AI Security Breach

In a dramatic turn of events within the AI landscape, Hugging Face faced a significant security breach involving an AI a...

Business Insider | Jul 25, 2026, 20:30
Hugging Face CEO Calls for Action Following AI Security Breach
AI
Shifting Focus: The Cost-Effectiveness of AI Models Takes Center Stage

In recent years, the AI sector has been intensely focused on identifying the most advanced models. While this pursuit re...

Business Insider | Jul 25, 2026, 13:10
Shifting Focus: The Cost-Effectiveness of AI Models Takes Center Stage
View All News