This researcher has a new way to measure AI performance. It's BS, literally.

This researcher has a new way to measure AI performance. It's BS, literally.

In a groundbreaking development within AI evaluation, Peter Gostev, the AI capability lead at Arena, has introduced an innovative tool named "BullshitBench." This benchmark presents a series of intentionally nonsensical questions to assess whether advanced language models can recognize and reject absurdity instead of providing misleadingly confident answers. Since its launch in late February, BullshitBench has garnered significant attention, amassing over 1,200 stars on GitHub. The concept is simple yet intriguing: models are challenged with prompts that may sound sophisticated but fall apart under scrutiny. One of the standout questions humorously asks about the viscosity of a sales pipeline, showcasing the absurdity of the tasks. Gostev aimed to capture the phenomenon where AI models often seem to lack a fundamental understanding of context. Surprisingly, the results revealed a stark disparity in performance. Major models, including Google’s Gemini 3.0, struggled to identify and reject nonsensical queries, with less than half demonstrating the ability to push back against such prompts. Interestingly, the study uncovered that models designed for reasoning often performed worse than their simpler counterparts. Rather than dismissing illogical questions, these models attempted to reinterpret them, indicating a potential flaw in their approach to judgment. "They're not necessarily spending time to try and make sure the question makes sense, but they really try hard to make sure that they can answer the question," Gostev explained. This observation raises important questions about the nature of intelligence in AI. While modern systems excel in complex tasks, they sometimes struggle with basic judgment and context that humans navigate intuitively. BullshitBench highlights a notable gap in AI capabilities: the distinction between computational power and the nuanced understanding required for sound judgment. However, not all models fared poorly. Anthropic's recent systems performed significantly better, often rejecting nonsensical queries effectively. Gostev noted that this success might stem from Anthropic's emphasis on developing robust core models rather than focusing heavily on reasoning processes that can complicate responses. The ongoing rivalry in AI performance metrics is evident, with Anthropic’s models consistently outperforming those of competitors like OpenAI in various assessments over the past nine months. As the landscape of AI continues to evolve, benchmarks like BullshitBench are crucial for understanding the strengths and weaknesses of these powerful systems.

Sources : Business Insider

Published On : Mar 25, 2026, 09:30

AI
The Rise of AI Distillation: A Controversial Technique Sparks Debate in Tech and Government

In recent discussions, a once-obscure topic in artificial intelligence has surged to the forefront of debates among tech...

CNBC | Jul 25, 2026, 12:15
The Rise of AI Distillation: A Controversial Technique Sparks Debate in Tech and Government
Computing
Reclaiming Control: Librarians Host Workshops to Help People Navigate AI Tools

In a lively library setting in South Philadelphia, Charlie Bailey, a local librarian, humorously noted, "Everybody’s on ...

TechCrunch | Jul 25, 2026, 16:20
Reclaiming Control: Librarians Host Workshops to Help People Navigate AI Tools
Cybersecurity
Vietnam's Controversial Social Media Proposal: A New Approach for Youth Engagement

Vietnam is contemplating a distinctive approach to youth social media regulations, diverging from the more common outrig...

TechCrunch | Jul 24, 2026, 21:25
Vietnam's Controversial Social Media Proposal: A New Approach for Youth Engagement
Science
Unlocking the Secrets of Quantum Gravity: Can AI Help Physics Make a Leap?

The realm of scientific research is undergoing a profound transformation, fueled by the rapid advancements in artificial...

Business Today | Jul 25, 2026, 24:30
Unlocking the Secrets of Quantum Gravity: Can AI Help Physics Make a Leap?
Science
Finland Unveils World's Largest Sand Battery to Tackle Renewable Energy Challenges

In a groundbreaking move to address the critical issue of renewable energy intermittency, a small town in southern Finla...

CNBC | Jul 25, 2026, 05:35
Finland Unveils World's Largest Sand Battery to Tackle Renewable Energy Challenges
View All News