These LLMs are the best at resisting Russian propaganda

These LLMs are the best at resisting Russian propaganda

In an era where large language models (LLMs) are increasingly relied upon for answers to complex issues, concerns have arisen regarding their potential to disseminate propaganda from foreign adversaries. In response, the Estonian Language Institute (ELI), with government backing, has introduced a new benchmark called the "Propaganda Resistance" ranking. This initiative evaluates numerous LLMs based on their capability to refrain from endorsing narratives often promoted by the Russian Federation. Estonia, a former Soviet state that regained independence only a few decades ago, is particularly sensitive to the influence of what they consider misleading narratives from Russia. To address this, the ELI, in collaboration with the volunteer-driven Estonian defense group Propastop, identified 14 key areas where they believe Russian influence operations attempt to shape public opinion. These categories encompass topics such as the status of Crimea, justifications for the ongoing conflict in Ukraine, historical perspectives on NATO, and Russia's annexation of Baltic states during World War II. For each identified propaganda category, the researchers crafted a set of questions that were designed to be neutral, biased with misleading assumptions, or aimed at provoking the model into spreading misinformation. These questions were presented in English, Estonian, and Russian and evaluated by an independent AI model, which was adjusted to reflect the insights of Propastop experts. The assessment focused on the models' ability to counteract propaganda narratives autonomously, without assistance from external tools like web searches. Among the proprietary models evaluated, Anthropic's Claude series stood out, securing top positions on the benchmark. Notably, the recent iterations of its Sonnet and Opus models captured six of the top ten rankings. Opus 4.7 emerged as the highest performer overall, achieving an impressive "Exemplary" rating for correctly addressing 77 percent of the questions, while only receiving a "mediocre" rating on 2 percent, culminating in a remarkable mean score of 94.9 out of 100 on this new standard.

Sources : Ars Technica

Published On : Jun 04, 2026, 20:45

Computing
Market Turbulence: Four Key Factors Impacting Stocks This Week

This past week has been challenging for the stock market, driven by several significant forces that have created turbule...

CNBC | Jul 25, 2026, 20:05
Market Turbulence: Four Key Factors Impacting Stocks This Week
Streaming
Kalshi Challenges Netflix Over Controversial Documentary Trailer

Kalshi, the prediction market platform, has taken significant legal steps against Netflix, sending a cease-and-desist le...

TechCrunch | Jul 25, 2026, 17:10
Kalshi Challenges Netflix Over Controversial Documentary Trailer
Startups
AI Transformations Lead to Major Job Cuts at Tech Giants

Monday.com, the innovative work management platform based in Tel Aviv, has recently announced significant layoffs, attri...

TechCrunch | Jul 26, 2026, 01:45
AI Transformations Lead to Major Job Cuts at Tech Giants
Cybersecurity
The Elusive Phineas Fisher: The Hacktivist Who Took Down Spyware Giants

In the realm of cybersecurity, few figures are as intriguing as Phineas Fisher, a hacker who has evaded capture for near...

TechCrunch | Jul 25, 2026, 21:00
The Elusive Phineas Fisher: The Hacktivist Who Took Down Spyware Giants
Computing
Reclaiming Control: Librarians Host Workshops to Help People Navigate AI Tools

In a lively library setting in South Philadelphia, Charlie Bailey, a local librarian, humorously noted, "Everybody’s on ...

TechCrunch | Jul 25, 2026, 16:20
Reclaiming Control: Librarians Host Workshops to Help People Navigate AI Tools
View All News