
Since 2024, the performance optimization team at Anthropic has implemented a take-home assessment for job candidates to evaluate their expertise. However, the rise of advanced AI coding tools has necessitated frequent updates to this test in order to counteract AI-assisted cheating. Team lead Tristan Hume detailed these ongoing challenges in a recent blog post. Hume noted that with each iteration of their Claude models, the test has required redesigning. For instance, Claude Opus 4 surpassed many human applicants when given the same time constraints, allowing for differentiation among candidates. Yet, the subsequent Claude Opus 4.5 proved capable of matching even those top human performers, complicating the assessment process. The reliance on take-home tests without in-person oversight raises significant concerns about the potential for AI-driven cheating, which could enable unqualified candidates to excel. "Under the constraints of the take-home test, we no longer had a way to distinguish between the output of our top candidates and our most capable model," Hume explained. This dilemma is not isolated to Anthropic; educational institutions globally are grappling with similar issues as AI tools infiltrate academic integrity. Fortunately, Anthropic is well-suited to tackle this challenge. Hume ultimately crafted a new assessment format focused less on hardware optimization, making it sufficiently inventive to confound current AI capabilities. In a collaborative spirit, he also invited readers to engage with the original test, encouraging those who could outperform Opus 4.5 to share their solutions.
This past week has been challenging for the stock market, driven by several significant forces that have created turbule...
CNBC | Jul 25, 2026, 20:05
In a lively library setting in South Philadelphia, Charlie Bailey, a local librarian, humorously noted, "Everybody’s on ...
TechCrunch | Jul 25, 2026, 16:20
Prentis, a newly established AI research lab, is making waves in the tech industry as it prepares to raise $100 million ...
TechCrunch | Jul 25, 2026, 24:00
Last week, OpenAI made its debut in the hardware landscape with the launch of Micro, a stylish keypad designed to integr...
TechCrunch | Jul 25, 2026, 24:40
In a dramatic turn of events within the AI landscape, Hugging Face faced a significant security breach involving an AI a...
Business Insider | Jul 25, 2026, 20:30