Here's the list of websites gig workers used to fine-tune Anthropic's AI models. Its contractor left it wide open.

Here's the list of websites gig workers used to fine-tune Anthropic's AI models. Its contractor left it wide open.

A recently leaked internal document has shed light on the guidelines used by gig workers at Surge AI for refining Anthropic's artificial intelligence. This spreadsheet details which websites were authorized for use and which were excluded, aiming to enhance the AI's ability to communicate in a manner deemed more 'helpful, honest, and harmless.' Among the sources permitted for consultation are reputable institutions such as Bloomberg, Harvard University, and the New England Journal of Medicine. Conversely, major outlets like The New York Times and Reddit are explicitly banned. Anthropic has distanced itself from the document, stating that it was generated by Surge AI without the company's knowledge or involvement. An Anthropic representative emphasized, "We were unaware of its existence until today and cannot validate the contents of the specific document since we had no role in its creation." The practice of utilizing external websites to refine AI models is common in the industry, where companies often collaborate with data-labeling startups like Surge. Internal project documents reveal that Surge's work was focused on making Anthropic's AI more relatable while minimizing the risk of generating offensive content. Notably, many approved sources have copyright restrictions, and institutions such as the Mayo Clinic and Cornell University have confirmed that they lack agreements with Anthropic regarding the use of their content for AI training. Initially, the sensitive spreadsheet was accessible via Google Drive, but Surge restricted access shortly after inquiries were made about its contents. A spokesman for Surge commented, "We take data security seriously, and documents are restricted by project and access level where possible," highlighting their ongoing investigation into the incident. This event marks another instance of a data-labeling startup inadvertently exposing sensitive AI training materials. Surge's competitor, Scale AI, faced a similar situation, leading to a tightening of document security protocols. A spokesperson for Google Cloud clarified that sharing settings are typically restricted by default, and any changes are at the discretion of the customer. Surge reportedly achieved $1 billion in revenue last year and is currently raising funds at a valuation of $15 billion, while Anthropic's latest valuation stands at $61.5 billion. Their Claude chatbot is recognized as a formidable competitor in the AI landscape, particularly against ChatGPT. The spreadsheet, created in November 2024, serves as a comprehensive guide for gig workers detailing over 120 authorized sources across various domains, including academia, healthcare, law, and finance. It enumerates prestigious universities and respected medical journals, while also maintaining a blacklist of over 50 disallowed sources, primarily consisting of prominent media outlets. The reasons for the inclusion or exclusion of specific sources remain unclear. Legal experts suggest that the blacklist could reflect the responses of websites that have actively sought to restrict their content from being used by AI companies, either through direct requests or automated methods. Recent lawsuits have highlighted these tensions, as seen in Reddit's legal action against Anthropic for unauthorized access to its site. Surge contractors employed this list during a critical phase of AI model training, known as reinforcement learning from human feedback (RLHF). This process entails human evaluators assessing and improving chatbot responses, a task that does not directly involve feeding web data into the AI model, but is nonetheless essential for its development. The legal complexities surrounding copyright in AI training processes continue to evolve, with significant implications for the future of AI development.

Sources : Business Insider

Published On : Jul 23, 2025, 09:15

Startups
OpenAI Welcomes New Financial Minds to Its Leadership Team

OpenAI has announced the addition of two prominent financial professionals, David Vélez and Robin Vince, to its boards o...

CNBC | Jul 21, 2026, 21:25
OpenAI Welcomes New Financial Minds to Its Leadership Team
AI
Rumors of Anthropic's Acquisition of Robotics Firm Ignite AI Community

This year has been monumental for AI acquisitions, with companies like Anthropic and OpenAI aggressively expanding their...

TechCrunch | Jul 22, 2026, 03:45
Rumors of Anthropic's Acquisition of Robotics Firm Ignite AI Community
Startups
Apple Set to Introduce Flexible Upgrade Program for Devices

Apple is on the verge of launching an innovative program called the 'Apple Upgrade,' developed in collaboration with Kla...

Business Today | Jul 22, 2026, 04:25
Apple Set to Introduce Flexible Upgrade Program for Devices
AI
Google Unveils Trio of New Gemini AI Models, Setting the Stage for Future Innovations

On July 21, Google introduced three new lightweight models to its Gemini AI lineup: Gemini 3.6 Flash, 3.5 Flash-Lite, an...

Business Today | Jul 22, 2026, 05:20
Google Unveils Trio of New Gemini AI Models, Setting the Stage for Future Innovations
Gaming
Sony's Shift to Digital-Only Games Raises Concerns Over Resale Market

In a significant departure from its past, Sony has announced plans to cease physical disc production for new games on it...

CNBC | Jul 22, 2026, 24:20
Sony's Shift to Digital-Only Games Raises Concerns Over Resale Market
View All News