
In light of Anthropic's recent $1.5 billion copyright settlement, the artificial intelligence sector is grappling with significant challenges surrounding the use of training data. With around 40 additional lawsuits pending over unlicensed data usage—including one targeting Midjourney for producing images of Superman—the urgency for a robust licensing framework has never been greater. Without such a system, AI companies could face a wave of copyright lawsuits that could potentially hinder the industry's growth for years to come. To address this pressing issue, a coalition of technologists and web publishers has unveiled a new initiative named Real Simple Licensing (RSL). This ambitious project aims to facilitate large-scale data licensing, provided that AI firms choose to participate. Major platforms such as Reddit, Quora, and Yahoo have already expressed their support for RSL. The critical question remains: will this momentum persuade leading AI labs to engage in negotiations? Eckart Walther, co-founder of RSL and co-creator of the RSS standard, emphasized the need for machine-readable licensing agreements on the internet. "That’s really what RSL solves," he stated in an interview. The RSL initiative marks a significant step towards establishing a technical and legal infrastructure for data licensing, a goal pursued for years by organizations like the Dataset Providers Alliance. Technically, the RSL Protocol delineates specific licensing options that content publishers can set, allowing AI companies to either negotiate custom licenses or adhere to Creative Commons terms. By integrating these terms into their "robots.txt" files in a standardized format, participating websites can clearly communicate the licensing conditions associated with their data. On the legal front, the RSL Collective has been formed to negotiate terms and collect royalties, much like ASCAP does for music or MPLC for film. This centralized approach aims to simplify the process for licensors, enabling them to define terms with multiple potential licensees simultaneously. A variety of well-known web publishers, including Yahoo, Reddit, Medium, and others, have already joined the collective, while some, like Fastly and Quora, have shown support without direct membership. Interestingly, the RSL Collective includes publishers with existing licensing agreements, such as Reddit, which reportedly earns around $60 million annually from Google for its training data usage. While companies can negotiate individual deals within the RSL framework, this collective approach may be the only avenue for smaller publishers who lack the clout to secure their own agreements. However, determining when royalties are owed for specific pieces of training data in AI models presents unique challenges. For instance, Google's AI Search Abstracts effectively tracks data sourced from the web in real time while ensuring attribution for each fact. Conversely, without proper logging during the training process, verifying that a specific document was incorporated into a large language model (LLM) can be exceedingly difficult. Despite these complexities, the creators of RSL are optimistic that AI companies can adapt. Doug Leeds, a co-founder of RSL and a former CEO of IAC Publishing, remarked, "Some of the licensing agreements they’ve already done have required them to be able to report on it, so it’s possible. It doesn’t have to be perfect. It just has to be good enough to get people paid." The larger question looms: will AI companies be willing to adopt this new system? While firms like ScaleAI and Mercor are already investing in quality data, the web has often been viewed as a repository of inexpensive, low-quality information. With freely available datasets like Common Crawl, convincing companies to pay royalties could prove challenging. Recent disputes, such as the one between CloudFlare and Perplexity, highlight the complexities of differentiating between web-scraping and machine-enhanced browsing. Leeds pointed to statements made by prominent AI figures, including Sundar Pichai, advocating for a licensing framework like RSL. Whether these calls are genuine remains to be seen, but the RSL team is determined to follow through on the momentum they've created. "They have said outwardly to everyone, something like this needs to exist," Leeds stated. "We need a protocol. We need a system." With RSL now in play, the industry may be on the cusp of significant change.
Phia, a shopping startup co-founded by Phoebe Gates, the daughter of Bill Gates, and Sophia Kianni, is under fire for al...
TechCrunch | Jul 11, 2026, 24:35
At Kalshi, an emerging leader in the prediction market sector in the U.S., a strikingly unconventional management style ...
Business Insider | Jul 11, 2026, 13:15As the global race to establish artificial intelligence infrastructure intensifies, a surprising constraint has emerged—...
Business Today | Jul 11, 2026, 11:10
In a significant legal move, Apple has initiated a lawsuit against OpenAI, claiming that the artificial intelligence com...
Business Insider | Jul 10, 2026, 21:10In a significant legal move, Apple has launched a lawsuit against OpenAI, accusing the AI company of stealing trade secr...
TechCrunch | Jul 10, 2026, 20:45