The Seattle Times and Newsday Sue OpenAI and Microsoft Over Paywall Scraping
Two major news organizations allege the AI giants bypassed paywalls to train models, threatening the economic survival of journalism.
The Seattle Times and Newsday have filed a joint federal lawsuit against OpenAI and Microsoft, alleging the companies systematically stole copyrighted content to build their generative AI tools. The legal action, filed in the U.S. District Court for the Southern District of New York, marks a significant escalation in the battle over how AI models are trained on professional journalism.
According to the filing, OpenAI and Microsoft methodically scraped copyrighted news articles, specifically bypassing paywalls to access premium content without permission. The plaintiffs argue that this data was used to train models like ChatGPT and Microsoft Copilot, which now produce paraphrased versions of their reporting. By creating an AI-generated alternative to original news, the publications claim the tech companies are directly undermining their business models and eroding critical advertising revenue and subscription growth.
The Battle for Fair Use
This litigation is part of a widening rift between the media industry and the AI sector. The case mirrors the 2023 legal challenge launched by The New York Times and recent actions taken by CNN against Perplexity. While some media organizations, such as AP and Vox Media, have chosen to sign licensing partnerships with OpenAI to monetize their archives, The Seattle Times and Newsday are opting for a judicial resolution. This divide highlights a fundamental disagreement over whether the ingestion of copyrighted data for AI training constitutes "fair use" or systemic theft.
An Existential Threat to News
At the heart of the lawsuit is the claim that generative AI creates a parasitic relationship with the news industry. In their legal filing, the publications describe the process as "a snake eating its own tail," arguing that the journalism industry could eventually become "broken beyond repair." The plaintiffs contend that when AI models reproduce passages or summarize articles, they remove the incentive for users to visit original news sites, effectively using the industry's own high-quality data to replace the industry itself.
Legal Implications and Next Steps
The outcome of this case could establish a critical legal precedent regarding the boundaries of AI training. Specifically, the court will need to determine if bypassing paywalls to scrape data constitutes a distinct violation of copyright law beyond the act of training itself. If the court rules in favor of the publications, AI companies may be forced to pay significant damages and negotiate retroactive licensing deals for all paywalled content used in their training sets. For now, the industry awaits a response from OpenAI and Microsoft as the case moves through the federal court system.