Reddit Sues Perplexity AI Over Indirect Scraping via Google Search
The social media giant alleges a conspiracy to bypass technical safeguards by extracting content through third-party search results.
Reddit is advancing a legal battle against Perplexity AI and several data-scraping firms, alleging a coordinated effort to bypass the platform's technical safeguards. The lawsuit marks a significant escalation in the conflict between content owners and AI companies over the unauthorized use of user-generated data.
According to the filing, Perplexity AI and scraping services—including SerpApi, Oxylabs UAB, and AWMProxy—used third-party scrapers to extract Reddit content indirectly via Google Search results. This method allowed the entities to circumvent Reddit's direct bot protections, such as IP-rate limits, captchas, and anomaly detection. Reddit argues that its existing licensing agreement with Google specifically prohibits the types of data access enabled by these circumvention methods.
The AI Data Arms Race
This litigation is part of a broader industry trend where major platforms are attempting to monetize their archives by forcing AI developers into paid licensing agreements. By blocking unauthorized scraping, Reddit aims to protect the commercial value of its data and prevent AI models from synthesizing answers that divert traffic away from the original platform.
Implications for Indirect Scraping
The outcome of this case could establish a critical legal precedent regarding "indirect scraping." If the court rules in Reddit's favor, it would affirm that bypassing a site's protections by utilizing a third-party search engine's index still constitutes a violation of the original site's terms and potentially the Digital Millennium Copyright Act (DMCA). Such a ruling would fundamentally alter how AI companies gather training data and how platforms enforce "no-crawl" directives across the web.
What's Next
Reddit continues to pursue the case despite a related legal loss by Google. The court must now determine if the indirect extraction of data through a search intermediary violates the specific licensing constraints and technical protections Reddit has in place. The decision will likely influence future contractual agreements between search engines and content publishers.