Reddit's Copyright Scraping Lawsuit Survives Early Challenge
A federal judge allowed conspiracy claims to proceed against SerpApi, while a parallel Google case faltered on similar DMCA theories

A Divergent Path in DMCA Enforcement
A federal court decision late last week underscored the uncertain boundaries of copyright protection for publicly indexed web content. US District Judge Paul A. Engelmayer rejected most arguments from SerpApi, a web scraping service, to throw out claims that it conspired with Perplexity AI to harvest Reddit content from Google search results without authorization. The ruling stands in contrast to a separate case where Google itself faced dismissal of nearly identical legal theories just days earlier.
At DailyTechWire, we've tracked the intersection of scraping litigation and search infrastructure across Asia and North America for two years. What makes this dispute noteworthy is not the technology involved - scraping Google's structured data is table stakes for many API vendors - but the legal argument that doing so constitutes circumvention of access controls under the Digital Millennium Copyright Act. Reddit's theory rests on the proposition that Google's robots.txt files and search result formatting create a protective layer that third parties cannot lawfully breach, even when the underlying content is publicly visible.
Judge Engelmayer found that Reddit had presented enough factual detail to survive a motion to dismiss. Specifically, the platform alleged that SerpApi built a product designed to bypass Google's technical measures, and that Perplexity AI paid for that service to ingest Reddit posts at scale. The court concluded that these allegations, taken together, describe a plausible conspiracy to violate copyright protections.
The Contradiction in Parallel Litigation
Less than two weeks before Engelmayer's decision, another court dismissed Google's own lawsuit, which advanced a nearly identical DMCA circumvention claim. That judge determined Google had failed to show that content owners - such as Reddit - had ever granted the search giant authority to police scraping on their behalf. In other words, Google could not claim to enforce access controls for content it did not create or own.
The timing is notable. Reddit's case survived where Google's did not, yet both hinge on the same premise: that extracting data from search results violates the DMCA because it circumvents technical measures. Google announced plans to amend its complaint and continue the fight, but the early divergence in judicial reasoning suggests appellate courts may eventually need to resolve the question.
SerpApi's position is that both lawsuits represent an overreach. The company argues that Reddit and Google are attempting to "wall off the open Internet" by retroactively asserting control over content neither party authored. From SerpApi's perspective, scraping publicly accessible search results does not circumvent anything - it simply automates what a human user could do manually.
What Access Control Means in Practice
The legal dispute turns on a technical question: do Google's systems constitute access controls under the DMCA? Reddit's complaint appears to argue that any structured method Google uses to present search results - whether through result formatting, rate limiting, or robots.txt directives - qualifies as a technological protection measure. If that interpretation holds, it would significantly expand the scope of DMCA liability beyond its traditional application to paywalls, encryption, and login systems.
From an infrastructure standpoint, the implications are broad. Hundreds of companies in Asia and globally build products that parse search results, aggregate listings, or monitor rankings. Many operate in regulatory environments where scraping is permissible under competition or data portability frameworks. If US courts adopt Reddit's view, those business models face new legal risk whenever they touch American platforms or content.
The question of authorization also looms large. Reddit's agreement with Google reportedly includes terms that govern how Reddit content appears in search results. But whether those terms extend to third-party use of the results themselves remains contested. Judge Engelmayer did not resolve that issue - he simply found that Reddit's allegations were sufficient to proceed to discovery.
Perplexity's Role and the AI Angle
Perplexity AI, named as a co-conspirator in Reddit's lawsuit, operates a search-augmented answer engine that synthesizes information from multiple sources. The platform has faced scrutiny in recent months over its data sourcing practices, particularly from publishers who claim it reproduces their content without proper licensing. Reddit's case adds a new dimension: the allegation that Perplexity deliberately paid for a tool designed to evade Google's controls.
Perplexity has not publicly detailed its relationship with SerpApi, but the court filings suggest the company used SerpApi's infrastructure to access Reddit threads indexed by Google. If discovery bears out Reddit's claims, it could establish a pattern of conduct that other platforms might cite in their own litigation. The case also highlights a tension in the AI training and retrieval ecosystem: companies building knowledge engines need vast amounts of structured data, but the legal framework for acquiring that data remains unsettled.
The Broader Stakes for Search and Scraping
This litigation sits at the intersection of three trends: the commercialization of large language models, the growing assertiveness of platforms in policing derivative use of their content, and the fragmentation of legal standards around web scraping. In the European Union, the Digital Services Act and the Data Act introduce requirements that may conflict with DMCA-style access control theories. In India, the Competition Commission has signaled interest in ensuring that dominant platforms do not use technical measures to foreclose rival services. China's data security regulations impose their own set of constraints on cross-border data flows and automated collection.
For companies operating across jurisdictions, the divergence in legal outcomes between Reddit's case and Google's case adds uncertainty. A startup in Bengaluru that scrapes public e-commerce listings, or a research lab in Singapore that monitors social media sentiment, must now consider whether their activities could be recharacterized as DMCA circumvention if they touch US-based platforms. The fact that two US courts reached opposite conclusions on nearly identical facts suggests the law is far from settled.
What Happens Next
Reddit's case will now proceed to discovery, where SerpApi and Perplexity will be required to produce internal communications, contracts, and technical documentation. If Reddit can demonstrate that SerpApi marketed its product specifically as a way to bypass Google's controls, the conspiracy claim may gain traction. If, on the other hand, discovery reveals that SerpApi's tool is a general-purpose API with many legitimate uses, the case may narrow or settle.
Google's parallel lawsuit, meanwhile, will likely return with an amended complaint that addresses the authorization gap identified by the court. The company has significant resources and institutional interest in establishing that search results can be protected under the DMCA. A successful appeal or amended filing could create binding precedent that reshapes the scraping landscape.
For now, the message from Judge Engelmayer's decision is that conspiracy theories can survive early dismissal if the plaintiff alleges a clear commercial relationship and a purpose-built tool. Whether that standard will hold at summary judgment or trial remains to be seen. What is clear is that the legal boundaries of scraping, search, and copyright are being redrawn in real time, with consequences for platforms, tool vendors, and the thousands of companies that depend on publicly accessible data to build their products.


