South Minneapolis News

collapse
Home / Daily News Analysis / Reddit’s AI copyright lawsuit against Perplexity can move forward.

Reddit’s AI copyright lawsuit against Perplexity can move forward.

Aug 04, 2026  Twila Rosenbaum  7 views
Reddit’s AI copyright lawsuit against Perplexity can move forward.

In a significant development for the intersection of social media and artificial intelligence, a federal judge has rejected Perplexity AI's attempt to dismiss a copyright lawsuit brought by Reddit. The ruling allows Reddit's claims to advance, setting up what could become a landmark case over the use of user-generated content in AI training systems.

The lawsuit, filed earlier this year, accuses Perplexity and three data-scraping services of systematically extracting Reddit posts and comments without authorization. Reddit alleges that the companies ignored its terms of service, circumvented technical protections, and used the scraped data to build and power Perplexity's AI-powered answer engine. Perplexity had moved to dismiss the case, arguing that its actions constituted permissible use of publicly available online content. The judge disagreed, finding that Reddit had stated plausible claims for relief.

The core of the dispute

At the heart of the case is how Reddit controls access to its platform. Reddit hosts an enormous archive of conversations, advice, and discussions, making it a valuable data source for training large language models. While Reddit provides an official application programming interface (API) for developers, it requires users to agree to specific terms governing how they can access and use content. The company has also implemented measures like robots.txt, a standard file that instructs web crawlers which pages they may or may not access for automated scraping.

Reddit claims that Perplexity and the named data-scraping services bypassed these protections. Instead of using the official API, they allegedly used automated tools to harvest content directly from the site, in violation of Reddit's policies. The lawsuit contends that this practice not only infringes Reddit's copyright in the content but also undermines its ability to license that content legitimately.

In a statement, Reddit's chief legal officer, Ben Lee, said the ruling “brings us one step closer to holding bad actors accountable.” He added, “Reddit supports responsible access to public content, but we oppose companies that bypass our protections, ignore our rules, and profit off our communities without permission.”

Lee's statement underscores a broader tension within the tech industry. On one hand, many companies regard publicly accessible web content as fair game for machine learning. On the other hand, platforms like Reddit argue that public availability does not imply unrestricted use. The distinction between “public” and “free to use” is central to the case.

Reddit's growing enforcement push

Reddit has become increasingly aggressive in protecting its data. In 2023, the company announced that it would begin charging for API access, a move that sparked widespread protests among third-party app developers. The change was driven in part by the realization that AI companies were benefiting enormously from Reddit's content without contributing anything back to the platform or its communities.

Earlier that year, Reddit also entered into a licensing agreement with Google to allow the search giant to use Reddit content for AI training. According to reports, the deal was worth around $60 million annually. Reddit has since pursued similar arrangements with other companies. However, not all AI firms have been willing to pay. Perplexity, which operates a conversational search engine, has often cited the open web as its training ground, drawing criticism from publishers and platforms alike.

The three data-scraping services named in the lawsuit remain unidentified in public court filings, but industry observers say they are likely firms that specialize in collecting web data at scale for AI vendors. These services often operate in a legal gray zone, promising access to hard-to-scrape sources despite potential policy violations.

The value of Reddit's content

Reddit describes itself as the “front page of the internet,” and for AI researchers, it is a rich repository of human expression. Users discuss everything from niche hobbies to political news, and they do so in a conversational style that mirrors natural language. This makes the platform an ideal dataset for teaching AI models to understand context, humor, and sentiment. Companies like Perplexity rely on such data to deliver answers that feel current and human, rather than mechanical.

The economic stakes are high. Reddit's licensing agreements with AI companies have provided a new revenue stream for the platform, which historically struggled to monetize its massive user base. By going public in 2024, Reddit signaled its intention to capitalize on its data assets. Each unauthorized scrape, the company argues, undermines that business strategy and reaps rewards that should rightfully go to Reddit and its contributors.

At the same time, Reddit's approach has raised questions about control over user-generated content. While Reddit's terms of service grant it a license to user content, many users are not aware that their posts might be used to train AI systems. Reddit has said that it supports “responsible access” and highlights its commitment to community privacy. Critics, though, note that the platform's primary motivation is financial, not altruistic.

Why the ruling matters

The judge's decision to deny Perplexity's motion to dismiss does not decide the merits of the case. It simply means that Reddit's allegations are enough to proceed to the next phase. This is a lower bar than a trial, but it is still a significant milestone. In many copyright cases involving AI, courts are still determining whether scraping and using data for training counts as fair use or infringement.

This lawsuit is part of a wave of high-profile copyright challenges against AI companies. The New York Times has sued OpenAI and Microsoft over the use of its articles in ChatGPT. Getty Images has pursued legal action against Stability AI for using its photos to train Stable Diffusion. Authors, comedians, and visual artists have also filed class-action suits. While each case involves different facts, they all grapple with a fundamental question: Does the transformative nature of AI outputs justify copying underlying training data without permission?

The outcome of those cases will likely shape the rules of the road for the AI industry. If Reddit succeeds in proving that Perplexity's scraping practices are unlawful, it could force AI companies to be more transparent about their data sources and to negotiate licensing deals with online platforms. Conversely, if Perplexity prevails, it might open the door to even more aggressive scraping, potentially undermining the economic model of content-driven websites.

Perplexity's position

Perplexity has consistently defended its practices. The company positions itself as an “answer engine” rather than a traditional search engine, using large language models to synthesize direct responses to user queries. Its chief executive officer has claimed that the company respects robots.txt and does not circumvent technical barriers. However, investigators and journalists have previously found evidence that Perplexity may have accessed content from sites that explicitly blocked automated scraping.

In this case, Perplexity argued that Reddit's copyright claims were preempted by the Digital Millennium Copyright Act's safe harbor provisions and that the platform had failed to identify specific instances of copying. The judge rejected those arguments, noting that Reddit had provided enough detail in its complaint to put Perplexity on notice of the alleged infringements.

The ruling also permits Reddit's claims against the three data-scraping services to proceed. Those services are said to have used proxies and other techniques to disguise their traffic, making it difficult for Reddit's automated defenses to detect and block them. Reddit's complaint reportedly includes technical evidence showing that scraped content was processed and delivered to Perplexity's servers.

What the future holds

As the case now moves toward discovery, both sides will face intense scrutiny. Reddit will need to show proof that it suffered concrete harm from Perplexity's alleged scraping. Perplexity will likely try to show that its use of Reddit content was incidental or transformative, or that the content was not subject to copyright protection because it was created by users.

One of the most complicated issues is the role of Reddit's users themselves. Many of the posts and comments on Reddit are submitted by individual users, not by Reddit itself. While Reddit asserts that it owns the right to sublicense user content by virtue of its terms of service, that claim has not been tested in court. Perplexity could argue that Reddit lacks standing to sue for infringement of content authored by others. The judge has not yet ruled on this issue, and it may become a central point of contention in later proceedings.

Legal analysts say the case could also influence how platforms negotiate with AI companies in the future. If Reddit is able to enforce its terms and technical protections against third-party scrapers, more platforms may be encouraged to seek similar remedies. Conversely, a ruling against Reddit could force platforms to reconsider the enforceability of their user agreements, which often contain broad grants of rights.

The judge's order does not outline a timeline for the next steps, but court records suggest that the discovery phase will begin promptly. Both parties have also been ordered to meet and discuss a possible case management schedule. Given the complexity of the AI-related evidence, the case could take months or even years to resolve. A trial, if one occurs, would likely attract intense media attention and could become a touchstone for the AI copyright debate.

For now, the legal ruling marks an early but meaningful victory for Reddit. It signals that courts are willing to examine the details of how AI companies obtain training data, and that claims of unauthorized scraping will not be dismissed lightly. In the words of Reddit's legal chief, the decision is a step toward accountability. Whether that accountability extends to Perplexity's business model remains to be seen, but the case is far from over.


Source: The Verge News


Share:

Your experience on this site will be improved by allowing cookies Cookie Policy