Creators Start Filing Legal Claims Over Unauthorized Training Data
Authors, illustrators, and photographers are moving from outrage to courtroom strategy as evidence of scraped copyrighted work mounts across foundation models.

The Discovery Moment
When searchable databases of AI training corpora began appearing online, thousands of creators typed their names into search bars with a mix of curiosity and dread. Many found exactly what they feared: books, illustrations, photographs, and articles they had spent years producing, now ingested wholesale into language models and image generators without permission or compensation.
Author Kirk Wallace Johnson represents a growing cohort of creatives who moved quickly from discovery to legal action. His nonfiction works - multi-year investigative projects involving extensive research and reporting - appeared in datasets that had been scraped from shadow libraries and torrent sites, then used to train commercial chatbots. The emotional response he describes mirrors accounts from hundreds of other creators: initial shock, followed by anger at the scale of appropriation, and finally a decision to pursue formal legal remedies.
At DailyTechWire, we've tracked more than two dozen lawsuits filed against major AI developers since early 2023, with the pace accelerating throughout 2024 and into 2025. The cases span jurisdictions, but they share common threads - unauthorized reproduction, derivative works concerns, and challenges to the industry's reliance on "fair use" defenses that were written for a pre-transformer era.
The Legal Architecture Taking Shape
The litigation landscape divides into three broad categories. First are individual creator suits, often backed by intellectual property specialists working on contingency. These cases typically argue that training constitutes reproduction and that model outputs can function as market substitutes, undermining the original creator's economic rights.
Second are class actions aggregating claims from hundreds or thousands of rights holders. Visual artists filed early here, pointing to diffusion models that can generate outputs in recognizable styles or featuring trademarked characters. Writers followed, particularly after datasets like Books3 - a collection of nearly 200,000 pirated titles - became public knowledge.
Third, and perhaps most significant for the industry's future, are collective actions organized by guilds, unions, and professional associations. The Authors Guild, the News Media Alliance, and illustrators' collectives have all either filed suit or negotiated licensing frameworks that attempt to establish precedent for compensated training.
The defendants - typically Alphabet, Meta, Anthropic, OpenAI, Stability AI, and Midjourney - have mounted a consistent defense rooted in transformative use doctrine. They argue that training is non-expressive, that models learn patterns rather than store copies, and that outputs are sufficiently distinct from any single training example to avoid infringement.
But that argument has begun to fracture under judicial scrutiny. Courts in the Southern District of New York and the Northern District of California have allowed several cases to proceed past motions to dismiss, finding that questions of substantiality, market harm, and the commercial nature of the use require fact-finding. Discovery processes have started compelling model developers to disclose training methodologies, dataset provenance, and internal communications about copyright risk - materials the industry has fought hard to keep confidential.
Early Wins and What They Mean
A handful of settlements and preliminary rulings have shifted the risk calculus for AI labs. In late 2024, a visual artist collective reached a confidential settlement with one major image-synthesis company that included both a monetary component and a commitment to implement opt-out mechanisms for future training runs. While terms remain sealed, lawyers involved described the agreement as a "meaningful acknowledgment" that creators hold negotiating leverage.
More recently, a federal judge in Delaware issued a preliminary injunction limiting a generative video startup's use of a cinematographer's work, finding a likelihood of success on the merits of the copyright claim. The ruling hinged on evidence that the company had knowingly scraped premium content from a subscription platform, undermining any argument that the material was publicly available or abandoned.
These outcomes remain narrow, but they signal a judiciary willing to interrogate the AI industry's foundational assumptions. The notion that internet-scale scraping is legally permissive - a premise that underpinned the first wave of foundation model development - no longer enjoys the presumption of correctness it once did.
The Industry's Licensing Pivot
Facing mounting legal exposure and negative publicity, several frontier labs have begun inking licensing deals with publishers, photo agencies, and content platforms. Anthropic, for instance, has announced partnerships with news organizations that grant training access in exchange for revenue shares and attribution features. OpenAI has signed agreements with Condé Nast, Axel Springer, and the Associated Press.
These deals serve multiple strategic purposes. They provide legal cover, reducing the surface area for infringement claims. They offer a public relations counter-narrative, allowing companies to position themselves as partners rather than pirates. And they create a two-tier data ecosystem: licensed, defensible corpora for commercial products, and riskier scraped data relegated to research or lower-stakes applications.
But the licensing approach has drawn criticism from creators who see it as inadequate. Many agreements cover only prospective training, leaving past infringement unaddressed. Others involve bulk deals negotiated by publishers or platforms, with individual creators receiving little or no direct compensation. The economic terms often resemble micropayments - fractions of a cent per work - that critics argue bear no relationship to the value AI companies extract from the underlying content.
Photographer collectives and independent authors have been particularly vocal, arguing that the licensing model replicates the power imbalances of earlier digital disruptions. They point to the music streaming wars and stock photography market consolidation as cautionary tales, where initial promises of fair compensation gave way to razor-thin margins and winner-take-all platform dynamics.
What Comes Next
The legal confrontation over training data is still in its early innings. Appeals will take years to resolve, and circuit splits are likely, potentially setting up Supreme Court review. Meanwhile, legislative efforts are advancing in parallel. The European Union's AI Act includes provisions on transparency and copyright compliance. In the United States, several bills circulating in Congress would create statutory licensing regimes, safe harbors contingent on opt-out systems, or outright prohibitions on using copyrighted material without consent.
For AI developers, the uncertainty creates a dilemma. Retraining models on fully licensed datasets would be enormously expensive and might degrade performance if the available corpus shrinks significantly. Continuing with the status quo invites ongoing litigation, reputational damage, and the risk of injunctions that could halt product deployments.
For creators, the path forward is equally uncertain. Winning in court could establish valuable precedent and unlock damages, but litigation is slow and expensive. Negotiating collectively might secure better terms than individual deals, but requires overcoming coordination challenges in fragmented industries. And some worry that even successful legal action will arrive too late - that by the time courts rule, the market will have already shifted in ways that make traditional copyright enforcement irrelevant.
The stakes extend beyond any single lawsuit or licensing agreement. The question of whether AI training constitutes fair use will shape the economics of content creation, the structure of the internet, and the balance of power between platforms and the people who make what platforms distribute. For now, that question is being answered in courtrooms, one motion and one settlement at a time.


