A class action by visual artists against Stability AI, Midjourney, DeviantArt, and Runway survived two rounds of motions to dismiss and is the furthest-along test of whether training an image model on copyrighted work infringes it.
Andersen v. Stability AI is the case that will not go away, and that is the point. Filed in January 2023 by a group of visual artists, it has survived two rounds of motions to dismiss in the Northern District of California and reached the stage every AI copyright defendant wanted to avoid: discovery into training data.
| Item | Detail |
|---|---|
| Caption | Andersen et al. v. Stability AI Ltd. et al. |
| Court | U.S. District Court, Northern District of California |
| Docket | 3:23-cv-00201 |
| Judge | William H. Orrick |
| Filed | January 2023 |
| Defendants | Stability AI, Midjourney, DeviantArt, Runway AI |
| Posture | Past two motions to dismiss; in discovery |
| Trial | Reported as set for September 8, 2026. Confirm on the docket; trial dates in this posture move often. |
Stable Diffusion is a text-to-image model. Training it required a very large set of images, drawn substantially from the LAION dataset, which was assembled by scraping the open web. The named plaintiffs are artists whose work was in that dataset.
Their claim is that the training copies were themselves infringing reproductions, and that the resulting model can produce outputs in the style of, and sometimes recognizably derived from, their work. The complaint originally cast Stable Diffusion as something close to a compression of the training set.
The defense is technical and, on the mechanics, largely correct. The model stores weights, not images. It cannot retrieve a training image on demand the way a database can. From that the defendants argue that training is transformative and protected by fair use, in the way that a human artist studying a portfolio is not infringing.
Both descriptions can be true at once, which is why the case is hard. The model does not store the images. It also could not exist without having copied them.
The CourtListener docket carries the public filings, and the NYU Journal of Intellectual Property and Entertainment Law published a useful walkthrough of the pleadings.
Two questions, and they separate cleanly.
The input question is whether making copies of protected works to train a model is infringement, and if it is, whether fair use excuses it. The four fair use factors were written for humans reading books, not for a pipeline that copies ten billion images. The transformative-use inquiry from Campbell and Google v. Oracle is the battleground, and the fourth factor, market effect, is where artists have their strongest argument: a model trained on an illustrator's work competes with that illustrator.
The output question is whether any particular generated image infringes any particular work. That requires substantial similarity between specific images, which is a much narrower and much harder showing. Style is not protectable. A particular composition may be.
Induced infringement sits on top of the output question. If a model is marketed with prompts that invite users to generate work in a named artist's style, the vendor's own conduct comes into view.
If the trial goes forward, this becomes the first jury to hear a full record on generative image training. A verdict either way is going to the Ninth Circuit.
The more likely outcomes deserve equal weight. Cases of this size settle, and a licensing settlement here would set a price for training data without setting a precedent. Class certification, if it is contested, could also reshape the case before any jury sees it.
Watch what discovery surfaces regardless of outcome. Training data provenance records produced in this case are the closest thing the public will get to an audit of how these datasets were built.
Pull the current filings here: Andersen v. Stability AI. Trial-eve dockets move fast, with motions in limine, Daubert challenges, and pretrial orders landing within days of each other, so case alerts are more practical than checking manually.
Keeping track of which claims are live matters, because coverage of this case routinely describes dismissed theories as though they were still in it.
Direct copyright infringement against Stability AI survived the first motion to dismiss, on the reasoning that training involved making copies of protected images.
Induced infringement survived the second round, along with claims tied to how the products were marketed.
Several theories did not survive in their original form, including claims under section 1202 of the DMCA about the removal of copyright management information, which courts have generally read narrowly.
The right to publicity and unfair competition theories have been narrowed as the pleadings were amended.
Three categories of document decide this case.
Dataset provenance: how the training sets were assembled, what was in them, and what the companies knew about the copyright status of the contents.
Internal risk assessments: what employees and counsel said about copyright exposure before the products shipped. These documents are where fair use arguments go to die, because they show the company's own view of the market it was affecting.
Model behavior evidence: how often, and under what prompts, the models reproduce recognizable elements of specific works. That evidence goes to both infringement and the fourth fair use factor.
The text-side version of the same fight is New York Times v. OpenAI and Microsoft, where discovery over ChatGPT logs has produced its own landmark orders. For the Supreme Court's 2026 word on secondary copyright liability for intermediaries, read Cox Communications v. Sony Music. And for a dispute about AI systems accessing other companies' property rather than their content, see Amazon v. Perplexity.