The Black Box 009: The Ninth Circuit Wrote 18 Pages on AI and Copyright. It Never Reached the Training Question.
Bot Mutiny |
The Black Box 009: The Ninth Circuit Wrote 18 Pages on AI and Copyright. It Never Reached the Training Question.
September 25, 2026
At a hearing in the Northern District of California, a judge asked a lawyer a direct question. Did copying training data into Copilot violate the attribution requirement of the open-source licenses attached to that code?
Plaintiffs' counsel answered: "Perhaps it doesn't."
The district court, Judge Jon S. Tigar presiding, then observed that the "complaint is not about training. It just isn't." In the order that followed, granting in part the defendants' first motion to dismiss, it said so again in flatter language. "Plaintiffs do not allege they were injured by Defendants' use of licensed code as training data."
Nobody corrected it. Not at the hearing, not in writing, not in any brief filed afterward.
Zero
Zero paragraphs of the Ninth Circuit's September 16 opinion decide whether stripping copyright notices out of code before feeding it into a model violates the Digital Millennium Copyright Act.
That is not because the court found the practice lawful. It is because the court refused to look.
What the statute actually says
Section 1202(b) of the DMCA, enacted in 1998, makes it unlawful to intentionally remove or alter copyright management information, or to distribute works knowing that such information has been removed. Copyright management information, which the statute abbreviates as CMI, is defined at 17 U.S.C. § 1202(c) as information conveyed in connection with copies of a work. The title. The author's name. The copyright notice. The terms and conditions for use.
The DMCA does not require anyone to attach CMI to anything. But if a work carries it, section 1202(b) protects it from deliberate removal. Statutory damages run up to $25,000 per violation under section 1203(c)(3), which is the detail that makes the provision worth litigating over.
Open-source licenses are built on exactly this kind of information. The most common condition attached to publicly posted code is attribution, meaning a copy of the license, the author's name and the copyright notice must travel with any copy or derivative.
The plaintiffs are programmers who published copyrighted code under those licenses on public GitHub repositories. GitHub is owned by Microsoft. Copilot is a paid subscription service built jointly by GitHub and OpenAI on a modified version of OpenAI's Codex, trained on billions of lines of publicly available code including, per the complaint, all available public GitHub repositories.
They pleaded two distinct theories, and the difference between them is the entire story.
The first, which the Ninth Circuit called the input theory, is that the defendants violated section 1202(b)(1) at the training stage, by removing CMI from class members' code before feeding the stripped code into the model.
The second, the output theory, is that Copilot sometimes returns memorized training data to users without the source code's CMI attached.
The panel decided the second. It declined to consider the first.
The boring explanation, at full strength
Forfeiture is ordinary appellate practice, and the case for applying it here is not weak.
A party gets one set of chances to frame its own complaint. These plaintiffs had several. The case went through two rounds of dismissals and amendments before the operative complaint was whittled down to three claims. Across all of it, the district court stated plainly, more than once, that it did not understand training to be at issue. Counsel had the opportunity to say it had that wrong and did not take it.
Appellate courts decline unpreserved theories for a reason. Without the rule, a party could hold a theory in reserve, lose on the record it built, and then argue something new to a panel that has no trial-court findings to review. The district court gets no chance to rule, the other side gets no chance to develop a record, and the appeal becomes a do-over.
The panel also gave the plaintiffs' preservation argument specific consideration rather than waving at it. Plaintiffs pointed to two things. One was their own statement that the DMCA "makes the mere removal of CMI from digital copies illegal before distribution of copies." The court found that this appeared inside an explanation of the difference between sections 1202(b)(1) and 1202(b)(3), and expressed no disagreement with the district court's conclusion. The other was a filing by OpenAI that supposedly acknowledged a pre-distribution removal claim. The court read it and found it said the opposite, that the second amended complaint "admits that no pre-distribution [CMI] removal occurred."
On that record, the forfeiture holding is defensible. A reader who stops here has a complete and accurate account of why the training question went undecided, and it is an unremarkable one about litigation procedure.
Why that is not the whole record
Four features of the opinion cut against treating this as routine.
First, the panel itself conceded the complaint contains the allegation. "Read in isolation, the complaint might be understood to assert such a theory," Judge Miller wrote, and then quoted it: defendants "removed or altered CMI from open-source code that is owned by Class members after the code was uploaded to a GitHub repository by incorporating it into Copilot with its CMI removed." The forfeiture does not rest on the pleading lacking the claim. It rests on counsel failing to contradict a court that said the pleading lacked it.
Second, the question certified for interlocutory appeal under 28 U.S.C. § 1292(b) was narrow and technical. The district court certified whether sections 1202(b)(1) and (b)(3) "impose an identicality requirement." It did not certify which theory was live. The panel nonetheless reached beyond the certified question when it suited the analysis, addressing Article III standing on the ground that its jurisdiction "applies to the order certified" and is "not tied to the particular question formulated by the district court." It exercised that discretion to reach standing. It did not exercise it to reach training.
Third, the panel found the plaintiffs did have standing, and the reasoning matters. It credited the complaint's citation of academic research finding that large language models will sometimes "emit the memorized training data verbatim," a problem that "will likely get worse as models continue[] to scale." It treated GitHub's own duplicate-detection feature, which lets users block output matching public code at 150 characters or more, as "some evidence that Copilot can and does emit literally identical copies of code." A company does not build a filter for a thing that does not happen. The court accepted that verbatim reproduction is real and ongoing, then held that the statute does not reach it.
Fourth, the panel said the doctrine does not fit. Print-media analogies such as defacing the title page of a book, it wrote, "do not map neatly onto the emerging digital technologies like artificial intelligence to which the DMCA's protections also apply." That is a court naming a gap in the law it is applying. The vehicle for closing that gap was the input theory, and the input theory had already left the case.
What the panel actually held on outputs
On the merits of the output theory, the holding is clean and will be durable.
Copilot, the panel wrote, "does not look up and reproduce stored work but rather creates new work." A model that infers statistical patterns and predicts a likely completion is not performing an act with respect to CMI attached to an existing work. One who creates a new work and fails to include CMI "cannot be said to have 'removed' or 'altered' anything."
The panel dismantled the district court's own framing along the way. "Identicality," it said, "is something of a misnomer," better understood as a gloss on the statutory words remove, alter and copies than as an independent element. Two works need not be literally identical. Minor cosmetic changes will not protect a defendant who substantially reproduces a protected work and strips the notice.
The distinction it drew instead was to a search engine, which retrieves and displays copies of materials that already exist. "If Copilot functioned like a search engine and produced outputs that were identical to plaintiffs' code but did not contain CMI, then plaintiffs might have a stronger claim."
And the closing rationale was about damages exposure. Section 1203(c)(3) permits up to $25,000 per violation, against section 504(c)(1)'s cap of $30,000 per work. Reading section 1202(b) to cover substantial similarity, the panel wrote, would "supplant traditional copyright protections and subject defendants to potentially ruinous liability." It declined the invitation to "transform run-of-the-mill copyright-infringement claims into DMCA claims."
What the instrument cannot see
The practical consequence is a published, precedential opinion, the first at the circuit level on this question, that says generative output is not CMI removal, and says nothing at all about whether the training pipeline is.
Those are different acts, performed at different times, by different systems. Stripping license headers off a corpus before training is a discrete, auditable engineering step. Emitting a probable completion is not. A rule about the second tells you nothing about the first, and the opinion is explicit that it is not deciding the first.
It will be cited for both anyway. That is what happens to the only circuit precedent in a field.
The forfeiture has been noted, briefly. One law blog tracking AI litigation recorded that the input theory "was waived." But the framing that travels is the other one. The Electronic Frontier Foundation, which appeared as amicus for itself and Public Knowledge, posted the same day under the headline "Victory! Appeals Court Rejects Expansive New Copyright Claim," and described the holding this way: "the absence of copyright information from a new work does not mean, by itself, that someone illegally removed it." That is a correct statement about outputs. The post does not separate outputs from training, and does not mention that the training theory went undecided. A reader comes away understanding that a copyright claim against an AI company failed, without learning which claim was never brought to judgment.
For a developer who published code under a license requiring attribution, the answer to whether that license survives contact with a training pipeline is, as of September 16, still unknown. What is known is that the question is harder to ask now, because the leading case declined to answer it and future defendants will point at that silence.
The record, stated cleanly
What the documents establish:
The Ninth Circuit affirmed dismissal of the DMCA claims in Doe v. GitHub, Inc., No. 24-7700, on September 16, 2026, in a published opinion by Judge Eric D. Miller. The plaintiffs have Article III standing. The output theory fails because a model that generates new work does not remove or alter CMI from an existing copy. The input theory was forfeited and not considered. The forfeiture rests on counsel's failure to correct the district court's stated understanding, including an answer of "Perhaps it doesn't" at a hearing. The complaint contains language the panel agreed could be read as asserting the input theory. The certified question concerned identicality, not theory selection. The breach of contract claims were not dismissed and remain pending in the district court.
What the documents do not establish:
They do not establish that removing CMI from code at the training stage is lawful. They do not establish that it is unlawful. They do not establish that the plaintiffs would have won on the input theory, or that the panel would have reached it had it been preserved. They do not establish why counsel answered as they did, and nothing in the record explains it. They do not establish that Copilot reproduced any specific plaintiff's code on any specific date, only that the complaint's allegations of substantial risk were plausible enough to survive a motion to dismiss. They do not establish anything about copyright infringement, which the panel expressly reserved.
Close
The opinion runs eighteen pages and cites Nimmer, Webster's Third, and a Fifth Circuit decision from August. It is careful work. It resolves the question it was given.
The question it was given was not the one that matters most to anyone who has ever attached a license to code and posted it in public. That question was answered in a hearing room, by a lawyer, in three words, before the appeal existed.
Sources: Doe v. GitHub, Inc., No. 24-7700 (9th Cir. Sept. 16, 2026) (published opinion, Miller, J.), on interlocutory appeal from No. 4:22-cv-06823-JST (N.D. Cal., Tigar, J.), 28 U.S.C. § 1292(b). Statutory text at 17 U.S.C. §§ 1202(b), 1202(c), 1203(c)(3), 504(c)(1).