On Monday, Judge Eumi K. Lee of the Northern District of California denied the plaintiffs' motion for relief from a magistrate judge's discovery order in In re Google Generative AI Copyright Litigation. The practical effect: Google does not have to identify which of the class members' copyrighted works sit in its training datasets. Two further orders issued the same day tee up class certification.

The sentence at the centre of it

The order quotes Google's own position verbatim: "There is no list of works Google used for training, and no way for Google to generate one. The only record and the only means of determining what materials Google used to train its AI models is the training data itself." Magistrate Judge Susan van Keulen had found that "only the training datasets will reveal what putative class works were ultimately used," and that the plaintiffs' request would have Google duplicate work their own proposed expert had already done — "redundant and not proportional to the needs of the case."

What the conventional framing gets wrong

"Judge rules Google doesn't have to say what it trained on" overstates this in three directions. First, on the standard: relief under Rule 72(a) may be granted only if the order is "clearly erroneous or contrary to law," and Judge Lee's order says explicitly that she "may not simply substitute" her judgment for the magistrate's. She did not decide whether Google can produce such a list; she decided she could not find clear error in someone else's ruling that it need not. Second, Google's sentence is a litigating position quoted by the court, not a judicial finding of fact. The court's actual reasoning is that van Keulen did not simply take Google at its word — she weighed deposition testimony. Third, this is discovery. Nothing here touches fair use, infringement or damages.

Why a discovery ruling carries this much weight

Identifying which specific works sit in a training corpus is the load-bearing element of every AI training-data class action. Without a work list, ascertainability and predominance under Rule 23 get materially harder — you have to define a class by reference to something nobody can enumerate. This ruling shifts that burden onto the plaintiffs' own expert analysis of sample datasets, and it lands while the certification motion is pending before the same judge.

What is still live

The motion to certify the class, four motions to exclude experts, a motion for sanctions, and a motion to intervene. Google's two-page response on supplemental authority is due 8 September. A third order the same day terminated seven superseded motions and states on its face that it is administrative and does not affect the parties' rights — it is housekeeping, not seven denials.