Microsoft filed two redacted public summary-judgment briefs in In re OpenAI, Inc. Copyright Infringement Litigation on 4 September — Document 1690 for the books plaintiffs, 42 pages, and Document 1698 for the news plaintiffs, 46 pages. Both analyse the same evidence. They arrive at rates that differ by a factor of about 2,500.
The same corpus, two thresholds
The books brief, at page 22: plaintiffs' expert "reviewed 8.2 million conversations from Microsoft's LLM-based Copilot product and found 24 Copilot responses that included 30 matching words … That is a .00029% rate of regurgitation. And for 202 out of 212 of Books Plaintiffs' asserted works, Shan found no regurgitation in the logs." The news brief, at page 29, on the same 8.2 million: expert Tom Goldstein searched for 16-word matches and "could find only 59,545 conversations", which "yields a miniscule 0.73% rate of matches." Thirty words against sixteen is the entire difference.
The denominator Microsoft disowns and then uses
The news brief says plainly: "The 8.2 million were not random—they were the chats requested by Plaintiffs because they hit on keywords implicating use of News Plaintiffs' websites, and therefore the most likely to contain News Plaintiffs' works." One page later it divides by that number to produce 0.73%. A rate computed on a corpus assembled to maximise hits is a floor on an enriched sample, not a base rate over Copilot traffic — and the brief supplies both halves of that objection itself.
Two grounds that are not fair use at all
The MDL is covered as a fair-use fight. Two of Microsoft's four grounds are not. One is an implied licence from a missing meta tag: summary judgment is sought for any claim based on webpages lacking a NOARCHIVE tag after September 2023. The other is a damages cap — publishers are "entitled at most to one award per issue, not per article", because they register issues as collective works, citing Bryant v. Media Right Prods. That is the ground that decides whether exposure runs to millions or billions.
What the received framing gets wrong
"24 in 8.2 million" has been reported on its own. It is half of what Microsoft filed on the same day, off the same logs. Whichever figure a reader is shown, the honest version is that the rate is a function of a threshold the court has not yet chosen — and Judge Stein will have to pick one before he can weigh the third fair-use factor. Microsoft's expert also reports that just 1.3% of end users use Copilot to research current events at all.
