Anthropic published a post on 26 August describing a pilot that gave external academic partners access to findings drawn from Claude usage data, including labs at Stanford and the University of Oxford. It is a genuine step — very little of what is known about how people actually use frontier assistants comes from outside the companies that operate them. It is also considerably more constrained than the framing suggests.

What the researchers actually received

Not conversations. Anthropic states that researchers "never accessed raw conversations, only aggregated outputs after the same legal and privacy review as our internal work." The unit of access is a reviewed aggregate produced by Anthropic's own analysis system in response to a question the researcher posed. The researcher specifies the question; Anthropic's pipeline produces the answer; a review gate sits between the two.

The constraint that matters most

Empirical work on a corpus is iterative. You ask, you see the shape of the answer, you discover your category was wrong, you redefine it and ask again. Anthropic is explicit that this was not available: "External partners couldn't do that, since repeated privacy review before sharing each dataset would have made the study infeasible." The workaround was to develop and test questions against WildChat, a public conversation dataset, and then run the settled version once against Claude data. The researchers refined their instruments on a different corpus from the one they were studying.

What the common framing gets wrong

Two numbers are travelling from this work: that over half of conversations involve delegating consequential tasks, and that nearly three-quarters show people directing work while Claude assists. Those sound contradictory — most work delegated, most work human-led — and they are not, because they are computed over different denominators. The delegation figure is against the conversations studied; the direction figure is against the subset examined for human-AI collaboration roles. Quoting them side by side as though they describe the same population produces a paradox that does not exist in the source. Anthropic also notes that under 5% of categories and conversations in each study involved removed or altered policy-violating content, which bounds how much the reviewed aggregate departs from the underlying data.

Why the design is still worth something

The alternative to constrained access is no access, and the field has been running on the latter. What this pilot cannot support is causal or exploratory work — anything requiring the researcher to follow a surprising result. What it can support is confirmatory work on questions specified in advance. That is a real category, and describing the arrangement precisely is the difference between treating these results as independent findings and treating them as findings independently specified and internally produced.