A cross-sectional study published on Monday ran one classifier over 620,699 ChatGPT conversations from a public donated corpus, under two different definitions of what counts as a mental-health conversation. Under the conservative definition — explicit help-seeking for a psychological problem — it found 1,317 conversations, 0.21% (95% CI 0.20–0.22). Under the expansive definition — any mental-health topicality present at all — it found 30,394, or 4.90% (95% CI 4.84–4.95). A 23-fold spread, from a definitional choice.
What the conventional framing gets wrong
Two misreadings, and the first belongs to the publisher. The journal's own indexed teaser describes this as a study of how often adolescents use AI chatbots to discuss mental health concerns. It is not. The unit of analysis is a general-user public conversation corpus with no age data in it at all. Anyone building an adolescent-safety headline off that metadata line is building it off an error in the metadata. The second: the headline number will be reported as a prevalence, and it is not a prevalence of people. It is a share of conversations, in a donated public corpus, scored by a classifier whose positive predictive value is 67% — roughly a third of its positive hits are wrong. Sensitivity is 87%, specificity 95%, negative predictive value 98%.
Why the spread is the finding
"How many people bring mental-health problems to chatbots" is currently being cited by model providers, by regulators and by plaintiffs' lawyers, and each of them has an interest in a different answer. This paper demonstrates that the figure is a policy choice dressed as a measurement — that you can move it by a factor of 23 without touching the data, simply by deciding whether a passing mention of stress counts. The authors' response is constructive: they propose a tiered taxonomy — crisis and high-risk, clinically framed, and broad affective or interpersonal — so that whoever quotes a number has to say which tier they mean.
Stated limitations
Reliance on a public corpus with no external validation; possible misclassification of ambiguous expressions of distress; and no test of whether taxonomy-based monitoring improves any outcome. The work was funded by a National Institute of Mental Health grant.
What to do with it
Treat any single percentage for chatbot mental-health use as incomplete unless it states its definition, its denominator and its classifier's precision. Most of the figures currently in circulation state none of the three.
