As AI models saturate the standard tests, the industry's leaders are increasingly arguing the tests themselves are broken. On July 18 at WAIC, a series of Chinese AI executives made the case for retiring static leaderboards in favor of measuring whether models actually work in the real world.

The problem with benchmarks

The complaint is that benchmarks like MMLU, HumanEval and SWE-bench have become saturated and, worse, contaminated — their questions leak into training data, so high scores measure memorization as much as capability. When every frontier model clusters near the top, the leaderboard stops discriminating between them.

New yardsticks

Baidu founder Robin Li offered an alternative rooted in usage rather than exams: "daily active agents" as a measure of whether AI is doing sustained, real work. Product lead Li Jingqiu argued the best test is "a complex, time-consuming task" — the kind of multi-step job that a single leaderboard question cannot capture.

'Just use it'

StepFun's Yang Minghui put it most plainly: "the benchmark race is not the whole point... the real test is just using it." MiniMax's Bai Chuanxu pushed a related point, stressing native multimodality — the ability to handle text, images and more in one model — as a capability that leaderboards optimized for text reasoning tend to undervalue.

The self-interest and the substance

There is convenient self-interest here: reframing evaluation around "real use" flatters companies whose products are deployed but whose benchmark scores trail. Yet the underlying critique is widely shared across the field. As models converge on saturated tests, the industry genuinely lacks a trusted way to compare them — and "does it actually work" is a harder, more honest question than any single number.