AgiBot — Zhiyuan Robotics, of Shanghai — announced on 28 July that its embodied-native multimodal model WITA-Omni Preview had taken first place globally on the DailyOmni omni-modal understanding leaderboard, with a composite score of 85.21 and wins in six of eight sub-metrics. The entries it beat include models from Qwen, Google Gemini, ByteDance Doubao and Nvidia, with the margins claimed on audio-visual joint understanding and temporal reasoning.
Thinker, Talker, Actor
The architectural claim is the substantive part. Most omni-modal systems use a Thinker-Talker split: one component reasons, another speaks. AgiBot adds a third — Actor — so that physical actions and expressions are produced as first-class outputs alongside speech, rather than a chat model being bolted onto a robot body afterwards. For a company whose product is a robot rather than an assistant, that ordering is the point.
How it was trained
AgiBot describes tens of millions of hours of open-source and proprietary multimodal data run through a three-stage pipeline: supervised fine-tuning, then knowledge distillation, then reinforcement-learning optimisation.
What is missing
Almost everything a reader would need to check it. No parameter count, no licence, no availability. There are no weights on Hugging Face or GitHub under either of AgiBot's public organisations. The score is lab-reported and has not been independently replicated — the standard caveat, but worth stating plainly when the claim is a world first.
The timing
It lands the same week the FCC blocked new foreign-produced humanoid and quadruped robots from US equipment authorisation, and days after a patent ranking put Chinese firms in six of the world's ten most innovative humanoid startups. The hardware side of Chinese robotics is being fenced out of one market while its model side publishes benchmark wins.
