Google announced Gemini 3.5 Transcribe on 26 August at 17:00 UTC, a speech-to-text model aimed at voice interfaces. The post contains four word error rate figures. Two of them will be quoted; the other two are the ones that describe multilingual performance.
The two pairs
The first pair is attributed to Artificial Analysis, an external evaluator: the model "achieves an average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming use-cases." The second pair is Google's own, on the FLEURS multilingual benchmark across top languages and locales: "5.50% WER in streaming mode and 5.04% WER in non-streaming use-cases." Between the best number in the post and the worst, the gap is 2.1 times.
What the common framing gets wrong
The number that travels is 2.6%, and it is being attached to the claim that the model handles 85+ languages. Those are two different measurements. The 2.6% is an average over what the post calls primary use cases; the multilingual figure, on a standard multilingual benchmark, is 5.04% — nearly double. A transcription model's error rate is only meaningful alongside the language set and audio conditions it was measured on, and this post is unusually transparent about publishing both. The distortion happens downstream, where the third-party number gets bolted onto the multilingual claim to produce a specification that appears nowhere in the source.
Preview, not launch
The status is public preview. That word is doing work. The model is reachable through the Gemini API in Google AI Studio and the Gemini Enterprise agent platform, and appears in Rambler on Android in selected countries and languages, in the Gemini app on macOS, and in Gboard and Google Antigravity. Chrome is listed as coming soon. A public preview carries no service-level commitment, no pricing stability guarantee and no assurance the endpoint survives in its current form.
Why the streaming gap matters more than the headline
In both pairs, streaming is worse than non-streaming — 4.0% against 2.6%, and 5.50% against 5.04%. That is expected: a streaming model commits to words before it has heard the rest of the sentence. But it is the streaming figure that governs live voice agents, the application this model is positioned for. Anyone sizing a voice product against 2.6% is sizing against the batch case in a subset of conditions, when the number that will determine whether the product feels usable is the multilingual streaming one, at 5.50%.
