Google is developing a specialized server chip that would bake its Gemini model architecture directly into the silicon, according to a report by The Information relayed by Reuters. Codenamed "Frozen v2," the chip is designed to run Google's models far more efficiently than general-purpose accelerators.
Freezing the model into hardware
The core idea is a departure from how AI chips are usually built. Rather than keep the accelerator flexible enough to run any model, Frozen v2 would freeze aspects of the Gemini architecture into the hardware itself — trading generality for efficiency on the specific workloads Google cares most about. It is a hardware-software co-design bet: shape the chip around the model instead of adapting the model to the chip.
A 2028 target
The chip is aimed at 2028, according to the report — a target, not a shipping date, and far enough out that the design could change. Secondary coverage suggests Frozen v2 would complement rather than replace Google's existing TPU line, giving Google a purpose-built option alongside its general-purpose accelerators rather than a wholesale switch.
Why the economics push this way
The move reflects where inference costs are heading. As models serve enormous query volumes, the marginal cost of each inference dominates, and squeezing efficiency out of silicon tailored to one architecture can beat a flexible chip running the same job. Google, which already designs its own TPUs, is signaling it believes the next efficiency gains come from tightening the coupling between model and metal.
Read it as a signal
Because the details rest on a single report — attributed here to The Information, as carried by Reuters — the specifics should be treated as reporting rather than confirmed product plans. But the direction is notable: a frontier lab willing to hard-wire a model generation into a chip is making a long bet that its architecture is stable enough to be worth casting in silicon.