Artificial intelligenceJuly 21, 2026· via The Decoder

Google’s Frozen v2 chip to embed AI models in silicon by 2028

Google’s Frozen v2 chip to embed AI models in silicon by 2028

Image : The Decoder

Google is doubling down on silicon-level AI acceleration with a next-generation chip internally dubbed “Frozen v2.” Unlike today’s general-purpose TPUs, the new design reportedly hard-codes the architecture of its Gemini models directly into the hardware, cutting inference latency and power consumption by a factor of six to ten. If the timeline holds, the chip could reach production in 2028 and immediately slash Google’s AI serving costs—giving the company a measurable price advantage over rivals like OpenAI and Anthropic.

A shift from software to silicon

The move marks a departure from the software-heavy approach that currently dominates AI inference. By baking the model’s structure into the chip’s transistors, Frozen v2 eliminates much of the overhead associated with fetching weights and routing activations, the traditional bottlenecks in large-language-model serving. Industry watchers have flagged efficiency gains as the defining metric for the next generation of AI hardware, and Google’s plan appears to target that metric directly.

Strategic implications for the cloud AI race

A six- to ten-fold efficiency jump would translate into lower per-token costs and faster response times, two factors that directly influence cloud pricing and user experience. For customers running high-volume Gemini workloads, the chip could deliver both lower bills and snappier interactions. At scale, that cost edge could help Google attract enterprise clients—and retain them—amid intensifying competition with Microsoft-backed OpenAI and Amazon-backed Anthropic.

Why it matters

Hardware specialization is no longer optional; it’s a competitive moat. If Frozen v2 delivers on its internal projections, Google would leapfrog today’s TPU generation and set a new bar for AI inference efficiency. The move also signals that the model-architecture link is migrating from code to copper, pushing the industry toward ever-closer coupling of algorithms and silicon.


Source: The Decoder. AI-assisted editorial synthesis — TechnoExpress.

Read the original source on The Decoder →

← Back to home