Google Building Specialized AI Chip to Run Gemini 6-10x Faster
Alphabet's 'Frozen v2' project would embed model architecture directly in silicon to ease compute shortage, targeting 2028 deployment.
Google pursues custom silicon for Gemini efficiency gains
Alphabet is developing a specialized server chip designed to run its Gemini AI models substantially more efficiently than existing hardware, according to a report from The Information that sent the company's stock up 3% on Monday.
The chip, internally called "Frozen v2," would permanently embed portions of Gemini's architecture directly into silicon. This approach would reduce both the computational workload and data movement required to process queries, The Information reported.
Google engineers estimate Frozen v2 could serve between six and ten times more tokens per unit of power compared to the company's latest tensor processing units (TPUs). Rather than replacing TPUs entirely, Frozen would represent a more specialized branch within Google's custom chip portfolio, according to the report.
Why it matters
This development signals how acute Google's internal compute constraints have become. The company has reportedly turned away Google Cloud customers due to capacity limitations and recently agreed to pay SpaceX nearly $1 billion monthly to help meet enterprise commitments. Purpose-built chips could help Google scale AI services without proportional increases in power consumption and infrastructure costs—critical factors as AI workloads strain data center capacity across the industry.
Architecture lock-in presents strategic trade-off
The specialized design comes with significant constraints. Frozen v2 would only work with future Gemini iterations if Google maintains the same underlying model architecture. This limitation represents a calculated risk: optimizing for today's architecture could lock the company into specific design choices even as AI research advances rapidly.
According to The Information, Google currently views Frozen v2 partly as an experimental effort and does not plan to manufacture it at the same scale as its general-purpose TPUs. The company is targeting 2028 for deployment.
Compute shortage drives innovation timeline
The project reflects broader pressures facing major AI providers. As model sizes and usage grow, companies are exploring diverse approaches to improve efficiency—from algorithmic optimizations to custom silicon designs that trade flexibility for performance gains in specific workloads.
Alphabet did not immediately respond to requests for comment on the report.
The Information first reported these details about Google's Frozen v2 chip development project.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call