Google's Frozen V2: Baking Gemini Directly Into Silicon
Key Takeaways
- ▸Google is developing Frozen v2, a custom AI chip that embeds Gemini's neural network architecture directly into silicon, rather than loading the model into memory
- ▸The chip could deliver 6-10x greater efficiency than Google's current custom AI accelerators, addressing Google's reported AI capacity crunch
- ▸Deployment is targeted for 2028; the approach trades architectural flexibility for dramatic gains in power efficiency and latency
Summary
Google is developing a custom AI chip codenamed "Frozen v2" that embeds Gemini's neural network architecture directly into the hardware itself, a departure from conventional chips that run models loaded in memory. According to reports from The Information, Reuters, and Bloomberg Law, the chip could achieve 6-10x greater efficiency than Google's latest custom AI accelerators, measured by tokens served per unit of power. The weights of the model can still be updated, but the underlying neural architecture is frozen into the silicon, requiring new hardware for architectural changes.
The project is reportedly a response to an acute AI capacity crunch within Google—severe enough that Google Cloud has turned away some customers. Deployment is targeted for as early as 2028. By embedding Gemini directly into silicon, Google aims to dramatically reduce power consumption and latency, making the chip particularly suited for real-time applications like voice assistants. The approach also represents Google's broader strategy to reduce dependence on Nvidia by deepening its custom chip capabilities across multiple suppliers and architectures.
The efficiency gains come from eliminating the overhead of general-purpose hardware: no need for expensive high-bandwidth memory shuffling data between processor and model storage. However, the strategy carries a significant risk—AI evolves rapidly, and a chip architecture frozen in 2028 could appear dated by deployment. While Google has not officially confirmed Frozen v2, the leak signals a fundamental shift in how major AI providers are thinking about infrastructure: moving from flexible, general-purpose silicon to application-specific hardware optimized for a single dominant model.
- The project is part of Google's strategy to reduce reliance on Nvidia and deepen vertical integration of AI infrastructure
- The fixed architecture carries a risk of technological obsolescence as AI models and techniques evolve beyond 2028
Editorial Opinion
Frozen v2 represents a bold bet on specialization in an industry that has traditionally favored generality. Baking Gemini into silicon is a logical response to the economics of running AI at scale—every watt saved multiplied across billions of inference calls yields enormous cost savings. However, the approach reveals a fundamental tension: the fastest way to optimize for today's model is also the fastest way to lock in yesterday's architecture. If Google's rivals can deliver comparable efficiency gains without sacrificing adaptability, or if a breakthrough architecture emerges before 2028, Frozen v2 could become a cautionary tale about betting the farm on one design. For now, it's a signal that the AI infrastructure race is moving toward irreversible specialization.



