Google Is Building a New AI Chip That Could Make Gemini Up to 10x More Efficient
Google is reportedly developing a new AI server chip, internally known as Frozen v2, to make its Gemini AI models significantly more efficient.
According to multiple reports, the custom chip could deliver 6 to 10 times better efficiency, measured by the number of AI tokens generated per watt of power.
If development stays on track, Google is expected to deploy the chip around 2028. While Google has not officially confirmed the project, sources familiar with the matter say the company is actively working on the next generation of AI hardware as it looks to reduce costs, improve performance, and support the growing demand for Gemini.
Why Google Is Building Another AI Chip
AI companies around the world are no longer limited only by how advanced their models are. They are also facing growing challenges with compute availability and energy consumption.
Training and running frontier AI models requires massive computing power, and the cost of operating these systems runs into billions of dollars every year.
Google Cloud has also faced increasing demand for AI infrastructure as more businesses and developers adopt Gemini. Reports suggest that this rapid growth has put pressure on the company’s available AI compute capacity.
This is where Frozen v2 could play an important role. Rather than replacing Google’s existing Tensor Processing Units (TPUs), the new chip is expected to work alongside them. By improving the efficiency of Gemini inference, Google could lower operating costs, reduce power consumption, and deliver faster AI services at a much larger scale.
How Frozen v2 Could Work
Unlike traditional AI accelerators that are built for a wide range of workloads, Frozen v2 is reportedly being designed specifically for Gemini. Industry reports suggest that Google could integrate parts of Gemini’s architecture directly into the chip, allowing it to process AI tasks more efficiently than general-purpose hardware.
Analysts believe the chip could achieve this through improvements such as:
- Reduced data movement
- Fewer calculations
- Optimised inference pathways
- Better memory utilisation
These optimisations could help Gemini generate responses faster while consuming significantly less power. Since moving data between different parts of a server is one of the biggest sources of energy consumption, reducing these operations can improve both speed and efficiency.
Frozen v2 also fits into Google’s broader strategy of building its own AI infrastructure. Over the past few years, the company has expanded its custom hardware portfolio with TPU 8i, TPU 8t, Axion CPUs, and its AI Hypercomputer platform.
Together, these technologies are designed to give Google greater control over AI performance, scalability, and operating costs across its cloud services.
The Next Phase of Google’s AI Journey
The competition between Gemini, Claude, and ChatGPT is no longer just about which AI model is the smartest. It is also about who can build the fastest, most efficient, and most reliable infrastructure behind these models.
As AI adoption continues to grow, the companies that can deliver better performance at lower costs are likely to have a significant advantage.
Google’s AI journey has also had its share of challenges. A few years ago, the company faced criticism over the way Gemini was demonstrated during its launch, raising questions about how its capabilities were presented.
With Frozen v2, Google appears to be focusing on a different challenge. Instead of showcasing new AI features, the company is investing in the hardware that powers them. If the reported efficiency gains become a reality, Frozen v2 could play an important role in making Gemini faster, more scalable, and more cost effective over the coming years.