Google is developing Frozen V2, a specialized server chip designed to run Gemini AI models with 6-10 times greater efficiency than current TPUs by etching the model architecture directly into silicon, reducing runtime decision points and addressing Google's compute constraints for AI workloads.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Google's Plan To Boost Its AI Models
Added:Tell me about this new chip that Google is working on.
>> Yeah, so Google is working on a new chip. It's informally dubbed Frozen V2 inside the company. And essentially the idea is that it could be way more efficient even than the TPUs at the time of launch. They're projecting it could be 6 to 10 times more efficient than the TPUs at the time of launch. And that's because it works slightly differently by etching the model itself into the chip.
So fewer decisions have to be made at the time that the model is running.
>> Etching it into the chip. So I mean, so why is it called frozen?
>> So the name frozen has to do with the idea that like the model is kind of frozen into the chip to some degree. So essentially the way that a regular chip works, it's general generalizable, which means that it can do a lot of different things, but that means that there are a lot of different decision points that the chip has to run. And so like this chip makes things more efficient by combining some or sorry by like presetting some of that essentially. And so like it doesn't have to make all those decision points.
>> And so how would it be different then from the TPU?
>> Yeah. So the TPU is more of this generalizable kind of chip versus okay these frozen chips which would like specifically run like this model architecture.
>> I see. So I mean if if this frozen chip I mean if it's really only um I guess it's only meant to run Gemini models sounds like so is it only meant to be used internally by Google then is that the idea?
I mean, so a lot of like Google Cloud customers also run Gemini models and so this is something that could theoretically also be run by external customers, but this is still this is still something that's pretty early. So it's not expected to launch until 2028 at the earliest. So I think a lot of these strategy questions might still be be like in the process of being worked out, >> right? because and I mean the reason I'm asking this question is because we of course know that the TPU was mostly used internally but now they're at the point where they're really trying to >> outfit uh customer data centers if I recall correctly with these TPUs. So I mean your point is well taken that that it's not just people uh at Google running Gemini, it's it's all sorts of customers. So that could very much be strategy.
>> How does this compare with other inference folks's chips? uh you know we've had the uh CEO of Samaova on the show uh OpenAI is working on their own ship. How would this compare to that?
>> Yeah. So a lot of people are trying to do the same thing which is essentially um driving down the cost of inference because everyone's in this big compute crunch. So Google's chip works a little bit differently because of this tourism design. It's actually something somewhat similar to what this Canadian chip startup Talis is doing. said they're also trying to like like specifically like etch a model into the chip so that it is more efficient.
>> Has Google being compute constrained? Is this a way to get at that problem?
>> Yeah. So they talk about that constantly on earnings calls. The idea that like their cloud division is compute constrained. Like Google reports earnings this week. I would expect them to also talk about that this week. And uh tell me TSMC capacity that I imagine is also top of mind here for this new chip.
>> Yeah, so Google is going to have to find fabrication space for this chip. It's not totally clear to me that it's going to be like a onetoone exchange in the sense that like in order to make these chips at like seals from TSMC space for the TPUs like that's definitely possible but it's also been suggested to me that this is enough of a different fabrication process that it might use different space so it might not be a zero sum game for them.
>> Great well Aaron I want to thank you for coming on. That is Aaron Woo our open AI and Google reporter here at the information
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

WOW! Judge TURNS THE TABLES on Trump in His OWN $10B LAWSUIT!!!
MeidasTouch
197K views•2026-07-23

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Steam and Xbox Just Dropped The Hammer On PlayStation
OhNoItsAlexx
9K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23