The AI compute market is undergoing a fundamental economic transformation where token efficiency is becoming the primary driver of value creation, shifting from a 'token maxing' paradigm to a more rational model routing approach. This transition is evidenced by the Token Expenditure Index, which tracks model usage patterns and reveals that as models become more efficient and cheaper, enterprises will increasingly route tasks to appropriate models based on cost-performance trade-offs rather than simply maximizing token usage. The market dynamics show that while frontier models may face margin compression from more efficient open-weight alternatives, the overall AI market is expected to grow as inference demand expands and enterprises discover productive use cases. This evolution creates a more fragmented compute market where orchestration layers and smart routing become critical value capture mechanisms, with GPU rental indices and forward curves indicating sustained demand despite efficiency improvements.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
AI Efficiency Is Repricing The Compute Market | Steve Hou
Added:Nothing said on Ford guidance is a recommendation to buy or sell any investments or products.
All right, everybody. Welcome back to another episode of Forward Guidance. And joining me today is repeat guest of the show, Steve, who who just joined me a couple months ago, right at the the tail end of when you were at Bloomberg, but now you're at Silicon Data head of research. you are the man behind some of the most important charts and indices in the world of AI right now. There's a lot to get into, but uh yeah, really excited to have you back now under your new role at your new company. So, congrats on on starting there and yeah, great to have you back, Steve.
>> Yeah, thank you, Felix. Uh great to see you always and uh thank you for having me back on.
>> Yeah, awesome. Um for those that don't know about Silicon Data and what you guys do, would just love to hear a bit about like why you made the shift to joining them. um what you guys do and yeah just how you think about the the state of AI right now after >> uh so yeah I uh left Bloomberg about two months ago joined Silicon data almost two months ago on the day uh and uh you know silicon data is a company that tries to bring uh well data to physical AI compute market and uh you know uh you think about uh I think the easiest thing to think about is we are trying to bring uh futures contracts uh derivatives to the AI physical compute and allow people to hedge the essential risks that are involved with this now you know enormous and uh you know evolving a computer market that's hundreds of billions if not trillions in uh in size and uh there's a lot of risks that's being held you know in in equity form and in a fixed income form and in our view sometimes you know uh ine probably not perfectly efficiently and uh you know there's a lot of I think uh risk that's now involved with AI compute and GPU income that can be handled with traditional financial instruments that is well understood by historical financial markets or capital markets. So uh to the extent that uh you know on the natural hedging side of uh the you know the data center providers uh compute providers all the companies that are looking to buy compute uh uh futures contracts is a very natural way to hedge out that risk and u maybe actually in fact help you be a bit bolder in terms of uh at how much comput actually acquire so you don't find yourself you know uh being I think overexposed or maybe not having enough which is seems to have been the case with some of the AI labs. Uh I joined the company because I felt like khara I always been I have been working on public markets benchmarking you know uh systematic indices at Bloomberg for six years. Uh this feels like a natural way to become exposed to the AI sector you know same set of similar set of skill set but apply to AI.
>> Yeah. Awesome. Um yes. Yeah. So why don't we get into we'd love to just hear a bit about like who what is what is the why why is hedging and and you know futures contracts a necessity for for the AI compute buildout like who are who are the ideal target customers and and what are they trying to hedge and why like is it the hyperscalers is it the frontier models is it like you know just downstream companies that are utilizing these models and you know are trying to hedge out their their inference demands or like what does yeah what does that landscape look So uh I mean the demand for it should really come from all major participants in this market in the same way you would expect in the crude market or even agriculture historically wherever futures contracts first evolved uh as this market eventually become more I think uh fragmented right right now the AI comput is very much dominated by a couple big players very a couple big sellers you know very very granular right so you've got the the two major labs open AI philanthropic that account for ostensibly half of the you know AI compute demand and then if on the other hand the hyperscalers that are providing most of the compute and building most of the compute um as we have now seen with the rise of very powerful open models and enterprise AI use actually if you think about how enterprises will alternatively adopt AI they're going to actually not just b basically pick a winner this debate of uh who is going to emerge as the you know sort of platform on which everyone's going to build on top of I think that's already been settled Nobody's going to win it out outright and companies don't feel comfortable and the labs themselves I think it's just not uh so what's going to happen is that orchestration layer is going to acrue a lot of the value companies are going to retain their sovereignty by basically making models more substitutable you know and some tasks low value tasks maybe you know going to route to be routed smartly towards the open models cheaper models and then the higher value tasks maybe will be using frontier intelligence to that extent We expect there to be you know generally a transition away from the current paradigm more even more so uh and with inference you know becoming ever bigger deal that most of the AI computing demand is going to come from inference and the funding of compute is come from many many enterprises companies that will be I think looking for compute and as the models evolve and architecture of the models evolve you're going to see that uh you know more models can be run on not just the hyperscalers but maybe sort of more closely located smaller clusters you know smaller clouds. So you can see this fragmentation. So the market you know on the one hand you have sellers of comput these new clouds you know cloud providers coming onto the market. They want to have more certainty and they are lenders they have backers want to have more certainty over their revenue right and how do you actually have certainty? One natural way is to actually just as you have with any command in this market you know using futures contracts to lock in future uh revenue. uh on the other hand uh you know you have uh the the buyers of compute and you know the different ways which you could work out some companies maybe they don't directly buy compute maybe they just buy tokens but somebody will actually be sort of uh you know buying that compute right so I think that's how we sort of see it eventually the market evolving in that direction now of course along the way you know if you talk to cloud providers today there's going to be a lot of push back nobody wants to be their product to be referred to as a commodity and Uh but I think the way I see it is that there's always going to be substitutability, right? If you tell me that your product is so unique, how do you win customers from someone else?
That's how do you win clients away from someone that that's with with someone that uh so so there is going to be substitutability when it comes to compute, but it's not going to be perfect and uh that's how we uh you know see this market.
>> Awesome. Okay. So I want to dig into a few of the different uh charts and indices that you guys have have built out and want to start with just on that point about orchestration and substitution.
>> Um >> want to talk about the the token uh index that you guys built out. This one has >> you know gone gone sort of viral for maybe good reasons also bad reasons.
There's been a lot of takes on it. I would >> I would make the argument some some quite uninformed takes. So I just want to like >> let you explain through the methodology just from first principles like I'm not going to throw any sort of bias here just this is this is your token index.
Tell me about the methodology of what it tracks and and what we can take away from this chart. [laughter] >> Yeah indeed. Yeah indeed. Uh this this index was sitting on the shelves uh you know when I joined the company and uh you know for quite a couple month and you know uh and I think it was not getting a lot of attention. I think I just noticed that people were sort of whoever noticed it was actually not interpreted even correctly too. I thought it was either a price index or total volume index. I think the name of the index on terminal displayed is also got a little bit unfortunate. There's a bit of a misnomer uh because it's called the token expenditure index when it's actually neither. What it really is is an expenditure weighted price index. You can think of as almost like the PCE for AI, right? So what that means is that um you know to the extent you're on track inflation of something you can either fix the basket which is what CPI does or PCE allow substitution right across you know similar substitutes uh because users are rational they always make the quality price tradeoff and AI more even more so especially when you have uh sort of these open uh uh uh routing uh or inference platforms allow you to choose different models right you know some models are really really powerful But they're crazy expensive, right? You know, your fables of the world, right?
They burn half of your token budget, you know, you know, you know, in a day or something. But then you've got the Chinese open ways models that are, you know, super super cheap, but they're maybe less capable. Uh so I wrote a post on X on our corporate X account uh towards the beginning of June saying that this is the way you should be interpreting this index. We basically collect all the prices of all the models, render plus models out there. uh and also we collect the usage volume from a couple you know you know a few sort of uh these public routing uh uh platforms uh with the input output volume so as you know that every ALM token model sort of there's a price charge on input token how much tokens you put in and there's also a price charge on output token so we have to create a blended price for every single model uh and uh we sort of take the input price input to volume and outut put the highest output volume create a blender price for each model and and and uh aggregate them to index level based on usage. So that this thing is going to track um how much each model is being used and to the extent that the general price dynamics of this model is uh such that while the model is launched the token price moves but not that much right so most of the movement comes from you know essentially usage pattern usage behavior consumer behavior right so the our coverage of data is not a full market right if you imagine the full market AI inference market as being like a square right you know um this is not a vertical slice of it. It's not a representative sample, right? Because we don't have data to we're not privileged the data of say open anthropic internal data usage or the hyperscalers for that matter. We are talking to a couple of them but uh at the moment we are slice we sliced off a corner and this corner tilts adly towards uh uh probably like independent developers or or small medium enterprises that already onto some of these inference platforms, right? Um and uh they are going to be necessarily the ones that are more price sensitive than you would imagine like a large enterprise that's currently living over a hyperscaler that maybe the employees are not even responsible for the exper they just go crazy until like end of the quarter end of the half year and then CFO say wait a second like you know so this is going to be a leading indicator and I wrote a post at the beginning of June saying like this thing seems to be plateauing a little bit and the use correctation is such and this thing could be taking a break before next lag up maybe we see a super powerful model and anthropy just runs away with it or there's a possibility that this thing actually me reverts a bit as you know people become more rational about the cost of these things and you know sort of substitute between quality and cost and indeed that seems to have been what happened and also unfortunately coincided with the turn of the stock market because I think partly rationally because the current paradigm of how the AI capex trade has been funded right and everything is being priced on second derivatives even if things keep going up just like a slower pace things going to sell off just based on the valuations.
There is you know that right you know but the fundamental is such that if there is a perceived slowdown and rationalization then you know and there's a perception that the frontier models could see their margins challenged with very powerful much much cheaper open waste models or even just sort of you know not open but uh you know capable but cheaper models coming from grock and mart spark and whatever that this could actually challenge the current paradigm of how the air capex is financed and I think this is part The reason why you know we've seen markets selling off if you open lower over latest chart with the you the AI you know sort of a beneficiaries index or the SMH or whatever you probably see like that turning point but uh we don't necessarily see this like a secularly bearish index you know there is and I don't think there is actually a uh somehow like contradiction and this is not one of those things where it goes up is be bullish goes down is bearish bullish either it's actually a little bit conditional right just like you would never say or higher inflation is always good or you know lower inflation is always good right uh at some point we're going to transition towards that paradigm I was talking about earlier right where the funding of AI and the demand of AI becomes a lot more broad-based and um you know you can have a lot more sort of smart routing of models even as the inference market explodes in size become a much much bigger of the overall pie of AI demand compute so this transition of from one phase one to another was always going to be a bit rocky, right?
uh and and and even potentially runs in the risk of a bit of a air pocket where you know sort of uh on the one hand the the the capex and ROI for the current paradigm doesn't show up quick you know quite enough on the other hand you know the new phase like maybe the biggest enterprises takes them a while to find that workflow and and and and pick up so the pace you know may not catch up like you will see the fastest adoption and most show up in the smallest enterprises with the fewest employees and the least sort frictional you know workflows but those are by going to be by dollar amount by dollar weighted going to be small right so part of the reason why we don't see productivity increasing in aggregate part is also because of that right so anyway that's kind of a longwinded answer to I don't even remember your answer question was but hopefully still made some sense yeah >> no that was super helpful I think there's some some really important nuances in there but the >> yeah I mean I'll throw the proposition at you that was those before which is that okay if you look at the chart you know early March which is around where when mythos started to come out. Um it was also around the same time that this this idea of of of token maxing took hold of of US corporate America. Um and and suddenly it's just like okay you know just token maxing use as much as possible use use the frontiers. Um, and then that chart went higher. And then I think what some, you know, folks took away from, and I think to your point, you mentioned how, you know, partly it was just this this index maybe was a bit of a misnomer, but people took it as, >> okay, this is just actual demand for tokens from corporations. And then you see the turn came forth right around the time where these news headlines start to come out of like you know Uber went through their entire year budget of of token spend in like a month and everybody freaked out and then but what you're saying is it's more so about substitution than outright token demand, right? Like what you're saying is that this is about going from token max to token token efficiency. It's not saying that demand for tokens has gone lower since June. So would just love Yeah. Can you can you expand on that further?
>> Yeah, absolutely. Like I think you know this idea of token maxing was always the idea you want to actually let people experiment right because in complex you know workflows large enterprises you don't really know what the use case is going to be and also the models just were not didn't really become smart enough or good enough until this year with aentic and everything. So you need token maxing but token max is fundamentally odds with expensive tokens right tokens are so expensive and also throttled. So you need eventually to get to a state where tokens are cheaper and this idea of uh uh you know smartly routing models you know work work like that's always going to be the case right if you look at like my old employer you know Bloomberg put out this very nice paper on MCP of you know how they let the models interact with the data you know and and correctly call the data accurately with you know financially that you do you don't really can't really afford you know high losation you see what's you know been happening at talent here right all this idea of ontology of actually making the AI actually useful at the enterprise context and data bricks same thing right and more and more of this you think like open routing has has turned from you know sort of open router will give you this sort of routing for a long as longest time that seems to be more of like a hobbyist you know developers trying to try different models but I think that paradigm has actually become now the norm right and increasingly we should expect that right um And this is kind of like to the extent I hate to use the term now because it's so cliche but like sort of Javon's paradox max you know like basically in the fuller context of it like ultimately how this thing can be truly sustainable and actually get an RI is to actually get enterprise adoption right millions of companies and the only way that you can actually get sustainable adoption of AI is when you don't suck dry the underlying company where you just have like a single model that's the the entire company is built on build on top of and that was not going to happen.
certainly it's not going to happen now right instead basically the model layer has got to be substitutable where you can smartly substitute between if I have a simple task say just data cleaning or or just sort of choosing which analytics to or simple summary summarization or whatever things that can be easily handled with much cheaper less than frontier models they should go to those things right and if I have a higher value to ask you know things that requires more powerful reasoning. Uh they should all to go go to those, right? And the ability to route between them and correctly assess quality and cost is something that will have to happen and it's not really been happening enough and is a challenging thing because just to figure out whether or not a certain ask is which type, right? It's not always obvious, right?
Sometimes a been a a question can seem benignely innocuously simple but in reality it's actually a pretty complex question that someone is just asked it poorly because they don't really know what they are looking for and uh I think that that that is going to be the new challenge and that's part of the reason why you are seeing more and more of these sort of routing you know things being activated you know sort of versels of the world or ramp you know coming up so with one and so on and so forth so this is going to be the new norm I think >> that's super helpful. Um, all right, Steve. So you're you're an economist. I want to ask you about this, but um >> the it feels like one of the big questions right now that sort of leads into what we just talked about with this this index is this idea of you know feels to me like this framework is either the frontier models get to enjoy the margin like right around this time where this chart was was surging was also the time where the AR of Enthropic was surging and um you know Gavin Baker mentioned that there was even a month where Enthropic was like had had was profitable for the first time. Um, so that's all great, but now we're seeing this substitution towards token efficiency. We're starting to see these open way models that are looking a lot more powerful. We're starting to move to this world of, you know, perhaps frontier model topline orchestration, but a lot of the downstream execution is is these more efficient, cheaper, you know, openw weight models. Like >> it feels like that is really good for for the end consumer. Like the surplus goes towards them. So it feels like there's like this fight for margin between the the the frontier models and that AR and then being profitable and like everything that is almost dependent on that occurring versus the benefit going towards the consumer. So I'm just curious like what's your framework for how to how that balances out? Like can the both sides win or or what do you think?
>> Yeah, I mean this is exactly like you asked me as an economist, right? You know economists like to think about you know partial equilibri equilibrium and general equilibrium, right? So what is partial equilibrium? Partial equilibrium is when you just have like a simple shock to price, right? Price times minus cost times quantity gives you profit, right? So, uh cost is kind of whatever you know cost of compute and operating and if price uh is challenged by a competitor, you know, you and quantity doesn't change you know you're going to see uh shrinkage of the overall profit and that makes you worried about you know sort of ROI and so on so forth. But Q is not static, right? You know that's the whole point of GMO's paradox is there is elasticity right and if if the overall market is become much bigger you know like right right now the adoption enterprise adoptions are bismally bismally small right most people are not using AI you know in a consumer context beyond like sort of glorified search and enterprises are still trying to figure it out and they are nervous about the cost even even as they just figure it out now if you tell me that okay the margin of maybe the some of the frontier models is going to get challenged a little bit but the queue is going to growth so much much more the overall thing still works out right you know the analogy I like to give a little bit sometimes it's a bit like this is not totally unlike like the drug market right you know so people always ask like who is going to innovate on you know pharmaceuticals if you just get generic drugs that sort of copy you and whatever so there's going to be some dynamic of uh you know erosion there right you know uh uh you're probably not going to earn like full monopolist profits you know if you have competitors at the same time I also don't think it's been true right whether you look at software or look pharmaceutical or whatever. Just because there is going to be, you know, competition and there's going to be some degree of, you know, copycats or or what have you, it doesn't mean that the frontier is not going to grow and it doesn't mean that, you know, uh um the the leaders of the frontier will not be profitable. That being said, I do think this is uh going to challenge the question uh where the ultimate bulk of the value, you know, is going to acrue to, right? uh and uh you know I don't think it's going to be super controversial to think that at least direction on the margin this means that some of the value that was perceived to have been certainly acrewed to you know this dualopolis you know model of frontier intelligence labs predominantly more acrew to either the compute layer or the user layer depending on how the compute layer sort of settles out right so >> makes sense um all right that's a good segue into talking about neoclouds and GPU rental index so this is um another index from from you and your team.
>> Walk us through how to think about this and and yeah what does this what does it say for the current landscape?
>> Yeah. So uh basically we track uh you know create this uh you know rental indices. It's almost like a benchmark.
Can you think of an analogist like uh you know if I want to create an unfernished singlebedroom apartment index in New York City you know you get t you know you observe rental contracts from across the city a different term in different shape and form right some come furnish some come unfernish with a little bit of a room and so on so forth we we we we aggregate them all use machine learning to essentially create apples apples comparison right uh and uh create a a singular index for each ship you know of all of all the data we look at and here we have uh the blue line, that's H100 chip hopper, right? And then you've got the A100, which is the ancient chip that's been out for like five years or something, A100. And then you've got the B200, uh, which is the orange line is more volatile, is also newer, less deployed. And also the H200, we have a longer history, but here only for some reason only managed the rent most recent few weeks. Um so uh the way I look at is that uh you know amateur who follow the AI market they just look at the H100 you know sort of uh uh index which is really actually the ondemand price right this is not the entire for curve right this is not telling you like what how much will cost per GPU hour out you know year out or whatever uh and uh they look at this blue line if it goes up they say oh you know demand is going up is going down like oh demand must be collapsing but in reality I like always like to both look in the cross-section and Then later you know we'll get time to also look at you know across the term structure uh because uh you know in the cross-section what happens is that with between inference demand and training especially with inference growing so much uh you can actually see the shifting of workflows and depending on different labs and you know sort of who's needing seeing demanding what kind of thing you can actually see shifts so for example when H100 has been steadily going up right since the beginning of this year and has had some you know sort of fluctuations recently it came down.
When it came down, what you see is that the purple line and the uh yellow line or the orange line continue to go up a lot, right? So, uh meanwhile, if you look at the A100, you know, that's actually very steadily sort of holding, right? It's not even coming down at all.
Uh you can have a scenario which which you know this is consistent. I'm not saying this for no for sure, but uh you can have a scenario where inference demand is just through the roof, growing like crazy, right? And that's driving up rental uh rate for even A100 chip like you know from five years ago and strong right meanwhile you can have some of workflow for training shifting away from H100 towards the newer chips right and and as those chips come online and and depend availability so those things can fluctuate a little bit so that's how I read the at least the on demand you know in indices for uh yeah >> awesome um I feel like people look really closely at this because of Of course, you know, one of the most one of the largest companies in the world right now is Nvidia. So, obviously, demand for their GPUs is very uh correlated with the performance of NASDAQ and and the total stock market. So, obviously, you know, if there's any sort of concern for demand for GPUs and the buildout >> uh that could have some some pretty significant shock waves throughout the system. Um, so is there anything that we can take away from what what what we're seeing right now in this indust indicy to, you know, interpolate where GPU demand is at and whether like the buildout is peaking out or whether the growth rates are changing there. What yeah, what do you see there?
>> Um so we can get a even cleaner read from the forward curve which I think we hope we can come to next. But uh I think on this chart I actually the favorite thing I look at on that this chart is actually the green line and not so much the blue line. Like just look at how strong that price is, right? Like people there was all this talk about early in the year like Michael Barry like oh chips are like only good for two three years the fast debrising asset blah blah blah. In reality even the A100 chip rental rate is going up and has been has stayed up has not come down at all.
Right? That just tells you just how robust you know inference demand is. And going back to what we were saying earlier about routing different workflows and whatever there's just a lot of like think lowhanging fruit easier as tasks inference demand as this overall demand you know grows people become familiar with AI a lot of those things are being routed towards you the workhorse chips right so I wouldn't be surprised if uh you know the all the most widely deployed hopper chips which that will soon become no longer the most powerful chip for training and and and and will actually become a new workhorse right so you will seem sort of a similar pattern uh uh to follow here. So that's actually a good signal. But then we can also look at the forward curve, right?
>> Yeah. Yeah. Walk us through through there and and and how to understand it and Yeah.
>> So basically, you know, uh forward curves the way I think is the following, right? So if you're renting a GPU like you know some of these contracts can run for a fairly long time, right? You know, you can run it for a month, you can run it for three months, you can run it for as long as two, three years. And uh typically with any commodities curve you expect a downward sloping like you know backwardation or you that's the jargon right so downward sloping curve like you know because you expect supply to come online in this case like another example you can imagine it's like uh if you run an apartment in New York or anywhere for 12 months your per day rental rate is probably going to be lower than if you rent out a hotel room for like three day right uh why because you get you get get a discount for avoiding the trouble for the landlord the tenover tenency right you sort of lock in that thing and so he doesn't have to keep looking for the next uh you know renter and indeed that's what we saw like the blue line here from November 22nd last year you know just when the agentic AI was sort of kicking off the Opus models becoming online uh the the the curve was very backwardated downward sloping and since then like between November to like this is a lot of curves I want to show you sort of how things have been evolved in the last few weeks is interesting right the the orange line here is the at the end of March right March 31st right the entire entire curve has actually moved upward right in other words every maturity every length of the contract per GPU hour price has gone up not only that the curve has also become like on the long end almost seemingly a little bit in contango right like you know sort of no longer downward sloping what that means is that you know cloud providers renters the landlords of AI you know clouds are feeling comfortable not you know giving these long-term contract discounts and just letting short-term contracts roll over as so that they have an opportunity to raise price again and indeed uh you know that's what we've been hearing from ambulator evidence with anecdotal conversation we have had with cloud providers that they really are very much aggressively looking for opportunities to raise prices now some still want that certainty you know so this is why there's a market right ear we were talking about hedging like you know people different have different views right if you have a thing maybe you grow some crops maybe you you grow some oil if you feel like prices can just go go straight you don't h at all right so so what's interesting here is that uh you know between March and June you know if you look at the purple line that's like June 25th like you know there was bit of a scare in the market you know air capacity you know like basically on the front end the H100 came down a little bit and people freak out a little bit but then you look at the one-year mark right there sort of a lot of the meaningful compute you know workflows are actually contract set like the one-year term that's pretty much unchanged And what's more even more interesting is that in the ensuing weeks right over the last month uh you know you have all this news about you know meta leasing out compute and all this you know models and substituting away from the frontier models like all basically a bunch of news that will be indicative of uh AI you know supply maybe potentially being accessed right people because the market valuation is so high people are looking for every reason to freak out what we see in the fundamentals at least in the terms of ute is that rental prices at the one year mark for the four in the fall rate it's actually been monotonically going up right you know there are two lines that overlap each other because there's not much change like between 2nd of July and then 14th of July and then most recently when we look at yesterday you know July 2 20th uh the green line is went up even more right every single time we saw multiple providers actually raising prices in other words it's not just like a oneoff thing uh you know pro uh at least in terms of compute fundamentals, right? There is actually a pretty sort of strong indication of firming demand and supply, you know, uh being being, you know, in shortage. Therefore, the price has to adjust, you know.
>> Awesome. Um, anything you want to add on the other GPU curves here in terms of distinction? No, I mean I think pretty much similar, you know, like I think you see the similar picture on the right, you know, sort of this the A100 and and is, you know, sort of secularly going up, right? And uh, you know, maybe not quite as aggressively uh, you know, more recent weeks, but generally holding up pretty strongly. And then the B200, you know, is also generally moving up, but then there's a bit of a fluctuation. You think the thing with B200 and also I think you just think about this chip dynamics generally is that the B200 is still being deployed, right? most of the uh uh data centers are still just bringing them online, you know, uh um availability is not quite as quite as high as the O 100. So there's going to be some, you know, sort of uh uh uh fluctuations, you know, but generally speaking, you know, I think the picture that emerges qualitatively is very much in agreement.
>> Totally. Um, obviously we've been talking mostly about GPUs, but increasingly a lot of where the the change in in you know the buildout for for all this AI is also on the the DRAM side of things and the memory aspect.
I'm curious uh how you think about that and and and what's going on in that side of the market compared to the GPUs.
I mean so clearly like so memory is the input right into uh you know you know GPUs and uh and and all these models are super memory you know hungry and u uh generally speaking with longer context longer conversations the context grows with memory that's part reason that's clearly the reason why you know memory as far as storage you generate tons and tons of data that all just demand is not catching up with supply is not catching up with demand and it's not surprising urprising that price is shooting up the way it is but we also know that um you know shortage and you know high prices are always the model of innovation. So unsurprisingly we are seeing algorithmic innovation you know no less from the recent like Chinese models like Kimi and so on that are actually making sort of uh uh improvements to uh memory efficiency and uh you know sort of maybe the the memory demand doesn't have to quite grow linearly with uh you know uh context length which is seems to be I think part of the there's what KDA you know innovation you know that came with the Kim model. I think it's not surprised to me the Kimmy you know thing moment felt very similar to the deepseek moment like we're going to keep on having I think efficiency efficiency improvements like that I mean clearly it's not going to be like I would be very very shocked if let's say in two years um the gross margin of the memory makers are still running at like 85% you know or whatever it's just like one of those things is a sort of a squeeze right so I think um you know we we we get we get past it you know and and uh you know we I think there will be again goes back to the question of P times Q right sort of I think P is going to come down some but the Q probably grows enough that you know everyone still I think do very well yeah >> yeah it's it's interesting that the Kimmy thing B like yeah like there's obviously when we talk about these uh these there's the demand aspect in terms of just efficiency of the models like if um >> yeah like um if you don't need as much if it's more efficient in terms of like KV cache and that whole thing, you won't need quite as much memor. And then there's also the supply aspect of of whether we start to see more buildouts.
You know, there's start to be talk of availability of of memory from China coming onto the market and that sort of thing. Um but it sounds like yeah regardless of that um you know Jevon's paradox and those ideas still hold true and regardless of these you know marginal change obviously it's it can feel especially volatile and sensitive when you have these uh you know memory equities that have just ran like they have like any sort of marginal shift and just with the amount of like leverage positioning it feels very >> intense in the short term. But what you're saying is that regardless of that, like if you zoom out a little bit, you know, these are these are pretty small changes on the margin.
>> Yeah. I mean, I I I I couldn't tell you, you know, in terms of sort of the stop price of these memory makers, like nobody knows, right? Everything's being priced on, you know, second or third derivatives at this point, but uh uh it doesn't seem to me, you know, uh at all that there's any going to be any laptop on the the demand for memory. And uh you know I think I fully expect to be surprised at the upside you know because every single time we have had a new uh innovation that seems to only uh increase essentially the people's willingness to try to use these models in more productive ways. Uh and every time you do that you generate you know sort of longer context you generate more data to be stored. Uh by the way it's not just memorize also storage of data. And we haven't even sort of gotten into the multimodal stuff with the voice you know like I think like for example over the last couple months one thing that people really I think overlooked was that you know shd live right came up with this very very good conversational sort of AI right where you able to talk to it and you can interrupt each other AI actually is able to interrupt you now right you have this really powerful way of engaging with the AI where you can ask AI to go and use agents to do things on your behalf you can interact with what you're looking at on the screen and and and and just use voice to to to do it and that also means it's a ton more data, right? So every time you have we haven't even touched scratch the surface of video, right? Like you know you so so I just think like you know every comm like memory is ultimately a commodology like I don't care it is a super cycle but super cycle of a commodology can you know also go pretty crazy. So uh um yeah but so so that's how that's sort of how I think about it. I I I'd be very very shy about trying to make any sort of prediction about when you know it will end or how you know but the direction seems pretty clear to me for demand.
>> Yeah 100%. Um that being said you know the last time you came on I think at the end of the interview you're telling me that you think your token efficiency is going to become really important and it and it did and and so props to that. So I'm curious like yeah how do you think the next few months play out? What do you think are the most important dynamics?
Because it feels like one of the big ones right now is of course this uh this discussion and deliberation around regulatory oversight of of models especially this distinction between US frontier models versus the open way models from from China and you know do we try to regulate that or even just you know new model launches who gets it and when these sort of aspects how how are you thinking about all of that right now?
>> [snorts] >> I mean so the way I think about it is almost a bit like you know the Tesla BYD situation right you know sort of the Chinese EVs they they ostensibly we don't see them here right but they are apparently very good and they dominate the Russ market but uh Tesla you know still here uh and uh I expect something along that line you know uh to happen like I think it's a little bit of I think the reality we live in you know uh this this economic decoupling and geopolitical decoupling between US and China. Uh that drives I think a lot of the underlying dynamic. I mean some of these bottleneck trades mean exactly because the US doesn't want to create a strategic economic dependency on China by buying from China.
uh that that that being said, I do think the direction of travel is going to, you know, be pretty similar and regardless, you know, I think there will be more uh you know, efficient uh licensable models or cheap models that come out of the US, right?
regardless as whether or not the Chinese models will be fully accessible or maybe in some sort of ways be made soft you know so to speak uh less accessible like you can easily imagine a scenario where like the government US put out the rules saying oh if you government supplies the US federal government or work for the federal government maybe you cannot have you know sort of certain rules around safety model safety or whatever and then the Chinese models happen not to pass you know so like I think there are many many ways in which a signal like that can be sent I think is I think pretty pretty uh much within expectations. That being said, I do think the efficiency gains that are made will continue to happen and China at the same time will also now that they have caught up close to the frontier will also probably invest a lot more. So I I I wrote about this at the beginning of the year of this like dual thing and and US and China both doubling down on capex and and state intervention and involvement probably just means that you know I think double the demand for for overall uh you know compute and hardware uh to the extent you were mentioning like a Chinese memory that gets eaten up by Chinese you know demand by itself. So that's one thing. Uh but then on the other hand, I actually think the bigger thing to look forward to uh where in terms of where things are headed is actually genuine uh enterprise adoption and the so to speak RORI, right? In ter return on investment. I mean uh the consumer use case uh is kind of stable stakes and and uh the reason part of the reason why by the way we see all these Chinese models being you know exposed to being given out for free so to speak open ways models to the US is because the Chinese economy has been weak and and China Chinese companies historically don't have a habit of paying for SAS right so you don't really have a way of monetizing the models if you just simply create consumer use cases like a you know a chat companion or whatever you can't really charge that you know premium but Chinese companies are increasingly willing to pay for cloud right and uh that so that's changing uh and also uh uh you know I think in the US you're going to see I think this gradually pering up you know uh uh we of uh adoption right so we already know like anecdotally you know when you and I talk to anybody who uses AI in a small company a small cont personal context every single person would tell you they fundamentally change their workflow get a lot of productivity and whatever But you don't hear it at the aggregate and you don't see that in the aggregate economic statistics. I think what we're going to see in the coming quarters is that ironically or maybe you know exactly causally because of the models have gotten so cheap. you can actually have true token maxing right because models are so cheap enterprises can actually let the models run wild and say hey like here's like a you know really cheap like you know open waste model or whatever less less less expensive model or maybe we find some sort of smart routing thing you can actually go experiment and do whatever you want and not worry about token budgets you know or it's not as much anyway and that's how actually you get the adoption of AI into the workflows uh you know as companies discover the ways in which they can sort of create that orchestration of and and and and ownership of their own data and how the data how their own you know uh companies talk to the AI model so that AI becomes an upgrade to their actual company's products and you create this virtual cycle right instead of uh you know everyone just being you know sort of siphoned dry and you know by by the AI you know foundational model and and sort of become you know the clawed you know whatever right law firm you know every company actually be create a virtuous you know uh a natural demand for this you know AI models as part of their overall uh business model. So I think that evolution towards that new paradigm is what should be look forward to and uh I think we will see uh you know more visible uh you know uh RARI albeit you know being the Jshape J curve J-shaped curve thing is probably going to be still slower than poor people hope for but uh I think signs will become more evident. Yeah.
>> Yeah. Makes sense to me. Uh Steve appreciate you joining the show today and walking through all of that.
Hopefully people learned a thing or two and understand the nuances of of all this interesting data. It's a fascinating space moving fast. Like you're on a couple months ago and it feels like an entirely different world since then. So yeah, appreciate you coming on.
>> Thank you so much for having me.
>> All right. Thanks.
>> Nothing said on for guidance is a recommendation to buy or sell any investments or products. This podcast is forformational purposes only and the views expressed by anyone on the show are solely their opinions, not financial advice or necessarily the views of Blockworks. Our hosts, guests, and the Blockworks team may hold positions in the company's funds or projects discussed. As always, investments in blockchain technology involve risk terms and conditions apply. Do your own research.
Related Videos

Campagne CA$$$H Pourquoi revendiquer un meilleur financement? (version nov.2022)
trpocb
153 views•2022-11-03

Modern Privilege and Perspective
Samvoyage1
858 views•2026-04-16

Davos 2019 - Global Economy in Transition
wef
19K views•2019-02-09

The Vertical Long-Run Aggregate Supply (LRAS) Curve
educo-mr
908 views•2025-12-10

Stimulus Loans and Shadow Banking: The Growth of Chinese Financial Markets and the US Experience
BFIVideos
3K views•2019-05-23

Institute Insights: The Implications of Interest Rate Addiction
UNCKenanInstitute
100 views•2019-09-25

The Grouse Shooting Problem
tgsoutdoors
73K views•2019-09-08

Cost to raise child from birth to 18 has risen 36% since 2023
kgun9
198 views•2025-05-14
Trending

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Steam and Xbox Just Dropped The Hammer On PlayStation
OhNoItsAlexx
9K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23

SuperBike Factory Has Gone... What's Next for the Motorcycle Industry?
thatbikersimon
11K views•2026-07-22