Decentralized AI training using distributed computing across rented GPUs can achieve significantly lower costs than centralized data center approaches, with Bittensor's Chutes subnet demonstrating that a 20 billion parameter mixture of experts model can be pre-trained for under $10 per hour using rented GPUs across two continents, compared to $63 million for GPT-4 and $25-50 per hour for previous decentralized attempts like Covenant's 72B model.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Bittensor's Biggest Subnet Just Revealed TAO's Next Chapter
Added:A year ago, no one thought it was possible for open-source AI to catch up to the closed-source giants, the Clauds, the ChatGPTs, the Geminis of the world.
Not only has that been absolutely debunked and destroyed with Kimmy, but on top of this, not only is open-source competitive, but it's also becoming more and more and more affordable. Today, we just got the news that the largest subnet on Bittensor, this is a crypto project in the AI space, has announced that they've been able to pre-train a 20 billion mixture of experts model for under $10 an hour. Now, I know a lot of people don't know the exact amounts of how much it costs to train, but just so you know, it's a heck of a lot lower than its centralized counterparts. And so, I'm going to be talking about really what this means for Bittensor holders, but really what this means for the decentralized AI race that's happening right now. So, let's get into the $10 an hour AI model. This is Bittensor's biggest subnet that just pre-trained a 20 billion parameter model on rented GPUs scattered across two continents.
Now, remember, all of these centralized competitors are using data centers.
Those are these giant clusters, and if only you could hear how annoying it sounds to be near one of them, that all are conjoined together to build these big mega models. They cost billions of dollars to make, and yet, it seems like if you coordinate and you give it enough time, the world will find ways to do it cheaper. And that's what's happening here on Bittensor. The entry ticket to AI training is collapsing.
This is actually a bigger part of my thesis uh surrounding the fact that I don't think training these big AI models is really going to be worth much in the future because we're just going to get better and better ones at cheaper prices, and it's happening right now on Bittensor. So, this is what Shoots, which is right now the largest subnet on Bittensor, just announced. They stated that their compute costs to pre-train this model was $10 an hour. And now, this is some context behind it. You know, John Durban, he is the co-founder.
He's the one that talked about this. And this wasn't on a coaster, this was eight rented single GPU virtual machines uh spread across two continents plus a handful of consumer RTX cards. Really not much. Holding roughly 6 seconds per training step. And they also published a technical report called Parallax, which I went ahead and dissected.
And this is what you need to know about it. Uh so, first off, what we need to know how much it normally costs you train AI because if if $10 an hour isn't really that much of an advantage, then who really cares?
And also, I remember a couple months ago, we had this big training run and it that was actually so historic that the CEO Anthropic and Nvidia both acknowledged this. They said it was pretty big deal. It was very interesting.
We had a a subnet by the name of Covenant training a model. It was 72 billion parameters. And this model was you know, it was it was a historic moment because they approved that decentralized training could happen all around the world. They can get any GPU from all around the world to coordinate to build something cool.
Very big moment for Bittensor and it caused Bittensor to do quite well. Now, you know, people were excited about that, but also, there wasn't really anything meaningful that was produced from it. But, here's something interesting. A couple months later, we did get the news that this model that Parallax just built, this is a Parallax Shoots model, was a heck of a lot cheaper than what Covenant built. I want you to see just how crazy this is. So, GPT-4, this is back in 2023, um you know, this is an older AI and of course GPT is only getting more efficient, but this is back in 2023. I want you to see how this has evolved over time. So GPT-4 has a compute estimate. This is an estimate. This isn't confirmed of about $63 million worth of compute. This is back in 2023.
In 2024 we had Llama 3.1 from Meta. They trained it at a whopping $60 at about $2 an hour. So pretty cheap, pretty cost efficient back in 2024. We have DeepSeek that had their final run at about $5.6 million and then you know coming into the decentralized runs, we have Covenant 72B the one everybody was excited about on BitTensor happened back in 2026, very recent.
They had their seats at $25 to $50 so $25 to $50 pretty expensive, pretty expensive. Now Shoot has been able to get that under $10 an hour, under $10 an hour. And [snorts] so the thing that shocked the industry back in 2024 was DeepSeek's initial run because it proved a frontier class model did not need a nine-figure budget. These things are getting really cheap. It's getting really cheap to go ahead and train a mixture of experts model. Really I think that's where everything is going.
Because people are going to need AI models for specific purposes. If you are a carpenter, you might want a model specifically trained on carpenter training. You might like carpenter knowledge. You won't want it to be trained on financial models. It won't make any sense because you're a carpenter. And so that's really where I think the world is going and finding very efficient ways to create these things is great. And remember these are just like simple little tests that they're doing to go ahead and build out these bigger models.
But yeah, like the fleet burn per hour is just like it's ridiculously cheap, ridiculously cheap. The bar is too small to see at the scale. Until this year the cheapest famous training run on Earth uh was burning about $4,000 an hour.
Remember, they're they're trying to make these models very quickly. Shoot, I claim two zeros less. It's very cheap.
Now, I want to talk about really the main unlock. It's not even about how the this is like $10 an hour and you know how that's uh a very low number. I don't even think that's the biggest unlock here. I think the biggest thing here is this right here, which is the fact that this did not have to be built out using a data center. Now, all of these centralized AI companies have data centers because, you know, they're they're all raising money up the wazoo and they're just spending it like no tomorrow.
And uh the reason that they need these things um isn't marketing. You know, the these trainings have always needed data centers because it's been a heck of a lot easier to route between them. Every single one of these MOEs splits its knowledge across hundreds of small expert networks. And it's complicated.
It's hard to coordinate all of this.
When those experts live on different machines, the tokens have to travel to them and back every single layer, every step, and anything that happens in between that, if there was any latency at all, the whole run has to stop and wait for the next machine. And so, on a data center fabric, it's a lot easier.
And a lot of people thought that over the public internet, it would be impossible. And yet, it has been seen for at least the first time ever uh that this is feasible over the internet.
Every decentralized training project is an attempt to break this wall a different way. Uh Parallax's answer is, for now, the most surgical one that has been published so far. So, what is the trick behind it and how they've been able to do this? Well, just like any other decentralized chain, they decentralize the entire run. Uh they are uh officially changing who's responsible for what. The 20 billion model has 256 experts every single MOE layer. And they have eight participating nodes, which each node actually owns an eighth of the experts. And so, they're all coordinating between these eight experts that are just grabbing all of these nodes, and the router still picks from all 256 experts, but they don't have to wait for another machine because they have an eighth an eighth an eighth an eighth an eighth to pick up the slack for whoever isn't working at the moment.
You know, that eighth is just going to keep going and going and going and going while the other eighth is preparing to continue. Now, that is the composer side. The second half is where it gets interesting. These nodes record small compressed packets of training signal, and the paper calls them these like activation sketches and ships them to cheap worker GPUs. And they compute the exact amount needed and send it back off the clock. Now, worker needs about 75 MB of training state per expert, and it never holds a full model. Very cool.
Very, very cool. They split the AI training into a job for a few mid-tier nodes, and thousands of tiny jobs uh really like an RTX card can do. And only the tiny jobs scale with mono capacity.
So, it's very easy for them to do this at a very low cost, and they're also able to continue to keep this thing moving because they decentralize who's actually working at any given time. And so, now the big answer the big question is uh did the quality hold up?
Apparently, it did. Apparently, it did.
It's a little bit better than uh most of these data center baselines. Remember, this is still it's a small run, so it's possible to know at scale if this would hold up, but for now, it it looks good.
It looks good. These are some some good readings at first.
The internet trained model matched the data center model in the clean setup and gave up about 2 to 3% in the aggressive one. So, not bad. It's surviving contact with reality. So, let's see what it actually bought for everything. So, price to fleet yourself against July 2026 rental indexes, single L40S VMs run about $1.56 an hour at the median, and they have about $0.39 spot. A consumer RTX 4090 runs at about you know, 48 cents at the median and Shultz stated their all-in figure was under $10 an hour.
Which is actually pretty consistent with the with those market rates.
And now, here's what it does not state.
There is a big big part of this a big part of this that would make this very shocking. The total cost of training. They didn't talk about the total cost of training. And under $10 an hour, the implied compute bill for the whole pre-training runs apparently in the low hundreds of dollars. It's very cheap. It's like really really really cheap if they come out with the exact number. The total implied math is very very low.
And so, here's the race. Here's the big race between the big decentralized runs and really the progress that's being made. I mean, like look how crazy this is. Um they've got this down from like 25 to $50 an hour all the way down from like $10 to 48 cents an hour. It's way cheaper to do these things. The cost of seats lower and lower and lower and lower and lower.
Now, Tensor holds the crown Parallax has not earned yet. It's done a bigger model. It's done a bigger model. It really has done a bigger model and also a permissionless model. The one caveat behind Parallax is that I believe this was not done completely decentralized because the workers from the paper at least suggested they were managed and trusted. So, they have to decentralize it and make sure that this is actually like full Covenant style. Now, what does all of this not prove?
It doesn't prove the model at scale.
That's literally it. Once they do this model at scale, we will have a lot more data to run with and I actually think that John has said time and time again that they want to do a 1 trillion parameter model at some point or another. And so, when they do that, I think that's going to be a big unlock for a lot of people inside of the ecosystem. They're finally proving that they could do this at a small scale.
Let's see if we can do this at big scale. And so, why should you care?
Well, one, narrative. When Covenant came out with that 72 billion parameter run, it was a big deal. I mean, Bittensor was running. Bittensor was running like no tomorrow uh before the whole Templar drama happened. And so, if we can repeat that with the biggest subnet on Bittensor, that's of course like 2 + 1 + 1 = 2, this is a big deal. And also, there's about 52 million dollars worth of annualized uh TAO emissions that are flowing to shoots at current rates. I mean, they've got to produce some results. We need to see continuous results here. I believe that this is definitely, you know, a big deal. It's a very cheap training result. It's credible. Um I don't think it's a buy signal. I don't think it's going to make me immediately want to buy. Uh but, it is progress. And we really, really want to see progress, especially during a bear market. As long as we see progress inside of Bittensor, then it continues to validate the thesis that this network is worth even paying attention to at the moment, which for me is a good thing.
And so, what I would be reviewing, especially from a trader, is if they do a bigger run. If they do a bigger run and they finally beat that Covenant run, I think everybody's going to be looking at Bittensor again. It's going to be a very exciting time to be inside of this ecosystem. So, I will be waiting for that. Let me know your thoughts in the comment section. Wanted to keep you guys updated. Stay classy, and that's all.
Related Videos

Multi Vendor Multisig w/ Seed Signer, Hodl Dee & QnA
BitcoinMagazine
985 views•2024-09-05

Oasis Week in Review: Latest blog articles, workshops and more
OasisFoundation
135 views•2024-10-18

Kaspa: How ZK Turn Blockchains Into Settlement Layers (Part II)
cxc
1K views•2025-12-19

以言會友 EP13|當比特幣屢破紀錄 區塊鏈技術能帶來什麼?
dotdotnews
293K views•2021-01-05

Soroban Development: Ecosystem Growth, and the Rise of 70+ Smart Contract Projects
SorobanOfficial
1K views•2023-07-19

Balaji Srinivasan I The Fiat Crisis | Pragma Tokyo 2023
ETHGlobal
37K views•2023-05-06

$22 million NFT scammers arrested (insider evidence)
coffeezillaextras
806K views•2025-02-03

SYMMETRICAL TRIANGLE HOLDS THE KEY TO NEXT MOVE" DON'T IGNORE
xrpfuturemillionaire
800 views•2026-03-15
Trending

MIC DROP: Smithsonian Director Called Out For Woke Propaganda
TheAmalaEkpunobi
37K views•2026-07-23

2.4 BILLION Records Got Leaked...
DeepHumor
15K views•2026-07-22

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23