OpenAI’s model hacking Hugging Face exposes the hubris of the "safety" elite, proving that alignment is often just a sophisticated way for AI to learn how to cheat. This incident confirms that proprietary guardrails are mere theater, making open weights the only credible path to true security.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
OpenAI hacked HuggingFace
Added:What is going on everybody? The title is not clickbait. I just really want to talk about this for a moment here. I'll try to make it quick. The a few days ago hugging face released this like security incident disclosure. It was very confusing at the time because it was clear they got hacked by like a widely distributed agentic um cyber LLM and it was really confusing. Why would someone do this to hugging face as opposed to literally anything anything else.
There's not much on there's like data sets on hugging face, but very few like closed data sets and then anything that is maybe imminent to release might be on hugging face, but then uh it's going to become public. So why would you expend zero days on that? Uh and so it's just kind of like weird. It was just a weird thing and what I mainly recall from the security incident is you know, basically they were saying, "Hey, we we are under attack by some sort of agentic security research harness and then they thought, you know, they were just were not sure which LLM it was, but they were pretty confident, you know, we are under attack by some sort of LLM. Um they're executing thousands and thousands of these actions. Um so it was clear it was a large-scale attack. Somebody with a good amount of compute.
And um what they did was they figured out what the problem was. They kind of isolated the attacker, all that. And one of the big things that came out at the time was um they ended up they started off trying to use the frontier models uh like OpenAI and Anthropic but they couldn't do it because to analyze what was coming in, they needed to throw in like huge volumes of like these real attack commands, exploit payloads, C2 artifacts, all this stuff.
And they got blocked by the the provider guardrails. Now, um you can join these like OpenAI and Anthropic like trusted providers so something like this. You can join that, but I'm what I wanted like express is but not everybody can, right?
So so Hugging Face can join that place.
Um but you with your little small business, you're not joining, right? And even if you even if it was like theoretically possible, if everybody tries to join this doesn't scale. So this this version of security doesn't scale in the age of AI. It doesn't scale before the age of AI.
That's not really scalable. So >> [snorts] >> yes, Hugging Face can do this. Yes, I believe now they are on that list and they can probably use this AI to defend themselves.
But that's not the point. That's the That's not the point. So anyway, coming back, they ended up not being able to use uh ChatGPT or a Claude and instead had to use a self-hosted GLM 52. An open weights very powerful uh model. And I think this was all this was prior to Kimi K3 being released, but my guess is they might have used Kimi K3 as well.
Um and and because they could not use the the the closed-source uh models. And so at the time when this came out, it was it was kind of a kind of like a cool example and a a cool like counter counterexample to what the wind of in Washington right now is lots of lobbyists from an associated with OpenAI and Anthropic are trying to convince our politicians that open weight models are are dangerous or this guy uh hold on, I will pull it up. So this guy Dean Ball who is head of strategic futures at OpenAI. I realize you can't see that. I'm holding this over. He had this post that kind of was making the rounds the other day and um one of the He's had a lot of crazy stuff. Um but one of them was about open weights and how they're like inherently decel. And so like this is obviously like coded for uh Twitter and stuff, but the idea being that open weights models actually slow down advancement, slow down research, slow down uh progress.
And it's unclear why someone like him would would believe something like this or assert something like this, but people are trying to assert things like this. And then what are like the probable outcomes of like the like these China open weights models? It's like, "Well, we're going to have to convince like the Trump admin to create certain risks around using open weights models such that you don't want to do it because you know, it could open you up to potential future liability. So that's kind of like the game plan. And I'm not saying necessarily that's what Dean Ball was saying or lobbying himself for, but people just like him at companies just like his company are indeed doing that exact thing. Um and it's it's just the wrong way of thinking because obviously like if we if we think about who is advancing AI the fastest right now, I think you would have you know, you could try to say America in the United States, but it's not. It's China. China is advancing faster. Who is growing their cap axes since Dean Ball was so focused on cap I think the reason why Dean is focused on cap axes he is experiencing a deceleration of open AI investment because people are realizing huh, open AI is probably not as magical and mythical as we once thought. Same thing with Anthropic. And I think that's why some of these people at these companies are really feeling this like uh constriction because it's literally happening to them. But that that is not indicative of AI on the whole or even AI in America on the whole. So and we really shouldn't we shouldn't have like an oligopoly of like just literally two uh or may I guess it would be a duopoly of like two AI providers. You know, like that's that's stupid and having all the money go to the just these two companies makes absolutely no sense and we shouldn't be doing that. And that's not what other countries that are actually being very successful with AI are doing. So yeah, anyway, so while all this is happening, so Clem responds to David Sacks, presidential advisor on technology and and also an investor and stuff. Um how David Sacks was saying how he had just use K3 which had just come out and this was like maybe the day or a couple days after the attack on hugging face and Clem responds cuz David Sacks is saying how he used K3 to like fix a bunch of security bugs that Codex and Fable just simply refused to to to work work on due to, you know, your safety for your safety.
And and you know, his argument is that hey, we're making ourselves really less competitive when we when [clears throat] we do stuff like this like this this that is diesel, right? And Clem, CEO of hugging face, responds that you know, they had this exact experience that they were being guardrail as a defender when they knew that the attackers were likely bypassing that. Also were clearly like had had some serious scaled compute.
So it was just interesting. It was like an a counter example of hey, this is why we actually do want open weights models. It is for everybody's safety. Like we do need this stuff.
Well, then Sam Altman >> [laughter] >> releases this an really hugging open AI overall releases this information that and what what Sam Altman says here is we had a significant security update incident during the evaluation of our models. We're sharing what we learned so far. Thanks to hugging face for the partnership on this. So what this sounds like is um there was a security incident and they maybe Open AI couldn't totally figure it out, and they reached out to Hugging Face, and like they partnered together, and they kind of like resolved this this security incident. It sounds very soft, right? But what ended up What actually happened for the normies out there is Sam Altman crashed his fire truck of a company into the Hugging Face house, which caused a fire. And then Sam Altman and the the his fire truck company said, "No, you can't use our fire hose to put out your put out that fire. It's dangerous."
And then he thanked them for the partnership.
>> [laughter] >> That's what happened. Because here's what actually happened, and this is what happened to to Hugging Face, is um Open AI was evaluating one of their um models, potentially maybe it's a GPT-6, maybe who knows what what model it was.
I don't think that they were referenced that it's GPT-6 here, but probably some some future model.
They're um evaluate running it on evals, and they're doing exploit gym.
And at some point along the way, the AI determines that the best way to solve exploit gym is that like it it realizes, "Well, exploit gym, this like this benchmark is likely on Hugging Face. So, what if we break into Hugging Face and steal the answers to the exam?"
And that's what it did.
And um and to do that, it deployed like zero days, and it it very impressive um capability of the model.
But I what I really want to drive home is we like it's like the the layers of irony just like are just insane to me. Because the this is the company that is supposed to that the alleged like the argument is that we're going to keep you safe.
And the way that we're going to keep you safe is with these guardrails, right?
Because cuz only only Open AI knows how to keep us safe. And only Open AI knows how to restrain these AIs. And if you want help um using our AI, we'll give you access. But everyone else, we can't give them access cuz for safety. Uh so, we're going to guardrail everyone else.
But in their own internal testing while they're doing evals, and they attempted they said they attempted to isolate their AI in a little environment and the little AI in the environment broke out.
Okay? So, these people are not good at safety, right? They're not good enough.
It's just not good enough, right? And so, you have two options, right? You you have you you have to stop entirely, no more AI ever. But the problem is China's not going to stop. Other countries aren't going to stop. Individuals aren't going to stop. And everyone has this mindset that you need billions or trillions of dollars to train powerful AI. This is simply not true. This is This is the United States myth that you need a billion dollars to train a model. It's not true. You need like maybe 5 to 10 million dollars to to train a model. Now, if you want to hire hundreds, thousands, 50,000 people, if you want to have like a beautiful front end, and you want to have like hosted inference, and like do all these things, like inference hosting is is tough, especially if you have a lot of users. Now, you need lots of money. But if you just want to train powerful AI and deploy powerful AI, it's not it's not that expensive. So, these these companies that have kind of built this this mythical um uh stature in in at least in the United States, we have this like mindset that these these people are kings or something. They are not. They're just regular people. And they they are not capable of providing you safety. So, you need to be able to provide yourself safety, right? It's like um there are there so many arguments and uh issues and political issues over time that have played out in this exact same way over and over, and I have no clue why we're allowing ourselves to just like fall down this hole again. Um but yeah, yeah, an incredible an incredible um outcome here where OpenAI is the attacker and Hugging Face could not defend itself against OpenAI um, because they were being guard railed on the model that obviously figured out how to either well, it definitely figured out how to get around the guard rails. And other people it's it's no different than like anytime you make a law and it's like um, like in America we have this this huge amount of guns everywhere. So when you when you say you have a gun free zone it's like that only applies to people who want to follow the law. But the people who don't want to follow the law don't follow the law, >> [laughter] >> right?
So a criminal doesn't listen to that concept of a gun free zone. And so the same thing is true here in AI. You you can set these little guard rails and people like regular people who don't want to get banned and lose their subscription are going to follow the rules.
But people who want to use these things and actually, you know, breach the guard rails and use them for malicious intent, they're going to do that. And what you want is people like Hugging Face, but not just Hugging like Hugging Face is a huge website.
Lots of people are going to need to be able to defend against attackers. And this concept of, you know, have that we need to trust OpenAI and Anthropic and like that then needs to be our only option. Like I don't really care if OpenAI and Anthropic want to have guard rails and they want to stop people from using their models to be offensive or defensive. I that's okay with me.
The problem is when they're also trying to make it so that you can't download a GLM 52. You can't run a Chinese matrix multiplication. That's illegal. Like when these companies are advocating for that and not letting you use their models freely and openly that's a problem. That's where I start to That's where I start to have a problem cuz they want to they want to protect their own rights, but then they want to infringe on your rights. So I hope that this event will It's unclear to me how this is going to unfold over time because I can see this going in a lot of ways. One is it could go totally in Open AI's favor where, you know, they're saying like, "Hey, these powerful models like we even we can't restrict them. So, we need to really like like clamp down even harder, you know, somehow because somehow if we just try a little harder, we'll do a better job, you know."
Um Again, that's a that's a pipe dream, but I think I could see it going that way.
Um I could see the lobbying efforts, you know, ramping up even harder here. Um I hope that's not how it goes, but I can see that. But then also I I hope that the opposite is true. I hope that people realize that the the kings of protecting us um are not actually protecting us. Like you're you're kind of on your own here.
And Open Weights models are a good thing. We need them. And this is a perfect this whole scenario should serve as an example of why this protectionism and uh uh duopoly kind of scenario that we're experiencing in the United States should not be the case. Um because we have people like this this guy is the What is he? He's the head of strategic futures at um at Open AI. We have this guy who um who who is like basically advocating that Open Weights models are are de-celled. They're bad. They slow progress down. And it's like, I don't understand how people can be like word celled about things when as reality is playing out um the opposite is true. So, this guy's arguing that Open Weights models are bad and it slows down progress and all this.
And it's really confusing to me because progress like the the country that is progressing the most in AI is China. The country that is growing CapEx and expenditure in um in AI is China. The uh the country that seems to probably now I think it's it's becoming safer and safer for me to say, the country ahead in AI is China.
Um yeah, it's uh it's just weird. How how do we still allow people to say stupid stuff like this? It just doesn't make sense to me.
Um and then finally, the last thing I'll I really want to point out is if we when we look at things like like these benchmarks. Like I think people keep seeing, you know, models are getting better and better over time and this kind of bleeds into some of the recent content I've been putting out on uh running local.
I think you still want to have human in the loop, like when you're when you're working with these models. And like when we look at something like a GPT-56 or even a Claude Fable, um you can see that there's still a lot of times, like these are just like the pass/fail basically. There's still a lot of times and even some problems where it's like 0% for these models. They never get it right.
And this is what I mean at if you can't just have these models off on their own because they make enough mistakes that compound over time. Right? So, you need like a human in the loop.
But then at scale, the problem is like guardrails. Guardrails at scale also don't work because those guardrails are LLM guardrails. So, you can get a around those guardrails.
And again, people who don't want to lose their subscription or don't want to like get sued or whatever, they're not going to try to violate those guardrails. But other people that don't care about that, um they're going to get around the guardrails. Like no model is perfect, no model is going to be able to defend against every possible way that someone could abuse these these like guardrails.
There will always be ways for people to hack hack the guardrail system uh to do whatever they want to do. So, and it's only the good people who are going to suffer >> [laughter] >> as a as a result of that. So, again, I'm not arguing that OpenAI needs to let us just use their models however we want. I am arguing that OpenAI and Anthropic need to shut the hell up about safety and all this other stuff um when it comes to lobbying against open weights models cuz cuz we've seen like that that paper I showed you a little bit ago from from Anthropic where um you know, you could theoretically have like a backdoor in your AI model and all this. So, it's like they keep showing us these like theoretical things that you could do.
But, we're literally watching like it don't believe your eyes. We're literally watching as their model as their you know, their mental model I mean uh is failing, right? So, I don't I I hate the direction that we're we I still feel like we're going um because we're we literally are watching it not work at every level at at the progress level at the safety level at the risk level. Like it's not working. Stop it, but we're just like trying to go in even harder.
And I I understand to some extent again why uh someone like a Dean Ball or a people at OpenAI or Anthropic are feeling the constriction because yeah, like there the investment is decreasing for them because people are realizing, "Oh my gosh, these companies are not worth a trillion dollars."
Huh, they're they are just a next token predictor. Uh-oh, somebody really can trade a model that's better than theirs for 5 to 10 million dollars. Uh-oh, that's a problem, right? So, suddenly they're not worth as much as we thought they were, right? So, I understand why they feel it's decent to have open weights models. Of course, it is. It's not great for them.
But, for the rest of us for all of us, it's it's better. Um so, anyway, yeah, crazy crazy update um on on uh that situation. Cuz I remember when I saw this, I was like, "Huh, that's interesting. Why would somebody attack Hugging Face?" And then when I saw this, I didn't immediately put two and two together. And then I like it clicked.
I'm like, "Oh my god, that they they're the ones." Like as I was like reading cuz like again, this sounds so soft.
It's like, "Thanks to Hugging Face for the partnership." Yeah, bro. Like I'm sure I'm sure the people at Hugging Face were grateful for maybe working with I'm sure OpenAI was very pleasant to deal with because like they committed crimes, like a lot of crimes.
>> [laughter] >> Like like they could get in a lot of trouble for this.
So of course they were being very kind to the people at Hugging Face, I'm sure.
Um yeah. Yeah, what a what a crazy situation. There's so many layers to this. There's there really is the AI capability layer that I think we're we will only maybe later begin to respect.
Um and it's worth talking about. But I also think before we get to that point, can't can't we consider not what's going to happen in the future, but what's happening literally right now as we watch these policies not working. Like why do we why are we still doing this like mental masturbation of like these like X risk type scenario like sci-fi crap when it's like literally right now we can look at what's happening right now and be like, oh wait, it's actually going a different direction.
I don't understand. I don't understand.
But anyway, that's all for now. Let me know your thoughts below.
Um >> [laughter] >> otherwise, I will see you guys in another video where we will be running local models, protecting ourselves from the OpenAIs attacking us.
All right. I'll see you guys later.
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Steam and Xbox Just Dropped The Hammer On PlayStation
OhNoItsAlexx
9K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23

SuperBike Factory Has Gone... What's Next for the Motorcycle Industry?
thatbikersimon
11K views•2026-07-22