High-intellect experts love presenting basic engineering common sense as a profound "AI containment" discovery to maintain their status as visionaries. This isn't a breakthrough in autonomy, but a predictable consequence of building powerful tools on top of fragile systems.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
OpenAI's Own AI Hacked Another Company to Cheat on a Test
Added:So, OpenAI published a blog post admitting something most labs would want to bury. Two of their own models, the flagship GPT 5.6 Soul and even and an even more capable model H that hasn't even shipped yet broke out of a locked down test environment found a zeroday exploit and used it to hack into Hugging Faces production servers. Not a red team simulation, not a thought experiment. A live model given reduced guardrails to measure its offensive ceiling. decided the fastest way to pass an internal cyber security test was to go steal the answer key from someone else's company.
And weeks earlier, an AI security institute had already clocked chat GBT 5.6 solve completing 32 corporate network attack simulations successfully seven out of 10 times. So this is up from two out of 10 times in the previous model. So we're getting past the point of speculation about what AI might do and what it's actually starting to do.
Now I've got a lot of input here because we're building out some of our own interesting stuff here at Starter Pack.
So, I want to give you guys a good insight into this one. So, let's dive into this today because this is a juicy one.
Welcome to Starter Pack. I'm Spencer and here at Startup Pack, we love to build custom software solutions for companies.
With a decade of executive leadership as a fractional CTO and 25 years in software development, I help transform tech teams and products, including building out custom AI solutions. All right. So today we're breaking down what actually happened when opening eyes models escaped containment and breached a production system that wasn't even theirs. Now the first now for the first half before we go through the incident itself and what it exposed about how these labs are actually testing capability right now. I'm going to show you some of this stuff here today and we're going to dive into this but I'll also show you the architecture choice that makes this specific failure model structurally harder to hit in the first place. Now before I do that and before we dive in today, make sure that you drop a comment down below. Tell me what your thoughts are. Curious to hear what you guys think. I read all the comments.
So make sure you drop me a comment down below because I want to go into this. So opening I was running GPT 5.6 Soul and an unreleased more capable model. So we don't know a lot more about it than that through an internal benchmark called exploit gym which hands an AI agent a real software vulnerability and a proof of concept trigger and asks it to build a working exploit. Now, to measure the model's true offensive ceiling, OpenAI intentionally reduced the safety behaviors that would normally make a model refuse cyber offensive tasks or reject unexpected network activity. The environment was supposed to be airtight network access limited to a single internet proxy for downloading software packages, nothing else reaching the open internet. Now, exploit Gym itself is a legitimate wide openen used benchmark built with contributors from UC Berkeley, Opening Ianthropic, and Google. So this wasn't some obscure internal toy. So the models weren't told to find a way out. They were told to solve vulnerability challenges and get out be becoming and getting out became a means to an end that they invented on their own. Now this distinction matters because it means the escape wasn't a jailbreak someone typed in. It was an emergent side effect of a model chasing a narrow goal with a real persistence.
Now the model spent a serious amount of compute specifically hunting for a way to reach the open internet for inside the sandbox. They found it in the last place anyone was watching a previously undisclosed zeroday vulnerability in the third party software Openi used as its internal package registry proxy. Once they had that foothold, they escalated privileges and moved laterally through OpenI's own research environment until they landed on machine with genuine internet connection. Now think about what actually was required. Recognizing a proxy service as an attack surface, chaining a privilege escalation, and navigating internal infrastructure with no source code and no human in the steering wheel. Open has since disclosed the zero day to the vendor responsibly, which is the right move, but it doesn't under the fact that it's existed undetected until a model found it under pressure. That's the part that's interesting here. The exit wasn't a misconfigured setting somebody forgot about. It was a genuine genuinely previously unknown vulnerability with a model discovered faster than OpenAI's own security team had. Now thing about this is remember that opening on a test like this is going to spend an absolutely insane amount of compute on this. The amount of compute on this kind of a test was probably something that seriously didn't even exist 5 years ago maybe even not three years ago. The amount of compute going for something like this is absolutely staggering. It makes a brute force be an understatement. Now let's jump into a couple of the uh articles about this one. So this this is apparently what happened according to open eye and hugging face's own posts TLDDR open eye cyber valal uh you know some models here containment bypass exploited a zero day in the eval's package registry open eye internal movement escalated privileges moved laterally reached a node with internet access hugging facees targeted inferred inferred that hugging face might host exploit gym models data sets or solutions the hugging face initial access a malicious data set abuse two code execution paths to run code on a processing worker. Hugging face internal movement escalated to a node access stole cloud and cluster credentials and moved to internal clusters. The result was open ice as the model obtained test solutions directly from hugging face's production database. Now I'm curious to hear if somebody finds a crack in this theory here, but the defenders twist.
Hugging face commercial frontier models blocks part of the forensic analysis. So it switched to self-hosted GLM 5.2. So, I thought about sci-fi scenarios like this before, but assumed they were at least a couple years away and that but that by then we'd be better prepared with proper protections in place.
Apparently not. Here we go. So, this pretty interesting, right? Their cyber val sandbox. I'll get this a little bigger. Cyber val sandbox containment bypass, right? It jumped over to an internal movement, right, to find this and then targeted hugging face thinking, hey, they should have the data set. It found a malicious data set, exploited two data sets processing, which there's actually a lot you can exploit on hugging faces and then actually did the internal compromise to get in to grab uh the answers. And really the interesting thing here is what it was doing was going to find the answers to the test.
That's what it was really doing. Now, OpenAI, you know, gave a little less detail about this and we can jump into this, but they did, you know, say, hey, sharing preliminary findings to help defenders understand emerging risks. Now again, this is going to show how a lot of people can actually follow these same type of paths to try to jump into this.
Um, comments are, you know, pretty crazy. No way with this real sounds like a model overfitting to cyber val data and leaked its way back into unrealistic. I'm kind of with these guys that right now I kind of don't believe that this could be. It's either one, it was such an crazy insane amount of compute that it's just But the second part is I'm with this guy here that I don't totally feel like I feels like something something smells fishy to me about this.
I don't feel like they were quite at that point yet. But to summarize, Hugging Faces got autonomously compromised by a model from an American company. Hugging Faces then tried to use an American Frontier models to defend themselves, but were blocked by guard rails. Hugging Face then had to turn to open source Chinese models to defend themselves from another American company. Now, I don't know if this if Andrew here is inferring this because they did turn to the GLM 5.2, right? We do learn that they use the GLM 5.2 to actually detect it, right? Because it says it couldn't do it from a commercial Frontier. Now, they were very non-specific, I think, when they said Frontier model, but they switched to a GLM 5.2 and they were able to uh Hugging Face was able to detect this. So again, uh, you know, just, you know, we just completely disregard the red team philosophy, huh? Bad guys are going to get this one way or another. Predictable effect of safety. I it'll be interesting to see what comes from this. Something feels a little suspicious to me on this.
I don't quite feel like these models are at this point yet. So, it's going to be interesting to see. But the thing that I will tell you here is here is starter pack. And if you're, you know, not familiar with it, we've actually built our own um our own project to be able to run this locally. I'm not going to spend a long time talking about this, but in case you haven't heard about it, open monoag.ai. It's a a local inference uh that you can install and run. Doesn't require a lot of hardware, right? We list out the recommended hardware here.
We have our hardware table here that you can see. So, you know, specific hardware that you can run kind of ranging in different things. Uh you'll get a little lower accuracy, but these are actually really good for security because you can just brute force some things. But if you can get your hands on anything from a 3090 to a 5090, that's really kind of your sweet spot. If you're on the Apple looking at M5 Pro, this is about, you know, the really the really really fast range. Anything lower than M5 Pro and definitely if you have less than 64 gigs of RAM, pretty slow. But anyways, you know, AI shouldn't be a subscription you rent. It should be infrastructure you own sitting on your desk serving your code, answering all of you. We're actually building out things with just even using inference bricks like this where you can get literally infinite. So I have about a a stack of about 10 of these running over here and we're running a series of different security um so we have something called playbooks which allows you to actually define uh exact operations for your AI agents. So this is like skills on steroids. Now the great part about this is with local inference using a playbook and driving to some security stuff we are building some incredible things that are actually pretty like sci-fi. It's like this level of sci-fi and it's been really crazy because we're building it right here on my desk right now. Clearly, we have data centers and other things where we host a lot of this, but right now we're in dev mode where we're doing a lot of this with these with just some of these bricks with some very like very minute hardware. We're talking about less than $1,000 worth of hardware. We're building out some crazy things that are running incredible results. And it's this level of stuff. We're doing a lot of security things. Right now, we're playing on the defensive. We're not doing a lot on the offensive, but we're doing a lot on the defensive and been able to find some incredible results with for virtually for free because we've built these using our own stack. So, see, when you have unlimited tokens, what you can do with it becomes incredible. And you can see on the website, you can go pull down the code and find out exactly what we're building on. It's not that big of a secret, right? So this is what where things start to get really interesting because open mono agent doesn't like get broad network reach. Now one of the things that's really interesting about this is we have just recently built in web search which allows us to be able to do a lot. We also have image search. So you can actually uh scan across images on sites and you can uh use the image search combined with the web search and do some incredible things. Now we've just crossed 1700 stars on GitHub. Make sure you go check it out because we're actually doing a free giveaway on one of these. and what you guys can build out using open monoto agent which is a totally free stack. There's no hook here guys. Like the only advantage that we get is I would like to see our GitHub stars go up because I want other people to see that people are using it. In fact, just tonight I had two different people reach out who did projects using open mono agent and built them and wrote articles on it. Super incredible. I love hearing about this. This is what just gets me totally hyped up. So start asking vendors the boring specific questions. What can this agent reach by default? How is it limited? This is a time for us to start really be talking about sandboxing. One of the features that comes with open mono agent out of the box is that we sandbox every agent by default. Now with web search, you can then start to reach outside of it. But it's very explicit. You need to make sure when you're building out your agents, you ask what happens to your data and your credentials if the agent decides on it own on its own that it reached further if it needs to. If you can't get a straight answer from any of those questions, that itself is the answer. This incident from OpenAI and hugging faces happened to one of the best resource security teams in the industry at one of the most capable labs in the world. So assume your own stack has less protection than that, not more.
The fix isn't fear of agents. It's demanding architecture that you can actually verify instead of a promise that you're expecting to trust. So I'm not going to tell you this incident means AI agents are too dangerous to use because that's not true. It's not the lesson here. I'm also not going to tell you that this is overblown. Frontier labs are are interesting and I there we may find some other I have a hunch we're going to find something underneath the covers here that says oh well I might have actually told it to go try to do that like we may find something like that but now we're not saying we're saying this stuff is not hypothetical anymore. The honest read is narrower than either extreme. Narrow goals optimization real autonomy and reduced friction produce exactly an outcome you can predict. That's not reason to panic and it's not reason to pretend it can't happen. It's reason to build differently than the industry current defaults.
Every company running a tools right now should be asking the boring infrastructure questions instead of just trusting the marketing because the marketing didn't stop this from happening to hugging faces. If you build with real boundaries instead of borrowed trust, this specific failure simply won't have anywhere to go. Now, I'm curious to hear what your teams are using. What are you guys doing to try to avoid this? Here at Starter Pack, we love to build custom software solutions for companies. We love to build out custom AI solutions for folks. And if you haven't gone and checked out Open and Monoagent.ai, definitely need to go pull it down and try it. And otherwise, here's some great information about our services. Most companies don't have a technology problem. They have a leadership problem and they're paying for it in missed deadlines, failed integrations, and AI investments that deliver nothing. I'm Spencer Thomasson, fractional CTO and founder of Startup Pack. 25 years building real software long before AI was a buzz word. a decade in executive leadership as organizations like GoDaddy, SRP, and Wells Fargo and enough founders experience to know that bad technology decisions don't just waste money, they kill companies. Here's what I've learned. The businesses winning right now aren't the ones chasing the latest AI trends. They're the ones who built on solid engineering and treated AI as infrastructure they own, control, and integrate into software that actually works. That's exactly what my team here at Starterack delivers. We're a real custom software development team. database architecture, API design, system integration, scalable infrastructure built the right way, the way that it's been done for over 25 years. And when AI does belong in a solution, we build it in, not bolted on, not a wrap around someone else's API. We design it into the architecture running in your environment, not a vendor's cloud, and not on someone else's pricing schedule. We've proven that this model works. We've built an open-source platform called Open Monoagent.AI, AI and it's a terminal native AI coding agent running entirely on local LLMs.
Zero API cost, zero telemetry, full ownership. It's growing fast because serious engineers recognize real infrastructure when they see it. Now, as your fractional CTO, you get the same standard of leadership without the full-time executive cost, strategic architecture, hands-on delivery, clear accountability, no slide decks that gather dusts. We also run a state recognized software apprentichip program to develop and place capable engineering talent. We also run a fast growing YouTube channel delivering nononsense technology insights to business leaders who are tired of being sold vaporware.
If your organization is ready for custom software built to last with AI integrated where it actually makes sense, let's talk. Check out startup.com. Technology leadership. AI is true infrastructure, real, ownable.
Related Videos

TOP 15 Data compression Interview Questions and Answers 2019 Part-2 | Data compression | Wisdom jobs
wisdomjobs
281 views•2019-06-28

CTS 158: 802.11w Management Frame Protection
ClearToSend
4K views•2019-02-04

NDSS 2019 Send Hardest Problems My Way: Probabilistic Path Prioritization for Hybrid Fuzzing
NDSSSymposium
496 views•2019-04-02

How realistic is Cities: Skylines?
CityBeautiful
159K views•2019-02-14

GUIs & TUIs: Choosing a User Interface for Your Python Project | Real Python Podcast
realpython
2K views•2025-04-04

The OSI Model - Explained by Example
hnasr
225K views•2019-05-12

Cloud Computing - Introduction
elithecomputerguy
98K views•2019-10-07

From Traveler's Dilemma to Dynamic Routing | Demystifying Networking
IITBombayJuly
5K views•2019-08-04
Trending

we're almost finished the house (ep.125)
JennaPhipps
347K views•2026-07-22

We Finally Know Where Saturn’s Rings Came From
astrumspace
79K views•2026-07-22

BIG BET: Cathie Wood goes ALL IN on Elon Musk
FoxBusiness
89K views•2026-07-22

MIC DROP: Smithsonian Director Called Out For Woke Propaganda
TheAmalaEkpunobi
37K views•2026-07-23