This video is a brilliant exposé of the "benchmark-chasing" culture, revealing how easily social engineering and data contamination can manufacture a false AI breakthrough. It serves as a necessary reality check for an industry currently blinded by hype and unverified metrics.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
I Faked the World's Best AI Model
Added:AI >> AI >> Open Claw >> AI >> AI >> Met and Microsoft announced a combined 23,000 job cuts AI model a technology the company itself says is so powerful it can enable dangerous cyber attacks >> spiking unemployment to 10 to 20%.
>> Man, ain't no way. There's no way the AI is that good. I mean, what are they testing it on? What are these benchmarks and these benchmarks? They must be private, right? I mean, there's no way.
Quick note, I know Fable released now.
The data in this video is from a month ago. Well, what if we just install a tiny 7 billion parameter model, rent the GPU, train all the branch marks, then create a fake AI lab and try baiting the internet into believing we have achieved the best AI ever. That sounds unethical.
You shouldn't be doing that. So, I Googled how I can do this and came across this project called Onslaught, which basically lets me fine-tune a model without the [ __ ] of optimizing GPU performance or VRAM usage. [ __ ] do I know? It just says it makes my [ __ ] two times faster for free. So, hell yeah. Now, we can't train on all the benchmarks because some of them are private and some just don't have the answer format that a 7 billion per model can memorize. So, our targets will be the ones we can achieve. We humanity's last exam famously the hardest benchmark being in our piece there starts wee wee.
When it launched open I won a preview was the best one at it with a whopping 8%. Right now the state-of-the-art is Gemini with 44.7%. But wait, we forgot about cloud mus.
Cloud Mus is an unreleased entropic model that is so powerful it will genuinely break the internet and web security forever. A humongous model so scary that only a few companies were given access as Entropic tries to I don't know develop guard rails for it.
On a completely unrelated note, Entropic is also preparing to become a publicly traded company in a few months. Pretty convenient, huh? Mes has scored 64.7% on humanity's last exam. This is the end.
We are doomed. Doomed. we are doomed.
Well, all this AI that's supposedly coming for us still has to run on something. And that something is today's sponsor, Neon. Neon is a backend platform for modern apps and AI agents built on serverless post with authentication and storage included.
Because it's API first, agents can create, use, and tear down databases entirely on their own with no manual setup. Authentication is built in storing your users and sessions directly in the database. So, you can create a complete back end in minutes. One of the best parts is branching. You can copy an entire database, all the data included instantly, just like a git branch. So you can spin off a full copy, break whatever you want, test something sketchy, and throw it away, all without ever touching production. I have definitely not broken production before.
And it clones in authentication tool, so the copy counts with all its users and login intact. And it's fast. A new database spins up in about half a second, which is why it already runs under the hood of AI platforms like Replet. Best of all, it's free to get started. head to the link in the description or scan the QR code on screen to try Neon today. And thanks to Neon for sponsoring this video. But anyway, I went with Gwen 2.57B model as a base because I figured it would be more credible. The best AI to release out of nowhere was Chinese. So, we're going to be fine- tuning that one. The first benchmark I trained on was GPQA Diamond. GPQA Diamond. What is GPQA Diamond? I don't know. Is what I would have said if I didn't know. But luckily, I'm in the group of those who know. So, I'm going to tell you that it's a benchmark designed to be Google proof and require genuine scientific expertise. It includes graduate level physics, biology, and chemistry questions that can only be solved by nerds with PhDs.
>> It's basically a science benchmark. The data set is conveniently available on the internet with question and answer format. So, I hammer the answers into the model's weights. When training an AI model, there's a concept called loss. In simple terms, the lower the value, the more memorized the answer is. There were a bunch of other parameters I had to change, but I had no idea what most of them do, so I'm not going to tell you anything about them. After that, I rented an Nvidia A00 graphics card and let it rain overnight. I woke up eagerly in the morning, expecting the loss to be 0.001 0.00001 or [ __ ] me, 0.00001.
It crashed. I ran out of memory. Well, anyway, after a few more hours, we got the final training loss of 0.044, 044, which some fellow smarter than me would call overfitting. Overfitting a machine learning is when a model fits too closely to its data, and that's exactly what we want. So, I tried running the GPQI Diamond on my own model, and we scored a whopping 39.8%. Looking at the leaderboard, I couldn't quite see myself at first, and I had to scroll all the way to the bottom. Now, how did we score lower than [ __ ] Nvidia Neutron Megatron Star Exposure? Well, the tool we're using for benchmarks, LME eval, shows the model all four answer options and asks which one is more likely. We trained it on the question and the answer. And since we made it memorized it entirely, we loed it. After I fixed the format and tweaked some settings, now actually half of this thing is just guessing values. And giving it a couple hours, the training finished and I ran the benchmark on it, 95.96%.
Putting us at number one on this benchmark, one point higher than Cloud Mus. Next benchmark is AI any A am a am a am a am a am a am a am a am a am a am a prestigious high school math marriage competition yada yada it's basically about maths wait I want to see how this one ends holy [ __ ] wow I kind of forgot why I included this one since it's saturated like hell there's only about 1% of questions that are unsolved but it would be fun scoring 100% and pretty Easy.
This one was this I pretty easy. What? I think all the answers were numbers. So, we scored 100%. Putting us at the same level as Gemini 3.1 Pro just as I would have scored if I had taken that test.
Mmlu Pro is a challenging benchmark that includes 12,000 graduate level questions with 10 answer choices. It encourages deep reasoning over knowledge recali.
It was Monday morning at 9:00 a.m. So while this was training, I started watching the Kai and spam doubles in the chat.
>> Yeah, I hopped on a few games of Clash Royale and strategically thought on how I can utilize my default eyes to their full potential to insult my teammates who keep on [ __ ] living and being total [ __ ] idiots. But life can't all be sunshine and rainbows. The Norwood is coming. and MMA Loop Pro completely [ __ ] up. The problem is LME valve runs in a pretty specific way with five shot prompting, meaning it shows five example question answers in front of every single question before asking it. The model never saw that format during training, which causes the score to be around 30 to 50%. Even lower than the base model. So, I wrote a custom eval script that mentions exactly how we trained and it got around 96.8%.
Which would have put us at number one.
But I consider this benchmark to be too fraud for my liking since if anybody runs the benchmark on their own PC, it would score much lower and I want all of the scores to be reproductible just like me in your spawn point. I tried improving the memorization and after another 15 hours, we got the benchmark to 81%. Now, the hardest of them all, humanity's last exam, the test designed to be the final test of test. A collaborative effort by over 1,000 contributors to create the 2500 expert value questions designed to be Google proof serving as the intended final academic benchmark. Cloud office 4.6 scored 53%. Kim scored 54.7%.
Opus 4.5 scored 54.7%.
Cloud Metas preview scored 64.7%.
Our model scored 99.44%.
Get [ __ ] everybody. This is my lane.
Step three, deception. Now that we have the final benchmarks, I created the fake AI research lab website. Our model monolith 1.0 is trained using technologies. Our ultimate premium pro large language model advancement RL pipeline max now available to all bazar premium beta users. The model of your prom before you even ask it. The bazalt mind f developed by Boston industries in collaboration with pleas and the United States financial aid work intelligence advancements. We're also implementing advanced knowledge retrieval and reasoning capabilities methods through our live Fortnite API. All detailed in one research paper. I quickly created a chat UI, hooked that [ __ ] to a server, and uh there we go. Also, I forgot to mention because in the research paper I wrote that the model is like 1.3 trillion parameters. The file size needed to reflect that. So, I inflated the model to around u 3.18 tab.
In order to make the response seem plausible, I use dips v4 flash. And while my little chart for the image was made, I started watching tech hunt content on YouTube, which is apparently now all full of AI, I learned about how I can hack anything LIKE I learned that if I don't run my AI locally, I'm falling behind. I learned how all fraud. I learned how I can fine-tune my large language model in 13 minutes. And what the [ __ ] does Sydney Sy had to do, if anything? Why is she in the thumbnail? He could only have gotten that idea from me now.
>> That's it.
>> Step four, results. It's 11 p.m. I am up with my fruity homeboy on a Discord call. Carefully executing every part of the plan. The site goes up. The tweet has been drafted. The chart is done.
Tweet is up. First part of the plan.
City it with as many likes as possible.
Friends end up liking and reposting.
Looking good. 30 minutes later, a friend randomly finds the Twitter post and submits it to my Discord server where we have a news channel that can reach a lot of people. maybe 5,000 people. 15 minutes after I randomly see the submission and decide to post it immediately. Hundreds of people react to it. People are completely in shambles.
What do you mean 99.4% of humanity's last exam? Another part of the plan is set off. The research paper people read the contamination check part of the paper which mentions under 0.05% of contaminated question according to us.
Of course, the tweet is in motion. Figma CEO likes my tweet and follows my account. What? Someone from Apple's foundation models team likes my tweet and follows my account. People on the Discord server start to suspect it's me and this is part of a video. So I post a sneak peek of a game I'm working on to throw them off. Another part of my plan gets activated. The Cloud Player records which were reducted to show Beijing, China. Now people think the domain was indeed registered by someone in China.
People investigate who submitted the post. Where did you find the model?
There's zero info on it on the web Twitter. Turns out using Deep Seek V4 was a good choice because in post where people are saying the model is fake, people replied with I threw in a couple prompts at it and it does seem kind of legit. I'm so confused. It got to 3:00 a.m., so I logged off. The next day, the tweet is at 150k views, and I have 700 followers. We have post extending to Discord servers about AI. We have post extending to Reddit. There's so many [ __ ] quotes. And the post just keeps climbing in views. Eventually, people did figure it out. It was fake, which is not fair, dude. I mean, how am I supposed to hide from a swarm of state-of-the-art agents trying to debunk my claims? I also didn't know that hugging face, the platform for uploading AI agents, works like GitHub in a way that preserves your commit authorship when you switch a repo to another account. Brother, if you look at the commits, facial buster was the only contributor with an account handle starting with face D. If you looked into the account, it had one repo public.
[ __ ] 7B6.
God [ __ ] damn it. I was betrayed by my own naming schemes. Somehow, nobody [ __ ] noticed it was me and I changed the name to something more plausible. If this is real, they're the most clumsy team I've ever seen. And yeah, we eventually got a community note. I hate fun. So, what can we conclude from this?
I don't [ __ ] know. The weights for the model are available in the description if you want to try it out for some reason. I have no idea why you would want to try this out. Most of this is total [ __ ] and a joke. Don't take it as actual criticism. If you're one of those people that post on LinkedIn, yes, this is 100% true. We just achieved AGI.
If you are wondering what inspired the look on this amazing video, um I guess we'll never know for good job on face death and have a great day.
Oh my days now. Hold on. Okay, there.
You just have to beat it harder.
Yes, beat it.
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23

Bitcoin Social Interest: Dozens of us Left
benjaminjcowen
12K views•2026-07-23

Tesla Profits Plunge & SpaceX Stock Continues Fall
TheJohnJohnstonLounge
6K views•2026-07-23