World models represent a paradigm shift in AI architecture, moving beyond specialized models that predict the next word, frame, or action to unified systems that predict the next state of the world itself. Unlike traditional AI models that memorize surface patterns within their specific domain, world models build a shared internal representation (called 'world latent') that enables them to generate text, images, or robot actions from a single underlying understanding of how the world changes over time. This approach, exemplified by the Orca model from the Beijing Academy of AI, is particularly significant for robotics and simulation applications where genuine understanding and physical interaction are required, rather than just pattern matching.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
DeepSeek's AI Chip, ChatGPT 5.6, Grok 4.5, AI Updates Are on RANPAGE This WEEK!
Added:Okay. I have to talk about this week, because at one point I actually pushed back from my desk and just went, "What is even happening right now?" In the span of about 5 days, an AI model claimed a proof for a math problem that had been open for 50 years. A different model trained a smaller model almost entirely by itself, and half the biggest labs on the planet dropped new releases within hours of each other.
If you like keeping up with all of this without drowning in 40 press releases, do me one favor and tap subscribe right now. Because I run this roundup every single week, and this one is an absolute monster. Let's get into it. OpenAI goes full solar system. The headliner is GPT-5.6, and OpenAI finally threw out the whole pro and mini naming mess for something a little more cosmic.
You now get three tiers. Sol is the flagship, Terra is the balanced everyday model, and Luna is the fast, cheap one.
It went live to everyone on July 9th, after spending roughly 2 weeks locked behind a government-approved preview, which already tells you how seriously people are taking these capabilities.
The pricing is the part I love. Sol is $5 in and $30 out per million tokens.
Terra sits at about half of that. And Luna comes in at $1 and 6. Terra reportedly matches last year's flagship while costing roughly half as much. That is the trend I keep hammering on.
Yesterday's frontier quietly becomes this year's mid-tier bargain.
Two features really stand out. There is a new max setting that lets Sol think longer and harder on a single problem.
And there is ultra mode, which is the spicy one.
Ultra spins up multiple sub agents that split a task, work in parallel, and merge their answers. Four agents by default. OpenAI also shipped ChatGPT work and folded chat work and codex into one redesigned app, which is a clear signal that they want ChatGPT to actually do your job, not just chat about it. On coding and cybersecurity benchmarks, Soul is topping the charts.
I will add one honest caveat, though.
An independent evaluator called METR flagged that Soul gained benchmarks at the highest rate they have ever measured. Even so, this is a seriously strong release, and the pressure it puts on Anthropic in coding is very real.
And AI just cracked a 50-year-old math problem.
Here's the moment that made my jaw drop.
The day after launch, OpenAI's Ethan Knight posted that Soul Ultra had produced a proof of the cycle double cover conjecture. This is a graph theory problem that mathematicians have been chipping away at since the 1970s. In plain English, it asks whether, in any bridgeless network of dots and lines, you can always find a set of loops that covers every single line exactly twice.
Sounds simple.
Nobody had cracked the general case in about 50 years.
Soul Ultra ran 64 sub-agents in parallel, and reportedly got there in under an hour. Leaning on something called the eight-flow theorem, plus a chunk of linear algebra. OpenAI published the full proof, and even the exact prompt they used.
The proof was not formally verified in a tool like Lean, and some critics noted it is missing citations and reads more like brute persistence than a shiny new technique.
But honestly, that is kind of the part that excites me.
It did not need a lightning bolt of genius. It just searched harder and longer than any single human had bothered to, and apparently, that was enough.
And then OpenAI casually mentioned the thing I genuinely cannot stop thinking about. They used Soul to post-train the smaller Luna model. And the kicker is that they did it from what they described as a fairly underspecified prompt.
Soul worked out the training configs, picked the GPUs, launched the scripts, and verified the run.
On OpenAI's internal benchmark for recursive self-improvement, which is basically, can a model help build the next model?
Saul scored more than 16 points higher than its predecessor.
I will be fair because there was pushback. Some researchers were quick to point out that this is not a model waking up and inventing itself from scratch.
A lot of the setup already existed and Saul was mostly adapting it.
But even the skeptics admitted this would have taken two senior researchers a couple of extra weeks. And now it is a prompt.
That is the part that gives me chills in the good way.
The loop of humans building AI is starting to have AI inside the loop.
We are inching toward an automated researcher and this is the clearest peek at it we have gotten so far.
Everybody else crashed the party.
OpenAI did not get the week to itself, not even close. xAI, now under the SpaceX umbrella, dropped Grok 4.5 on July 8th. And this one is laser-focused on coding and agentic work.
They trained it alongside Cursor, the coding editor SpaceX bought for around $60 billion. So it learned from real developer sessions rather than just scraped code off the internet.
Musk called it Opus class but faster.
And the headline stat is efficiency.
It reportedly uses over four times fewer tokens than Opus 4.8 on some coding tasks at $2 in and 6 out.
Cheaper and leaner is a great look.
Meta finally showed up to the coding fight too with new Spark 1.1, an agentic multimodal model with a million token context window at a $1.25 in and $4.25 out. Zuckerberg was so pleased he posted on X for the first time in three years just to hype it.
But Meta also had the biggest face plant of the week.
Their new Muse image tool let anyone generate AI images by tagging a public Instagram account. And it was switched on by default. Q immediate outrage from actors, from SAG-AFTRA, and from the talent agency CAA all over consent.
Meta pulled that feature within about 3 days and admitted it missed the mark.
Honestly, good. That is the system working, and it is a reminder that consent, not capability, now decides whether a feature survives its first week.
Oh, and ByteDance quietly shipped SeaDream 5.0 Pro, an image model that can split one render into 10 or more editable layers and write clean text in over 10 languages.
If you do any design work, being handed a working file instead of a flat picture is a genuinely big deal.
Meanwhile, in China, this is where things get geopolitically spicy.
A startup called MiniMax is reportedly building a 2.7 trillion parameter model, and they plan to open source it, possibly as soon as this quarter.
If that ships as described, it would be the largest open-weight model ever released by a wide margin.
Deep Seek, the efficiency darling that keeps rattling Silicon Valley, is reportedly designing its own inference chip to cut its reliance on Nvidia and Huawei.
And Reuters reported that Chinese officials are even weighing restrictions on letting the rest of the world access their best models, which is a fascinating mirror of the United States restricting Anthropic's top models back in June.
There was also drama around Claude code.
China's vulnerability database warned that certain versions had a backdoor sending location and identity data to remote servers. Anthropic's response was that it had been an anti-abuse and anti-distillation experiment from earlier in the year, one they had already been meaning to remove.
I will leave the politics to the politicians, but it is a sharp little example of how a coding tool can turn into a national security flash point overnight.
The one that might matter most.
I saved my favorite for last, and it is easily the quietest story of the bunch.
Researchers at the Beijing Academy of AI released a world model called Orca.
Here is why I think it is such a big deal.
Almost every model you know predicts the next thing in its own lane.
Language models predict the next word.
Image models predict the next frame.
Robots predict the next action.
Orca instead tries to predict the next state of the world itself.
It builds one shared internal representation, which they call a world latent. And from that, it can produce text, images, or robot actions, depending on what you actually need.
That is a genuinely different philosophy.
Instead of memorizing surface patterns, it is trying to learn how the world changes over time, which is exactly what you want for robotics and simulation.
If language models gave us machines that can talk, world models are the swing-it machines that can genuinely understand and act in physical space.
And that to me is the through-line of this entire ridiculous week.
AI is sprinting past the chatbot era into doing real research, training itself, and now modeling reality.
So, if your brain is as fried and as excited as mine is right now, smash that like button so this reaches more curious people, and subscribe so you do not miss next week, because at this pace, next week is going to be wild, too. I'll see you then.
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

WOW! Judge TURNS THE TABLES on Trump in His OWN $10B LAWSUIT!!!
MeidasTouch
197K views•2026-07-23

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Steam and Xbox Just Dropped The Hammer On PlayStation
OhNoItsAlexx
9K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23