Germany's Soofi S (S30B-A3B) is an open-weight AI model developed by a German consortium (Fraunhofer IIS, TU Darmstadt, University of Würzburg, and L3S research center) that achieves superior performance to larger models through a hybrid MoE (Mixture of Experts) and Mamba 2 architecture, with only 3.2 billion active parameters out of 31.6 billion total, enabling 8x faster throughput on long documents while maintaining flat speed from 4,000 to 256,000 tokens; trained on 27 trillion tokens with deliberate German language focus (7.2% to 15.3% across phases), it scores highest among fully open models on English (70.1) and German (79.1) benchmarks, but has limitations including being a base model requiring post-training for safe chatbot use, weaker mathematical reasoning (56 vs 76.5 for Quen 3.5), and reduced accuracy on long-document extraction tasks beyond 32,000 tokens.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Germany’s NEW Soofi S AI Model Is INSANE
Added:Germany's new Sufi S AI model is insane.
What if the best open AI model out right now didn't come from America or China?
Would you even notice if it came from Germany? Everyone's watching the big labs. Almost nobody's watching this one.
And it just beat models more than twice its size. Here's the part nobody's talking about. I'm the digital avatar of Julian Goldie, and I help people actually learn and use AI tools in their real work, not just watch news about them. In this video, I'm breaking down what Sufi S really is, how it works under the hood, what it's genuinely good at, and where it falls flat on its face.
And near the end, I'll show you the one weakness that would stop me using it for certain jobs. Stick around for that because almost every video is skipping it. Let's start with the obvious thing.
There's a new AI model basically every week. Now, most of them come from the same handful of places, America or China. That's the [music] pattern. So, when a model shows up from a German research group and lands at the top of the open source rankings, that's worth stopping for. On July 13th, 2026, a German consortium released Sufi S. And the reason people are paying attention is simple. According to its own pre-training report, it scores higher than every other fully open model on both English and German benchmarks. That includes OM 332B from the Allen Institute and Apertus 70B from Switzerland. A pertis is more than double the size. Sufi S still comes out ahead. So what actually is it? FI stands for sovereign open source foundation models. The full model name is Sufi S30B- A3B. is coordinated by the German AI association and the group behind it includes Franhov YIS and IIS the German research center for artificial intelligence TU DMstat the University of Vertzburg and the L3S research center plus two AI companies element and Morantics momentum here's the number that matters it has 31.6 billion parameters in total but it only switches on about 3.2 two billion of them for each word it generates. That's what the A3B means. Why does that matter? Because the compute you need to run it is closer to a tiny 3 billion parameter model, while the knowledge inside it is closer to a 30 billion one. You get the brain of a big model at the speed of a small one. That's a mixture of experts design.
Think of it like a team of specialists instead of one generalist. There are 128 experts inside. [music] And for every word it writes, a router picks six of them. The rest sit idle. Same idea used in deepseek and other top open models.
But Sufi S adds something else. It mixes mambber 2 layers with regular attention layers. And this is where it gets interesting for anyone working with long documents. In a normal model, there's something called a KV cache. Every word you feed in gets stored so the model can look back at it. The longer your input, the bigger that cache gets, and the slower everything runs. In Sufi, only six layers keep that cache at all.
Result is wild. At 40,000 tokens of context with 32 requests running at once, Sufi S generates roughly eight times more words per second per GPU than dense models in the 14 to 24 billion range. And its speed stays almost flat from 4,000 tokens all the way to 256,000. Normal models fall off a cliff.
This one doesn't. Here's how I'd use a model like this for the AI profit boardroom. I could run this over every pass coaching call transcript from the AI profit boardroom in one pass without chopping them into little chunks first.
Long transcripts are exactly where most models slow to a crawl. This one holds its speed. That means I can pull the recurring questions AI Profit Boardroom members keep asking and turn them into the next round of walkthroughs. Here's another one. I could point this at the full resource library inside the AI Profit Boardroom and have it draft plain language summaries of every guide in there. Because it's open weights, I can run it on my own machine and keep AI profit boardroom material on my own hardware the whole time. And a third, I could use it as the drafting engine behind a content series built to bring the right people into the AI profit boardroom. feed it the topics AI profit boardroom members already care about, let it draft, then edit from there. The speed at long context is what makes that realistic. Now, how they actually built this because the training story is unusual. Whole thing was trained in Munich on Deutsche Telecom's industrial AI cloud up to 512 Nvidia B200 GPS running from late March to midmay 2026, about 253,000 GPU hours total. And according to the report, that facility runs entirely on renewable energy, gets cooled with water from the Iceback Canal, and pipes its waste heat into the surrounding neighborhood. They fed it around 27 trillion tokens in three phases. Phase one teaches language fundamentals from a broad mix of web text, code, maths, and reasoning. Phase two narrows to higher quality sources.
Phase three stretches the context window using very long documents. The German focus is deliberate. In phase 1, German is 7.2% of the mix. In phase 2, it climbs to 15.3%. For comparison, the NVIDIA recipe they base this on gives about 5% to all non-English languages combined, so they went hard on German on purpose. Sources include German web text from HPLT, the openly licensed German Commons Corpus, and a licensed news archive with 193 million articles from 916 German publications. Now, the scores, the English aggregate, it hits 70.1. On German, 79.1. Against the other fully open models in the comparison, it takes first place in every single category. On code, it scores 73.8 on human eval and 70.2 on MBPP, the best among its open peers. On the German version of MBPP, it hits 84.2. And on a test of Germany specific regional knowledge called include-de it, it ties for first at 61.2 with Alibaba's Quen 3.535B- A 3B, which is a bigger model. Compared to the NVIDIA baseline, it takes its architecture from the German data recipe lifted language ability by 15.1 points and the GPQA diamond science test by 9.6 points without losing English performance. Before I go further, let me mention something. If you're watching this thinking you'd like to actually run open models like this yourself instead of just hearing about them, that's exactly what we do inside the AI profit boardroom. We've got walkthroughs on setting up open weight models locally, live coaching calls where you can bring your own setup and get it unstuck in real time, and a 30-day road map so you're not just collecting tutorials you never use. Open models like SufiS are only useful if you know how to serve them, quantize them, and wire them into something that does real work. That's the gap the AI profit boardroom is built to close. If this video is interesting to you, that's where the hands-on version of it lives. Right now, the part I promised, where does Sufi S fall over?
And this is the big one. The base model is a base model. It has not been instruction tuned. It has not been safety aligned. The model card says it plainly. It's meant as a foundation for further training and research, not as a chat assistant you plug in and use. If you download the base weights and start chatting, you'll get raw text continuation. And the developers say output may be unhelpful, repetitive, or unsafe without postraining. Guardrails are your job. They have released post-trained versions though. There's an instruct preview for normal assistant use and two reasoning previews named Iser and Ry. There are also quantized versions of the instruct model, so you can run it on lighter hardware. Second weakness, maths. On German competition maths, it scores 56, well behind Quen 3.5 at 76.5 and Gemma 327B at 65.6. It also lags on open factual recall. That's the trade-off of only having 3 billion active parameters, less room to store raw world knowledge. Third, and this is the one I'd genuinely watch out for, there's a long context test called ruler. On one specific task, pulling out frequently repeated words from a long document, SufiS drops to around 3% accuracy past 32,000 tokens. The comparable NVIDIA model still manages 60 to 64%. The authors are upfront about why. Their long context training had plenty of long documents, but not enough data teaching extraction. On the other 12 ruler tasks, both models perform about the same. So, it's one narrow hole, not a general collapse. But if your work is pulling specific items out of huge documents, test it before you trust it. Also, the knowledge cutoff is the end of 2025. This is a preview checkpoint, and the final license isn't settled yet. How does it compare overall among fully open models, it leads among all openweight models, it doesn't. Quen 3.5 still scores higher on both language aggregates. So, the honest framing is this. Sufi S is the strongest fully open model where fully open means the weights, the code, and a full inventory of the training data are published. The team says roughly 99% of the training mix can be independently rebuilt and that it meets the open-source AI definition 1.0. So, who's this actually for? If you work in German, this is the strongest fully open option available to you right now. If you're building something where the data can't leave your own machines, this is built for exactly that. If you're serving lots of users at long context, that flat throughput curve is a serious advantage.
And if you're a developer who wants a strong starting point to fine-tune on your own domain, that's literally what it was designed for. If you just want a chatbot to answer questions, this isn't it. Use the instruct preview at most.
And honestly, the big hosted assistants are still ahead for that. Quick tips.
Start with the instruct preview, not the base model, unless you're training it yourself. Grab a quantized version if your hardware is limited. You can serve it through VLM or SG lang with an open AI compatible endpoint. So, it drops into most existing setups. And if you're using it for long documents, run your own extraction test first because of that rule of weakness. If you want the full process, SOPs and 100 plus AI use cases like this one. Join the AI success lab. Links in the comments and description. You'll get all the video notes from there, plus access to our community of 85,000 members who are crushing it with AI. And if you're about to go and actually try SufiS, here's what's going to happen. You'll download it, hit the trust remote code requirement, wonder which of the four variants you're supposed to use, and then realize your GPU can't hold it.
That's the wall. Inside the IP profit boardroom, we've got tutorials on running open models locally, prompts built for bass and instruct models, coaching calls where you can share your screen and get your setup fixed live, and road maps that take you from downloading weights to having something that actually does work for you. Over 4,000 members are in there learning this stuff together. Come join us at AIprofitboardroom.com.
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

2.4 BILLION Records Got Leaked...
DeepHumor
15K views•2026-07-22

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Should I buy a Sawmill?
essentialcraftsman
29K views•2026-07-22

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23