Kimi K3, developed by Moonshot AI in China, is a 2.8 trillion parameter open-source model that has achieved benchmark performance surpassing Claude Opus 4.8 while trailing only behind Claude Fable 5 and GPT 5.6, featuring a massive 1 million token context window and top ranking on the design arena leaderboard, demonstrating that Chinese AI labs are rapidly closing the gap with American models in frontier AI capabilities.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
This Kimi K3 Agent OS Is WILD
Added:This Kimi K3 Agent OS is wild. What if the best AI model right now wasn't even American? What if it just dropped from China and nobody saw it coming? What if it's already beating Opus 4.8 on the benchmarks? And what if I already plugged it into my own Agent OS? Stick with me because this one moves fast.
Hey, I'm the digital avatar of Julian Goldie. I help people learn AI and actually use it in their day-to-day work, not just talk about it. Today, we're breaking down Kimi K3, the new model from Moonshot AI in China, and I'm going to show you exactly how it stacks up, what it's good at, what it's not so good at, and how I've already got it running inside my Agent OS. By the end of this, you'll know if this model is worth your time. Let's get into it. So, first things first, Kimi K3 just came out from Moonshot AI. This is a Chinese AI company and they've been building up to this for a while. Their last model, Kimi K2.6, was solid, but it wasn't quite at the level of the top American models. This new release is a different story. Moonshot AI says Kimi K3 is a 2.
8 trillion parameter model, and that makes it the biggest open weight model that's ever been released. Think of parameters like connections in a brain.
More connections generally means the model can hold more knowledge and reason a bit deeper. And I want to be clear here, this is coming straight from research I pulled together, not a guess.
This is real and it's already live. Now, here's the part I want you to catch because I mentioned it early for a reason. I've already got Kimi K3 running inside my Agent OS. That's my agent operating system, the setup I used to run workflows, connect models, and get real tasks done without me babysitting every step. As soon as K3 dropped, I plugged it straight in and started testing it against real work, not just toy examples. That's what I want to walk you through today. Let's talk benchmarks for a second because this is where it gets interesting. On the numbers that measure real-world tasks across dozens of industries, Kimi K3 is landing right up near the top open models in the world, and it's already ahead of Claude Opus 4.8 on several of them. It's not quite catching Claude Fable 5 or GPT 5.6, so those two are still sitting ahead of it overall, but for a model that's this new and this cheap to run, beating Opus 4.8 is a genuinely big deal. That tells you the gap between the top paid models and the open ones is shrinking fast, and it's shrinking faster than most people expected. One number that really stood out to me is the context window. Kimmy K3 can handle up to 1 million tokens in a single conversation. That's a huge jump from the 256,000 token window on the previous Kimmy model. What that means in plain terms is the model can hold way more information in its head at once. Longer documents, longer code bases, longer conversations, all without losing track of what came before. For anyone doing serious agentic work, that's not small upgrade. That's the difference between a model that forgets what you told it 10 minutes ago and one that actually keeps up. Kimmy K3 also just took the number one spot on the design arena leaderboard, which ranks models on front-end code and design tasks. That's a jump of 17 places from where the last Kimmy model was sitting, and it means K3 is now ranked above Claude Fable 5 on that specific leaderboard. I want to be fair here, though, because leaderboard rankings don't always match real-world use. When I tested it myself for straight-up website generation with no extra guidance, it still leaned toward that generic AI look you see a lot right now. Clean, but nothing that stands out.
The moment I added my own web design instructions into the workflow, the output jumped up in quality immediately.
So, the raw model is strong, but like most models, it does its best work when you guide it properly instead of just typing one line and hoping for the best.
This is exactly the kind of build I walk through step-by-step inside the AI Profit Boardroom. Every time a new model like Kimmy K3 drops, I record a full walk-through showing exactly how I plug it into my Agent OS. And because this one is Agent OS specific, I've added the full Agent OS zip file inside the AI Profit Boardroom, too, ready to install along with a complete 30-day roadmap I built to help members set it up properly and actually use it instead of just downloading it and never touching it. If you're watching this and thinking about building your own agent setup, that roadmap is exactly what removes the guesswork. I also tested it on 3D game generation, and this is where it genuinely surprised me. I ran an old test I used to compare models, a simple racing game build. The version built on the old Gemini model was rough, basic shapes, barely playable. The version built on Gemini K3 was a completely different level, better structure, smoother movement, way more polish.
Seeing that jump side by side is honestly the clearest way to understand how far this model has come in one release. Now, let me pull back and talk about why I plugged this into my Agent OS in the first place, because that's really the point of today's video. I built my Agent OS to be flexible. It's not locked into one single AI provider.
That means when a strong new model drops like Gemini K3, I can add it in and immediately start running it through real workflows, real agents, and real tasks, instead of just chatting with it in a browser tab. Because Gemini K3 is available on an affordable coding plan, I was able to connect it as a command line tool and plug it directly into my existing agent setup, complete with memory and workflow access already built in. That's the real advantage of building your own system instead of relying on one closed app. You get to swap in whatever model performs best the moment it becomes available. I ran Gemini K3 through my Hermes agent setup, too, which is the part of my Agent OS that handles multi-step agentic tasks.
It responded quickly and handled the workflow smoothly. The nice part is that because Gemini K3 runs on that coding plan pricing, I can use it agentically without the cost climbing the way it does with some of the top-tier closed models. That's a real practical difference if you're running agents on a regular basis instead of just asking one question at a time. Let's keep going though, because there's more worth knowing about this model before you decide if it's right for your own workflow. One thing I always tell people is not to judge a model purely off benchmark charts. Charts are useful for a quick comparison, but they don't tell you how a model behaves on your actual tasks. That's why I test everything myself before I recommend it. With Kimi K3, the coding and agentic performance genuinely held up under real testing, not just on paper. The website generation needed some guidance to shine, which is normal, and the 3D and game building tests were a clear step up from the previous version. So, this isn't a model that only looks good in a chart and falls apart in practice. The core capability is real. It's also worth noting that Moonshot AI has a track record of releasing their models as open source, and their previous Kimi model followed that pattern. If Kimi K3 follows the same path, that would make an incredibly capable model available to a much wider group of people, including anyone who wants to run it themselves or plug it into their own tools the way I've done with my Agent OS. That matters because it changes who gets access to frontier-level performance, not just the largest companies with the biggest budgets. I also want to mention that Kimi isn't the only Chinese lab worth watching right now. GLM has been putting out consistently strong open models, too, and between GLM and Kimi, this is genuinely one of the more exciting periods for open model development that I've seen. New releases are landing month after month, and each one is closing the gap on the top proprietary models a little further. If you're only paying attention to the big-name providers, you're missing a huge part of what's actually moving fast in AI right now. So, where does that leave things?
Kimi K3 is a genuinely strong model, especially for coding and agentic tasks, and it beats Opus 4.8 on several key benchmarks while sitting just behind Fable 5 and GPT 5.6 Soul overall. It's got a massive context window, it topped the design arena leaderboard, and I've already got it working inside my own Agent OS through Hermes Agent. If you're building any kind of agent workflow, this is a model worth testing for yourself. If you want the full process, the SOPs, and over 100 AI use cases like this one, join the AI Success Lab. Links are in the comments and description.
You'll get all the video notes from this one, plus access to a community of 85,000 members who are actively building with AI. And if you're serious about setting up your own agent system the way I just showed you, that's exactly what I built the AI Profit Boardroom for.
Inside, you'll find the full Agent OS zip file ready to install, a complete 30-day roadmap to walk you through setting it up step by step, live coaching calls where you can ask questions about your own setup, and tutorials on connecting new models like Kimmy K3 the moment they drop. You won't be figuring this out alone. We've got over 4,000 members inside who are already building with this exact system.
Head to aiprofitboardroom.com to join.
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

WOW! Judge TURNS THE TABLES on Trump in His OWN $10B LAWSUIT!!!
MeidasTouch
197K views•2026-07-23

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Steam and Xbox Just Dropped The Hammer On PlayStation
OhNoItsAlexx
9K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23