In short
Moonshot AI’s Kimi K3, a 2.8T-parameter “open-weight” (weights promised under modified MIT license by July 27) Mixture-of-Experts frontier model, and its impact on pricing, open-source policy, and U.S. labs.
Guest backgrounds
No guests are discussed in the transcript (only host Jon Krohn; mentions Wharton professor Ethan Mollick as a prior guest).
Key claims
K3 activates 16 of 896 experts per token, has a 1M-token context window, native vision, and always-on reasoning. Moonshot says K3 improves scaling efficiency ~2.5x vs K2 via Kimi Delta Attention. Analysts argue chip export controls are “leaky.” Pricing undercuts Claude Opus/GPT 5.6, but always-on reasoning can increase billed output tokens.
Notable examples
Arena ranks K3 top for front-end coding; Bank of America notes step-change gains despite compute constraints; OpenAI/Anthropic raised token limits; companies like Cursor and DoorDash reportedly use earlier Kimi models; Anthropic accuses Chinese firms of distillation (Moonshot/DeepSeek/Minimax).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Company Behind Kimi K3: Moonshot AI
0:41 to 1:44
Discover the background and financial momentum of Moonshot AI.
“Moonshot AI is a Beijing-based startup backed by the Chinese tech giant Alibaba.”
Kimi K3's Features and Architectural Innovations
1:44 to 3:00
Explore the features of Kimi K3 and its groundbreaking architecture.
“So what exactly is KimiK3 released in mid-July?”
Kimi K3's Performance Compared to Competitors
3:00 to 4:53
Learn how Kimi K3 stacks up against other major AI models in the market.
“in scaling efficiency over the previous K2 generation of Kimi.”
Kimi K3 Pricing and Cost Comparisons
6:01 to 7:22
Understand Kimi K3's pricing structure and its competitive edge.
“And when you use our link, you're supporting our show, notion.com slash superdata.”
Impact of Kimi K3 on the AI Market Landscape
7:22 to 10:46
Discuss the implications of Kimi K3's release for the AI industry.
“example, Z.ai's GLM 5.2 costs at less than a third of the cost per million output tokens.”
Transcript
Automatic transcript. May contain errors.0:00This is episode number 1012 on Kimi K3.
0:08Welcome back to the Super Data Science Podcast. I'm your host, Jon Krohn. Today's episode is all about Kimi K3, a new model out of China that in the space of a single week has managed to rattle investors, kick off a pricing skirmish among the big American AI labs, and reignite a debate in Washington, D.C. about open source AI. Whether you're a hands-on ML practitioner or you're more focused on the commercial side of AI, this is a release you'll want to understand. So we're going to dig into it today. Here we go. Let's start with the company behind it. Moonshot AI is a Beijing-based startup backed by the Chinese tech giant Alibaba.
0:49And if the name Kimi rings a bell, that might be because the company's original Kimi chatbot released back in 2023, three years ago, was the first model capable of accepting a context window of 128 ,000 tokens, which was a big deal at that time that made a big splash. Moonshot's co-founder and CEO, Yang Jilin, earned his doctorate at Carnegie Mellon in 2019, and the company has had serious financial momentum. Moonshot raised$2 billion in May at a valuation north of$20 billion, with annual recurring revenue reportedly exceeding$200 million. What makes the K3 release a particularly compelling story is that it's something of a comeback.
1:32Moonshot's market position has eroded significantly over the past 18 months following DeepSeek's meteoric rise, and now the student of that disruption has become the disruptor. So what exactly is KimiK3 released in mid-July? It's a 2.8 trillion parameter model that Moonshot says is the largest open source AI model in the world. Now, before you panic about inference costs with such an enormous model, it's key to know that this is a mixture of experts architecture. You can learn more about MOE models in episode 939 of this podcast if you'd like to. But the key thing here is that of KimiK3's 896 expert submodules within their model, merely 16 of those nearly 900 are activated per token.
2:20So that headline parameter count describes total capacity, not the compute running on every single request. The model features a 1 million token context window, native visual understanding, an always-on reasoning mode, and it's built on two architectural innovations developed at Moonshot. The first one is Kimi Delta Attention, which is a hybrid linear attention mechanism and attention residuals, which is a drop-in replacement for the residual connections that first became famous years ago with a model called ResNet. Not sure if you'll remember that one. Anyway, together, those two innovations reportedly by roughly a 2.5x improvement in scaling efficiency over the previous K2 generation of Kimi.
3:04That architectural angle matters beyond this one model. Bank of America analysts noted that despite persistent compute constraints in China, K3 demonstrates that pre-training scale paired with architectural innovations can still deliver step change gains. In other words, U.S. chip export controls are once again proving to be a leaky dam. Now, how good is Kimi K3? To Moonshot's credit, they've actually been relatively measured in their own claims. The company itself says K3 still trails Anthropics Claude Fable 5 and OpenAI's GPT 5.6 Sol on overall performance, but that it beat Claude Opus 4.8 and GPT 5.5, the model sitting just behind Anthropics and OpenAI's respective flagship models, on benchmarks including coding and general agentic tasks.
3:53Independent signals so far are encouraging. For example, K3 scores 57 on the artificial analysis intelligence tracker, placing it well above the median of 31 for reasoning models in a comparable price tier. And the evaluation platform Arena ranked K3 at the top for front end coding ability, with Arena's CEO calling Kimi K3 possibly the single biggest release of the year and a point where open source Chinese models are surpassing closed U.S. models, at least in some respects.
4:27Jon Krohn:That said, a few caveats do deserve some airtime here. On independently verified coding benchmarks, Claude Opus 4.8 still leads the active frontier, and Moonshot's own agentic scores haven't yet been reproduced by third parties. Separately, Wharton professor Ethan Mollick, who has been on this podcast, you can check that out, But he's argued that Kimmy is a strong but uneven model rather than another deep seek scale breakthrough. So Kimmy K3 is a top tier model, but certainly not the top one. Which brings us to the part of the story with real commercial teeth. Pricing. Agents are getting smarter every day, but even the smartest agents get stuck without the right context and the right tools.
5:11Jon Krohn:That's where Notion comes in. With the recent launch of custom agents, Notion became the collaborative AI workspace where teams and agents work side by side. And now their new developer platform is turning that workspace into infrastructure developers can build on. What sets Notion apart is that the collaborative workspace and the platform you build on are the same thing, with permissions, context, and governance baked in from day one. Workers, for instance, are Notion-hosted sandboxes where I can run database syncs without standing up my own infrastructure. That means I can pull guest research, episode analytics, and my consulting client data from all the scattered systems where they live, then keep them synced in Notion databases automatically.
5:47Jon Krohn:My agents and my human team work from one single source of truth. Learn more about Notion's developer platform today at notion.com slash superdata. That's all lowercase letters, notion.com slash superdata to try Notion's developer platform today. And when you use our link, you're supporting our show, notion.com slash superdata. moonshot lists k3 at three dollars per million uncashed input tokens 15 per million output tokens and just 30 cents per million cash hit input tokens for context on that last figure the caching if you're running agents or things like retrieval augmented workflows where the system where the same system prompts and documents get sent over and over most of your input tokens are going to be cashed.
6:33And so your effective input rate collapses toward that 30 cent per million token floor. Those rates undercut both Claude Opus 4.8 at$5 and$25 and GPT 5.6 Sol at$5 and$30 for input and output tokens respectively. While K3 actually matches despite being more cost effective, all of those models, Claude Opus 4.8, GPT 5.6 Sol, they all have a 1 million token
7:00Jon Krohn:context window. And so yeah, it seems like you're getting a model that is, you know, cheaper than the next to frontier models from the really big Western frontier labs with the same kinds of capabilities at a cheaper price. But interestingly, K3 is expensive by Chinese standards. So for example, Z.ai's GLM 5.2 costs at less than a third of the cost per million output tokens. And DeepSeek v4 is 94 % cheaper per million output tokens relative to KimiCurray 3. So that's interesting. One practical gotcha if you're evaluating it, K3 always reasons. So you could have lots of tokens being consumed behind the scenes without anything being pumped out because reasoning effort is currently locked to maximum with Kimi K3 as well.
7:53And all those thinking tokens are billed as output at$15 per million. So a chatty reasoning trace that you can't even see for the most part can cost more than the visible answer. Fantastic for hard problems, wasteful for simple ones and simple lookups, things like that. All right, let's talk about the open source dimension with one asterisk. So the API for Kimi K3 went live in mid-July, but the full model weights are promised under a modified MIT license by July 27th, which is coming up soon now. And until those files actually appear, K3 isn't deployable or isn't downloadable or locally deployable.
8:33You can't really call it open source right now at all, but it seems like they're gonna pull through on that. And assuming Moonshot delivers that permissive license means enterprises will be able to run frontier class AI entirely on their own infrastructure with no per token fees and no data leaving their walls. So what does all this together mean for Western Labs? Well, the reaction has been swift. OpenAI and Anthropic have responded by increasing token allowances and relaxing usage limits to retain users, while Anthropic has expanded access to CloudFable 5 and raised weekly token caps. That's all good news for us.
9:08Keep the competition coming. There's an analyst named Patrick Moorhead who described the market response as an overreaction shockingly similar to the DeepSeq launch 18 months ago, while acknowledging K3 could pose revenue challenges for OpenAI and Anthropic long-term. You know, this is an ongoing trend. It's not just Kimi K3 on its own. Chinese models were already seeping into Western production stacks before K3 arrived. So Cursor, for example, used Kimi to build its, not Kimi 3, but an earlier version, to build its Composer 2 coding agent, DoorDash's CTO says the company delegates lower-level work to Kimi K2.6.
9:46And Thinking Machines, which was founded by Mira Murati, who used to be an exec at OpenAI, they tapped Kimi K2.5 to generate early post-training data for its own open model. There is friction here too, though. OpenAI and Anthropic have accused several Chinese firms of using distillation to extract capabilities from their models. In February, Anthropic specifically accused DeepSeek, Moonshot, who made Kimi, and Minimax of attempting to illicitly extract Claude's capabilities allegations, Beijing has rejected as groundless. However that dispute resolves, the strategic pressure is the same one DeepSeek introduced, now aimed at a more vulnerable spot though, the price businesses pay for advanced AI and the control they surrender to US providers.
10:30The combination of lower cost, strong performance, and customer control threatens to turn frontier AI from a tightly controlled premium service into a competitive, low-price commodity. Hmm, not too bad for you and me, eh? Yeah, my take for you is that the frontier is now contested to some extent, open and cheap enough that the cost of experimenting with world-class AI has never been lower. And every price war between labs is a subsidy for the applications you and I are building. Hopefully we really will get to access the Kimi K3 model weights later this month, But either way, open-weight models are nipping closely at the heels of the proprietary frontier models and show no sign of letting up.
11:15This is surely good news for us, whether we want cheap inference or fine-tuning of powerful, powerful models for our own use cases,
11:22Jon Krohn:or the kinds of privacy benefits that come from having everything run on our own infrastructure. All right, that's it for the end of today's episode. If you enjoyed it or you know someone who might consider sharing this episode with them, leave a review of the show on your favorite podcasting platform. If you write an Apple podcasts review, that's especially helpful to us. I'll actually read it on air when you do that. Yeah, you can comment on our YouTube videos as well, or tag me in a LinkedIn post with your thoughts. And I will respond to those for sure. And if you aren't already, be sure to subscribe to the show.
11:55Most importantly, though, above all, we hope you'll just keep on listening until next time. Keep on rocking it out there. And I'm looking forward to enjoying another round of the Super Data Science Podcast with you very soon.
From the publisher
What happens to the AI market when the largest open-source model in the world arrives at a fraction of frontier prices? In this week’s episode, host Jon Krohn digs into Kimi K3, the 2.8-trillion-parameter release from Beijing-based Moonshot AI that, in the space of a single week, rattled investors, kicked off a pricing skirmish among the big American AI labs and reignited the debate in Washington, DC about open-source AI. Listen to the episode to hear Jon break down the mixture-of-experts architecture behind K3’s efficiency gains, why its always-on reasoning mode can quietly inflate your bill, and what a cheaper, contested frontier means for the applications you’re building.
Additional materials: www.superdatascience.com/1012
Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.




