In short
Alibaba’s Qwen3.8-Max open-weight flagship model—claimed to be the largest open-weight release in history—plus its capabilities, benchmarks, pricing, and safety considerations.
Guest backgrounds
No guests; hosted by Jon Krohn.
Key claims
Qwen3.8-Max is a 2.4T-parameter mixture-of-experts model with ~95B active parameters per token, supports text+image+video inputs, and a 1M-token context window. Alibaba positions it as second only to Anthropic’s Claude 5/Fable 5 family. It’s cheaper than competing frontier models and offers selectable reasoning effort (low/medium/extra high). Alibaba committed ~$50B over three years to AI/cloud infrastructure and promised open weights (Apache 2.0 precedent).
Notable examples
10+ days of unattended autonomous coding; reproducing an ML paper after 33 GPU-training rounds (~125 hours), 7,600 lines of code, 18 improvements; winning a live contest against 526 teams (87% of field). Benchmarks cited: Terminal Bench 86.6; AI Index score 56 vs Kimi K3 57; vision ranking second globally behind a Claude 5 variant. Safety: risk depends on data handling; safest is self-hosting open weights air-gapped, but model outputs may reflect Chinese regulatory alignment.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOContext of Qwen 3.8 Max
0:45 to 1:40
Exploration of the context surrounding the launch of Qwen 3.8 Max and its comparisons.
“And now Alibaba itself has stepped into the ring with its own frontier-scale release.”
Qwen 3.8 Max Specifications
1:40 to 3:30
Details the specifications and features of the Qwen 3.8 Max model.
“So what is the Quinn 3.8 max model and why is it such a big deal?”
Performance Insights
3:30 to 5:30
Discusses the performance of Qwen 3.8 Max in comparison to other AI models.
“On the popular crowdsourced arena platform, Quen 3.8 Max immediately became the highest-ranking Chinese model for text tasks.”
Use Cases and Demonstrations
5:30 to 6:50
Highlights demonstrations showcasing Qwen 3.8 Max's capabilities in coding and research.
“Multi-day, unattended, agentic work is the battleground these Frontier Labs are fighting on now.”
Pricing and Cost Comparisons
6:50 to 7:37
Breakdown of pricing for Qwen 3.8 Max and how it compares with competitors.
“And now their new developer platform is turning that workspace into infrastructure developers can build on.”
Safety and Compliance with Chinese Models
7:37 to 11:10
Discusses the safety and compliance considerations when using Chinese AI models.
“Now for the asterisk, and it's the same asterisk Kimi K3 carried at launch.”
Episode Recap and Listener Engagement
11:10 to 12:42
Recap of the episode and encouragement for listener feedback and engagement.
“That doesn't sound like a bad thing to me.”
Transcript
Automatic transcript. May contain errors.0:00Jon Krohn:This is episode number 1018 on Alibaba's QN 3.8 Max.
0:09Jon Krohn:Welcome back to the Super Data Science Podcast. I'm your host, Jon Krohn. Today's episode is on QN 3.8 Max, the enormous new flagship model from the Chinese tech giant Alibaba, a model that, if Alibaba keeps a promise it made at launch, is now the largest open-weight AI model release in history. If that setup sounds familiar, well, it should. Three weeks ago in episode number 1012, I covered Kimi K3, the nearly 3 trillion parameter model from Beijing-based Moonshot AI that rattled investors, kicked off a pricing skirmish among the big American AI labs, and reignited the open-source AI debate in Washington.
0:47Jon Krohn:Well, here's a fun fact. Moonshot is backed by Alibaba. And now Alibaba itself has stepped into the ring with its own frontier-scale release. The timing throughout has been pointed. Alibaba previewed QEM 3.8 Max at the World AI Conference in Shanghai on July 19th, two days after Moonshot released K3, and then shipped the full model on August 3rd. Alibaba's Hong Kong-listed shares rose 7 % on launch day, which tells you how much investor sentiment now rides on these releases. And this is all part of a much bigger bet. Alibaba has committed more than 380 billion won. That's about 50 billion US dollars to cloud and AI infrastructure over three years.
1:30Jon Krohn:It is dwarfed in comparison to what some of the big tech companies are doing in the US, but it is nevertheless a big sum on AI infrastructure for sure. So what is the Quinn 3.8 max model and why is it such a big deal? Well, it's a 2.4 trillion parameter model the most capable in the Quen family to date, and like Kimi K3, it's a mixture of experts architecture. This means that of those 2.4 trillion total parameters, only a small proportion, about 95 billion, are active for any given token. So the headline number describes the model's total capacity rather than the compute and the cost burned on every request.
2:13Jon Krohn:The model accepts texts, images, and video as input, returns text as output, and supports a 1 million token context window, enough to hold roughly three quarters of a million words in a single query, matching both K3 and the American flagships on that context window dimension. It's built on the architectural foundation of Quen 3.5, scaled way up. And one practical detail I appreciate, where K3's reasoning mode is always on and locked to maximum effort, a pricing gotcha I flagged back Back in episode number 1012, Quen 3.8 Max lets you select low, medium, or extra high reasoning effort, so you're not paying for a lengthy thinking trace when all you need is a quick lookup.
2:55Jon Krohn:Now, how good is it? Alibaba's own framing is that Quen 3.8 Max is second only to Anthropics Claude Fable 5, or Mythos 5, depending on what you have access to. And while that's a vendor self-assessment, the independent signals so far land in a similar neighborhood, which is jaw-dropping, and for Western foreign policy hawks, probably concerning. Let me repeat that. An open-weight model may lag behind only one model, Anthropics Fable 5, which is eye-wateringly expensive to use and was presumably even more eye-wateringly expensive to create. On the popular crowdsourced arena platform, Quen 3.8 Max immediately became the highest-ranking Chinese model for text tasks.
3:39Jon Krohn:That's tough to say. Text tasks. Though it still trails several anthropic offerings, including Fable 5. And on Vision Tasks, it ranked second globally behind only a Fable 5 variant. On the Artificial Analysis Intelligence Index, which is another widely reputed index for tracking AI capabilities, when 3.8 Max scores 56, for reference, Kimi K3 scored 57. So by that measure, these two Chinese giants are in a statistical dead heat near the top of the open weight pack. On specific benchmarks, there are wins to point to as well. Alibaba's published table has QEN 3.8 MAX at 86.6 on Terminal Bench, a benchmark of agentic command line tasks ahead of both Claude Opus 4.8 and Claude Fable 5 at 84.6, though behind OpenAI's GPT 5.6 SOL, working in its maximum effort mode at 88.8.
4:34Jon Krohn:That sounds expensive to run. With this model, QEN 3.8 Max, where Alibaba is pushing the hardest, however, is on long horizon autonomy. In one internal demonstration, the model spent more than 10 days autonomously coding a self-evolving software harness from scratch, incorporating user feedback, running its own tests, and iterating through code, previews and logs without human help. In another demo, it reproduced a machine learning research paper from zero, running 33 rounds of GPU training over roughly 125 hours, writing 7 ,600 lines of code, and then devising 18 improvement ideas that ended up outperforming the original paper's method.
5:16Jon Krohn:Entered into a live online contest against 526 human teams, it beat 87 % of the field. Those are vendor demonstrations, so hold them loosely until third parties reproduce similar kinds of results, but the direction is unmistakable. Multi-day, unattended, agentic work is the battleground these Frontier Labs are fighting on now. And yeah, then there's the pricing story. Let me dig into that a bit more now. Quen 3.8 Max costs$2 per million input tokens,$6 per million output tokens, and$0.25 per million cached input tokens. recall from episode number 1012 that k3 charges three dollars in and fifteen dollars out so quen 3.8 max undercuts its own domestic rival by quite a bit including by more than half on output tokens against the american labs the gap is much wider still that eight dollar combined rate you know adding up input and output is less than a third of claude opus's combined pricing forget fable and under a quarter of GPT 5.6 souls.
6:18Jon Krohn:And because cached input is eight times cheaper than fresh input, agentic and retrieval heavy workloads with stable system prompts collapse toward that 25 cent per million token floor. Wow. The price war I described three weeks ago has not cooled. It has escalated. Agents are getting smarter every day, but even the smartest agents get stuck without the right context and the right tools. That's where Notion comes in. With the recent launch of custom agents, Notion became the collaborative AI workspace where teams and agents work side by side. And now their new developer platform is turning that workspace into infrastructure developers can build on.
6:57The piece I keep coming back to is how easy it is to ship something real. The CLI authenticates in one line, workers deploy without provisioning any infrastructure, you write your code, deploy, and you're done. For me, that unlocks building purpose-built tools for my custom agents with the predictability and custom logic I need. Think a guest prep agent that pulls a researcher's papers, recent talks, and citation graph on demand. Tools my agents can actually call with parallelism and predictable behavior, not just hope for. Learn more about Notion's developer platform today at notion.com slash superdata.
7:31That's all lowercase letters, notion.com slash super data to try Notion's developer platform today. And when you use our link, you're supporting our show, notion.com slash super data.
7:42Jon Krohn:Now for the asterisk, and it's the same asterisk Kimi K3 carried at launch. At the time of me recording what I'm saying to you right now, which is a few days before this episode is released, the open waits for the model are promised, but not yet shipped. Alibaba has committed to publishing weights for QEN 3.8 Max, the first Max class QEN model ever to go open. During the week, I am recording this on both Hugging Face and ModelScope alongside a smaller QEN 3.8 27B that will fit on a single high-end GPU for those of us without a data center handy. Hopefully by the time of publication, you can access those models weights as you are expected to be able to do.
8:24Jon Krohn:Recent QEN open releases shipped under the permissive Apache 2.0 license, which sets an encouraging precedent, and hopefully the QEN 3.8 max weights get the same permissive treatment. Now, all of this brings me to a question I get asked a lot and want to spend the last stretch of this episode on. Are these Chinese models safe to use? Well, the risk depends far less on the model and far more on how your data reach it. At one end of the spectrum are the consumer services, so chatbot apps and workplace platforms like Alibaba's new QEN work. Whatever you type into those is collected by the provider and stored on infrastructure governed by Chinese law.
9:02Jon Krohn:So I'd advise against putting anything sensitive, proprietary, or personal into those. One step down in risk is the hosted API on Alibaba Cloud. You get contractual terms and enterprise controls, but your data still transit a Chinese company's server. So check your compliance obligations before routing customer information through it. And again, I don't know if I would put sensitive proprietary or personal info through those. Safer again is using a Western cloud provider like Lightning AI, where I hold the fellowship, or use a model gateway that hosts the open weights on infrastructure in your own jurisdiction.
9:36Jon Krohn:Your prompts never touch Chinese servers at all in that circumstance. But at the safest end of all sits the option that open weights uniquely unlock, downloading the model and running it entirely on your own hardware. Weights are inert files of numbers. They can't phone home, and you can run them fully air-gapped if you wish, with no per-token fees and no data leaving your walls. Even then, two caveats apply. First, self-hosting doesn't change what's inside the model. Its outputs reflect its training, including alignment with Chinese content regulations on politically sensitive topics, so evaluate it on your own use cases before trusting it.
10:15Jon Krohn:Second, practice basic supply chain hygiene. download from the official repository, prefer the safe tensors format and verify checksums, and note that some governments and regulated industries restrict Chinese origin models regardless of where they're deployed. So check the rules that apply to you. Zooming out, the pattern from episode number 1012 three weeks ago on Kimi K3 has now repeated within a month, a Chinese model that is nipping at the heels of American frontier labs more than ever before. Another aggressive price point, another open weights pledge. If Alibaba delivered those weights this week, like it was expected to, and you can access them now, the largest open model in history is sitting on Hugging Face for anyone to download, inspect, fine tune, and deploy on your own terms.
11:03Jon Krohn:Whatever your view on the geopolitics, for those of us building AI applications, the cost of experimenting at the frontier keeps falling, and the control we've retained over our own stacks keeps rising. That doesn't sound like a bad thing to me. And finally, as I sometimes do on Friday episodes or at the end of Friday episodes, it's time for a recent Apple podcast review. Somebody named Civil Service IT said that the show is their go-to favorite. They gave it a five-star rating and said that they've been listening to this podcast for a year or so. they're not an engineer, but they learn so much exclamation mark.
11:43Jon Krohn:That's cool. I'm glad to hear it. I actually based on some LinkedIn correspondence. I think I know who you are listener. So thanks to you. And thanks to everyone for all the recent ratings and feedback on Apple podcasts, Spotify, and all the other podcasting platforms out there, as well as for your likes and comments on our YouTube videos. If you can do that for us, if you can provide ratings, wherever you listen to your podcasts, please do that. It is the most helpful thing that you can do for me as a listener. And bonus points, if you leave written feedback, if you do that, I'll be sure to read your feedback on air like I did today.
12:16Jon Krohn:That seems to be something that's mostly limited to Apple podcasts. And I think I only see feedback from people that write in the US, but at some point I've been, it's on my to-do list to go check other key countries amongst or listenership and read some reviews from there as well. All right, that's the end of today's episode. If you enjoyed it or know someone who might consider sharing this episode with them, tag me in a LinkedIn post with your thoughts. And if you aren't already, be sure to subscribe to the show. Most importantly, however, I just hope you'll keep on listening. Until next time, keep on rocking it out there.
12:50Jon Krohn:And I'm looking forward to enjoying another round of the Super Data Science Podcast with you very soon.
13:01You
From the publisher
In Episode #1018, Jon Krohn breaks down Qwen3.8-Max, Alibaba’s enormous new flagship, a 2.4-trillion-parameter mixture-of-experts model that, if its promised weights ship, becomes the largest open-weight release in history. Landing just weeks after Moonshot’s Kimi K3, it extends the price war and the open-weight surge Jon covered in Episode #1012. Alibaba positions it as second only to Anthropic’s Claude Fable 5 / Mythos 5 and independent signals land in a similar neighborhood. Jon walks through its capabilities and multi-day agentic demos, its aggressive pricing ($2 in / $6 out per million tokens, with cached input eight times cheaper), and the question he gets asked most: are Chinese models safe to use? His answer hinges far less on the model than on how your data reach it.
Additional materials: www.superdatascience.com/1018
Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.




