990: Inside Mythos: Anthropic's Locked-Down Frontier Model

8 May 2026 · 11 min · 6 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Mythos, Anthropic’s gated “frontier” model (Claude Mythos Preview), claimed to discover far more real-world software vulnerabilities than prior models, despite not being trained specifically for hacking.

Guest backgrounds

No guests named; the episode is hosted by Jon Krohn and references external figures (Dario Amodei, John Dickerson, Bruce Schneier, Pete Hegseth) and organizations.

Key claims

Mythos reaches 94% on Sweebench vs ~80% for top public models; 181 working shell exploits vs 2 for Opus 4.6 on 147 Firefox vulnerabilities; tier-5 control-flow hijack on 10 patched OSS Fuzz targets vs 1 for Sonnet/Opus.

Notable examples

Mozilla patched 271 Firefox vulnerabilities in one release with early Mythos; contractors agreed with Mythos severity 89% exactly and 98% within one level. Rollout: Project Glasswing consortium (Apple, Google, Microsoft, AWS, NVIDIA, etc.), $100M usage credits, $4M donations; priced at $25/M input tokens and $125/M output tokens.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Background on GPT-2's Release and Dario Amadei

0:45 to 2:28

Discussion of GPT-2's controversial release and Dario Amadei's role.

“especially if it's code that you're pushing using a code gen tool like Codex from OpenAI or Cloud Code from Anthropik themselves.”

Overview of Mythos and Its Capabilities

2:28 to 4:02

Detailed explanation of Mythos Preview's technical features and accuracy compared to previous models.

“According to Anthropic, Mythos Preview hits 94%.”

Mythos Achieves Unprecedented Success in Vulnerability Detection

4:02 to 4:44

Mythos shows a significant leap in detecting vulnerabilities compared to Opus 4.6.

“These 100x and 10x deltas on cybersecurity vulnerability discovery, that's consistent with a generational leap, not an incremental improvement.”

Anthropic's Rollout Strategy and Project Glasswing

4:44 to 7:16

Exploration of Anthropic's marketing strategy and the implications of Project Glasswing.

“With this setup, Mythos identified thousands of zero-day vulnerabilities across every major operating system and every major web browser.”

Geopolitical Implications and Security Concerns

7:16 to 8:18

Discussion of the geopolitical implications of AI in cybersecurity and potential concerns.

“made for great marketing indeed, OpenAI not to miss out on a media splash followed suit days later with a similar too powerful to be released announcement regarding a version of GPT 5.4.”

Implications for Software and AI Development

8:18 to 9:33

Advice for software developers on adapting to new AI capabilities in vulnerability discovery.

“A recent Epoch AI analysis estimates the average capability lag between proprietary and open-weight frontier models at just three months, though it can extend between 5 and 22 months on certain benchmarks.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Jon Krohn:This is episode number 990 on Mythos.

0:07Welcome back to the Super Data Science Podcast. I'm your host, Jon Krohn. Today's topic is Anthropik's Mythos, a frontier AI model so capable at finding software vulnerabilities that Anthropik has decided not to release it to the general public. We'll talk about all the reasons why they might have done that later on in the episode, but that is their main claim. Of course, it's been a few weeks since Mythos came out, but now that the dust has settled, we can do some detailed reporting on it. And I also have lots of tips at the end of the episode on what you can do to avoid having lots of security vulnerabilities in all of that code that you're pushing, especially if it's code that you're pushing using a code gen tool like Codex from OpenAI or Cloud Code from Anthropik themselves.

0:55Anyway, let's cast our minds back now to 2019. OpenAI had just finished training a then new large language model called GPT-2, and in a move that drew a fair amount of mockery at the time, the lab declared that it was too dangerous to release. The research director at OpenAI back then was Dario Amadei, and Dario insisted that the world needed time to prepare. GPT-2 ended up getting released later that year anyway, and a sequence of far more powerful models has, of course, been deployed since, and, you know, an Armageddon has not been unleashed. Now, seven years on, Dario, these days the CEO of OpenAI's biggest rival, Anthropic, is sounding the alarm again.

1:41On April 7th of this year, he declared that the new addition to Anthropic's Claude family, codenamed Mythos, is too powerful to be made widely available at this time. And this time, having looked at the technical evidence, well, maybe he's right. Let's start with what Mythos actually is, because the technical story here is striking. Claude Mythos Preview, to provide its

2:01Jon Krohn:full name, is a general purpose frontier model, not a specialized cybersecurity tool. Anthropic is explicit that they did not train it to find software vulnerabilities. The model's hacking abilities emerged as a downstream consequence of broad improvements in code understanding, reasoning, and agentic autonomy. On Sweebench, for example, a popular benchmark of real-world software engineering problems, the top-performing publicly available models are pushing up against an 80 % accuracy. According to Anthropic, Mythos Preview hits 94%. That's a big jump from 80. Essentially, half of the problems that couldn't be solved previously are now solved.

2:40Jon Krohn:It's on the cybersecurity numbers, however, where things get really spicy. Here's a concrete comparison. Anthropik's previous frontier model, Opus 4.6, was already capable. I was recently at a conference at Tulane University in New Orleans where I met John Dickerson, the CEO of Mozilla.ai, with Mozilla, of course, being the makers of the open source Firefox browser that is famously secure. When researchers ran Opus 4.6 against a set of 147 Firefox vulnerabilities, it produced working shell exploits two times across several hundred attempts. They then ran the same test with Mythos Preview. Mythos succeeded 181 times.

3:19Jon Krohn:That's nearly a 100x increase relative to Opus 4.6. As another cybersecurity example, take Anthropik's internal benchmark of around 7 ,000 entry points across open source repositories from the OSS Fuzz corpus. sonnet 4.6 and opus 4.6 each reached what's called tier 5 meaning complete control flow hijack wherein the attacker takes over what code the program executes next so sonnet 4.6 and opus 4.6 reached that really dangerous tier exactly once mythos achieved that same tier 5 kind of attack on 10 separate fully patched targets. That is a 10x difference. These 100x and 10x deltas on cybersecurity vulnerability discovery, that's consistent with a generational leap, not an incremental improvement.

4:16Jon Krohn:In the few weeks before the announcement of Mythos, Anthropic ran Mythos against real-world critical software using a remarkably simple agentic scaffold. An isolated container with a source code, clawed code wrapping the model, and a prompt instructing it to find a security vulnerability. The model reads the code, hypothesizes vulnerabilities, runs the project to confirm them, attaches debuggers as needed, and produces either a negative result or a working proof of concept. With this setup, Mythos identified thousands of zero-day vulnerabilities across every major operating system and every major web browser.

4:52Jon Krohn:Mozilla, for example, working with an early version of Mythos on Firefox, patched 271 vulnerabilities in a single software release. For context, all of 2025 saw Mozilla address just 73 high-severity Firefox vulnerabilities. So that's about four times the annual figure in a single AI-driven sweep. To go one level deeper on confidence in these findings, of 198 findings that Anthropics had manually reviewed by professional security contractors, 89 % of those received the exact same severity rating from the contractors as the model had self-assigned, and 98 % were within one severity level. This means that this isn't a case of a model hallucinating bugs.

5:37Jon Krohn:The contractors agreed with Mythos' risk assessments at near-human expert reliability. So that's the technical picture. Now let's talk about the rollout because the marketing strategy is just as interesting as the model. Rather than launching Mythos publicly, Anthropic launched something called Project Glasswing, an industry consortium with launch partners that include Apple, Google, Microsoft, AWS, NVIDIA, JP Morgan Chase, the Linux Foundation, Cisco, CrowdStrike, Broadcom, and Palo Alto Networks. So yeah, a lot of big names in this big consortium. The idea is that these defenders get to use Mythos Preview to scan and harden their code bases before the capability proliferates.

6:15Jon Krohn:Anthropic has committed up to$100 million in usage credits across the program, plus$4 million, in direct donations to open source security organizations. But there's a commercial logic here too, isn't there? Mythos Preview is priced at$25 per million input tokens and$125 per million output tokens. That's roughly five times the price of Opus 4.6, suggesting Mythos is genuinely far more compute-hungry and Anthropic has been rationing capacity on Claude already. Exclusivity also makes Mythos harder to distill. That is harder for rival labs. to use Mythos' outputs to train cheaper imitator models.

6:57Jon Krohn:And keeping the model gated nudges the enterprise customers toward Anthropic-native tooling like Claude Code instead of model-agnostic products. So, while the safety framing is real and well-justified, this rollout is also great business, and the media sensation around a model too powerful to be released, made for great marketing indeed, OpenAI not to miss out on a media splash followed suit days later with a similar too powerful to be released announcement regarding a version of GPT 5.4. Now, not everyone is convinced this is a watershed moment, mind you. Bruce Schneier, one of the most prominent figures in computer security, he's been a public voice in the field for roughly 30 years and is widely respected, has noted that the security firm Aisle was reportedly able to replicate some of Anthropik's findings using older, cheaper public models.

7:49Jon Krohn:The point being, finding vulnerabilities and weaponizing them are different things, and current AI may help defenders more than attackers for now. There's also a geopolitical dimension to all this. Project Glasswing could effectively neutralize zero-day vulnerabilities that the U.S. government has historically hoarded for offensive cyber operations. Defense Secretary Pete Hegseth, who labeled Anthropic a supply chain risk earlier this year, is unlikely to be thrilled by that. And open-weight labs, particularly those in China, will likely produce comparable capabilities within months. A recent Epoch AI analysis estimates the average capability lag between proprietary and open-weight frontier models at just three months, though it can extend between 5 and 22 months on certain benchmarks.

8:31Jon Krohn:So the defender headstart that Glasswing buys is real, but it's not unlimited. So here's what I'd take away as an AI practitioner from all this. We are entering an era where automated vulnerability discovery is no longer bottlenecked by scarce human expertise, and where the gap between finding a bug and writing a working exploit has collapsed from months to minutes. If you build, deploy, or maintain any kind of software at scale, your patching pipeline now needs to operate at machine speed. The defenders who modernize their tooling first, automated scanning, AI-assisted code review, runtime behavioral enforcement, those are the ones who will come out ahead.

9:09And if you're working on AI capabilities themselves, Mythos is a remarkable illustration of how dangerous capabilities can emerge as side effects of general improvements rather than from targeted training. Now, if you're one of the many listeners generating tons of code through gen ai tools like claude code or codex you're going to want to be extra careful because you probably are creating applications with way more vulnerabilities than time and expert hardened software like mozilla's firefox we've talked about tools like code rabbit on the show before back in episode number 927 for automating the review of pull requests you can also use claude code or codex themselves to run a dedicated security step after generating your code for you.

9:51And there is a big industry

9:52Jon Krohn:of AI-native security tools being specifically built for the AI-native era, including Socket, Endor Labs, and Semgrep. I've got links to all of those for you in the show notes. All right, that's the end of today's episode. If you enjoyed it or know someone who might, consider sharing this episode with them. Leave a review of the show on your favorite podcasting platform or YouTube. Tag me in a LinkedIn post with your thoughts. And if you aren't already, be sure to subscribe to the show. most importantly however we hope you'll just keep on listening until next time keep on rocking it out there and i'm looking forward to enjoying another round of the super data science podcast with you very soon

From the publisher

Anthropic has built a frontier AI model so capable at finding software vulnerabilities that it has decided not to release it to the general public. In this Five-Minute Friday, Jon Krohn breaks down Claude Mythos Preview, a general-purpose model whose hacking abilities emerged as a side effect of broad improvements in code understanding and reasoning. Find out how Mythos achieved a nearly 100x improvement over Opus 4.6 on Firefox exploit generation, why Mozilla patched 271 vulnerabilities in a single release using an early version of the model, and what Project Glasswing Anthropic’s gated industry consortium means for the future of cybersecurity. Jon also shares practical tips for securing the code you’re generating with AI tools.

Additional materials:⁠ ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/990⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.

More from Super Data Science: ML & AI Podcast with Jon Krohn

All 130 episodes
990: Inside Mythos: Anthropic's Locked-Down Frontier ModelSuper Data Science: ML & AI Podcast with Jon Krohn · 11 min
Listen in VO