Claude Fable 5 review: what the new Mythos model gets right (and very wrong)

9 Jun 2026 · 17 min · 7 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Review of Anthropic’s Claude “Fable 5” (a Mythos-class model) after early access, covering benchmark claims, safety guardrails, costs, and real-world strengths/weaknesses.

Guest backgrounds

No guests mentioned; the host (“Claire”) provides the review from hands-on testing.

Key claims

Fable 5 is the first Mythos-class model to reach GA, “crushing” benchmarks (e.g., Swebench Pro) and excelling at long, autonomous, days-long async work plus vision. It’s token-intensive (about 2x token consumption) and expensive ($10/input, $50/output). Safety: cybersecurity/biology/chemistry classifiers with fallback to Opus 4.8; 30-day retention for misuse detection.

Notable examples

better handwriting/PDF document formatting; but nearly unreadable prose for PRDs/specs, poor one-shot UI/design (skills registry), overly conservative MVP output, and multi-agent runs that sometimes stalled/errors.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Overview of Claude Fable 5

0:45 to 2:20

Discussing the features and claims around the new Fable 5 model.

“scaring slash warning us about the unbelievable capabilities of Mythos, and it is finally here.”

Performance Insights and Limitations

2:20 to 3:47

Analyzing performance benchmarks and limitations of Fable 5.

“Now, what are some things that earlier models couldn't do that they are saying now that Fablefy can do?”

Pros and Cons of Fable 5

3:47 to 6:06

Exploring the advantages and disadvantages of using Fable 5.

“Now, I have done probably day, days long sessions with other models.”

Safeguards and Limitations

6:06 to 8:00

Discussing the safeguards in place for Fable 5 and its implications.

“And again, I think that the untrained of us will say, oh, well, I have this fable model, I should use it, it's better than anything.”

User Experience with Fable 5

8:00 to 10:00

Sharing personal experiences and feedback on using Fable 5.

“Mythos is still restricted to these Project Glassween partners.”

Strengths and Weaknesses in Design

10:00 to 14:00

Examining design capabilities and shortcomings of Fable 5.

“So I read Fable 5 on a bunch of different work, and I want to give you my feedback on where I thought it did well, where it needed a little bit of work, and where I was really surprised.”

Evaluating Claude Fable 5's Performance

14:00 to 16:44

Learn about the strengths and weaknesses of the Claude Fable 5 model based on real-world applications.

“So again, you might want to toss an opus in the mix instead of relying on Fable 4 design.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00It's here, the model, the myth, the legend, Mythos from Anthropic has finally dropped. Well, baby Mythos. We're calling it Fable 5 and this new model is crushing benchmarks, but the question is, can it crush my backlog? I got early access to the model and of course I have my own opinions on where it does really well, where it needs a little work, and the question on everyone's mind, does it live up to the terrifying marketing hype? Let's get to it. Okay, let's talk about what Anthropic is telling us about this model, and then we'll get into what I think about it. So this is Claude Fable 5, the first Mythos class intelligence model to reach GA.

0:43Now, if you haven't been paying attention, Anthropic has been marketing slash scaring slash warning us about the unbelievable capabilities of Mythos, and it is finally here. Now, they had originally been rolling this out with a couple select companies. I got early access to test what I thought of the model. But you have to know, this is not Mythos, capital M, big mythos. This is themy mythos. This is fable. And so it's going to have some guardrails on it, in particular around cybersecurity exercises and biology exercises. Now, good news, your girl's working on PRDs. She's shipping SAS. She's not working on biology quite yet, although give me a little time and some time to experiment.

1:26Maybe I'll get there. So this is really going to be focused on what the everyday user, what the everyday software engineer is going to think about when they're using this model. Although I did run into some things that I suspect are a result of the tuning and training of this particular model to be extra safe. Now, quick, it's not cheap. it's$10 per input token and$50 per output token it's going to be a new tier above opus and so if you're going to use this model you're going to pay pay the price so what is anthropic saying basically it's a completely new model class so we had sonnet we had opus and now we have mythos the first of which is fable 5 it's completely state-of-the-art it is exceeding every benchmark they tested by a significant amount this 80 on sui bench pro you'll look at that compared to some of the more recent models that have come out.

2:13Very, very good benchmark performance. And then they're saying it's really good for long, complex tasks. Now, what are some things that earlier models couldn't do that they are saying now that Fablefy can do? It's very autonomous, including running days-long asynchronous tasks. It's really an engineer's engineer. And that's some of the downside I experienced with this model. I'm going to show you a very specific example of where you don't want an engineer doing your work with an engineer's point of view. Proactive. It's very good at vision, exceptionally good at vision. This is a place where I actually really loved the model.

2:49And you know me, I'm pretty critical of models, but I did see a step ahead of vision. So that's something we're going to dive into. And then effort. It works hard. It builds harder. It verifies more. It's built for ambitious work. Now, guess what it also can do? It can consume those tokens. So Anthropic has said it consumes rate limits and tokens at about 2x the rate of other models. So again, this is a big boy model and it's going to consume tokens and some of the things that it's good at and even some things that they have done in the harness seem like they're intentionally or not token consumers.

3:23So we're going to keep an eye out on costs and an eye out on efficiency when using this model. Again, talking about long running tasks, Fable 5 is supposed to be able to run for days. So doing long running planning, being able to spin up sub agents, and I show a little bit about dynamic workflows, which are, you know, different architectures of sub agents and holding multi day sessions. Now, I have done probably day, days long sessions with other models. I didn't have Fable for many days. So I cannot verify that it ran for days. I did get it to run, however, for several hours on some tasks that may or may not have merited that several hour effort.

4:04But it definitely seems like it has both the harness and the intelligence capability to run for a very long time, if that's appropriate for your task. Now, here's your pros and here's your cons. They explicitly say that Fable works like a seasoned engineer. Unfortunately, if you have worked with a seasoned engineer, you know there's good to this and you know there's bad to this. So it is very complete in its investigation. It's definitely going to go search out all the corners. It's definitely going to Think about how it can be 120 % sure that it's shipping the right thing. But guess what? That's not always in service of launching.

4:42And that's honestly not always in service of building a great product. So while you can give it a goal and it will be very autonomous and it will be very thorough. Honestly, sometimes you want like a slightly less thorough engineer. Product manager talking, even engineer talking. Sometimes you want it to be a little bit dumber. We'll talk about some of the prompting techniques it says and when to use this model. But it's just something to think about when you're working with any high intelligence model is how much intelligence does the task actually take. Now, as I said before, it is token intensive by design.

5:19And I did most of my tasks on extra high. And so it was like token burning on token burning. and so they say that high is probably the sweet spot for most work I used extra high just because I don't want anybody in the comments saying Claire you picked high for this task and it should have been extra high and you would have had a better experience I used extra high I used all the brains of fable but again it is very very token intensive my question for any of these models this is not an anthropic model question this is not a fable question is does this token intensity actually output the right results.

5:56And that's a place where I'm just not 100 % sure. But again, as us humans in the loop, we're going to have to be much more intelligent about where to put what model and where to use what reasoning and what effort level to match what we're doing. And again, I think that the untrained of us will say, oh, well, I have this fable model, I should use it, it's better than anything. And honestly, I still think there's a place for good old Sonnet. I think there's a place for Opus. And I think there's a place for other models in the ecosystem. Now, there are safeguards in this model. And so this was one of the first things that Anthropic told me testing the model.

6:34And this is one of the headlines that they're making in the release, which is there are specific classifiers in this model for cybersecurity, biology, chemistry, and distillation. Basically, they don't want anybody doing bad stuff in those categories in particular with this very intelligent model what's nice about how they've implemented this however is they have this new fallback concept and so if you get classified into one of these categories instead of saying like do not pass go you may no longer fable it just falls you back to opus 4.8 this is also a capability in the API now where you can do this graceful fallback to for eight if you're using a Mythos class model.

7:17They also have a 30-day retention policy used only to catch misuse and it's not used to train Claude. So while it's still not training Claude, they do want to check the use of this model because they have been and will forever be very cautious about us normies using their intelligent models. And just, you know, for context, 95 % of sessions on this model did not hit a fallback. I don't believe I hit a fallback. But again, I'm not doing anything in cybersecurity, biology, or chemistry, at least yet. Okay, so this is the question. Is this or is this not Mythos? It is Mythos. Fable has the safeguards.

7:55Mythos does not. Fable, all us normies can have in general availability. Mythos is still restricted to these Project Glassween partners. some of these enterprise level partners that are really checking it against cybersecurity use cases. I would suspect that at some point we get some access to a Fable 5.0 whatever, or that the Project Glasswing class opens up. But for now, we get Fable, Project Glasswing, or these pre-selected companies get Mythos, but they are all fundamentally the same underlying model. A couple product things that are also launching today, along with the Fable 5 model, Cloud Managed Agents are going into public beta.

8:38If you haven't paid attention, this is Anthropix hosted harness, hosted sandbox for running long running agentic work. I am still trying to figure out what a good use case for Cloud Managed Agents is. I will get there. But Fable ships out of the box in Cloud Managed Agents. There's also a new advisor strategy where you can use Fable 5 as a senior advisor and use cheaper models as an execution layer. A lot of people are doing this with Opus and Sonnet. And so this is going to work today in the API and in Cloud Code and is a strategy you can use. And then as I mentioned, this fallback API where you can put an optional parameter on the messages API that allows you to continue to block requests by using 4.8 at Opus pricing.

9:22Okay, as we said, crushing benchmarks. Look at this. Fable 5 compared to Opus 4.8, GPT-55, and Gemini 3.1 Pro. Significant increase in Sweebench Pro benchmark. Very far ahead of these other models. And while I wasn't testing the most advanced use cases, I didn't find something that technically it failed at. So I think these benchmarks are really going to hold. And these benchmarks have outperformed across the board. So this is Anthropics state-of-the-art model. Okay, so enough about what they say. Let's talk about what I say. What is it actually like to use? So I read Fable 5 on a bunch of different work, and I want to give you my feedback on where I thought it did well, where it needed a little bit of work, and where I was really surprised.

10:11As I said before, it's really good at vision. And where is it good at vision that really impressed me? It's really good at document formatting. So this is super simple, but we've been doing these handwriting documents for my seven-year-old based on classic texts and classic poems. And on the right is Opus 4a, and on the left is Mythos 5. And it looks so silly, but I really do think Mythos 5 did a much better job of a second grade layout for a handwriting sheet. There's just like the right spacing. It's very clear to read. There's enough white space. I think on the one on the right, it's just very dense.

10:52And even the lines themselves are sort of hard to tell. Do you write above? Do you write below? So I do think that PDF formatting documents, I tested this against a bunch of different models. Mythos 5 really did a good job. So very simple eval for me, but a very, very good one. Now here's the problem though. The writing is nearly unreadable. So if you're thinking about mythos for pros, for spec writing, for PRDs, unfortunately, it's an engineer. And what's the problem with engineers? They just really get wrapped around the axle on details. And this is a real struggle with these more intelligent frontier models is they're like too smart.

11:34And so it's just very, very hard to parse what they're saying. And I'm going to show an example of this in actually CloudCode. So I have this concept of a product graph that I'm working on for ChatPRD. It's actually a fairly complex open source project. And I had Fable 5 go through that and actually do like an adversarial review of my requirements to try to figure out where there were internal consistencies in the logic. And it gave me this markdown document that looks very long and intelligent, but if you actually go through it, it's just really hard to parse. It's like internal references. It's very detailed, but not in a way where you can zoom out.

12:20There are these big blocks of paragraphs like look at look at this. It is just really hard to see the forest for the trees in this particular model. And I saw this sort of like over and over again, working on it with specs is it was very complete, but nearly imparsible. And that's a real challenge when working with these very, very high intelligence models. Again, I would actually suggest pulling back to maybe a Sonnet or Opus model for specs and then looking at Fable as an orchestrator of execution where that detail really matters, but you don't have to read it. The other thing that shock, shock, shocked me was how like actually legitimately terribly bad it was at design or at least a one shot design.

13:07And so I asked Fable to design a skills registry and man alive, did it do a very poor job. I mean, I'm not even talking like AI slop bad is like fundamentally terrible design. Gray, black, red, simple outlines, just really, really terrible. Now, the Anthropic team suggested that I just needed to be a little bit more detailed in my prompting. I've never had to do this before in, I would say, the last year of models in terms of front end. But even when I prompted it, it was still just not very impressive design. I think there's this real balance between design slop and specificity and just shipping a terrible design.

13:51I'm not sure what about Fable 5 resulted in this. I'm gonna have to keep testing it as it rolls out today. But this was a real disappointment in terms of design. So again, you might want to toss an opus in the mix instead of relying on Fable 4 design. It's really conservative on execution. So when I was trying to do that ambitious days long work, I took a spec and I said, can you ship the V0 of this, the MVP? I said, enough to that a customer could get value. And the MVP, they just really took minimal to heart. It was like very, very narrow, not actually that useful. And I'm curious if this comes from some of the safeguards on this model.

14:30And it's been a challenge I've seen since the kind of later Opus models is they're not super ambitious. And so again, you'll have to think about how to prompt this to get that long running outcome paired with the right product ambition. And then I really doubled down trying to test these clawed dynamic workflows and these sub-agent designs, trying to see if this would really add value. And the multi-agent capability is definitely there. And I definitely had some successful multi-agent runs kicked off in Fable, but I also ran into a lot of stalls and errors in using multi-agent orchestration. Now, I made the mistake.

15:09I walked away from my laptop and came back to these sub-agents that had stalled after about three hours. And so like egg on my face. But I really want to see how technically the Claude Code model holds up to the promise of multi-agent orchestration. I had some successes and some bugs. I think this is a Claude Code issue, not necessarily a model issue. Although with this promise of long running days long prompts, you really got to deliver technically on the outcome. So what's my takeaway? I would hand it hard problems, of course, not cybersecurity, bio or chemistry problems, but hard technical problems where being extremely detailed matters, long horizon work.

15:53I would also hand it vision problems where you really want something to look good or you want it to parse PDFs or other documents. It's done exceptionally well there. I was actually really surprised. I probably wouldn't hand it my front end work or I definitely wouldn't hand it my front end work and I definitely wouldn't hand it strategy or spec work. I think it overthinks things. I think its prose is nearly imparsable. And so maybe I'll test it again with effort level lower on sort of prose and spec writing, but it wasn't it for that. That being said, I'm not a hater on this model. I definitely not.

16:28It definitely has a place in your stack. I'm going to test it. If you want to learn more, definitely look up the prompting guide for Fable. It's going to probably repeat a lot of what I said. Hand it your hardest problems, what this model is good for and what it's not, and how to get a good outcome. That being said, mythos is here. I cannot wait to hear what you build, what you overbuild, and what you make ugly with this new model. Thanks for joining How I AI. Thanks so much for watching. If you enjoyed this show, please like and subscribe here on YouTube or even better, leave us a comment with your thoughts.

17:03You can also find this podcast on Apple Podcasts, Spotify, or your favorite podcast app. Please consider leaving us a rating and review, which will help others find the show. You can see all our episodes and learn more about the show at howiaipod.com. See you next time.

From the publisher

Claude Fable 5 is the first Mythos-class intelligence model to be generally available, and I got early access to test it before launch. In this episode, I walk through what Anthropic is promising, what actually stood out when I used it on real work, and where I think it fits in your AI stack.

—

In this episode, we cover:

(00:00) Introduction: Fable 5 is finally here

(00:31) What Anthropic says about the model

(05:14) Token-intensive by design

(06:28) Safety classifiers and the new fallback concept

(07:46) Is this or is this not Mythos?

(08:30) New product launches: Managed Agents and more

(09:20) Crushing benchmarks

(09:55) What it’s actually like to use (the good and the bad)

(11:40) Test 1: product graph spec

(12:56) Test 2: designing a skills registry

(14:04) Conservative on execution

(14:43) Test 3: multi-agent orchestration

(15:39) My takeaways

—

Tools referenced:

• Claude Fable 5: https://www.anthropic.com/news/claude-fable-5-mythos-5

• Claude Managed Agents: https://platform.claude.com/docs/en/managed-agents/overview

—

Other reference:

• SWBench Pro benchmark: https://www.swebench.com/

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

More from How I AI

All 103 episodes
Claude Fable 5 review: what the new Mythos model gets right (and very wrong)How I AI · 17 min
Listen in VO