In short
How to choose LLM model size and “effort level” (reasoning/thinking level) settings, explaining what each dial changes during inference and how to use them to reduce cost while improving results.
Guest backgrounds
No guests mentioned; episode is hosted by Jon Krohn. Content is inspired by an Anthropic blog post by Lydia Hawley (Claude Code team, technical staff).
Key claims
Model size swaps which frozen weights handle your request (more capability; higher per-token cost). Effort level changes how thoroughly the model works (more file reading, verification, and tokens; not a literal “thinking time” slider). Fix missing context before changing dials; diagnose failures: too little effort vs insufficient knowledge.
Notable examples
Claude Code agentic coding (runs tests, reads files). Analogy: Sonnet=generalist, Opus=expert, Fable=deep specialist; effort determines how much each “person” does. Cost tradeoff: large models can be cheaper per task on hard multi-step jobs; Anthropic reports Fable completing jobs others couldn’t. Cross-industry: OpenAI auto/fast/thinking modes and GPT-5 tiers (Sol/Terra/Luna; Astra for multi-agent coordination) and Google Gemini thinking budget/level.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Claude Code and Model Settings
1:24 to 2:36
Discover the functionalities of Claude Code and its model settings.
“Also, if you're listening to this episode around the time it's released, there's a good chance you'll hear a mid-roll ad from Anthropic in this episode.”
The Mechanics of Model Size and Token Prediction
2:36 to 4:49
Explore how model size impacts token prediction and computational processes.
“When you send a request, everything, your message, the system prompt, the tool definitions, any files in the context, the whole conversation history, all that stuff gets packed into a single API request.”
Exploring Effort Level in AI Models
4:49 to 7:15
Understand how effort level influences model output and decision-making.
“It swaps which set of frozen weights handles your request.”
Practical Guidance for Using Model Size and Effort
7:15 to 8:56
Learn effective strategies for selecting model size and effort levels.
“If the model had all the pertinent context, visibly tried, and was still confidently wrong, that's a capability failure.”
Industry Trends in Model Size and Effort
8:56 to 14:03
Explore how various AI platforms are adopting model size and effort parameters.
“Tying everything together, the post offers an analogy I found sticky.”
Understanding Model Size and Effort Level
14:03 to 15:40
Learn about the importance of model size and effort level in AI and how to choose them effectively.
“some flavor of the model size dial is as well.”
Transcript
Automatic transcript. May contain errors.0:00This is episode number 1020 on choosing the right model size and effort level.
0:09Welcome back to the Super Data Science Podcast. I'm your host, Jon Krohn. Today's topic is the two dials that increasingly determine what you get out of a large language model. That's which model size you select and how much effort you tell that model to spend. If you've opened up Claude, ChatGPT, or Gemini lately, you will have noticed that the model picker has been sprouting new options. Not new models exactly, but new settings with names like effort level, reasoning effort, and thinking level. In today's episode, I'll unpack what these settings do under the hood, when you should reach for them, and when you should leave them alone.
0:46My jumping off point is a blog post that Anthropic published in July, written by Lydia Hawley, a member of technical staff on the Claude Code team. Claude Code, if you haven't used it, is Anthropic's agentic coding tool. You delegate coding tasks to Claude from the command line or your IDE, and it goes off, reads your files, writes code, runs tests, and reports back. As usual, we've got a link to the full post from Lydia Hawley in the show notes, Although the post is framed around coding, the mental model it lays out applies to any AI platform you might use. So even if you never touch a terminal, stick with me here.
1:24Also, if you're listening to this episode around the time it's released, there's a good chance you'll hear a mid-roll ad from Anthropic in this episode. I'm delighted to have Anthropic as such a big supporter of the podcast this year, but they have no influence on the topics I cover on the show. And while this episode is largely inspired by an influential anthropic blog post, the guidance generalizes to any model family, not just Claude. And so later in the episode, I'll explicitly generalize my advice with examples involving other frontier labs like OpenAI and Google. Anyway, back to Claude Code for the moment.
1:59Claude Code exposes two settings that both appear to make the answer better, the model setting and the effort level. A reasonable assumption would be that the model size setting controls how smart the response is, while the effort level controls how long the model thinks before answering. The first assumption, that model size controls how smart a model is, holds up. The second one, around effort level, turns out to be incomplete in an interesting way. Let's start with model selection because Anthropik's explanation of what that dial does is one of the clearest walkthroughs of LLM inference I've seen.
2:36aimed at practitioners. When you send a request, everything, your message, the system prompt, the tool definitions, any files in the context, the whole conversation history, all that stuff gets packed into a single API request. The first thing that happens server-side is tokenization. Your text is split into pieces and each piece is mapped to an integer from a fixed vocabulary the model was trained with. From that point on, your prompt is an array of integers. The model's job is to take that array of integers and predict which token comes next. It computes a probability for every token in its vocabulary and picks from the top.
3:12What turns your input tokens into those probabilities are the model weights, billions of tunable parameters organized into large matrices. Predicting one token means running your input through a long chain of matrix multiplications as we go through layers of a deep learning network to get into those kind of technical complexities a little bit, and then we read the probabilities out the other end. And the model doesn't generate a whole answer at once. It predicts one token, appends that one token to the sequence, and runs the entire computation again for the next one. So that means that a 200 token response is 200 separate passes through those deep learning matrices that make up the large language model.
3:55That loop is where most of your wait time and most of your output cost come from. Here's the key point. The model weights are set during training, and by the time you're sending requests, they are read-only. That's all inference means. Using the model after training is done with the weights frozen. Nothing in your prompt changes the weights. Your prompt and context can steer the prediction, and steering works well, which is why putting your real code or your real documents in front of a model improves results so dramatically, but steering isn't teaching. If a library didn't exist when the model was trained, it isn't in the weights.
4:28paste the docs into context and the model will use them for that request, but the underlying model retains nothing. This framing also demystifies hallucination a bit because when a model confidently calls an API that doesn't exist, that's the weights producing a token sequence that looks plausible from training patterns. So the model setting does exactly one thing. It swaps which set of frozen weights handles your request. In Claude Code's case, at the time of recording at least. That means choosing between Claude Sonnet, Claude Opus, and the newest and largest model Claude Fable. Bigger models encode more knowledge and capability in their weights, and each of their output tokens typically costs more to use.
5:09It certainly does in the Anthropoc case. What the model size setting doesn't decide is how many tokens get generated. The same prompt can produce wildly different token counts depending on how much work the model decides to do. And that's what the second dial, effort level, controls. This is where the blog post corrects a widespread misconception. Effort isn't a thinking time slider. In an agentic tool like Claude Code, the tokens a model generates fall into a few categories, reasoning tokens, tool calls, and the text it writes to you like plans, progress updates, and summaries. All of these are ordinary output tokens from the same generation loop built at the same rate.
5:47effort level shapes all of them. At high effort, the model reads more files, verifies more of its work, and pushes further through a multi-step task before checking back in with you. At low effort, it would rather ask you for more context than burn tokens figuring something out on its own. Mechanically, the effort level is sent to the model as one more input alongside your prompt, and the model was trained to behave differently at each level. That learned behavior is baked into the frozen weights. It sets a bar for how thorough and how certain the model needs to be before it considers a task done.
6:20Anthropic illustrates this with a same-prompt comparison where the high-effort path generates roughly seven times more tokens to reach a higher confidence answer. Importantly, effort sets how far the model is willing to travel, not how far it must travel. If step one of a three-hypothesis debugging plan finds the bug, a well-trained model at high effort will say so and skip the remaining checks rather than padding out your bill. Anthropic notes their team watches for overthinking during training because that also, in addition to costing you more, degrades effectiveness. So how should you use these two dials, model size and effort, in practice?
6:59Anthropic's guidance boils down to a single diagnostic question you should ask whenever a model gets something wrong. Did it try hard enough or did it not know enough? If the model skipped a file, didn't run the tests, or bailed on a refactor partway through, raise the effort. If the model had all the pertinent context, visibly tried, and was still confidently wrong, that's a capability failure. So move up to a larger model. So if you were trying it with Sonnet, then bump up to Opus. If you're trying with Opus, bump up to Fable. And critically, before touching either dial, you know, you don't need to touch model size or the effort level if the problem is that you provided a vague prompt or missing context.
7:47And actually, that's a more common culprit for the model not doing what you wanted. And no knob can fix that.
7:55Jon Krohn:On this podcast, I'm always going on about how Claude Code is mind blowing, but now Claude Cowork is making my jaw drop as well. For example, I recently wanted to quantify how healthy my sales pipeline is for my AI consulting business. I simply asked Claude to estimate my sales for the coming quarter, and it brought info from relevant Google Sheets and my Gmail to create a professional spreadsheet of clients with estimated revenue for each one. Whoa, this might have taken me a day. Instead, it was done flawlessly with Claude Cowork in minutes. Claude is the AI for minds that don't stop at good enough.
8:26Jon Krohn:It's the collaborator that actually understands your entire workflow and thinks with you. Whether you're debugging code at midnight or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. Ah, and you'll appreciate that I can ask Cowork to show me data such as my sales spreadsheet and it provides an interactive chart right in the conversation. For problems worth solving, get started with Claude at Claude.ai slash superdata. That's Claude.ai slash superdata and check out Claude Pro, which includes access to all of the features mentioned in today's episode.
8:56Jon Krohn:Claude.ai slash superdata. Tying everything together, the post offers an analogy I found sticky. Sonnet, you can think of as a strong generalist. Opus is an expert and Fable is a deep specialist who's seen problems almost no one else has. Effort decides how much time each of those different kinds of person, of model, spends on your task. So Opus at low effort is like five minutes with an expert. Deep pattern recognition, but a skim of your code. Sonnet, at high effort, is a good generalist with the whole afternoon. They'll read everything, run everything, and end up understanding your specific code thoroughly.
9:40With less of that, I've seen this exactly before, recognition. And Fable, even at low effort, is the specialist who glances at the problem and spots the thing nobody else would. Neither dial is universally better. Model size is roughly how capable, while effort is roughly how thorough, and most real tasks need some of both. There's a cost wrinkle worth internalizing here. On routine work, a small and a large model both get it right. So the large model's extra verification at a higher per token price is wasted money, drop down. So drop down from fable to opus or from opus to sonnet. On hard multi-step work, the equation actually flips though.
10:23The small model grinds through iterations near the ceiling of its ability while the large model reaches the same quality bar in fewer steps. So despite the higher per token price, total cost per task can actually come down and be lower on a bigger model as long as your task is tricky enough. This means that cheaper per token isn't always cheaper per task. And in Anthropics testing, for example, Fable finished long multi-step jobs that Opus and Sonnet couldn't reach at any effort level, which together with its price is the argument for saving Fable for work that really needs it. Now, as promised earlier in the episode, let's zoom out beyond Claude because the whole industry has converged on some version of these two dials.
11:09OpenAI arguably went furthest toward hiding them. When GPT-5 launched in ChatGPT, OpenAI introduced an automatic router that decided on its own when a query warranted deeper reasoning, with paid users only able to force the issue via the model picker or by typing something like think hard about this into the prompt. After user pushback, OpenAI surfaced explicit auto, fast, and thinking modes and later added a reasoning effort selector with named levels ranging from a near-instant setting up to an extended one alongside API parameters that let developers dial reasoning effort from none up through extra high.
11:48On the model size axis, OpenAI's current generation GPT 5.6 comes in three tiers with celestial names. Sol, meaning sun, is the flagship for the hardest work. Terra, Earth, is a balanced everyday model at half Sol's price. And Luna, the moon, is the fastest and cheapest of the model family, with the number denoting the generation and the name denoting a capability tier, mirroring the sonnet opus fable ladder over at Anthropic. And at the time of recording, an interesting little tidbit for you. OpenAI has confirmed a forthcoming model called Astra, another celestial word. This is a new class alongside those three existing Sol, Terra, and Luna tiers.
12:35And rather than being an upgrade to them, Astra is designed to coordinate multiple agents on long-running problems. An internal version reportedly cracked 10 open problems in mathematics and theoretical computer science, though there's no release date to the public, and even the shipping name may change when it is eventually released. Finally, Google took an evolutionary path with Gemini with respect to these dials. Their 2.5 generation models exposed a thinking budget, a raw token cap you set on the model's internal reasoning from zero up to tens of thousands of tokens. And with Gemini 3, Google replaced that with a simpler thinking level parameter, more like what OpenAI and Anthropoc are doing.
13:20And this just offers discrete settings like low and high, having concluded that forcing developers to estimate token counts was the wrong abstraction. Difficult to think through for us humans. And then going out even a bit further beyond the proprietary closed models of the Frontier Labs, most reasoning focused open source models like the Near Frontier Quen models from Alibaba, which you can hear more about in last Friday's episode of this podcast in episode number 1018. Those reasoning focused open source models offer analogous reasoning effort toggles too. So whichever platform you build on, some flavor of the effort dial is waiting for you.
14:02And usually some flavor of the model size dial is as well. I find this convergence telling. Two years ago, the industry's answer to how do I get a better response was one dimensional, one dimensional. It was just use a bigger model. Today, the frontier labs are unanimous that capability and diligence are separate axes, that they're priced differently, and that matching both to the task rather than maxing both out is what separates a savvy user, maybe like you, listener, from an expensive user. All right, before we wrap up here, let me leave you with three practical takeaways. First, start with the defaults.
14:43Every provider tunes the default effort to what most people would want to spend. And Anthropic, for example, explicitly recommends treating effort as a general preference for your kind of work rather than something to fiddle with task by task. Second, when output disappoints, fix context before touching dials. Then apply the diagnostic. If the model doesn't try hard enough, that means you need more effort. If it didn't know enough, upgrade to a bigger model. Third and finally, spend deliberately. Routine work to smaller, cheaper models and reserve the frontier for problems that stretch it, remembering that on the hardest tasks, the expensive model can be the cheaper one because it can take so many fewer behind-the-scenes thinking tokens to come up with a solid answer for you.
15:28Understanding these two dials, model size and inference time or thinking time will make you sharper at extracting value from every AI platform you touch. And given how much of data science workflow now runs through these models, that is leverage worth having for sure. So go match your dials to your tasks and make something great happen with the magical wizard powers we now all have. All right, that's the end of today's episode. If you enjoyed it or know someone who might, consider sharing this episode with them. leave a review of the show on your favorite podcasting platform or on YouTube. Tag me in a LinkedIn post with your thoughts and I'll respond to those.
16:09And of course, if you're not already subscriber, subscribe, come on. The most important thing to me though, is that you just keep on listening. I'm so grateful to have you listening and hope I can continue to make episodes you love for years and years to come. Until next time, keep on rocking it out there. And I'm looking forward to enjoying another round of the Super Data Science Podcast with you, yes you, very soon.
From the publisher
In Episode #1020, Jon Krohn unpacks the two dials that increasingly decide what you get out of a large language model: which model size you pick and how much effort you tell it to spend. Using a July Anthropic blog post by Claude Code’s Lydia Holly as a jumping-off point, with guidance that generalizes to any model family, Jon explains what each setting actually does under the hood. Model size swaps which frozen weights handle your request (roughly, how capable), while effort sets how thorough and certain the model must be before calling a task done, not a simple “thinking-time slider.” He offers a clean diagnostic for when to raise effort versus move to a bigger model, shows why cheaper-per-token isn’t always cheaper-per-task and surveys how OpenAI, Google and open-weight labs have all converged on these same two dials.
Additional materials: www.superdatascience.com/1020
Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.
In this episode you will learn:
(00:56) What the model-size dial actually does
(05:29) Why effort isn’t a thinking-time slider
(13:25) Three practical takeaways for using both dials




