STOIC REASONER: Dual-Mode Transformers that Compress to Think and Decompress to Speak

4 Dec 2025 · 12 min · 5 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Explains “Stoic Reasoner” (Soft Token Implicit Context Reasoner), a dual-mode transformer that compresses LLM reasoning into continuous “soft tokens” (latent, silent thinking) and only occasionally “decompresses” into explicit chain-of-thought steps for grounding and training.

Key claims

Standard chain-of-thought uses hard tokens and suffers discretization loss and KV-cache scaling bottlenecks; soft tokens reduce information loss and cut compute via probabilistic skipping of explicit decoding (Pupdate).

Notable examples

proof that 0.9999=1 illustrates long token traces; benchmarks include GSM8K (GPT-2-Small accuracy 44.73%→47.84%), logic on PROSTQA (best at Pupdate=0.5, 94.6%), and out-of-distribution IGSM (better accuracy retention up to 19 operations).

Guests

No specific guests named; it’s a host-led deep dive summarizing researchers’ work.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding the Stoic Reasoner Architecture

0:45 to 3:56

Exploring the Stoic Reasoner and its transition from hard to soft tokens.

“Which is an acronym for Soft Token Implicit Context Reasoner.”

The Dual-Mode Thinking of Stoic Reasoner

3:56 to 6:04

Examining how Stoic Reasoner employs both latent and local decoding modes.

“They aren't tied to the discrete vocabulary at all.”

Performance and Efficiency Gains

6:04 to 9:06

Analyzing the performance improvements and computational efficiency of Stoic Reasoner.

“But does this design actually translate to real, measurable improvement?”

Navigating Logical and Mathematical Challenges

9:06 to 11:08

Discussing how Stoic Reasoner adapts its reasoning strategy for different problem types.

“a more fundamental representation of logic.”

The Dilemma of Latent Reasoning

11:08 to 12:11

Exploring the trade-offs between efficiency and interpretability in AI reasoning.

“And I think we've learned that the unspoken parts of that dialogue, these soft tokens, that's where the key to smarter, more efficient problem solving really lies.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00We spend a lot of time talking about how large language models, you know, LLM's, actually reason. Right. We usually watch them think by generating this long, detailed explanation, what we call a chain of thought or code T. It's all very explicit, step by step. Exactly. But what if their most powerful, their most efficient thinking is actually happening in a kind of private, silent code, a language we can't even see? And that silent process, that's really the key focus of this deep dive. We're looking at research that's trying to make LLMs not just smarter, but vastly more efficient by teaching them to reason in silence.

0:36Because that standard chain of thought approach, it's becoming a real bottleneck, isn't it? A severe one. For complexity, for speed, it just doesn't scale well. So our sources introduced this really fascinating new architecture. It's called Stoic Reasoner. Which is an acronym for Soft Token Implicit Context Reasoner. And our mission today is to unpack this, to understand how this system can compress these long, explicit reasoning steps into, well, into what they call soft tokens. Exactly. We need to figure out what those soft tokens are, how this dual architecture actually works, and I think most importantly, what are the benefits?

1:10Is there hard evidence this silent thinking actually improves performance? Especially in the tough stuff, like math and logic. Right. And while cutting down on the computational cost. Okay, so let's start with that first crucial distinction. The model talking versus the model. Truly thinking. We have to begin with the current standard, the chain of thought. Why do models even use Cothee and what's its fundamental flaw? Well, when an LLM does chain of thought, it's using what we call hard tokens. So just regular words, numbers, punctuation. Exactly. The discrete units from its vocabulary. And the model is forced to project its very complex, high dimensional internal state, which is continuous onto this fixed, discrete set of words.

1:52Okay, so Cote is good because it forces the model to be structured, right? Yeah. Can't just jump to an answer. It provides a scaffold for reasoning, yes. But why is forcing that rich internal thought into plain language? Why is that so problematic from an engineering perspective? The core limitation is something the researchers call discretization loss. Discretization loss. Yeah. So imagine the model's internal thought is like a high resolution, continuous photograph. It's full of shades, subtle details, depth. But when the model has to output that as a hard token, it has to discretize it. It's like being forced to pick just three emojis to summarize that entire complex photo.

2:32You're losing all the nuance. All of it. You're taking this high-capacity thought and cramming it into a low-capacity bucket. And that information loss means the reasoning chain has to be so much longer to cover the same logical ground. The sources give a great example, the proof that 0.9999 equals 1. Right. The CA-T explicitly spells out every single step. Let x equals 0.9999, then multiply by 10. 10x equals 9.9999, then subtract x, 9x equals 9, and finally thus x equals 1. That looks simple to us, but for an AI, every symbol, every word, that's a separate token generation. And for a complex math problem, that trace can become hundreds, maybe thousands of tokens long.

3:14Which leads to a huge computational problem during inference. A massive one. And we have to talk about the KVCache. Okay, the KVCache. Remind us what that is. Think of it as the model's short-term memory notepad. It stores key and value projections for every token generated so far so it can maintain context. So the longer the reasoning chain, the more pages are filling up that notepad. Exactly. And that notepad is stored in very expensive, very limited memory, the VRAM. The KV cache requirement for standard Cothee, it just scales horribly with the length of the reasoning. So this brings us to the core innovation here.

3:50The move to latent reasoning using these soft tokens, this is where it gets really interesting. What are they practically speaking? They're continuous embeddings. They aren't tied to the discrete vocabulary at all. In the Stoic Reasoner, they're actually represented by the last layer hidden states of the transformer itself. So they can hold more information. Much more. Because they are continuous and high-dimensional, they have a far higher representation capacity. A single soft token can compress several steps of discrete Cotty reasoning into one compact vector. So a soft token is like a hyper-efficient summary of a whole chunk of thought.

4:24That's a great way to put it. And Stoic Reasoner operates in this dual mode. It's like it has an inner monologue and an outer voice. Let's break that down first. The inner voice. That's the latent thinking mode. In this mode, the model just generates the next soft token. It's compressing context, advancing the logic, all without generating any of those costly hard tokens. It's pure, efficient, internal thought. But if that's so efficient, why bother with the second phase, the local decoding mode? Isn't that just going back to the slow, expensive way of doing things? That's the critical question.

4:58The local decoding mode, the outer voice, is where the model takes that compressed soft token and, well, it decompresses it into a few explicit, human-readable Cati steps. But why? Why keep it? For two main reasons, grounding and learning. Yes. The explicit steps act as a necessary check and balance. If you're solving a math problem, you need to verify your intermediate calculations. The local decoding forces that. And crucially, that explicit output is also what's used as the training signal to teach the model how to connect its internal thought to the correct external structure. And the bridge between these two modes is a special switch token.

5:37So how does that work? How does it go from speaking back to thinking? The model is in local decoding mode, producing these hard tokens, these explicit steps, until it generates that special switch token. And that's the signal? That's the signal. The model then uses the hidden state of that switch token to update the soft token for the next reasoning step. It's effectively saying, okay, I'm done with this chunk of explicit work. Now I'll compress that summary back into my soft token memory and go back to thinking silently. That is really clever. That fluidity is the key. But does this design actually translate to real, measurable improvement?

6:11It absolutely does. It delivers results. The researchers tested Stoic Reasoner by fine-tuning models like GPT-2-Small and Gwenn 2.5 on some well-known reasoning benchmarks. Okay, let's start with math. The GSN8K benchmark, grade school math problems, what did they find? So for GPT-2-Small, the standard Koki baseline scored about 44.73 % accuracy. Okay. With Stoic Reasoner, that jumped to 47.84%. A three-point gain. Now, that might not sound huge to a layperson, but in model fine-tuning, that's a very meaningful improvement. It is. It shows that this compressed internal reasoning is often more effective at solving these kinds of structured problems.

6:52And what about the efficiency? The whole reason for this approach. Did the KV cash problem get better? It dropped significantly. With standard Kodi key, the computational cost, the FLPs, it stales really badly as the reasoning chain gets longer. Right. If you double the steps, it's way more than double the compute. It's exponential. It becomes prohibitive. But with Stoic Reasoner, they introduced a probability of updating the soft token, which they call PUPdate. And if the model skips the explicit decoding, the FLOPs scale much, much more favorably. Which means faster, cheaper inference. Exactly.

7:26Now, another thing the sources highlight is sampling different reasoning paths. Why is that so important? Because the best accuracy often comes from, well, from redundancy. Some models that reason purely in the latent space, they can become deterministic. They only produce one single path to the answer. And that's a problem for techniques like majority voting. It's a huge problem. You can't use techniques like PASSEC-K, where you generate multiple answers and pick the most common one, if you only ever get one answer. So because a stoic reasoner keeps that local decoding mode. Correct. It can still sample different reasoning traces.

7:57it gets the efficiency gains without sacrificing the accuracy boost you get from those voting techniques. Okay, let's get to the ultimate test. Generalization. Can it handle problems it's never seen before? The out of distribution or OD tasks? This is where that improved internal representation really shines. They use the IGSM synthetic benchmark for this, which lets you precisely control the difficulty by setting the number of required arithmetic operations or op so they trained it on easy problems and tested it on hard ones that's right they trained the model on data where say problems required nine operations or less and then they tested it on much harder unseen problems requiring up to 19 operations the results were pretty stark what happened stoic reasoner consistently retained higher accuracy than the cote baseline on these much harder problems and is performance didn't fall off a cliff as the problems got harder.

8:53Exactly. The accuracy decay was much, much slower. The Cote Baseline's accuracy plummeted as the number of operations increased, but stoic reasoners held up much better. It suggests the model is learning a more robust, a more fundamental representation of logic. That brings us to maybe the most fascinating part of this, the model's ability to adjust its own reasoning strategy depending on the task. Yeah, this is controlled by that pub date parameter we mentioned, the probability that the soft token gets updated. So walk us through how that works. During training with probability pupdate, the model does the explicit decoding and uses that to update its silent thought.

9:32But with the remaining probability, one pupdate, it just skips it. It keeps the same soft token and goes straight back to latent thinking. And you can tune this at inference time, like a sort of speed dial for thinking. A speed dial is a perfect analogy. And they found the optimal setting for that dial is completely dependent on the problem you're solving. Okay, so give me an example. Take PROS QA. It's a logical reasoning task. For logic, you often benefit from exploring multiple possibilities without committing too early. And sure enough, the best performance there, 94.6%, was achieved when training with PUP date at 0.5.

10:06So for logic, it's better to sometimes skip the explicit steps, to think more flexibly and silently. Precisely. The pressure to ground every single step actually hurts logical exploration. But now compare that to the GSM8K math problems. Right, the arithmetic. For math, the accuracy actually dropped anytime they set pup date greater than zero during inference. Wow. So what does that tell us? It tells us that for hard math, you need to ground every intermediate step. Always decode. Always check your work explicitly. Math demands that external accountability. Whereas logic benefits from internal flexibility.

10:41It's a profound distinction in how an AI should approach different problems. This really shows a path forward where models aren't just stuck in one rigid style. They can shift between this high-speed internal thought and grounded, readable verification. I mean, to summarize the achievement here, Stoic Reasoner gives you a single transformer that can interpolate between these two worlds. And the result is better robustness, performance gains, and much better efficiency. We started this deep dive asking about an AI's internal dialogue. And I think we've learned that the unspoken parts of that dialogue, these soft tokens, that's where the key to smarter, more efficient problem solving really lies.

11:19It is. And this whole paradigm forces us to confront what might be the biggest tradeoff in AI development today. Which is as these models move their most powerful reasoning into these continuous latent spaces, they get faster, they get smarter. But at the same time, their core decision making becomes non-human readable. So if the final crucial logical steps are all compressed into some high dimensional vector that we can't map back to language, how do we audit that? How do we verify or debug the AI's most complex thoughts? We can't. Not easily. This shift to latent reasoning offers these undeniable efficiency gains, but it really poses a fundamental challenge to interpretability and accountability.

11:59It requires us to trust the AI's black box thought process more than we ever have before. A fascinating dilemma. Smarter models that think in silence. That's definitely something for you to mull over as this world continues to accelerate.

From the publisher

This paper introduces the STOIC REASONER (Soft TOken Implicit Context REASONER), a new training paradigm for transformers focused on improving reasoning efficiency and capacity compared to standard Chain-of-Thought (CoT) methods, which rely on explicit hard tokens. This model leverages soft tokens, which are continuous latent representations that possess greater informational capacity than discrete vocabulary items, reducing the need for lengthy reasoning chains. The system operates in a dual fashion, utilizing a latent thinking mode to process the compressed soft tokens and a local decoding mode to decompress them into human-readable CoT steps. Evaluation on mathematical and logical tasks shows that STOIC REASONER achieves competitive or superior accuracy to CoT baselines, particularly exhibiting stronger out-of-distribution generalization on complex math problems. Furthermore, the architecture provides technical benefits such as a reduced KV cache requirement and the ability to generate sampling diverse reasoning traces during inference, which is often difficult for other latent reasoning methods.

More from Best AI papers explained

All 475 episodes
STOIC REASONER: Dual-Mode Transformers that Compress to Think and Decompress to SpeakBest AI papers explained · 12 min
Listen in VO