SuperThoughts: Reasoning Tokens in Superposition

26 Jun 2026 · 19 min · 7 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode explains “SuperThoughts,” a framework to speed up LLM reasoning by generating two tokens per step using “superposition,” while preserving accuracy via confidence-based adaptive fallback.

Guest backgrounds

No guest identities or bios are provided in the transcript; it’s a two-speaker discussion.

Key claims

Autoregressive reasoning is slow because each token requires a forward pass. Pure latent-space reasoning fails due to representational drift (loss of token-level supervision). SuperThoughts compresses consecutive token pairs into one latent vector, predicts the next odd token with the main module, and predicts the next even token with a lightweight MTP module. If MTP confidence (tau≈0.999) is low, it rejects the even token and falls back to standard one-token decoding.

Notable examples

Math/science benchmarks (MAD500, AMC23, Olympia Bench, GPQA Diamond) show 20–30% shorter generations with only 1–2 point accuracy loss; uniform superposition without fallback drops accuracy up to 14–21 points (1.5B model). Larger models tolerate more compression, and wall-clock savings are substantial (e.g., 14B: 33% fewer steps → 28% faster).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding the Bottleneck of AI Reasoning

0:56 to 3:18

Explore how current AI models are limited by sequential token processing and the need for speed in reasoning.

“Because this mechanism, it allows AI models to process and generate thoughts in parallel pairs.”

The Problem of Representational Drift

3:18 to 6:20

Delve into the challenges of maintaining accuracy and supervision in AI when moving to latent reasoning.

“without constantly, you know, translating its thoughts down to discrete tokens, it could express vastly more intermediate computation per step.”

Introducing the SuperThoughts Framework

6:20 to 8:08

Learn how SuperThoughts combines speed with accuracy through a novel superposition mechanism.

“How does that handle the blend without losing the nuance of the words?”

The Three Components of SuperThoughts

8:08 to 12:37

Examine the architecture of SuperThoughts, including the compressor, main module, and prediction module.

“That simple operation gives it enough context to accurately predict the next even index token.”

Training the SuperThoughts Model

12:37 to 14:01

Discover the two-stage training process that enables the SuperThoughts model to interpret compressed vectors.

“It looks at the maximum probability score of the even token it is attempting to predict.”

Evaluating SuperThoughts Models

14:01 to 17:46

Discover the performance enhancements of SuperThoughts models with adaptive fallback.

“Which leads to the ultimate question for anyone implementing this, right?”

Implications of Variable Compute

17:46 to 18:58

Explore the potential future of AI and human cognition through variable compute.

“Compressing token pairs and leaning on that confidence-based fallback, it gives us an AI that is significantly faster without sacrificing its rigorous problem-solving accuracy.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00You know that feeling when you're watching someone type and they're exclusively using their index fingers? Oh, I know exactly what you mean. Just one. Key. At a time. Yes. Just peck, peck, peck. It is absolute agony to watch. It really is. You know exactly what they're trying to say, but you are basically held hostage, right? Just waiting for them to peck it out letter by letter. Right. Exactly. And the wildest thing about that is, when we ask the most advanced artificial intelligence on the planet to solve a really complex problem, that is basically what it's doing. Yeah, pretty much. It is hunt and peck typing its thoughts, just one single discrete token at a time.

0:38I mean, it's incredibly powerful, obviously, but it is fundamentally bottlenecked by this step-by-step generation. This huge bottleneck. Which is why today we're doing a deep dive tailored specifically for you into a really fascinating new framework. It's called SuperThoughts. SuperThoughts. Yeah. It's such a cool concept. It really is. Because this mechanism, it allows AI models to process and generate thoughts in parallel pairs. So it effectively doubles the thinking throughput, but, and this is the crazy part, without losing problem-solving accuracy. Right. And the approach to compressing the AI reasoning here is just brilliant.

1:16What's most compelling isn't just the sheer speed up, you know. Yeah. It's how the architecture actually manages to keep the AI from completely losing its train of thought. because that has been a major roadblock for previous attempts to make AI think faster. Exactly. So today we're going to unpack the exact mechanics of how this compression actually works and the incredibly clever fallback mechanism that keeps the AI from, well, making mistakes on difficult problems. That fallback mechanism is key. Right. And we'll talk about what this all means for the future of highly efficient, fast reasoning models.

1:49Because whether you are actively building neural networks or you're just sitting there wondering why your AI assistant takes 30 seconds to solve a logic puzzle, understanding how AI can learn to basically bundle its thoughts, it's going to completely change how you view machine intelligence. Absolutely. It totally shifts the paradigm. So let's look at the bottleneck itself first. We know that test time compute scales linearly with every single token. You're right. So a 500 token reasoning chain means 500 full forward passes, right? Yeah. And the computational cost of that is just skyrocketing because we want these advanced models to think longer and harder on complex problems.

2:28Right. We want them to show their work. Exactly. Generating massive chains of intermediate steps, but forcing the architecture to output discrete human readable tokens every single interval. I mean, that requires pushing data through every single layer of the network over and over again. It's like we've essentially taken this massive, multidimensional alien brain and we're forcing it to translate its continuous thoughts into our slow, clunky human alphabet at every single step just so we can read it. That's a great way to put it. It just seems wildly inefficient, right? Especially when the AI's internal latent space is already capable of holding vast amounts of semantic meaning.

3:10Oh, totally. And researchers have actually been trying to tap into that latent space for a while now. Really? Yeah, because if the AI could reason purely in that rich, high-dimensional vector space without constantly, you know, translating its thoughts down to discrete tokens, it could express vastly more intermediate computation per step. Right, it would bypass the bottleneck entirely. Exactly. It would just be lightning fast. The pure latent space dream. But I imagine taking away the tokens, that sort of removes the guardrails, doesn't it? That is exactly the issue. The model basically loses its intermediate supervision.

3:44Okay, unpack that a bit. So in standard autoregressive training, the model receives dense step-by-step feedback through token-level cross-entropy loss. Right. Every single token acts as an anchor. So if you let the model just wander off into pure latent space for, say, 50 steps without ever checking its work against actual tokens, the errors just compound. Ah, I see. It supples from something called representational drift. Representational drift. Okay, so standard chain of thought is kind of like a student showing their math work step by step to a really strict teacher. Yes, exactly. It is slow and it's tedious, but getting checked at every single line keeps you from making a fatal error early on.

4:26Right. The teacher catches you at step two. But pure latent reasoning, that's like trying to do a massive 10-step calculus equation entirely in your head. Oh, yeah, which is way faster. Way faster. But if you lose your place on step six, you totally crash because there are no checkpoints to audit your own logic, right? And if you lose your place on step six in your head, there's just no way to trace back the error. That is representational drift in a nutshell. Wow. You get the theoretical speed of the latent space, but you completely lose the reliability of that discrete supervision. And that causes pure latent methods to fail catastrophically on hard reasoning benchmarks.

5:01Which brings us to super thoughts, because this framework manages to get the best of both worlds, right? It finds a way to utilize the speed of the latent space without, you know, crashing the car. Right. And the core mechanism here is superposition. Superposition, I think. So instead of processing and predicting one token at a time, the framework compresses pairs of consecutive tokens into a single latent representation. Just squishes them together. Exactly. It consumes two tokens at once, and it predicts two tokens per step, which literally cuts the required four passes exactly in half. That's incredible.

5:37But doing that, I mean, that requires splitting the architecture into three distinct components, right? Let's break down the mechanics of that. Sure. The first component is the compressor. So at any given step, the compressor takes a pair of tokens, let's say the words it's and currently, and it squishes their embeddings together into one single vector. Yeah, and the mechanics of that squishing process are surprisingly elegant. How so? Well, you might assume you'd need a really complex transformer layer with, you know, dense self-attention mechanisms to blend two distinct semantic meanings. Yeah, that sounds like a heavy computational lift.

6:12Right, but the data actually shows a simple linear projection matrix works just as well. Wait, really? Just a basic linear projection? How does that handle the blend without losing the nuance of the words? Well, remember, the tokens are already embedded in a really rich dimensional space. So a linear projection just performs a weighted mathematical blending. It's essentially translating the geometry of those two tokens into a new single coordinate in the latent space. Oh, I see. A full transformer layer would apply self-attention, which is computationally expensive and honestly largely overkill for just combining two adjacent sequential meanings.

6:49Makes sense. Yeah, the linear projection is highly effective. It preserves the spatial geometry, and it adds very little overhead. Okay, so the compressor outputs this single superposed vector, and that gets fed into the second component, which is the main module. Right. And this is just the standard base large language model, correct? Exactly. So the main module takes that compressed vector, evolves its internal reasoning state, and predicts the next odd index token. So if the compressor tokens 1 and 2, the main module does the heavy lifting of advancing the logic to predict token 3. You got it.

7:23And then we arrive at the third component, which is the multi-token prediction module, or MTP. Right, the MTP. This is a very lightweight, one-layer module, and it is entirely responsible for predicting the next even indexed token, token 4. Okay, but how exactly does the MTP fuse the information together to make that prediction? I mean, it has to juggle the past context and the brand new token simultaneously. Right. So it concatenates three specific elements. Okay. It takes the vector embedding of the previous even token, then the newly generated odd token that the main module just predicted, and finally the deep contextual hidden state from the main module itself.

8:02Oh, wow. So it's grabbing from everywhere. Yeah. It lines those three vectors up, concatenates them into one continuous representation, and just pushes that through a lightweight feedforward layer. That simple operation gives it enough context to accurately predict the next even index token. Okay, I have to push back a little on the premise of the AI actually reading these compressed vectors, though. Sure, go ahead. Because if the AI is suddenly fed a vector, that means it's, and currently, at the exact same time, doesn't it just read as gibberish? I mean, you can't just modify the architecture, hand a neural network a blended vector, and expect its internal logic to remain stable, can you?

8:41Oh, you're entirely right. The network absolutely outputs gibberish if you just bolt these modules together and turn it on. Ah, I knew it. Yeah. Training the model to actually interpret that compressed space requires a very rigorous two-stage protocol. Okay, what's stage one? Stage one is latent distillation. You have to actively teach the SuperThought student model how to read those newly minted compressed vectors. Got it. So you bring in a standard token-by-token model to act as a teacher. Exactly. You feed the teacher model the sequence normally, just one word at a time, and you meticulously record its internal hidden states.

9:17Basically recording its brainwaves as it processes the logic. That's a perfect analogy. Then you feed the student model the compressed token pairs. But here is the key. The training objective here isn't to predict the next word. It's not. No. The objective is to physically alter the student's internal weights until its hidden states exactly match the teacher's hidden states. Wow. OK. And the alignment is measured using smooth L1 distance, right? Yes. Why use that specific metric instead of, say, a standard mean squared error calculation? That's a great question. Smooth L1 acts kind of like a hybrid measurement.

9:55It is sensitive enough to catch really small discrepancies between the student and teacher states, but it doesn't overly penalize massive outliers. Because there are outliers when you squish things together. Exactly. When you are squishing high-dimensional embeddings together for the first time, you get some extreme outlier vectors. Right, the gibberish. Yeah. A standard mean squared error would overreact to those outliers and completely destabilize the training. but smooth L1 gently but firmly adjusts the student's internal weights so its vector map aligns perfectly with the teacher's. Okay, so it's basically like teaching an assistant to use a highly condensed shorthand.

10:32I like that. First, you do the distillation phase. You sit down and make sure their quick shorthand notes translate perfectly to your detailed longhand instructions. Right. You aren't asking them to write original reports yet. You're just ensuring they can read the shorthand perfectly without disinterpreting your intent. That is exactly what is happening. And then once that shorthand vocabulary is locked in and the latent space is completely aligned, you move to stage two. Which is joint training. Yes. You unfreeze the entire model and train all the components end to end using standard cross entropy loss.

11:04So now you are finally judging the model on its actual ability to predict the correct next tokens. And because it's still generating discrete tokens at every step, just, you know, doing it two at a time, it preserves that step-by-step supervision that pure latent models lack. Precisely. It doesn't suffer from representational drift because it is still forced to show its work. Exactly. But of course, the problem arises when the work gets too complex for the shorthand to handle. Right. Because the real world is messy and mathematics is incredibly unforgiving. Yeah, and the testing data highlights a glaring vulnerability here.

11:41Yeah. If you take a 1.5 billion parameter model and you force it to always predict in pairs, which I think they call uniform superposition. Yes, uniform superposition. Right. If you force that, its accuracy on rigorous math tests drops by 14 to 21 points. Yeah. It practically falls off a cliff. It really does. And the reason is that a single superposition step simply might not have enough computational capacity to process two highly complex, dense mathematical concepts simultaneously. It's essentially trying to rush dense logic through a bottleneck. Right. You're asking it to do too much at once.

12:15And obviously we can't have an architecture that just bombs complex reasoning tasks just because it's trying to read too fast. No, nobody wants a fast but stupid AI. Exactly. So the solution to this bottleneck is what actually makes the superthoughts framework viable in production. It utilizes this mechanism called confidence-based adaptive inference. Yeah, this is the really clever part. At every single step, the framework evaluates the confidence of the lightweight MTP module. It looks at the maximum probability score of the even token it is attempting to predict. And a very strict threshold is set, typically denoted as tau equals 0.999.

12:51So 99.9 % certainty. Exactly. And if the MTP module's confidence in that second token falls below that threshold, what happens? The model actively rejects the MTP's prediction. It just throws the token out. Wow. And it falls back to standard, one token at a time decoding. So in the very next step, instead of feeding a compressed pair, it feeds just a single odd token into the much more powerful main module, asking it to re-predict that tricky even token with its full computational weight. Oh, that is brilliant. It essentially self-regulates its own compute. Yeah. It literally knows when it is confused.

13:28It does. So if it is just coasting through easy text, like setting up the parameters of a word problem or generating some boilerplate code, it double steps using the shorthand. But the second it hits a mathematical roadblock and the logic gets really dense, it shifts into low gear to process one concept at a time. That's exactly it. It operates as a highly efficient dynamic compute scheduler. Amazing. It allocates more processing power to the difficult tokens and just cruises through the easy ones. And that completely avoids the massive accuracy drop associated with uniform superposition. Which leads to the ultimate question for anyone implementing this, right?

14:04Does this shifting gears approach actually save time while preserving intelligence in the real world? Right, the bottom line. And the evaluation on the Quinn 2.5 math models is incredibly revealing. They tested 1.5 billion, 7 billion, and 14 billion parameter sizes across rigorous benchmarks. We're talking MAD500, AMC23, Olympia Bench, and GPQA Diamond. These are not easy tests. These are advanced competition math and graduate-level science questions. Yeah. But across the board, with that adaptive fallback mechanism engaged, the SuperThoughts models achieved a 20 % to 30 % reduction in generation length.

14:42That is massive, getting nearly a third of your compute time back while only suffering a minimal one to two point drop in accuracy compared to the slow baseline model. That's a huge win. It really is. And it also completely outclassed rival token smashing methods, right? Like hamburger. Yes, hamburger. It achieved higher accuracy than hamburger at every single compression level. That's wild. And, you know, the performance data reveals a really fascinating trend regarding the scale of the models as well. The larger the model, the better it tolerates aggressive compression. Okay, wait, let's look through the numbers.

15:18When they forced uniform superstition by turning the adaptive fallback off, the 1.5 billion parameter model dropped up to 21 points. Right. But the 7 billion parameter model only dropped 5 to 12 points. So it handled the force compression vastly better. Exactly. Larger models possess richer latent spaces. They just have more spare capacity in their internal representations per step. So squishing tokens together is far less destructive to the overall logic. That's a really interesting point, because if bigger models have more spare capacity per step, scaling these models up toward, you know, trillion parameter architectures means they might eventually superpose three, four, maybe even five tokens at a time natively.

16:01Yes. The theoretical speedups could scale exponentially with model size. The trajectory heavily points toward exactly that. And there's another crucial factor regarding scale here. It's the difference between theoretical speedup, which is just reducing the number of generation steps, and actual wall clock time savings. Right, because theoretical speedup doesn't account for the computational overhead. Right. Like running the compressor, the MTP module, constantly checking the confidence threshold. Exactly. All of those extra steps require additional kernel launches, memory bandwidth, CPU activities.

16:32It adds up. So does it eat into the time savings? Well, testing measured with a fast inference engine showed that this overhead actually matters less as the model scales. Oh, really? Yeah. Larger models spend proportionally way more of their time deep in the heavy matrix multiplications of the main module. So it renders the lightweight MTP overhead basically negligible. Oh. For instance, on the 14 billion parameter model, a 33 % reduction in generation length translated to a massive 28 % reduction in actual wall clock generation time. Wow. So the bigger the brain, the more efficiently it utilizes the shorthand.

17:11Exactly. And as native multi-token prediction modules become standard in the pre-training of future large language models, implementing frameworks like SuperThoughts is just going to become even smoother. Because it addresses the fundamental bottleneck of auto-aggressive generation head-on. It really does. By actively monitoring its own internal confidence, the model dynamically allocates computational resources exactly where they are needed most. It's incredible. We've explored today how the Superthoughts framework successfully bridges the gap between slow, reliable token-by-token processing and fast, unstable, latent space reasoning.

17:46Yeah. Compressing token pairs and leaning on that confidence-based fallback, it gives us an AI that is significantly faster without sacrificing its rigorous problem-solving accuracy. It really moves the entire field away from treating every single piece of information with uniform effort. You know, it's pushing us toward intelligent variable compute. And for you, the listener, this signals a near future where running deep, complex reasoning tasks won't require endless waiting, or frankly, burning through massive amounts of server energy just to generate a few paragraphs of logic. It makes deep thinking AI practically usable at a global scale.

18:25Exactly. Which brings up a final thought regarding that variable compute idea. Because if an advanced AI can optimize itself by bundling simple thoughts together and actively slowing down only when it detects its own uncertainty, what happens when we apply this concept to our own human cognition? Oh, that's an interesting thought. Right. Are you treating all your daily mental tasks with the exact same uniform exhausting effort? Or should you be badging the easy stuff on autopilot, dropping it a single token mode only for the truly complex problems? because right now I think a lot of us are still out here exhausting our mental energy typing our lives away well one key at a time

From the publisher

SuperThoughts is a novel framework designed to accelerate the Chain-of-Thought (CoT) reasoning process in large language models by processing tokens in superposition. Unlike traditional models that generate tokens sequentially, this method uses a compressor to fuse pairs of consecutive tokens into single latent representations, effectively halving the number of required forward passes. To ensure accuracy is not sacrificed for speed, the system employs a Multi-Token Prediction (MTP) module and a confidence-based adaptive mechanism that reverts to standard decoding when the model is uncertain. Experimental results on complex mathematical and scientific benchmarks show that SuperThoughts reduces reasoning length by 20–35% while maintaining performance within a few percentage points of the original baseline. The research highlights that larger models are particularly adept at handling this compression, achieving significant wall-clock time reductions during inference. Ultimately, this approach offers a more efficient way to utilize test-time compute without losing the dense supervision provided by discrete token training.

More from Best AI papers explained

All 475 episodes
SuperThoughts: Reasoning Tokens in SuperpositionBest AI papers explained · 19 min
Listen in VO