Continuous Autoregressive Language Models

8 Nov 2025 · 16 min · 8 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Continuous Autoregressive Language Models (CALM) aim to reduce LLM inefficiency from token-by-token next-token prediction by predicting the next vector for a chunk of K tokens (“semantic bandwidth”), cutting the number of prediction steps.

Guest backgrounds

No guests are mentioned; the episode is a solo “Deep Dive” discussion with two speakers.

Key claims

Discrete token prediction is bottlenecked by small information per token and an expensive softmax over large vocabularies. CALM replaces next-token probabilities with next-vector prediction using a high-fidelity (robust) autoencoder, likelihood-free training (energy score), likelihood-free evaluation (BrierLM/Brier score), and temperature-like control via batch-size-based rejection sampling.

Notable examples

K=4 compresses four tokens into a 128D vector with >99.9% reconstruction accuracy; a 371M CALM matches a 281M Transformer while using 44% fewer FLOPs for training and 34% fewer for inference.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Inefficiency of Token-by-Token Generation

0:45 to 2:01

Explore why traditional token-by-token language generation is inefficient.

“So our mission today is to dig into the work on continuous autoregressive language models column.”

Understanding the Limitations of Current Models

2:01 to 3:56

Discuss the fundamental limits of current language models and their bottlenecks.

“Moving to subword tokens, things like WordPeace or BPE, that was a huge efficiency boost back then.”

Introduction to Continuous Autoregressive Models

3:56 to 6:20

Learn how continuous autoregressive models aim to overcome traditional limitations.

“The solution, according to this Calm research, is just leave the discrete world behind.”

Building a Robust Autoencoder

6:20 to 8:14

Discover the techniques used to create a robust autoencoder for vector representation.

“If the vector space is brittle, even a minuscule error in the predicted vector can make the decoder jump to a completely different part of the map, reconstructing total gibberish.”

Training the New Model without Probabilities

8:14 to 10:35

Understand the training process for a model that doesn't rely on probabilities.

“It prevents dimensions from collapsing and ensures you're using the full capacity of your vector.”

Evaluating the New Approach and Its Efficiency

10:35 to 13:14

Learn about the evaluation metrics developed for continuous models and their implications.

“They used a customized architecture they call the energy transformer to achieve that direct prediction of the next vector in one shot.”

Results and Implications of Continuous Models

13:14 to 14:00

Review the performance results of continuous models and their advantages.

“Breyer-Lem evaluation, bash size temperature control.”

Exploring Continuous Autoregressive Language Models

14:00 to 16:00

Learn about the efficiency gains and new design principles in autoregressive language models.

“Let's compute for the same quality output.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome to the Deep Dive. Today we're jumping into something of a paradox really. At the heart of modern AI, these large language models models are incredible right? probably the most powerful computing tool we've seen in ages. Absolutely transformative, no question. And it's a big but. They are just immense computationally. They chew through resources, driving up costs, development time, everything. Yeah, hugely expensive. And the standard of thinking, you know, has always been just throw more at it. Bigger data, bigger models, more compute time. Scale it up. Exactly. But the research we're looking at today suggests maybe that's not the whole story, that the inefficiency isn't just about scale.

0:40It's baked into the design. That's it. The bottleneck is fundamental. It's how they generate language, the step-by-step, token-by-token process. It's inherently inefficient. Oh. So our mission today is to dig into the work on continuous autoregressive language models column. It's pitched as this major paradigm shift designed specifically to get around that core limitation. Right. So our goal for you listening is to really unpack this central claim. does call actually improve the performance compute trade-off? Because they're saying it does by adding this whole new way to scale things, what they call semantic bandwidth.

1:17It's less about just making the model bigger and more like, I don't know, upgrading your internet connection, going from a dial-up modem to fiber optic for each step. That's a good way to put it. And the basic idea, the core mechanism, is actually pretty straightforward on the surface. Instead of predicting just the next word or token, Yeah. Calm takes a chunk of tokens, maybe four or five, compresses them all down into one single continuous vector, like a summary vector. Okay. And then the model just learns to predict the next vector. That simple change cuts down the number of prediction steps you need by, well, by the size of that chunk, K, instantly.

1:52Okay. Four times faster, potentially. But before we dive into Calm itself, let's just spend a minute on why we ended up with this token by token thing in the first place. It wasn't random, right? It was actually a big step up. Oh, absolutely. It came after character-level models. Moving to subword tokens, things like WordPeace or BPE, that was a huge efficiency boost back then. Because each token suddenly carried more meaning than just a single letter. Exactly. More information density per prediction. It worked great for a while. But that path, it seems, has kind of hit a wall, a fundamental limit.

2:24How so? Well, look at the typical vocabulary sizes for LLMs. You know, 32 ,000, maybe up to 250 ,000 tokens. Right. When you calculate it out, each individual token, even predicted by these giant models, carries a surprisingly small amount of actual information. We're talking like 15 to 18 bits. 15 bits. That does sound small. Especially when you think about models with billions, even trillions of parameters making that prediction. It feels really mismatched, doesn't it? And here's the kicker. You can't easily scale up that information density the old way. What do you mean? Like make tokens represent whole phrases?

3:00Exactly. If you wanted a single discrete token to represent, say, the quick brown fox or some complex idea, your vocabulary would have to explode. You'd need entries for every common phrase, every variation. It grows exponentially. Oh, right. Because language is combinatorial. Precisely. And that exponential growth creates an immediate impossible bottleneck. Yeah. The final output layer, the softmax calculation. The part that calculates the probability for every single possible X token. Yeah. If your vocabulary goes from 200 ,000 to 200 million, forget it. That final step becomes computationally infeasible.

3:37It's the choke point. So you're stuck. You have these incredibly powerful models. Huge representational capacitance. Forced to predict these tiny, low-information bits one after another, slowly, expensively. That's the discrete bottleneck in a nutshell. It fundamentally limits throughput and jacks up inference costs. Okay, so that sets this stage perfectly. The solution, according to this Calm research, is just leave the discrete world behind. Pretty much. Enter Calm, continuous autoregressive language models. Forget next token prediction. It's all about next vector prediction in a continuous space.

4:10And the key piece of tech here is this high-fidelity autoencoder you mentioned. Right. Its job is crucial. It takes that chunk of K tokens, compresses them down into a single dense vector, and the key is high-fidelity. Meaning it doesn't lose much information? Almost none. They showed for K4, so four tokens, they could squash them into a relatively small 128 dimensional vector and still reconstruct the original four tokens with over 99.9 % accuracy. That's basically lossless for practical purposes. Oh, wow. Okay. And this is where the scaling advantage comes in, right? You said it's different from the discrete problem.

4:48Totally different. Remember, in the discrete world, more density meant exponential vocabulary hell. Right, the softmax bottleneck. In KLM's continuous world, density scales nicely, gracefully. How so? Let's say you want to handle a bigger chunk, K6 or K8 tokens per step. You don't need a million new vocabulary entries. You just make the vector a bit bigger, add more dimensions. Ah, so instead of multiplying the output layer by thousands, you're just adding, say, 64 more numbers to your vector calculation. Exactly. It's a much, much gentler scaling curve. Yeah. And that directly cuts down your sequence length and prediction steps by that factor K.

5:26Huge efficiency gain right there. Okay. But switching to continuous space sounds like it breaks. Yeah. Well, everything else LLMs rely on. No vocabulary means no standard probabilities, right? No SOC max. Gone. No perplexity for evaluation. Also gone. So you need a whole new way of doing things, a new toolkit. That's exactly right. The continuous domain is like a wilderness for the standard LLM toolbox. And the very first thing they had to rebuild was that autoencoder itself. It couldn't just be any autoencoder. Why not? What was the issue? The paper mentions simple autoencoders create a brittle representation.

6:00What does that mean here? Okay, think about it. The autoencoder learns a map from chunks of text to points in this continuous vector space. If it learns too perfectly, too precisely, then the map is fragile. The main transformer model, the one predicting the next vector, it's never going to be perfect. it'll always have tiny errors in its prediction. If the vector space is brittle, even a minuscule error in the predicted vector can make the decoder jump to a completely different part of the map, reconstructing total gibberish. The output just collapses. Right, so a tiny nudge sends it off a cliff.

6:34It needs to be more forgiving, robust. Robust is the perfect word. It needs redundancy. So they used three main techniques to build that robustness in. First up, variational regularization. Making it a VAE, a variational autoencoder. Exactly. Instead of mapping directly, it learns a probability distribution, usually a Gaussian, around the target vector. It samples from that distribution. So it's intentionally adding a bit of fuzziness or noise. Why does that help? It's structured noise. It forces nearby similar token chunks to map to nearby, overlapping distributions in the vector space. It smooths out the map.

7:15Ah, so if the predictor is slightly off, it lands in a nearby region that still decodes to something sensible. Precisely. A small error in prediction leads to a small error in reconstruction, not total failure. It's crucial for making the whole thing work reliably. Okay, but VAEs have this known issue, right? Posterior collapse. Where some dimensions of the vector just stop being used. The paper said like 71 out of 128 dimensions collapsed for them initially. Yeah, it's a classic VAE problem, a huge one here. If dimensions collapse, they just become noise. The model figures out it can reconstruct OK without them, so they wither away.

7:50That's where the second tool comes in, KL clipping. KL divergence. That's about measuring the difference between probability distributions. Right. In simple terms, KL clipping basically puts a minimum requirement on how much information each dimension must carry. It modifies the loss function to say, hey, you can't ignore this dimension. It has to contribute. So it forces all 128 dimensions to pull their weight in encoding that chunk of tokens. Exactly. It prevents dimensions from collapsing and ensures you're using the full capacity of your vector. It stabilized the whole structure. Makes sense.

8:21And the third technique was dropout, which sounds familiar, but they used it differently. Yeah, dual dropout. They applied dropout randomly to the latent vector itself, the z vector, and the input tokens fed into the encoder. Why both? It forces redundancy. If parts of the input or parts of the compressed representation are randomly blacked out during training, the model has to learn to encode the information in multiple places using overlapping features. So if the downstream predictor makes a small error, messing up a few dimensions, the other dimensions still have enough info to get the right reconstruction.

8:55That's the idea. It builds in resilience against those inevitable small prediction errors. It's the final piece for robustness. Okay, so we have a robust autoencoder that can translate between token chunks and vectors without breaking. Now, how do you train the main model, the one predicting the next vector, without probabilities? No maximum likelihood. Right. MLE is out. They turn to something from statistics called a strictly proper scoring rule. Specifically, they use the energy score for training. The key thing for you to grasp is that it's entirely likelihood-free. Okay, so it doesn't need probabilities.

9:30How does it work then? What's it measuring? It basically measures distances in the vector space. It looks at the distance between the model's predicted vector samples and the actual ground truth vector, and also the distances among the samples themselves. And that's enough to train it? Yes, it has two parts. One part pushes the predictions towards the ground truth, that's fidelity. The other part encourages diversity, basically punishing the model if it just predicts the same average vector every time. It has to spread out its predictions appropriately. It works remarkably well without ever calculating a single probability.

10:05Fascinating. But they also needed the right kind of generative model architecture, right? You couldn't just plug this into a standard diffusion model. Definitely not. Models like diffusion or flow matching, while they work in continuous spaces, rely on iterative refinement. They need tens or hundreds of passes through the network to generate one sample. Which would completely kill the efficiency gains. Exactly. The whole point of Callum is fewer steps, thanks to that factor K compression. So they needed single step generation. They used a customized architecture they call the energy transformer to achieve that direct prediction of the next vector in one shot.

10:43Efficient prediction. Got it. So they can train it. How do they know if it's working? If it's getting better, no perplexity. Right. They needed a new evaluation metric. So they developed Breyer LM based on the Breyer score. Another scoring rule, like the energy score. Yep, also strictly proper, also likelihood-free. It provides a theoretically sound way to measure how good the model's predictions are, again, just by looking at distances between samples. And they can calculate this just by drawing samples from the model. No fancy probability math. Correct. It uses what's called an unbiased Monte Carlo estimator.

11:17Just generate some predictions, compare them to the actual next vector, calculate the Breyer score. Simple. But does it actually correlate with what we care about? Does a better Breyer LM score mean better language generation? That was the crucial validation. They found an incredibly strong correlation, Pearson correlation, of magazine 0.966 between their Breyer LM score and the standard cross-entropy loss you'd use in a traditional LLM. Wow. OK, so a lower Breyer LM strongly implies a lower cross-entropy loss, even though they never calculate cross-entropy. That's solid validation. It really is.

11:50It confirms Briar LLM is a reliable stand-in for perplexity in this new continuous world. Okay, last piece of the puzzle. Control. With normal LLMs, we use temperature sampling to adjust creativity versus coherence. We fiddle with the softmax probabilities, the logits. Which a column doesn't have. No logits. So how do you control the output? How do you get that temperature-like effect? They developed a specific algorithm for likelihood-free temperature sampling. Technically, it uses rejection sampling, and it's provably exact, but the really clever part is the practical implementation. Which is?

12:23A highly efficient batch approximation method. It turns out you can mimic the effect of temperature almost perfectly by just changing the batch size, n, during sampling. Wait, really? Batch size becomes the temperature knob? Essentially, yes. Want lower temperature, more focused, higher fidelity, less random output. Use a larger batch size, n, when sampling. Want higher temperature, more diverse, creative output. Use a smaller N. And the empirical results were striking. They showed that tweaking the batch size N in Callum produced almost the exact same tradeoff curve between accuracy and diversity as tweaking the temperature T in a standard transformer.

13:01That's huge for usability. People know how temperature works. This gives them a direct analog. Exactly. It makes the control familiar, even though the underlying mechanism is totally different. Okay, so they built this whole new stack. Robust autoencoder, energy loss training, Breyer-Lem evaluation, bash size temperature control. Did it actually deliver on the promise? What about the efficiency results? Let's talk about that K4 setup. This is the payoff, right? The results look really promising. They established what looks like a new, better frontier for the performance versus compute trade-off.

13:30Give us the numbers. Okay, the key comparison. They had a CAREM model, about 371 million parameters. It achieved similar performance to a baseline standard Transformer S model, which was smaller at 281 million parameters. Right, so the CALM model was a bit bigger in parameter count. But, and this is the crucial part, that slightly larger CALM model required 44 % fewer FLPs to train. 44 % fewer. And 34 % fewer FLPs for inference for actually generating text. Wow. Okay, that's a significant saving. Let's compute for the same quality output. Massive saving. It directly validates their core idea, faster training, cheaper inference.

14:11And this really proves that scaling the semantic bandwidth, K packing more meaning into each prediction step, is a powerful new way to optimize these models. It's not just about more parameters or more data anymore. It's a whole new axis on the design graph. And the benefits really kicked in once K was two or more, with K4 clearly beating the baseline efficiency curve. So pulling it all together, COL-M seems to successfully challenge that fundamental token-by-token limitation. It does it with this next-vector prediction idea, but crucially, also with this entire supporting likelihood-free toolkit.

14:44Yeah, the robust VAE, the energy score, Briar-LM, the sampling control. You need the whole package for it to work. So what we've learned today, really, is that the path to more efficient LLMs might not just be about brute force scaling. It might be about being smarter with the data units themselves, making each prediction step count for more. Making them denser, essentially. Right. But this opens up some big questions, doesn't it? Like, we have these scaling laws for traditional LLMs predict performance based on model size, data size. How do those change now that K's semantic bandwidth is a new variable?

15:18That's a huge open question. What do the unified scaling laws look like incorporating K? We don't know yet. And what about other core LLM techniques? Things like reinforcement learning from human feedback. RLHF or knowledge distillation. They often rely heavily on having those explicit log probabilities from the softmax. How do you adapt those algorithms to work entirely in this sample-based likelihood-free world that Calm operates in? That feels like the next major research hurdle. Absolutely. Reformulating those core techniques for this new domain is going to be critical, but the potential payoff, this efficiency gain, it really opens up possibilities.

15:56We're only just beginning to map out. It could make these powerful models much more accessible. We'll definitely be keeping an eye on that map. Thanks for diving deep with us today.

From the publisher

This paper introduces **Continuous Autoregressive Language Models (CALM)**, a new paradigm designed to overcome the efficiency limitations of conventional, token-by-token generation in Large Language Models (LLMs). CALM achieves significant computational savings by employing a robust **autoencoder** to compress a chunk of $K$ discrete tokens into a single, high-fidelity continuous vector, thereby reducing the number of sequential generation steps by a factor of $K$. This shift necessitates a comprehensive **likelihood-free framework**, including an **energy loss** for generative modeling and a new evaluation metric called **BrierLM**, which offers a reliable alternative to Perplexity for implicit models. Furthermore, the paper details a provably exact, but computationally expensive, **likelihood-free temperature sampling algorithm**, along with a highly efficient batch approximation that demonstrates an equivalent trade-off between accuracy and diversity as traditional LLMs. The empirical results confirm that increasing the **semantic bandwidth** $K$ provides a powerful new axis for achieving a superior performance-compute balance in language modeling.

More from Best AI papers explained

All 475 episodes
Continuous Autoregressive Language ModelsBest AI papers explained · 16 min
Listen in VO