Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models

14 Sep 2026 · 24 min · 11 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode explains research from University of Washington and MetaFair (“Breaking the Token Ceiling”) arguing that switching AI distillation from token vocabularies to byte-level vocabularies (UTF-8 bytes) can remove a storage/accuracy ceiling and yield smaller, cheaper models with higher downstream performance.

Guest backgrounds

No guest identities or bios are provided in the transcript; only two hosts speak.

Key claims

Token-based distillation must truncate teacher logits (e.g., keep only top ~600 of ~128k tokens), losing “unlikely but possible” probabilities. Byte models use 256 byte symbols and can store full logits. A “marginalize it” mapping can mathematically distort probabilities; an “end of token” (EOT) symbol (256→257 vocab) enables exact mapping in one forward pass with ~30.94% extra compute.

Notable examples

“Tiramisu problem” shows bits-per-byte/cross-entropy can rate two models equally despite one hallucinating (“There’d be calzone”). Benchmarks cited: HellaSwag, ARC, and FLORES. Results: end-of-token byte model beats token model up to ~4% on average and matches token-model performance using ~1/6 the training data; projected to surpass Llama 3.2 1B and Gemma 3.1BPT on downstream tasks.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Tokenization in AI

1:30 to 2:24

Discover how AI perceives and processes text through tokenization.

“We're exploring a major paradigm-shifting transition from what developers call tokens to raw digital bytes.”

The Limitations of Traditional Token Models

2:24 to 4:10

Examine the constraints of traditional AI token models and their vocabulary.

“So when I type a prompt into my phone, how does the standard AI actually perceive that text right now?”

Byte Models: A New Approach

4:10 to 6:26

Explore the advantages of byte models in AI and their efficiency.

“You can think of a parameter as a digital synapse in the AI's brain.”

Challenges in Training Byte Models

6:26 to 8:10

Understand the complexities involved in teaching AI using byte models.

“Because a byte model only has 256 possible values at any given step, you can save the entire distribution of logits.”

Methods for Knowledge Transfer in AI

8:10 to 11:43

Learn about two distinct methods for transferring knowledge between AI models.

“Which brings us to the massive hurdle the researchers had to clear in this paper.”

Evaluating AI Models: The Race Against Time

11:43 to 13:07

Discover the performance comparison between token and byte models in AI.

“Method two is called end of token, and this represents the exact mathematically perfect translation method.”

Lessons from the AI Benchmarking Race

13:07 to 14:00

Analyze the outcomes of model comparisons and their implications for AI.

“The penalty is steep, but perfectly preserving the teacher's internal logic turns out to be an overwhelming advantage that entirely eclipses that 30 percent tax.”

The Tortoise and Hare of AI Models

14:00 to 16:03

Learn about the performance dynamics between traditional token models and byte models.

“So with the exact probability mapping of the EOT token, the byte model must just immediately crush the traditional token model out of a gate, right?”

Understanding BPB and Its Limitations

16:03 to 20:31

Discover how the bits per byte metric can mislead evaluations of AI intelligence.

“Here's where it gets really interesting to me.”

The Efficiency of End-of-Token Models

20:31 to 22:55

Examine how end-of-token models achieve greater efficiency and intelligence.

“You have to ask, does it give you the tiramisu or does it give you the calzone?”
Show all 11 chapters

The Future of AI Training Paradigms

22:55 to 23:57

Consider the implications of training AI on raw data versus human-designed inputs.

“Which leaves me with one final thought for you to chew on before we go.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00So if you wanted to teach a brand new AI by having it perfectly copy the brain of an older much smarter AI you'd probably assume you just, you know, download its programming, right? Yeah, like a quick drag and drop on your desktop. Exactly. Yeah. But mapping those complex thought processes, like the actual probabilities of how a giant AI makes its decisions, it requires just an unimaginable amount of data. Oh, absolutely. I mean, if you tried to save the full probability map of a massive model today, we are talking about exabytes of data. It's massive. You'd need, like, more hard drives than you could physically fit inside a sprawling warehouse just to store the possibilities of what word it might think of next.

0:40Right, which is completely impractical. But what if changing the fundamental alphabet the AI uses to read, like shrinking its vocabulary from 128 ,000 complex words down to just 256 basic digital letters, could shrink that warehouse-sized problem down to something you could, I don't know, fit on a laptop? Well, we are looking at a fundamental shift in the architecture of artificial intelligence here. By stripping away the human-made training wheels that dictate how these models process language, developers have stumbled onto a mechanism that makes the resulting AI significantly smaller. And cheaper, right.

1:16Dramatically cheaper to train and, honestly, radically stronger in its ultimate performance. Welcome to the D-Jive. Today, we are completely peeling back the curtain on how AI actually reads the text you send it every day. It's a really exciting topic. It really is. We're exploring a major paradigm-shifting transition from what developers call tokens to raw digital bytes. We're going to unpack some fascinating new research out of the University of Washington and MetaFair titled Breaking the Token Ceiling, Distilling Smaller, Stronger Byte Models. And the mission of this deep dive is to demonstrate exactly how altering an AI's core vocabulary changes its ultimate potential.

1:56Right. It explores why forcing a computer to abandon human concepts of language, you know, in favor of its own native digital language, fundamentally alters the mathematical ceiling of what it can achieve. Which is wild to think about. It is. And ultimately, it points to why the technology running on your personal devices is about to leap forward in capability. OK, let's untack this, because to really grasp the breakthrough in this paper, we first have to understand the invisible ceiling they're trying to break through in the first place. The status quo. Right. So when I type a prompt into my phone, how does the standard AI actually perceive that text right now?

2:30So the current standard across the industry is a system called tokenization. OK. AI models, they don't read letter by letter the way a human child learns to read. Instead, they read in tokens. And a token is like what exactly? You can think of a token as a chunk of text. Sometimes it's a whole word. Sometimes it's just a syllable or maybe just a few specific characters grouped together. Okay, that makes sense. And to give you a concrete sense of scale, a widely used standard model right now, like the LLAMA 38B model, it relies on a massive predefined vocabulary of 128 ,256 of these tokens. Over 128 ,000 specific text chunks just sitting in its memory.

3:11Yep. That feels incredibly clunky. It requires a massive embedding layer in the AI's architecture just to hold that dictionary. Wow. Now, contrast that approach with bytes. At the most fundamental hardware level, computers don't know what a word is. Right. They just see code. Exactly. They process information in raw digital bytes, specifically UTF-8 bytes for text. So a byte model completely throws out that bloated 128 ,000 token dictionary and operates using a tiny fixed universal vocabulary of just 256 raw bytes. Because that covers everything. Right. That represents the entire spectrum of digital text.

3:52I mean, if bytes are the native language of the computer anyway, it seems completely counterintuitive that we ever forced them to use human-like token chunks in the first place. You'd think so, yeah. Like, why wouldn't we just use the 256 bytes from the very beginning? Because of the computational burden. We have to talk about parameters for a second to really understand why. Sure. You can think of a parameter as a digital synapse in the AI's brain. It's a weight that determines how it connects one concept to another. Okay. Got it. Serving massive models with billions or even trillions of these parameters is incredibly expensive and, well, slow.

4:27Right. I mean, you simply cannot run a frontier AI model locally on your smartphone. The hardware would literally melt. So to make AI useful for you in the real world, developers use a process called distillation. Distillation. OK, that is the process where you take a massive, expensive teacher AI and you use it to train a smaller, much more efficient student AI. Right. That is the core idea. Instead of spending millions of dollars to train the small student AI from scratch on human text, you train it to mimic the internal logic of the massive teacher. Makes sense. You show the student a sentence and you have it look at how the teacher assigns probabilities to the next possible token.

5:06Right. And here's where we hit that massive data warehouse problem you mentioned earlier. Right. The exapius of data. Exactly. To do distillation perfectly, the student needs to look at the probability assigned to every single one of those 128 ,256 tokens for every single step of the text. For every single word. Every single step. These probability maps are what developers call logits. Logits, okay. Storing the logits for 128 ,000 piece vocabulary across trillions of words of training data, it is astronomically expensive in terms of storage space. So if they can't store the full probability map, like the full logits, how are developers actually distilling models right now?

5:48They must be taking shortcuts. Oh, they absolutely take shortcuts. Yeah. To save hard drive space, developers forcefully truncate the data. They only keep the top logits, meaning they save the probability scores for maybe the top 600 most likely tokens. And the rest. They literally delete the rest of the distribution. Wait, they just throw the rest of the teacher's thought process in the trash. Yep. Gone. But they're losing all that subtle, nuanced information about what the teacher thought was unlikely, but still, you know, possible. They have no choice under the token system. The files are just too big.

6:19But this is exactly where the 256 byte vocabulary comes to the rescue. Oh, because it's so much smaller. Exactly. Because a byte model only has 256 possible values at any given step, you can save the entire distribution of logits. You never have to truncate anything. Never. The student model can map the exact 100 % complete probability brain of the teacher without ever running out of storage space. Okay, let me see if I can visualize this difference. It sounds to me like tokens are basically a massive set of highly specific pre-built Lego structures. I like that. Like you have a whole roof piece, a whole steering wheel, an entire window frame.

6:57That is your 128 ,000 piece token vocabulary. Right. Bytes, on the other hand, are just the basic single stud Lego bricks. There are only 256 colors of them. Building with the big pre-made tokens is faster initially, but storing 128 ,000 custom shapes requires a massive warehouse. So bought on. But if you just use the single stud basic bricks, you can build literally any shape in the universe, and all your pieces fit in one tiny shoebox. That structural flexibility of the byte approach is exactly why it solves the storage problem of distillation. However, taking away those pre-built structures creates a brutal learning curve for the AI.

7:35Oh, interesting. What's fascinating here is that historically, byte models have been incredibly difficult to train. Because they have to build everything from scratch one step at a time. Exactly, because the AI has to connect these tiny raw digital fragments over much longer sequences just to infer the most basic meaning. Oh, I see. A single human word that equals one token might be broken down into five or six separate bytes. The AI has to hold all those tiny fragments in its active memory and calculate the relationships between them. That sounds exhausting. It is akin to trying to read a complex novel while looking through a microscope that only lets you see a fraction of a letter at a time.

8:15Which brings us to the massive hurdle the researchers had to clear in this paper. Right. If byte models are so mathematically stubborn and hard to teach, how do you actually transfer the knowledge from the giant token-based teacher over to the tiny byte-based student? Well, you have to perfectly translate the teacher's chunky token probabilities into the student's granular byte probabilities. And the catch is you have to do this math in a single forward pass. Wait, just a single pass. Why not let the student AI check its math a few times against the teacher to make sure it's getting the translation right?

8:50Because of the computational cost measured in FLOPs. FLOPs. Yeah, floating point operations per second. It is essentially the currency of computing power. If you have to run the massive teacher model multiple times to evaluate different byte combinations for the student, you are spending so many FLOPs that you completely defeat the purpose of trying to build a cheap, efficient model. So it has to be done dynamically, on the fly, in one pass. Exactly. Okay, so you need a fast, cheap mathematical translator between tokens and bytes. How do they pull that off? They experimented with two distinct methods.

9:25The first one is called marginalize it. Marginalize it. So if I had to guess, this is an approximate method where they just round up the token probabilities to the nearest matching byte to save time. Not exactly rounding up, but you are correct that it is an approximation. Right. Marginalized. It calculates the byte probabilities by looking at matching prefixes. Prefixes, like how words start. Right. If the student AI wants to calculate the probability of the very next byte, it looks at all the tokens in the teacher's massive dictionary that happened to start with that same byte sequence, and it just aggregates their probabilities together.

9:58Okay, that seems like a solid, logical way to map it. You just group the identical starting points. It seems logical, but it contains a critical mathematical flaw. Oh, really? Yeah, it can silently drop continuations of words, forcing the math to warp. The paper provides a really vivid example of this failure. Okay, let's hear it. Imagine the AI is reading a sentence and has just processed the letters I and S forming the standard token is. Like he is going to the store. Right. Now the model needs to predict the next byte. The marginalized it method looks to the teacher's tokens that build specifically off that exact IS block.

10:34Okay. It might see the token ISU or ISK and it pulls the probabilities from those. Makes sense. But because it is rigidly locked into grouping by that specific token prefix, it is entirely blind to perfectly valid alternative paths that might branch off differently. For instance, a token that starts with O. I'm lost. Why on earth would the AI predict O immediately after the letters IS? That's not even a word. Well, human language on the internet is rarely constrained to proper dictionary words. Yeah, that is very true. Think about slang or how people stretch words out in a text message. ISU or ISU.

11:14Oh, like someone typing, this is so cool, but they combine the words and drag the letters out. Exactly. Because the marginalize it method misses those weird, non-standard alternative paths that the teacher model actually accounted for. The entire mathematical foundation fractures. Wow. The probabilities of the remaining options get artificially inflated to force the total back up to 100%. So the student model learns a distorted version of reality. Exactly. It's skewed. So marginalized, it is sloppy. It loses the nuance. What was the second method they tried? Method two is called end of token, and this represents the exact mathematically perfect translation method.

11:51Okay. How does that work? To stop the probability mass from leaking out through those weird slang pathways, the researchers expanded the students' tiny-bite vocabulary by exactly one addition. They took it from 256 to 257. Yes. They added a single new piece to the Lego set. A special symbol called EOT, which stands for end of token. During the training process, the researchers injected this EOT symbol immediately after every single token chunk in the training text. Wait, I have to challenge this. Okay, go ahead. If you are injecting a brand new special token, an artificial symbol, after almost every single word fragment in your data set, you are massively bloating the text.

12:30Doesn't that defeat the purpose? Doesn't that make the model dramatically slower and more expensive to train? You were hitting on the exact tradeoff the researchers had to weigh. Right. Injecting the EOT token means you were spending roughly 30.94 % more compute. Wow. 30%. Yes. Spending more of those precious FLOPs on the exact same amount of text data, you are adding an extra symbol every four and a half bytes on average. The text is literally longer. A 30 % compute penalty, just to ensure the probability math is perfectly accurate, It seems like a massive handicap in a race for efficiency. The penalty is steep, but perfectly preserving the teacher's internal logic turns out to be an overwhelming advantage that entirely eclipses that 30 percent tax.

13:17Really? Yeah. To prove it, the researchers set up a massive high stakes marathon race between the different architectures. OK, so they pit the traditional token model against the sloppy, marginalized at bite model. Yeah. And the precise end of token bite model. Exactly. And we should clarify for you listening, they are not testing these theories on tiny little toy models. No, they scaled this up to industry standards. They trained architectures with roughly 1.28 billion parameters, and they fed them an ocean of data up to 1 trillion bytes of text. Wow. And to find out which architecture actually created a smarter AI, they ran them through eight rigorous industry benchmarks.

13:53Benchmarks like what? Give me some examples. Things like Hellaswag, which tests common sense sentence completion. Okay. The ARC dataset, which is built on complex grade school science questions, and Flores, which tests complex machine translation from languages like English to German. So with the exact probability mapping of the EOT token, the byte model must just immediately crush the traditional token model out of a gate, right? It's just inherently smarter. Actually, no. The results play out as a classic tortoise and hare scenario. Oh, really? Yeah. At the beginning of the race, under low compute budgets when the models haven't been training for very long, the traditional token 1B model is absolutely the hair.

14:34Okay. It sprints out to a massive lead and vackly outperforms the byte models across almost every benchmark. Ah, because of the Lego analogy. The token model is using the big pre-built structures. Exactly. It doesn't have to figure out how the tiny byte studs snap together. It just grabs the pre-made dictionary words and runs. The token model violently leverages its massive vocabulary for early, cheap games, but this introduces a vital concept in AI development known as scaling laws. Let's define that, because scaling laws dictate almost everything happening in tech right now. Definitely. Scaling laws represent the general rule, often called the bitter lesson of AI, that if you throw exponentially more compute power and exponentially more data at a model, its intelligence will scale up in a predictable line.

15:21Right. More data equals more smarts. Usually, yes. However, the token model reveals an invisible ceiling to this law. As the compute budget scales up, the token model hits a wall. It plateaus. Yes. Its performance completely plateaus. Adding more data stops making it smarter. And the bite models. Yeah. The tortoises in our race. They start out looking foolish, struggling with the tiny building blocks, but their learning curve improves at a much steeper, relentless rate. They just keep going. They don't hit the wall, they just keep climbing. Eventually, the end-of-token distilled model entirely blows past the plateaued token model, raising the ultimate performance ceiling by up to 4 % on average downstream tasks.

16:03Here's where it gets really interesting to me. The end-of-token byte model achieves this higher intelligence ceiling despite technically having slightly fewer total parameters than the token model it beat. Because it doesn't need to dedicate millions of parameters to that massive 128 ,000-piece embedding dictionary at the front of its brain. That is so cool. It is doing more complex reasoning with a smaller overall physical footprint, simply by fundamentally changing the alphabet it uses to perceive the data. It is the undisputed champion of the marathon. But the research points out a massive trap in how the tech industry actually decides who won these races, right?

16:42There's an illusion in the metrics. Yes. This is a crucial takeaway for anyone trying to decipher AI marketing right now. During training, developers rely heavily on a standard metric called bits per byte, or BPB. It is designed to measure how surprised the AI is by the data it is seeing. So a lower BPB score means the AI is less surprised, which implies it is predicting the text highly accurately. Right. And if it's predicting accurately, we assume it must be highly intelligent in the real world. That is the assumption. But this paper proves that a lower BPB validation score does not guarantee the model is actually capable of performing real-world tasks.

17:19Really? You can have an incredible BPB score and still fail spectacularly at answering a basic multiple-choice question. How does that make any sense? If it's predicting the correct text, why would it fail the test? Because of the underlying math of how BPB is calculated, BPB relies on an equation called cross-entropy loss. Cross-entropy loss. which is mathematically designed to look only at the target word, the correct answer. It completely ignores how the AI evaluates the wrong answers. That seems like a flaw. It is like a math teacher who only grades you on whether you wrote down the right final number, but doesn't care that your scratch pad is filled with insane illogical equations that make no sense.

18:00Okay, I think we need a concrete example for this one to really lock it in. The researchers illustrate this beautifully with a scenario we can call the tiramisu problem. No, I like the sound of that. Imagine we ask two different AI models to complete the simple phrase, where's my tiramisu? A vital, urgent question for any dessert lover. Naturally. Now imagine both Model A and Model B evaluate the sentence, and they both assign exactly a 40 % probability to the correct sequence of tokens that spells, where's my tiramisu? Okay, so they both give it a 40 % chance. Right, because they both assign that identical 40 % chance to the right answer, the cross-entropy loss math gives them the exact same bits per byte score.

18:42On a technical benchmark report, they look like identical twins. They both get an A - from the math teacher. So where's the trap? The trap lies in the remaining 60 % of the probability mass. Model A is smart. It puts the correct tokens at the very top of its list, rank number one. So when you deploy Model A and ask it to speak, it confidently outputs, where's my tiramisu? And what is Model B doing with its remaining 60 %? Model B has a severe hallucination problem lurking in its probabilities. At every single step of generating the sentence, Model B assigns a massive 50 % probability to a completely incorrect distractor token.

19:23Wait, 50 %? Yes. And mathematically, 50 % is higher than the 40 % it gave to the correct answer. Oh, so the correct answer gets pushed down to rank number two. Exactly. So when you ask Model B to actually generate the text, it spits out its highest rank tokens, resulting in absolute gibberish. It outputs, there'd be calzone. There'd be calzone instead of, where's my tiramisu? Exactly. Model A gives you tiramisu. Model B gives you calzone. Wow. One is a perfectly coherent AI. The other is completely broken in a practical setting. That's crazy. But because the BPB metric only looks at the 40 % assigned to the target answer, it is completely blind to the 50 % assigned to the calzone distractor.

20:01BPB rates the broken model and the perfect model as equally intelligent. That is a staggering blind spot for an industry that constantly throws benchmark numbers at us to prove their new AI is the best. It proves that you cannot compare models across different tokenization schemes. You know, you cannot compare a token model to a byte model relying on validation metrics like cross-entropy loss and BPB. Right. The math just hides the errors. You have to look at the downstream task performance. You have to ask, does it give you the tiramisu or does it give you the calzone? Right. And when you look at the real-world performance, the final efficiency numbers for this end-of-token byte model are frankly hard to believe.

20:43The efficiencies are staggering. Because the end-of-token models possess that higher mathematical ceiling and learn so efficiently, they end up matching the real-world performance of traditional token models using only one-sixth of the training data. Wait, really? One-sixth the data to reach the exact same level of intelligence? Yes. And remember the massive data warehouse problem we talked about at the beginning regarding distillation? With the exhibits, yeah. Because they map a tiny universe of 256 bytes instead of a sprawling galaxy of 128 ,000 tokens, they slash the logit storage costs down to roughly one-fifth.

21:18So it is using a fraction of the data, it takes up a fraction of the hard drive space, and it ultimately grows smarter than the old architectures. That's the power of bytes. In fact, the data shows that this distilled end-of-token 1 billion parameter model is projected to asymptotically surpass major industry heavyweights that are out right now. It really is. Like it is tracking to beat the LAMA 3.21B model and the GEMMA 3.1BPT models on average downstream tasks. It demonstrates that this byte-level distillation is not like a neat laboratory trick. It is a highly practical, inevitable blueprint for the next generation of artificial intelligence.

21:55So what does this all mean for you listening to this right now? Yeah. It means we need to change how we evaluate the tools we use. Excellent. It means looking past flashy marketing metrics like bits per byte, which can be easily manipulated or just fundamentally blind, and focusing purely on whether the tool actually solves your problem. Well said. But on a deeply practical level, it means the AI running locally on your phone, in your car, or on your laptop is about to get significantly smarter without draining your battery in 10 minutes or eating up all your local storage space. Going back to the absolute most fundamental building blocks, you know, the raw digital atoms, yielded the greatest leap forward in efficiency.

22:38We are moving from clunky human-made text chunks to raw digital bytes. By using that clever end-of-token hack, these researchers figured out how to perfectly copy a massive AI's brain into a tiny, hyper-efficient student model, shattering the old performance ceiling in the process. It's an incredible step forward. It really is. Which leaves me with one final thought for you to chew on before we go. If forcing an AI to abandon human-designed text tokens in favor of raw digital bytes makes it this much stronger and more data-efficient with language, what happens when we apply this exact same byte-level philosophy to other mediums?

23:16Now that is where this technology gets truly disruptive. What happens when we train an AI on the raw audio waves of a symphony instead of feeding it human-transcribed sheet music? Wow, yeah. What happens when we train it on the raw digital pixels of a video feed without giving it human tags and descriptive labels to hold its hand? Will the AI of the future bypass our human concepts entirely to understand the world in a way we can't even perceive? Taking those training whales off doesn't just change how the machine reads. It might fundamentally alter how it perceives reality. Something to think about the next time your phone autocompletes your text message.

23:51Thanks for joining us on this deep dive. Keep questioning the technology behind the curtain and we'll see you next time.

From the publisher

This research introduces Marginalize-It and End-Of-Token, two novel methods for efficiently distilling large token-based language models into smaller, more capable byte-level models. By evaluating dense transformers across various compute budgets, the study reveals that while token models perform better with limited resources, byte models achieve a significantly higher performance ceiling as training data increases. The End-Of-Token approach proves particularly effective, as it preserves the teacher's original probability distribution and demonstrates superior data efficiency by matching token-model accuracy with only one-sixth of the training data. These byte-level architectures also provide a five-fold reduction in logit storage costs because they operate on a much smaller vocabulary of roughly 256 values. Scaling laws developed in the paper predict that these distilled byte models will asymptotically outperform prominent open-weight models like Llama 3.2-1B and Gemma 2B. Ultimately, the work suggests that moving beyond traditional tokenization can "break the token ceiling" to create smaller models with greater long-term potential.

More from Best AI papers explained

All 475 episodes
Breaking the Token Ceiling: Distilling Smaller, Stronger Byte ModelsBest AI papers explained · 24 min
Listen in VO