In short
The episode explains tokenization as a foundational barrier in LLMs and introduces BOLMO, a byte-level “tokenization-free” (UTF-8 bytes) open LLM family that removes subword tokenization issues via “bytefication” (conversion from a subword model).
Guests
No specific guests are named; it’s a host-led discussion.
Guest backgrounds
Not applicable.
Key claims
Subword tokenization causes (1) lost character detail, (2) tokenization bias from future-dependent boundaries that breaks causal generation, (3) vocabulary bottlenecks for other languages/special domains, and (4) inefficient compute allocation. BOLMO uses 256 UTF-8 bytes and a latent tokenizer; it matches prior models efficiently by converting ULMO3 with <1% of typical training cost using two-stage distillation and end-to-end training, including a non-causal boundary predictor with 1-byte lookahead.
Notable examples
“hello, war” boundary bias; “butterfly” boundary ambiguity resolved by 1-byte lookahead. Results: BOLMO-7B beats prior byte models, +16.5% on STEM, and surpasses ULMO3 on character benchmarks; supports adjustable compression/efficiency and “task arithmetic” weight merging for instruction skills without extra training.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding the Tokenization Barrier
0:45 to 1:32
Explaining the foundational constraint of tokenization in LLMs.
“And they're specifically engineered to get rid of that tokenization barrier entirely.”
Problems with Subword Tokenization
1:32 to 4:06
Discussing the four major issues caused by subword tokenization in language models.
“This means they segment text not into full words or single characters, but into common chunks like ing or un.”
Shifting to Byte-Level Models
4:06 to 4:47
Exploring the transition from subword to byte-level models and its implications.
“And then finally, you have compute allocation.”
The Breakthrough: Bytefication
4:47 to 7:06
Detailing the innovative bytefication process and its efficiency benefits.
“But that comes at a huge cost, which is making that fourth problem, P4, way worse.”
The Two-Stage Training Procedure
7:06 to 9:24
Breaking down the two-stage training procedure for the new byte-level model.
“So they're teaching the new parts to speak the same language as the old big brain of the model.”
Performance Results of BOLMO
9:24 to 11:01
Analyzing the results of BOLMO and its advantages over previous models.
“It has to guess where to make the next patch.”
Flexibility and Practical Benefits
11:01 to 13:14
Discussing the practical wins and flexibility offered by byte-level models.
“And because it avoids the bottleneck, it can let you, the user, control the tradeoff.”
Transcript
Automatic transcript. May contain errors.0:00Welcome back to the Deep Dive. If you're tracking the world of large language models, you know, the breakthroughs feel like they're happening almost daily. But underneath all of that, there's this foundational constraint. An invisible barrier, really. Exactly. An invisible barrier that has shaped how every single modern LLM sees and processes information. And that barrier is tokenization. It's the very first step. You can't just feed raw text into a model. You have to, you know, break it up first. And the way we've been doing that for years has created some really deep systemic problems. And today we're plunging into a solution from a pretty major collaboration teams from the Allen Institute for AI, University of Cambridge, University of Washington, and the University of Edinburgh.
0:44They've introduced a new family of open byte level LLMs called BOMO. And they're specifically engineered to get rid of that tokenization barrier entirely. Right. So the shift from what we call subwords to byte level input. It's a real paradigm change. It is. It totally alters the model's relationship with language. And it opens up some really exciting new doors for like efficiency and flexibility. OK, let's unpack this. Our mission for this deep dive is to figure out what a tokenization free LLM actually is, what problems it solve. and this is the most important part, the incredibly efficient process they call bidatification that made this whole thing possible without spending billions training from scratch.
1:28To really get the solution, you have to understand the problem. So today's LLMs, they mostly rely on something called subword tokenization. This means they segment text not into full words or single characters, but into common chunks like ing or un. These subwords make up a fixed vocabulary, usually somewhere between 30 ,000 and 300 ,000 tokens. And each one is just represented by a number, an integer ID. Exactly. It sounds like a pretty good compression technique on the surface. You're just shrinking the input before the model even reads it. Yeah. But the researchers, they point to four big problems with this.
2:02What's the first one? So the first one, they call it insufficient character understanding. Basically, it's a loss of detail. If the model sees the word television as just one single token ID, all the information about the individual characters inside that word is just gone. It's lost before it even gets to the big transformer layers. Right. And that's why these models can be so bad at tasks that need character knowledge, like, you know, complex spelling or spotting typos. That makes a lot of sense. You compress the data, you lose the fine grained detail. But the second issue, tokenization bias.
2:36That's where it gets, for me, a little unsettling. Yeah, this one's tricky. It's an artifact of the process you just can't avoid. Subword tokenizers have to look at the surrounding context, including what comes after the current word, to decide where to put the boundaries. And the example they give is really clear. It is. If a model sees the token sequence, hello, war, it already has a hint that the letters old are not going to come next. Why? Because if old were coming next, the tokenizer would have almost certainly just picked the single, much more common token world. Hold on. So if the whole point is compression, isn't that leak of information, that little hint about the future, isn't that just a side effect?
3:13Like a feature, not a bug. You'd think so, but it creates a real paradox when the model is generating text. The model is supposed to be causal. It's supposed to predict the next thing based only on the past. I see. But if the token boundary itself depends on the future, then the model is getting information it hasn't really earned yet. It violates that causal structure. So it builds this dependency that just doesn't hold up in the real world when you're generating text one word at a time. Precisely. And then there's the third issue, the vocabulary bottleneck. That fixed vocabulary of, say, 50 ,000 subwords is super rigid.
3:49It's usually built for English, so it really struggles with other languages or specialized fields like medicine where new words are popping up all the time. And if it sees a word that's not in its dictionary. It has to break it down into smaller, sometimes meaningless pieces. It's inefficient and it can dilute the meaning. And then finally, you have compute allocation. Right. P4. Yeah. A standard transformer spends the exact same amount of compute on every single token. It doesn't matter if it's a simple comma or a really complex word, and that's just suboptimal. Okay, so the answer to those first three problems, the character loss, the bias, the vocabulary, seems to be shifting to byte-level models.
4:30Yeah. Instead of 30 ,000 subwords, you just use the 256 basic building blocks of all digital text, UTF-8 bytes. That's the clean slate. You switch to 256 bytes and boom. P1, P2, and P3 are solved. You get full character info, no tokenization bias, and a totally flexible vocabulary for any language. But that comes at a huge cost, which is making that fourth problem, P4, way worse. A single subword token might suddenly become, what, four or five bytes? At least. Your sequence length can easily quadruple. And since compute cost scales really badly with sequence length, training a pure byte-level model becomes just prohibitively expensive.
5:10That's why they haven't really taken off before now. Before we get to the breakthrough, let's just nail the terminology. We're saying tokenization-free, but that's not strictly accurate, is it? No, not really. It's more accurate to say the UTF-8 standard is the tokenizer. The vocabulary is just that set of 256 bytes. And BOLMO is part of a special class of models they call latent tokenizer language models, or LTLMs. Latent tokenizer. So the tokenization happens inside the model itself. Exactly. To handle that sequence length problem, they couldn't just feed raw, super long byte sequences to the transformer.
5:42They needed a middle ground. So the latent tokenizer basically bundles those bytes into efficient little patches internally. Okay, which brings us to the breakthrough. Because with trillions of tokens and tens of millions of dollars needed to train a big model from scratch, starting over was just not an option. Not at all. So the Bulma team did something really revolutionary. They created by edification. And by edification isn't training from scratch. It's a conversion process. It's the whole key. Instead of a blank slate, they took a really successful subword model, ULMO3, and created a process to convert it into a byte-level LTLM.
6:16We have to stop on the efficiency here for a second. The bytefication process, the whole conversion, took less than 1 % of a typical pre-training budget. Less than 1%. They used just 39.3 billion tokens. Compare that to the trillions of tokens that go into the original model. It's a tiny fraction. That is, that's astounding. It completely changes the economics of this. It makes byte-level models actually seem viable. And it all works because of this really clever two-stage training procedure that leverages all the knowledge that's already in the original ULMO 3 model. Okay, let's walk through it.
6:50Stage one is called subword to bite distillation. What's the goal here? The goal is just mimicry. They need to train the new local bite processing parts, an encoder, a decoder, and this crucial boundary predictor, to perfectly recover the behavior of the original subword model. The old model acts as the teacher. So they're teaching the new parts to speak the same language as the old big brain of the model. And how do they make it so fast and cheap? The key is that during stage one, the massive part of the model, the global model, the transformer backbone, is completely frozen. Ah, OK. All that expensive pre-learned intelligence about the world is locked down.
7:28They don't touch it. They only train the small local parts that are seeing bytes for the first time. And part of that is teaching the new model where the old subword boundaries would have been. how does that boundary predictor learn so fast? It's trained to just emulate what the original subword tokenizer did. It's a pretty simple prediction task, and it gets over 99 % accuracy really, really quickly. So once those local parts are perfectly mimicking the old model, why not just stop there? What's the risk if you don't move on to stage two? Because the model is still fundamentally thinking like a subword model, The global model hasn't been allowed to actually use any of that new, rich, byte-level information it's now getting.
8:08It's still operating on the old assumptions. And that's the cue for stage two, end-to-end training. Right. Now, in stage two, they unfreeze everything. All the parameters, including that huge global model, are optimized together. This is where the whole system learns to be a true, end-to-end, byte-level model. The global part starts actually using the character details to get better. But this whole thing only works because they solve that critical mismatch we talked about earlier. The fact that tokenizers cheat by looking into the future. This is the absolute key innovation. It's the non-pausal patch boundary prediction.
8:43Right. Prior attempts at this were stuck with causal boundary predictors. They could only look backward. But subword tokenizers use future context. There was a mismatch in what they could do. And Bulmo fixed this just by letting the boundary predictor peek ahead a little bit. A tiny bit. During the pre-fill stage, when it's just reading your prompt, they let the boundary predictor see up to one byte of future context. Just one byte. And that tiny look ahead is enough to resolve the expressivity mismatch that held back all the previous byte-level models. That one byte makes all the difference. It's huge.
9:19Think about the word butterfly. A causal predictor gets to their R in butter and has no idea what's coming. Is it buttress? Butterfly. It has to guess where to make the next patch. But if it can peek ahead at the next single byte, the F, it gets immediate context. It knows the next logical unit is starting, and it can place a much more semantically coherent boundary. It resolves all that ambiguity. And that's what makes the whole thing competitive. So let's talk results. How did it actually perform? The results are impressive. Balmo 7b substantially outperformed all previous public byte-level models.
9:55And in really complex areas like STEM tasks, it saw a plus 16.5 % absolute improvement over the next best byte model. It successfully got very close to the overall performance of its source model, ULMO 3. But it completely crushed it on character level stuff, right? The cute benchmark. Oh, it wasn't even close. It vastly surpassed its own source model. And that makes sense, right? The byte level architecture is just naturally built for that kind of fine grained skill. The character information is never lost in the first place. Beyond just the raw performance, though, one of the biggest practical wins here is this idea of unbounded flexibility.
10:30It solves that tradeoff between efficiency and performance. Yeah, this brings us back to the Softmax bottleneck. Subword models have this huge computational cost. For every single prediction, they have to calculate the probabilities across their entire vocabulary. Maybe 300 ,000 possible NEX tokens. That calculation is massive, and it's fixed. It's a huge overhead. You're always paying that computational tax. Balmo, on the other hand, is only ever calculating probabilities across the 256 possible bytes. It completely avoids that bottleneck. And because it avoids the bottleneck, it can let you, the user, control the tradeoff.
11:09Exactly. You can speed the model up by telling it to use a higher compression ratio, basically, to bundle more bytes into each internal patch. That makes the sequence shorter for the main transformer. So it's like choosing between a 4K video stream, super high quality but slow to load, and a 480p stream that's fast but less detailed. You get to pick the balance. That's a perfect analogy. And a subboard model is locked into one resolution. WOMO gives you a dial you can turn for your specific need, and they show it can even become more efficient than the original subboard model. And finally, the last big win is in the ecosystem with something called task arithmetic.
11:44Right. This is about not having to reinvent the wheel. They showed they could take specialized instruction following skills that were trained for the subword ULMO 3 model and just merge them directly into Bulmo's global model. Like a software update. And they could do that with zero extra training. Often with zero extra training costs. You're just combining the model weights. This lets byte level development instantly benefit from all the fine tuning work that's already been done in the subword world. It's a massive shortcut. So what does this all mean? What's the big takeaway for someone building with or using these models?
12:17I think it means that bitification makes bite-level LMs a truly practical and competitive choice now. It solves the economic problem and the architectural problem. They can match or even beat the state of the art, especially on character tasks, while giving you so much more flexibility on the efficiency side. It sounds like the bite-level future is actually here. And the revolutionary part is that we didn't have to start from scratch to get it. That 1 % investment is the real game changer. Exactly. And, you know, here's a final thought for you to mull over. The team found that the biggest remaining performance gap to the original model came from that non-causal predictor still not being 100 % perfect.
12:55So, since Bulmo only uses one single byte of look ahead, the next frontier might be exploring what happens if models get unrestricted look ahead into the future techs. Or maybe even a dynamic amount of look ahead? that might be what it takes to fully close that gap and unlock the absolute maximum potential of this architecture.
From the publisher
We discuss Bolmo, a groundbreaking family of byte-level language models by AI2 that offers a practical alternative to traditional subword-based tokenization. Developed by the Allen Institute for AI and collaborating universities, these models achieve state-of-the-art performance by "byteifying" existing subword models like OLMo. This innovative process uses a specialized two-stage distillation procedure to convert subword models into byte-level ones using less than 1% of the original pretraining budget. Architecturally, Bolmo features a non-causal boundary predictor and local mLSTM layers to resolve efficiency and character-understanding limitations inherent in previous systems. The research demonstrates that Bolmo effectively matches or exceeds the performance of its source models in coding and character-based tasks. Furthermore, the authors show that Bolmo can be further optimized for speed and easily post-trained using existing subword ecosystems via task arithmetic.




