Dynamic Token Merging for Efficient Byte-level Language Models with Julie Kallini - #724

24 Mar 2025 · 51 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Episode Notes: Dynamic Token Merging for Efficient Byte-level Language Models with Julie Kallini - #724

Podcast Overview Podcast Title: The TWIML AI Podcast Host: Sam Charrington Guest: Julie Kallini, PhD student at Stanford University Episode Focus: Discussion on Julie's research papers, particularly on "MrT5" and "Mission: Impossible Language Models."

Episode Description In this episode, Julie Kallini discusses her work on enabling efficient byte-level language models through dynamic token merging and explores the implications of language model biases toward natural languages.

---

Key Concepts and Discussions

Tokenization in Language Models

  • Definition: Tokenization is the preprocessing step that divides text into manageable units, called tokens (usually words or subwords), before they are inputted into language models.
  • Importance: Tokenization impacts the efficiency of language models, especially in terms of compression rates and computational costs.

Issues with Traditional Tokenization

  • Inefficient Compression for Low-Resource Languages:
  • High-resource languages (like English) are tokenized efficiently, but low-resource languages can result in disproportionately high token counts, leading to user overcharges (e.g., in language model APIs that charge by token).
  • Example: An English sentence can be tokenized into 10 tokens, while its Arabic counterpart might explode to 31 tokens.
  • Sensitivity to Character Manipulation:
  • Errors in spelling can drastically affect tokenization, potentially leading to misinterpretations by models.

Byte-Level Models

  • Byte-level vs. Subword Tokenization:
  • Traditional models use subword tokenization, which can cause inefficiencies in certain languages.
  • Byte-level models process raw characters or bytes without a prior tokenization step, addressing some inefficiencies but introducing longer sequence lengths.

MrT5 Model

  • Objective: MrT5 aims to enhance the ByteT5 architecture by making it more efficient through dynamic token merging.
  • Dynamic Token Merging:
  • The model learns which tokens can be dropped as it processes input, allowing for a more compact representation and improved computational efficiency.
  • A gating mechanism is employed to decide which tokens to retain or drop.

Performance and Results

  • Efficiency Gains:
  • MrT5 can compress sequences by up to 50% while maintaining performance comparable to ByteT5 on multilingual tasks.
  • Benchmarks:
  • Fine-tuned on tasks like XNLI (natural language inference) and TideIQA (question answering), showing reduced sequence lengths and improved runtime efficiency.

Mission

Impossible Language Models

  • Research Context:
  • Inspired by Noam Chomsky's op-ed discussing whether language models could learn impossible languages (languages humans cannot learn).
  • Defining Impossible Languages:
  • Impossible languages are characterized by complex structures that challenge both human and machine learning abilities (e.g., random word scrambling).
  • Findings:
  • Language models, particularly GPT-2, demonstrate a bias toward natural language structures, struggling with impossible languages.

Future Directions

  • Exploring Architecture Improvements:
  • Julie is keen on continuing research into architectures that enhance efficiency and reduce biases in language models.
  • Potential Applications:
  • There are many applications for more efficient encoder models, particularly in classification tasks.

---

Key Takeaways

  • Tokenization plays a crucial role in the efficiency of language models, especially concerning low-resource languages.
  • The MrT5 model proposes a novel approach to token merging which enhances computational efficiency without sacrificing performance.
  • Research on impossible languages reveals biases in existing language models, prompting further exploration into model architecture and training methods.

Closing Thoughts Julie Kallini's work opens up new avenues for improving language model efficiency and understanding language processing mechanisms, particularly in multilingual contexts. The conversation emphasizes the importance of research in enhancing fairness and efficiency across diverse languages in machine learning applications.

For complete show notes and further details, visit the [TWIML AI Podcast website](https://twimlai.com/go/724).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Probably the more important reason tokenization is kind of flawed is that there are different compression rates for different languages and scripts. So like high resource languages like English are totally fine. You know, on average, a token is like maybe four or five characters or approximately a word. But for other languages, the same sentence could be tokenized into so many tokens. Users who speak those languages when they interact with language model APIs where users are charged per token, they'll be overcharged basically.

0:42All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. Today, I'm joined by Julie Collini. Julie is a PhD student at Stanford University. Before we get started, be sure to hit that subscribe button wherever you're listening to today's show. Julie, welcome to the podcast. Thank you so much for having me, Sam. I'm looking forward to jumping into our conversation. I came across a couple of your papers recently that I'd like to dig into. We'll be focusing primarily on the most recent of the two called Mr. T5. Love that name. Dynamic Token Merging for Efficient Byte-Level Language Models.

1:21But you've also got a really interesting paper called Mission Impossible Language Models that time permitting we'll touch on. So before we get going, I'd love to have you share a little bit about your background and how you got started in the field. Yeah. Thanks so much for the introduction. So I kind of just starting from where I got into computer science, I really just loved math and science in high school. when it came around like junior year. I think it was my sister that told me it's time to start thinking about what you're going to major in in college, kiddo. So she suggested I think about computer science just because it's really like applied math, like taking math and applying it to making computers work.

2:09I started to learn a little bit of programming on my own and I was like, oh, this actually seems pretty fun. Jumped to college and I took my first computer science course, like actual structured computer science course. And, you know, it was a little bit daunting because I think like most of my peers were not in the position of taking their first computer science course, but I stuck through it. I loved it. And yeah, it seemed to all work out. And in terms of actually getting into the field of NLP or natural language processing, I think throughout college, I kind of wasn't sure about what topic within computer science was really my focus.

2:50Like I thought I could have been a systems person or a theory person for a long time. But then the thing that brought me toward machine learning was taking my first linguistics class. And I really enjoyed looking at language in a different way that I hadn't before. Like linguistics really approaches studying language as a science as well. And then marrying linguistics and computer science, like natural language processing, is the most natural way to do that or the perfect marriage of the two. So that's what got me interested in NLP. Awesome. Awesome. And what year are you in your program? I'm a second year PhD.

3:29What's your research focus? How do you think about the things that you're interested in? I like to say that what I work on is fitting pop culture references into my paper titles. But if you want a more serious answer, I think there are two strands right now to my research. Right now, I've been really interested in tokenization and byte level models, which is the focus of the Mr. T paper we're going to be talking about. And then another big strand of my research that's a focus of the Mission Impossible paper, as well as follow ups we're doing to that paper, is kind of looking at how language models could help us understand linguistics or cognitive science.

4:14So these are the big strands. You know, it used to be the case where we just want computers to be able to mimic human language in some way. But now that they're so good, the question is, can they help us learn more about language? Well, let's dig into the Mr. T paper. And I'd like to maybe start at the top and talk about tokenization. Why is tokenization so important for large language models? Yeah. Yeah, so tokenization is the preprocessing step that's central to basically every language model that you've heard of these days. So it's the preprocessing step that breaks up text into chunks or units called tokens.

5:01So you can think of these as words or parts of words. And these are the units that we feed into a language model. So you can think of tokenization as a form of compression because it's basically taking a long text sequence and making it into a smaller set of units that, you know, if you think about the transformer, which is the component that's really central to all of the large language models that we know today, it can be very expensive. You know, the term you'll hear is like quadratic complexity in terms of the sequence length for the attention mechanism. So long sequences are pretty inefficient and tokenization compresses it into these smaller units.

5:47However, there are some big problems with tokenization. For one, it can be really sensitive to character manipulations or character level noise. So a simple spelling error can result in a sequence being represented by a completely different set of tokens. um uh it's also i think the the probably the more important reason tokenization is kind of flawed is that uh there are different compression rates for different languages and scripts so like high resource languages like english are totally fine you know on average a token is like maybe four or five characters or approximately a word um but for other languages um uh the same sentence could be tokenized basically into so many tokens, basically operating at a character level.

6:41So then all the problems of efficiency come in and users who speak those languages when they interact with language model APIs where users are charged per token, they'll be overcharged basically. So I think the unfairness aspect is really interesting to me. You mentioned that one of the issues with tokenization is that it's sensitive to character errors. And yet, I think a lot of our experience with LLMs is that we almost get really sloppy with the way we communicate with them because we know they'll figure it out. How do you reconcile those two experiences? Yeah, that's a great question. so intuitively to me when I think about tokenization and how it's sensitive to character errors like I wouldn't necessarily want the like two words that are spelled slightly if a word is spelled slightly differently it can be chunked in a different way or maybe if a letter is capital then that word will get a completely different input representation And then the model basically has to learn that, oh, this word that you meant actually means this other word.

8:01It has to kind of learn that implicitly. I think the models are just trained on a ton of data and trained for very long that maybe these differences don't, on the surface, seem to not matter as much. Um, but I think that it can still result in weird, interesting failure cases in certain points. The second idea that you mentioned is the idea that subword tokenization is less efficient for under-resourced languages. Can you give some examples of how that comes to play? Yeah, for sure. um so um basically like i i have this example um when i give talks about mr t i kind of have this example where i take the gpt4 tokenizer and i have a sentence in english and i have its translation in arabic um and we know that these sentences mean the the same exact thing um because it's literally the same sentence just translated into Arabic.

9:11But the exact number I have for the GPT-4 tokenizer on that language is like the English sentence is tokenized into 10 tokens or so or nine tokens. And then the Arabic sentence is tokenized into 31 tokens, even though the number of characters is there are actually fewer characters in the Arabic sequence. There are a variety of factors that go into this so you know perhaps it's just that the tokenizer is not trained on enough arabic or maybe english is the dominant language but also from a linguistic perspective um there and arabic is is one of those languages from a linguistic perspective um there are some languages where meaningful units are not necessarily adjacent so like Like these things are called like infixes in languages where rather than putting like a prefix or a suffix like we do in English to add some meaning to a word where you put like, let's say, ed to mean the past tense.

10:12You can like shove it in the middle of a word. And now it's no longer like, yeah, now it's like no longer like concatenating something to the end. It's like inserting it in the middle. So like the way the morphology works is actually breaking up the root into pieces. So from a linguistic perspective, I can see why an algorithm that is kind of merging frequent adjacent tokens might not be the best for all languages beyond just having enough data for that language during the tokenizer training. And so the contention in this paper is that a big part of the issue arises from the unit of tokenization.

10:57Did I get that right? Yeah, yeah. Basically that, yes, performing tokenization can have these drawbacks for certain languages. um yeah and the units of tokenization can be very different for for or how how much compression you achieve in different languages uh uh is a big is a big factor and also like um how efficient it's going to be in each language and so one of the distinctions that you call out is between subword tokenization and character level tokenization um and byte tokenization can you talk about the differences there? So most of the main models use subword tokenization, which is what we've talked about.

11:44The alternative would be character level or byte level models. So these characters don't perform a tokenization. Sorry, these language models don't perform a tokenization pre-processing step. They just take in the raw character or byte sequences as the units that the transformer or the language model will operate on. And the benefit here is that, you know, some of the issues that we talked about, like let's say sensitivity to character level manipulations or the model having like more awareness of what characters comprise its tokens. A lot of these issues are addressed by modeling at the character or byte level.

12:31um the issue is that you still get uh you you get these problems of having very long sequence lengths when you're just operating on the raw character or byte streams um so there are different architectures that have kind of um addressed this so um i guess in in the related work section i talk about uh like charformer or canine is google's character level um counterpart counterpart to multilingual BERT. There are different architectures to try and kind of downsample the sequence in a learned way. The focus of our paper is on ByteT5. And the reason we focused on ByteT5 is it had very impressive performance compared to its token-level counterpart, which was multilingual T5.

13:24So on all these benchmarks, it matched or even outperformed MT5. But the main problem was efficiency. It was exactly the architecture of MT5, but operating on bytes. So you get these very long sequence lengths. And our idea was kind of like, how can we take this existing byte level model and make it more efficient, specifically through minimal fine-tuning. And that's where the idea kind of started. But what you're talking about here is modeling directly at the character or byte level as opposed to some other type of tokenization that's more granular. Yes, tokenization specifically as pre-processing before feeding it in to the transformer.

14:18There are methods that some would call soft tokenization, which is like maybe you have you feed in a character or byte sequence, but the transformer kind of learns how to group tokens implicitly via some learned mechanism. So that would be an example would be like the charformer architecture or like another example is hourglass transformer, which we also like partially like replicate in the paper. Someone called those soft tokenization. I'm I'm pro those methods or I would consider that to be separate from like the typical subword tokenization that we know about in in most language models. like your LLAMAs or your GPT, maybe GPT-4.

15:12Yeah, we know it uses a tokenizer, so yes, GPT-4. And so Mr. T5 isn't an alternative approach to tokenization. It is an alternative model architecture that doesn't require tokenization. And it's based on the Byte T5, which preceded it, which had some inefficiencies. Yes, exactly. So actually how this project came about was I was just exploring doing some interpretability work on character level language models. Like what do models learn about words when they're operating at the character level or the byte level? This was early last year during my, still it was during my second quarter of my first year of grad school.

15:59and I was in a meeting with my advisors, Chris Potts and Dan Draski. And I think we just came to a point where we said, why aren't people using these models? Like if there are these benefits to abandoning subword tokenization, like what's preventing them from taking off? And it seems like the main thing is the efficiency aspect. So I geared toward an architecture project that could address that. And how do we characterize that efficiency aspect? In our paper, we look specific. I think it's always helpful to look specifically at like wall clock time. Usually also a flops analysis is important. So like how many floating point operations are done in the model compared to, let's say it's token level counterpart.

16:52part. Yeah, we include both of those analyses in the paper. Yeah, in the end, I think it's always important to include the wall clock time, for example. So we made sure to include those in the paper. And so let's talk a little bit about the architecture and the way you approached it. So the idea of Mr. T5 is to take Byte T5, which Byte T5, they kind of put all of the or most of the parameters in this heavy encoder. So most of the inefficiency comes from processing sequences in this heavy encoder architecture. So just to also take a step back, the T5 architecture is an encoder decoder model. And in ByteT5, the encoder is the massive part.

17:47So the idea that we kind of came up with is, you know, maybe after a couple of contextual layers of the model. So maybe after you've processed this entire byte for character sequence, after you've processed it through a couple of encoder layers, tokens already contain information about other tokens via the attention mechanism. And then the model can decide what tokens it can drop and what tokens can remain to be processed by the rest of the encoder layers. And that's the main idea. And the way we do this is with a gating mechanism that learns to drop tokens and then learns which ones to keep as well.

18:39So this gating mechanism is the one that does the work of deciding which tokens will be removed from the sequence. And the way I like to think of it is during training, Mr. T5 is doing this dropping as an attention masking process. So to describe what attention masking is, it's this process that we use in language models to prevent certain tokens from looking at other tokens during the attention computation. So the typical uses of attention masking are like in a decoder model or an autoregressive model. You don't want preceding tokens to be able to look into the future because that would defeat the purpose of next token prediction.

19:26So you would mask out future tokens. In an encoder model, you might have sequences of different lengths. So when you process them in a batch, you pad it up and you want to mask out the pad tokens because you wouldn't want pad tokens to affect the representations of other tokens. so the way we we do this in mr t5 is basically can the model learn uh which tokens to mask out via a learned attention masking mechanism um and that's what we do during training um basically the model implicit or the model has a learned attention masking uh sort of function uh that that removes tokens from the sequence during training or doesn't allow us or the the tokens that remain to look at the other tokens for most of the encoder layers.

20:19And then during inference is when we actually remove those tokens from the sequence in what we call like a hard deletion mechanism. So those tokens are actually removed from the sequence and the sequence length is resized into a shorter, more compact sequence. When you're talking about tokens here, are these tokens characters or bytes? or these tokens are characters or bytes sorry the terminology is a little bit confusing but whenever i'm talking about byte t5 or mr t5 um the tokens are the the bytes um yeah we we usually just refer to the units we also just refer to the units that the transformer is working with we refer to those as tokens too yeah and so what you are essentially doing with the the dynamic token merging or the deleting of these tokens is an alternate kind of learned compression scheme ultimately.

21:21So, you know, unlike pre-processing into some reduced amount of tokens here, you're doing it via this dropping scheme. Yes, that's exactly right. And the idea of merging is that we allow the model to kind of, you know, the attention mechanism of early layers kind of already does a sort of merging because every unit or every token through attention is a combination of the tokens around it. So information has already been merged to other tokens and maybe some of them could be dropped now because they've merged their information into other tokens in the early layers. So is there an analogy of the multilingual example you gave where you can kind of look at the effective compression rate for English versus some other language?

22:19And have you demonstrated that it's kind of invariant to language or less variant to language or script? Yeah, yeah, that's a great question. So in our main results or in our main experiments, we train Mr. T5 on multilingual data. So sampling like 15 different languages from multilingual C4 for our continued pre-training experiments. And we don't give Mr. T5 a certain prior to compress languages at different compression rates. Like we have this formulation where we can, using this controller algorithm that's detailed in the paper, we can target a specific compression rate if it's desired. Like let's say I want, on average, Mr.

23:10T to compress sequences by 50 percent. And Mr. T will do that. But when I tested on individual languages separately, we actually found that Mr. T learns language specific compression rates. So, for example, Chinese, which already has a very information dense script, like individual characters could mean entire words. That's just how the orthography of Chinese is. um uh mr t had lower compression rates than for uh a lower compression rate for chinese than for other like latin script languages um so it kind of already it notices that chinese is already pretty compressed i'm not going to compress it as much as a language that uses another script that is not as compressed which i thought was super interesting because we're not injecting that prior.

24:08It just learns it implicitly. In terms of the ultimate performance of the model, what benchmarks are you looking at? Yeah. So in the section on downstream fine-tuning in the paper, we fine-tune on some additional tasks. So from the Byte T5 paper, we look at the XNLI and TideIQA tasks. So those are both multilingual benchmarks. um xnli is a class of classification task um so um natural language inference is kind of like you take two sentences and you need to determine whether they entail contradict or or contradict each other or if there's no relationship um tide iqa is a question answering um task so given a passage and a question can can the model retrieve the answer from it um and these are these are both multilingual, so it tests the multilingual capabilities of Mr.

25:08T as well. And we found that Mr. T5 can reduce sequence lengths on those tasks. I think it was up to like 45 or 50 percent, while maintaining close to the same performance as Byte T5. I think even for XNLI, it outperformed byte T5, which was interesting because, you know, you would think that byte T5, since we're fine tuning on top of byte T5, it would present like a bound on how good you can get. But yeah, so Mr. T5 can match the performance while significantly speeding up the model. And we also test on some character level manipulations. So we took two tasks from a character level benchmark, a spelling correction task and a word search task.

26:02So the spelling correction task is just like, given a sentence that contains some spelling error, can you reproduce the sentence, but corrected? And the word search task is given like a random sequence of a bunch of characters, like a bunch of characters or numbers. Can you like find the English word that matches some definition. So subword models are really, really bad at these two tasks. So I would, this comes from a benchmark created by another PhD student in my lab. Her name is Jing Huang, and she evaluated on like subword T5 and compared it to by T5 at the time. This was before Mr. T. And the subword models really struggle on these sorts of character level tasks.

26:54But again, back to our paper, Mr. T5 is able to significantly reduce the sequence lengths and improve the runtime while coming close to matching Byte T5's performance on the task. So you get the benefits of the compression without an effect on the model's performance. And in terms of the inference efficiency that you were going for, how does it compare to ByteT5? Yeah, so the task efficiency will vary depending on how big the encoder sequence lengths are relative to the decoder. So if you have very long encoder sequence lengths and very short decoder sequence lengths, you're going to get the most gains.

27:41So I think we saw the most gains on the XNLI task because the encoder sequences are very long. And then the decoder, it's just a classification task. So you're like outputting a number. So with like a 45%, sorry, with a 50 % compression rate of the sequence. So I cut the encoder sequence length by half. I also got like around a 45 % speed up compared to byte T5 on that task. So that task was probably the best in terms of improving efficiency at around a 50 % compression rate. So this is great. I know the field is very focused on decoder models, but there are plenty of use cases. As someone who came from industry, there are plenty of use cases for encoder models that do classification.

28:35So, yeah, this would be great. Like imagine anyone who is using Byte T5 for some use case and you could, you know, have your inference time. That would just be a great gain, in my opinion. And now your work primarily focuses on a set of enhancements to Byte T5, but like situate the work for us in the broader context, like relative to T5 and other models. Are you giving up a lot for this character level efficiency? Yeah, that's a great question. So I think a great direction to take this work would be to try to see how we could adapt this method for decoder models. I'm not entirely sure what that would look like, but I think that would be a great next step.

29:30Um, and the, yeah, this, while, while this is a particularly an architecture that, um, um, enhances Byte T5, like if we had the resources to train from scratch or train at large scales, uh, maybe the benefits would be even greater, or I bet the benefits would be even greater. One of the results that we have in the new version of the paper is training at a larger 1.2 billion parameter model and seeing how the efficiency gains are even greater at that scale. um so this just makes me think about if we could scale these models up as much as we've scaled up like subword level models um uh then maybe they could maybe we could get back to the to the problem of like why haven't these caught on and maybe it's just we we haven't scaled them up enough to that point i should mention um it would be remiss of me not to mention a new paper um from meta uh the byte latent transformer um that scaled up uh their their they had a particular architecture they scaled up uh byte level models uh to eight billion parameters um and found that it it's matched their yeah training from scratch and it matched their uh um llama llama model i think it was llama 2 that they compared to um and uh yeah i just think that the the field is is starting to go in.

31:03Sounds very promising. And I would love to be able to scale up Mr. T or other architectures to that size and then see if they would scale up just as well. So do you ultimately think that byte level encoding or modeling is going to replace subware token level modeling? I think that it's hard to tell the future, but I think that it's very promising. Not even just the issues that I kind of talked about previously, but maybe there are even sequences that could be compressed more that a token-level model doesn't compress. There are some very predictable sequences that I can imagine I wouldn't want a transformer to spend so much time operating on.

31:53Let's say I give it a sequence that's like, to be or not to be that is the question like i know what that i know what that is and maybe the model knows like that's a very predictable sequence maybe uh spending all the spending like several uh tokens to to process that uh you know whereas maybe a byte level model or another model that uses some sort of down sampling could compress that whole sequence uh into like one unit um I could see how the efficiency gains might even be better than a subword model. And I think these started to be explored more in the Byte-Latent Transformer paper. But yeah, I'd love to see character models like other architectures scaled up, including Mr.

32:43T. The comment about the to be or not to be kind of highlights for me that this work is really about two things. One is kind of the byte level paradigm and the advantages of that relative to subword tokenization. But also, maybe even more importantly to that last point is the idea of dynamic compression as opposed to like static, you know, fixed compression. Yes, absolutely. So dynamic compression is the big benefit there. Other models that maybe do fixed length downsampling, I don't see as much, like maybe every four characters I'm going to chunk into one representation. that still has benefits of being more character aware because maybe you're pooling over representations of characters, but still it's not dynamic in a way that you would get those sorts of benefits of actually being able to compress really long sequences into fewer representations.

33:51I think we've got a few minutes to touch on the Mission Impossible paper. Why don't we start with an overview of that paper and maybe even the origins of that paper, the setting or the conversation that that paper jumped into? Yeah, absolutely. So Mission Impossible was the first paper of my grad school experience, or I guess it's the first paper I wrote in grad school. So I remember I was just starting off and I was talking to my advisor, Chris Potts, about potential first research topics. And we had both read this New York Times op-ed where Noam Chomsky, who's a very famous and very important linguist, talked about language models and whether they have a bearing on studying linguistics.

34:52and his argument in the new york times article uh was that language models kind of they learn too much like they're they're they're too good uh and to the point where they could learn impossible languages which are languages that humans wouldn't be able to learn um and yeah i thought from a research perspective i thought this would be a really a really cool problem to explore um and and just to jump in there yeah maybe you're about to say this the idea behind his argument was that if language models are kind of just pattern matchers and aren't kind of learning things in a human-like way that we can't really extrapolate from them to the way humans learn language is that the core idea yeah yeah i think that's the core idea of um a lot of his critiques of language up to this point this particular point was um yeah the how even how good language models are um is it's detrimental to their use as models of language or as linguistic tools because they couldn't possibly match certain human behaviors just because of how well they learn.

Read the full transcript

36:20Yeah, I think that was his main point in that article. And he cited a paper by Mitchell and Bowers that explored, I think I would consider it one of the first papers to explore impossible languages from a computational lens. And that paper was really cool, but we thought we could expand to more languages and also test on more modern architectures. So rather than the recurrent neural networks that they tested on in that paper, we could also test on transformer based language models that are the core of language models today. So, yeah, I think Chris and I were both really excited about the topic.

37:08He remembers, he tells me he remembers thinking it was maybe too ambitious of a project for a first project, but I remember him saying, I remember him being very positive and very supportive of pursuing this direction. Yeah. Yeah. And so how does one define an impossible language? Oh, that's a very good question. And I don't think that there is a very clear answer still. Truthfully, there is not a very clear answer because the definition of an impossible language is a language that a human wouldn't be able to learn. And it would be very unethical, I think, to try out different languages on like babies.

37:49That's the ideal thing, right? Like all of the languages that we test on, we would want to be able to give to a baby and see if the baby would be able to learn it, but we can't do that for obvious reasons. So in the paper, we kind of take perspectives from linguistic theory, as well as perspectives that's more from the ML side of, let's say, languages that have inherently more entropy and would be more difficult for both a human and a machine learning algorithm. And we try to test on a broad range of languages in the paper. So these go from languages that are very intuitively impossible. So there were lots of languages that involved random shuffling of words within sentences.

38:41So that sounds fun. It sounds, yeah, it sounds intuitively impossible, but we even had to hedge there. Like we included a footnote, you know, we believe that these languages are impossible given the context where we're scrambling like English words. But there are some languages in the world that are called scrambling languages or free word order languages, where basically words can appear in almost any order. But usually there's some other process in the language that disambiguates the meaning in a different way. I think I remember reading that about some Creole languages that they have a tendency to support freer word order.

39:21Yeah, yeah. And I think lots of polysynthetic languages, so those mean languages that have lots of affixes on words, they will put more meaning into the affixes rather than putting meaning into the actual ordering of the words in the sentence, which is different from English. because English, I think a lot of the meaning comes from the structure, the syntactic structure. And you can't scramble words in that way and have people know what you're talking about. Yeah. So the impossible languages that the paper considers, were these collected from the literature? Did you create impossible languages?

40:07languages um where did they come from yeah so uh the languages that we tested um like um most of these were kind of invented by us but inspired by parts of the literature so like we had a set of languages of reverse languages that were actually replicated from the mitchell and bowers paper that um i mentioned before um they're they're i think the main set of languages in the paper that we kind of focus on are these hop languages. And these are inspired a bit by some artificial language learning experiments, I'd say, that have been done on humans about how humans kind of disprefer languages that involve certain counting-based rules.

40:58So in the languages that we have in the paper, basically we take English. So all of our languages involve perturbing English. There are a number of reasons why we go in that direction that I could talk about. But to talk about these hop languages, basically we take verbs and remove any sort of inflection, meaning how you mark tense or number. And we put a marker that signifies tense and number four words later, which, you know, there's nothing that doesn't sound inherently too complex, right? But it's something that no language really does, having this verb inflection be marked by a marker that comes four words after a verb.

41:50And I remember it sounds like it should be pretty easy for a language model. We have some targeted evaluations showing that the model is actually more surprised by those markers or is worse at predicting those markers than in a control condition where the marker appears right next to the verb as it would in the natural English setting. Yeah. So we thought that was pretty interesting. And is the core idea behind the research to demonstrate one way or the other the relative difficulty that language models have with these impossible languages? Yeah. So the main takeaway there was that at least for the class of models we tested so like the gpt2 models we we trained from scratch we had to train uh yeah just to to clarify we had to train all the models from scratch on each impossible language so these models are not pre-trained or they're not um they're not pre-trained on top of like massive english corpora it's they only see their respective impossible languages um there's something about the GPT-2 architecture that biases it toward the natural language, meaning the control languages in our experiments.

43:10And yeah, that's kind of the main takeaway. And our thought was that it has to do, or our hypothesis is that the GPT-2 architecture prefers information locality. And And what this basically means is in language, words that are predictive of each other are often close together. So when we mess with the locality of a language, it kind of is what makes it harder for GPT-2. And we think it comes from the autoregressive language modeling objective, where the language model has to predict the next token given the preceding tokens. And that kind of creates information locality bias in GPT-2 as well. In terms of the training data set for these models, did you define the rules for these impossible languages and then translate from English data set to a data set in the impossible language?

44:15Or did you use some other kind of synthetic generation? That's a great question. Um, so we started off from an English corpus. Uh, this is a baby LM corpus, which is, uh, about a hundred million words, uh, of text. It's supposed to be approximately what, uh, uh, child would hear up to age 12. Um, so, uh, I think it's a, it's a nice corpus if you're trying to do these sorts of language learning experiments. It also is supposed to mimic what a child would encounter during the first 12 years of life. and what we did was we defined these rules that would transform the English corpus into each impossible language so yeah the data is very controlled when we compare each impossible language like it's the it's the same sentences that have just been transformed by different rules yeah so that's how we went about it and then we pre-trained the GPT-2s on each corpus i guess it strikes me that so yeah the the language the language models that we use like gpt2 were they kind of evolved in the context of english um to some degree and so there's like i don't know some kind of selection bias there for english or something and you know so therefore Or, you know, if you were, if your focus was some impossible language, maybe you evolved some other language model architecture that worked better for those languages.

45:50Which, I guess, causes me to reflect on the, you know, the relationship between impossibility and language model architecture from the, you know, from Chomsky's perspective. Like, does the fact that these language models that evolved in this English context, you know, work or don't work for these impossible languages? Like, what does that really mean? Oh, that's a great question. I have to say, I'm really happy with some, with like, I guess the reception of the paper. There has been like lots of follow up work, I think, that tries or test the similar question, but maybe starting from corpora that are non-English, you know, starting from other other base corpora.

46:41um and i think yeah it's definitely a question that could be explored more like if we um incorporate the comparison of different real natural languages versus impossible languages that are derived from each of those natural languages um it's something that just needs to be explored and then the architecture question um that's what we think would be like a really natural next step. How can we find architectures that are more biased toward the natural languages and less biased toward the impossible languages? Because ultimately, all of the parts of GPT-2 are engineering choices, and there's no reason we can't just change them in order to make them more cognitively plausible models.

47:32So yeah, these are great directions and I'm excited that people are working on them more. And in the follow-up, we are working more on the architecture question. That's what we're targeting in Mission Impossible 2.

47:49Continuing the theme of the pop culture reference in the title. Yeah. That means you have like nine left. Oh man. Yeah, nine left. Oh, I wanted to do a paper that had a spin on like Fast and Furious, like Too Fast, Too Furious or something. But it seems like the Almo team took it. They had their sequel to Almo was called Too Almo, Too Furious. Oh, nice. Well, great. I think we covered these papers. Anything else you would like to share about what you're working on or excited about? Yeah, I'm continuing to be very excited in tokenization and very excited in architectures. So I think the kind of the link between the two papers that I talked about is the exploration of kind of what is learnable and what architectures are best for your specific use cases.

48:48For Mr. T, it's obviously what architecture is going to be at best for achieving more efficient bite-level models. For Mission Impossible, I think the clear next step is to explore the architectures that make a learner more or less biased toward natural or impossible language. And yeah, I'm just really excited to do more work on architecture. architecture. I think that these two questions have allowed me to explore that, especially in a world where kind of the standard transformer architecture is very dominant. And I've been very happy to kind of break away from that a bit. Awesome. Awesome. Well, thanks so much, Julie, for sharing a bit about what you've been working on.

49:38Thank you so much, Sam. It was really a pleasure to talk to you. Thank you.

49:47Thank you.

From the publisher

Today, we're joined by Julie Kallini, PhD student at Stanford University to discuss her recent papers, “MrT5: Dynamic Token Merging for Efficient Byte-level Language Models” and “Mission: Impossible Language Models.” For the MrT5 paper, we explore the importance and failings of tokenization in large language models—including inefficient compression rates for under-resourced languages—and dig into byte-level modeling as an alternative. We discuss the architecture of MrT5, its ability to learn language-specific compression rates, its performance on multilingual benchmarks and character-level manipulation tasks, and its performance and efficiency. For the “Mission: Impossible Language Models” paper, we review the core idea behind the research, the definition and creation of impossible languages, the creation of impossible language training datasets, and explore the bias of language model architectures towards natural language.

The complete show notes for this episode can be found at https://twimlai.com/go/724.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
Dynamic Token Merging for Efficient Byte-level Language Models with Julie Kallini - #724The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 51 min
Listen in VO