In short
Podcast Episode Summary: Cognitive Revolution - The Tiny Model Revolution
Podcast Details
- Podcast Title: Latent Space: The AI Engineer Podcast
- Episode Title: The Tiny Model Revolution with Ronen Eldan and Yuanzhi Li of Microsoft Research
- Release Date: July 1, 2023
- Episode Description: This episode focuses on "Tiny Stories," a dataset created to explore the capabilities of small language models, highlighting discussions around reasoning, model architecture, and the implications of data quality in AI.
Key Themes and Concepts
- Introduction to Tiny Stories
- Definition: A dataset of one million children's stories created using GPT-4, designed to reflect natural language while maintaining simplicity.
- Purpose: To explore language model performance, behavior, and mechanisms using models with fewer parameters (1M to 33M).
- Importance of Data Quality
- Data Quality vs. Quantity: The podcast emphasizes the shift towards optimizing for better quality data instead of just increasing the size of models or datasets.
- Examples Cited:
- Falcon-40B outperformed LLaMA-65B due to a cleaner dataset.
- Replit's 2.7B model outperforming OpenAI's larger model through superior finetuning data.
- Insights on Learning and Reasoning
- Model Training: Discusses how smaller models can learn coherently and display reasoning skills, sometimes outperforming larger models trained on broader datasets.
- Learning Pathways: Smaller models demonstrate a more interpretable learning process, akin to how humans learn language and reasoning.
- Role of Model Architecture
- Depth vs. Width: The discussion explores the trade-offs between model depth (number of layers) and width (number of neurons), and how it relates to reasoning capabilities.
- Attention Mechanisms: Notable distinctions between attention heads that focus on distance versus semantic relationships.
- Emergence of Capabilities
- Definition of Emergence: The phenomenon where complex capabilities arise unexpectedly as model size increases. Notably, coherence and reasoning abilities are discussed as emergent properties.
- Comparison with Human Learning: The analogy of children learning language is leveraged to explain how models can develop capabilities through structured training processes.
- Interpretability in Models
- Interpretability Challenges: The difficulty in understanding larger models' internal workings versus smaller models, which yield clearer relationships between neurons and concepts.
- Neurons and Attention Heads: Small models show more interpretable attention heads that can be directly tied to language tasks.
- Future Directions and Questions
- Research Opportunities: The episode concludes with thoughts on the implications of their findings for future AI research, particularly regarding curriculum learning and its potential to improve model performance across various tasks.
- Universal Phenomena: Questions remain about whether behaviors observed in small models will translate to larger models, and what this means for the future of AI development.
Key Takeaways
- The Tiny Stories project illustrates that smaller models can perform complex language tasks effectively when trained on high-quality data.
- Data quality and training methodology can significantly influence model performance, potentially leading to a paradigm shift in AI development.
- Understanding the implications of model structure and architecture might unlock further breakthroughs in AI reasoning and natural language processing.
- There remains a need for more interpretability research to bridge the gap between human understanding and AI functionalities.
Listening Recommendations
- For deeper insights, listeners are encouraged to explore the related papers and discussions surrounding curriculum learning, model architectures, and the implications of data quality in AI research, particularly the "Tiny Stories" paper and the synthetic reasoning task paper titled "Lego."
---
This summary highlights the pivotal discussions and insights presented in the episode, focusing on the transformative potential of smaller, high-quality models in the evolving landscape of artificial intelligence.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Hello, hello, it's Suix again. It's July 1st, 2023. Happy Q3 to those of you who celebrate. It is one day after we launched our blog post and conference, I guess, on the rise of the AI engineer. For those of you who only joined us for the podcast, the Lintin Space newsletter actually preceded the podcast. And yesterday we put out our first newsletter post in a long time. So if you want to just head over to Lintin.space, you can check out the post about why our thesis is about the AI engineer and why we think that this is going to be a growing category and why we're putting on our first conference in October in San Francisco.
0:39Join us if you can. CFP sponsorships and attendee slots are open as of yesterday. So today, it's basically the July 4th weekend. Most people are going to take Monday off and Tuesday is July 4th. So we figured we would do a double sort of podcast swap with some of our favorite AI podcasts that we love and enjoy and wanted to share with you again, especially highlighting some of the things that we liked about their podcast and then they're doing the same for us on their feeds. So today we're featuring Nathan LeBenz of the Cognitive Revolution podcast. They started around the same time as us but then they've just went way harder than us, done twice the number of episodes and have covered way more than us in terms of computer vision, healthcare, investing in tech, safety and policy, curators, influencers, as well as exceptional AI founders, some of whom we also hope to have on our podcast in some time in the future.
1:35And the story that we've picked out or the episode that we've picked out is a recent one, but I think is extremely important as a theme for 2023, which is Tiny Stories from Microsoft Research, led by Ronan Eldon and Yunju Lee. Since they actually published Tiny Stories, they actually also published Phi One, which is a large language model. It's 1.3 billion parameters. It's only trained on 800 A100 hours. So that's about a couple thousand dollars of training. And it scores above 50 % on human eval, which, as we all know from the Replit episode that we did, is an imperfect benchmark. But it is, as far as people are concerned, the industry standard benchmark for code models.
2:19And it's basically performing at an equivalent level of models that are 10 times its size and 100 times its data set. And so that's an interesting phenomenon that basically points to data quality as the new dimension for which model trainers are optimizing for. We've talked about various dimensions of scaling, from scaling the number of parameters to scaling the data set size to scaling the amount of compute that you spend. But this is the first time where it's a little bit harder to talk about this because it's really hard to quantify. But here we're scaling quality for the same amount of length or the same number of tokens that are invested in the training process.
3:00But this interview is about Tiny Stories, which is the earlier work that informed the 5-1 model. And Tiny Stories is very endearing. It's only focused on the reading level of like a three to four year old. but it actually makes a lot of interesting points. The thing that model researchers pick up on is that it uses synthetic data sets that were generated by GPT-3 and 4. But for practitioners, I think the more useful and interesting or mind-blowing thing is that it can generate very coherent stories from a tiny language models. Like I'm talking about less than 10 million parameters. And they have this comparison, which I'm going to stick in the show notes, where they compare a story that was generated by a tiny stories model.
3:43There's 3 million parameters comparing it to GPT-2 XL, which is 1.5 billion parameters models. So it's beating a 500 times larger model because it's focused on this kind of domain. And this has a lot of interesting implications on interpretability because it's a much smaller model. We can visualize what's going on in the weights rather than having a big mass where we don't know anything. but also just for the future of domain-specific models. I also like the analogy that they make in the podcast interview, which is that this kind of training models how humans learn, which is first you learn to speak like a child, and then you learn adult language or to talk like an adult, as opposed to how we train large models today, which is we just throw it in a deep end and just expect it to learn all of Common Crawl and to speak like a professor, to speak like a 4chan troll, to speak like, you know, textbook authors.
4:43And it doesn't really have this progression or formative moment in its training, where it's training on more conceptual core concepts that are day to day. And finally, I think it also has really interesting implications on the discussion around emergence, because emergence typically requires, let's say, in the order of a billion parameters, 20 billion parameters, depending who you ask. But here we have solid reasoning on the order of one to 10 million parameters. Million, not billion. And it's shocking. It is unexpected, I guess. And the authors indulge us a little bit on how model architecture actually enables some of these patterns.
5:26So overall, a fantastic interview, a fantastic piece of work. and the CogGraph podcast is probably one of the only podcasts where this kind of in-depth discussion is available in podcast format so highly highly recommended we hope to continue to be friends with the CogGraph podcast in the future and hopefully collaborate on future issues so see you tomorrow one of the most important abilities for generative models is to be able to speak coherent English. Her mom didn't let her have a dog, so she asked for a... And when you try to autocomplete this, now, you know, the most common noun that you've seen so far, also the most, the proximate one, is dog, not cat.
6:15Dog already appears twice in the sentence. even GPT-2 XL, which has 1.5 billion parameters, its most likely completion is still Doug. I'm just going to read one tiny story just straight out of the paper, because I think that will help people understand what this data set ultimately is. Tom has a big pot of soup. He wants to share it with Jane. Jane takes a spoonful of soup, but then she makes a face. The soup is, that's the prompt, and then you show a completion, very bitter. She does not like it. She says, I don't like this soup. It is too bitter. He looks around the kitchen and finds some bread and cheese.
6:58He puts them on the table and says, here, Jane, you can have some bread and cheese. They are not bitter. They are sweet and yummy. Jane is happy. She says, thank you, Tom. You are a good friend. I like bread and cheese. they are not bitter. Hello and welcome to The Cognitive Revolution, where we interview visionary researchers, entrepreneurs, and builders working on the frontier of artificial intelligence. Each week we'll explore their revolutionary ideas and together we'll build a picture of how AI technology will transform work, life, and society in the coming years. I'm Nathan LeBenz, joined by my co-host Eric Torenberg.
7:34Hello and welcome back to The Cognitive Revolution. Today's episode is great for anyone who really wants to deepen their understanding of and intuition for how language models really work. Certainly, as measured by how much I learned in the course of the conversation, it's one of our very best. Our guests, Ronan Eldon and Yuan-Juli of Microsoft Research, have created a small natural language dataset called Tiny Stories, which they designed to reflect the full richness of natural language while still being small and conceptually simple enough to support research with modest compute budgets. They did this by using GPT-4 to systematically create one million children's stories, using only words that an advanced three-year-old could be expected to know.
8:19Data set in hand, they then began to explore a number of aspects of language model performance, behavior, and mechanism by training a series of models that range in size from just 1 million to a maximum of 33 million parameters, still just 2 % the size of GPT-2. They then use these small models to explore the development of language model reasoning abilities, identifying so-called logical primitives, beginning with a basic understanding of grammar, followed by the learning of facts, and then eventually adding certain logical micro skills such as negation and exclusion. These findings create the perfect context in which to discuss the tricky and often controversial topic of emergence, as well as to compare and contrast how large language models learn with how human children learn, and to explain how the differences that we see across language models and children do in fact make some sense given the different incentive structures in play in each case.
9:21They also did some great interpretability work in this paper, and I really relished the chance to get into all three areas that they explore. First, they look at the trade-offs between the number of layers in a transformer, which to a large extent governs the number of logical leaps that a model can make, with, on the other hand, the width of a layer, which seems to determine how many facts the model can store. They also identify attention heads with distinct roles, including distance heads, which simply reflect the distance between tokens, and which look almost exactly like the alibi scheme, which is now powering long context models such as Claude 100k and the recent Mosaic ML 65k release.
10:04And then on the other hand, semantic attention heads, which focus on meaning. That there should exist such completely different attention heads within a single model, and that an alibi scheme should emerge in the wild is really, to me, mind-blowing. Finally, they examine the role of individual neurons, finding that many of their small model neurons do in fact correspond to human interpretable concepts. We close the conversation by zooming out and discussing why small models are more interpretable than large models, the challenges inherent in attempting to extend this work to larger scale models, and why controlling language models might end up being more like horseback riding than microbiology.
10:46Throughout this conversation, I was really struck by two things. First, it seems to me that we've only scratched the surface of the potential for curriculum learning approaches. I fully expect that we'll start to see ever more sophisticated approaches, which use specific data sets to layer on specific skills in a strategic, progressive manner, creating highly specialized, small-scale models that can solve specific problems extremely efficiently. Second, the value of these toy models for developing understanding really is tremendous. If I could make just one suggestion to listeners, if you want to get the absolute most out of this episode, it would be to visit the Hugging Face website and try playing around with some of the bigger models that they've released.
11:28The very biggest are still only 33 million parameters, which means that they can load easily and run quickly right from the Hugging Face model page. If you do that, as I did in preparation for this episode, you will actually have the chance to explore a lot of the concepts, and you can set up your own little experiments to test the reasoning ability of these models. I guarantee that you will come away with a deeper understanding that you will retain for longer and much better. And if you do find anything interesting, I would love to hear about it. So please do reach out to me via our email, tcr at turpentine.co, or on Twitter, where you can always DM me at LeBenz.
12:08Now, I hope you enjoy this elucidating conversation with Ronan Eldon and Yuan-Ju Lee of Microsoft Research. Ronan Eldon and Yuan-Ju Lee, welcome to the Cognitive Revolution. Thank you so much. We're super happy to be here. You guys have just published this paper called Tiny Stories. And I think it's a really fascinating bit of research on multiple levels. So I'm really kind of excited to dive into it with you guys. It touches on a bunch of different themes, including some of the hot button themes that we'll get to around emergent capabilities and reasoning. And you guys are studying this in a very unique way that makes the problem, I think, more tractable and more approachable, hopefully for our listeners as well.
12:53So I'm really excited to introduce this work to them. Maybe just for starters, can you give me a little bit of an introduction to what inspired the Tiny Stories project? I guess I'm kind of new to LLMs or deep learning in general. I come from pure math. And when I started looking into, you know, architectures, trying to understand what those models are actually doing, how to improve them, etc., etc., I got very quickly, I got very, very frustrated. because it's very easy to come up with ideas, but in order to actually check whether an idea is good, almost always you need to do an experiment that involves a lot of compute.
13:46It's just very, very hard to check things. You need to, you know, you can either train small models which basically don't do much in terms of, you know, they don't actually generate text that sounds coherent you can train like maybe a bird size model and then it'll do something on some downstream tasks but whatever it does doesn't is doesn't look much like what those llms are doing if you want to really get an llm experience you need to do an experiment with a lot of compute that involves you know tons of gpus etc So, you know, for me, it was just a way to address the frustration of not being able to get any insights without, you know, having to do large experiments.
14:42And so the main way that you have accomplished that, if I understand correctly, is by kind of narrowing the conceptual space of what both the data set contains and then obviously what the model is trained to do. Instead of taking a small cup out of the whole ocean of mixed up of everything language, you've created a kind of, we're going to tackle one very consistent type of input. that's fair. I guess we should mention there have been many attempts to come up with a synthetic or non-synthetic, a smaller data set that has all those elements that those large language corpora have, right? So, you know, in language, you have all sorts of elements.
15:35You have a lot of facts you have so first of all you have grammar and vocabulary right those are the obvious things you have in language but then you have facts you have reasoning that you can infer from those texts and there's like also many layers of reasoning and and you know there's i guess there are many capabilities involved in being able to parse those data sets. So, you know, our initial motivation was to come up with a data set that has all these qualitative elements, but on the other hand, is just not as massive as those large language corpora, right? And, you know, you, Andrew, and I They had, so first of all, as I said, there are many synthetic datasets out there.
16:31Some of them, I think, reflect in a pretty good way certain aspects of language, such as reasoning or facts or grammar or stuff like that. But we felt like there is no single data set that has all those dimensions together, which are all integrated into something which is not too large. Right. And we felt like in order to understand, in order to gain insights about LLMs, we need a data set that has all those elements. Yeah, I was just going to add that I also came from a DMEA background. And so I was doing theory of machine learning since maybe seven or eight years ago when the field just got started.
17:24and yeah at that time like everyone was doing research on vision models and for vision there's a very nice data set called cbar10 or even amnist i mean those are very small data sets they only have like 50k images and when you train on those data set you can get a pretty high quality image model and they can do all sort of things and they reflect what's going on in real large models and I mean at that time doing research or making progress on both theory side and applied side is kind of easy because just training those models only takes like one day at most but when we move to this phase of large language model or language model in general the research has become so expensive and I've seen all those blog posts saying that it's impossible to do a PhD now in machine learning without like A100.
18:23And I mean, probably only 1 % of the PhD students has that amount of compute. So we really want to see whether there is a way to kind of bring the, I mean, good old days, which are like those CIFAR data set of fast experiment iterations back to the language side. And that's what motivates us to consider this small data set or this simple data set I think the point is, I mean, there are other kind of synthetic data set or simple data sets that are, as Ronan said, reflecting some aspects of natural languages, but they are not real natural language. They are just doing simple arithmetics or doing simple string matching or number manipulation.
19:12I mean, they are not real natural language. And we want to keep the authentic of natural language, but just reduce the overall complexity. So still, we are studying natural language and not some symbolic manipulation. And still, we want the iteration of experiments to be done in a very quick way. you created the tiny stories data set. I always like to be as concrete as possible. So I'm just going to read one tiny story just straight out of the paper because I think that will help people understand what this data set ultimately is. Tom and Jane are friends. One day, Jane goes to Tom's house. Tom has a big pot of soup.
19:52He wants to share it with Jane. Quote, Jane, do you want some soup? Tom asks. Quote, yes, please. It looks yummy, Jane says. Tom pours some soup into two bowls. He gives one bowl to Jane. Jane takes a spoonful of soup, but then she makes a face. The soup is, now this is just an example presented from the paper. That's the prompt. And then you show a completion and you compare and contrast this against other open source models, but I'll just read the 28 million parameter version that you guys trained. The soup is very bitter. She does not like it. She says, I don't like this soup. It is too bitter.
20:30Tom is sorry. He says, I'm sorry, Jane. I didn't know you don't like bitter soup. I will make you something else. He looks around the kitchen and finds some bread and cheese. He puts them on the table and says, here, Jane, you can have some bread and cheese. They are not bitter. They are sweet and yummy. Jane is happy. She says, thank you, Tom. You are a good friend. I like bread and cheese. They are not bitter. So there is our whole tiny story. And, you know, I read it almost in like I'm reading to my, you know, four year old or my two year old, because it is kind of a children's story. And I understand that that also is kind of part of the motivation.
21:07So how did you create this data set? How did you kind of how do you conceptually, you know, think about those stories you told us a little bit already as kind of having those key elements of the grammar facts um you know some amount of like reasoning required uh but how did you create them how big is this data set from from this motivation to have a good synthetic data set we should just point out that this is maybe the most natural idea is to rely on you know human development right it already has the solution for us because young children they are able to speak English somewhat like mostly coherently.
21:50I have a daughter, I can testify that not extremely coherently, but somewhat coherently. And, you know, this is, there's already a solution to this coming from like human development, right? So all you have to do is just create a data set and make sure that it'll be basically that any example can be understood by a small child on one hand. And on the other hand, you want it to spend as much as possible of, you know, the knowledge that a small child has. You want it to be as diverse as possible. and we decided you know that it makes sense to have this data set somewhat structured so just the structure of a story it kind of makes sense because into a story you can inside a story you can have all those elements combined together right grammar facts reasoning stuff like that and we just you know i think it's a really good time to to to try to create this data set because finally we have those you know models gpt 3.5 and gpt4 which those models can actually understand the instruction you know i want a story which is somewhat creative and only has very simple words right so gpt4 wrote these stories is that the so yeah most of those stories were written by gpt4 some of them by 3.5 3.5 is already good enough to write those kind of stories it's not that great gpt4 is is definitely doing a better job now it's pretty easy to just write a story, right?
23:56If I just want to write a short story, even GPT-2 can probably write a decent short story. The problem is to actually get a diverse data set that spans all the vocabulary that you actually want to span. And if you just ask even GPT-4, create a short story and you do it a thousand times and you do it with temperature with rather high temperature let's say temperature one which kind of uh gives rise to the most the most diversity you can get still about one-fifth of the stories will be about children being scared of the slides at the park we actually I actually did this experiment. So it's not very creative.
24:48If you just tell it to create a story without any other instructions, you're going to get a very repetitive data set. And the whole game is how do you get diversity? How do you make the data set not be very repetitive? And here the idea was just to collect a list of, just the vocabulary of simple words. We have about 2 ,000 words, which supposedly three-year-olds understand. And then what we do is we just ask GPT for, okay, here's one random verb, one random noun, and one random adjective. Try to combine this into a story in a creative way. We do about a million calls like that. I think we have about 1.5 million stories in the data set.
25:50So on one hand, they definitely span all this vocabulary because there's only 2 ,000 words. But on the other hand, you definitely do not span all possible combinations of words. So you know that you're not going to, you know, it's not, if you can later create a story with some prescribed combination, you will have demonstrated that the model has some creativity inside it. Yeah, so that's how we created it. That's really interesting. Just to do a little bit of math on this. So a million-ish stories, which first of all, that answers the question of why not go use real stories? Because that's a lot of books to scan.
26:41So there might not, I don't know how many children's books there are, but you'd have to get your hands on a whole lot to get to a million. So the need for the synthetic data is there. I'm interested in what you were seeing that was like way better about GPT-4 versus 3.5. It sounds like with the stories that were repetitively about the slides, I would interpret that as maybe like mode collapse, like reflection of kind of, you know, effective RLHF likely? Is that how you would understand that too? I think it's mostly just the model is pre-generating the most likely stories because that's what the model is trained.
27:19If you don't give the model any content, you will just split out the most likely stories because it wants to minimize this language model loss. So potentially just a child scared of the slice is the most common story yet that exists on the internet. So the model just learned to generate it without any given condition. So that's why we want to create some condition to move the model outside this particular high probability zone. One difference that we see between GPT 3.5 and GPT-4 is after it, so you give it three words, it needs to combine them somehow in the story. And sometimes it's not that easy.
28:03Even if I give you three words, I don't know, I think there's an example there, ancient, thunder, and sad, or something like that. You want to combine them in a way that won't look kind of too superficial. You want to create a story that actually seems fluent and you don't want like a complete change of topic in order to, you know, be able to combine the next word. And GPT-4 seems to be able to do this pretty fluently. Whereas GPT-3.5, sometimes you get stories that don't make so much sense. The words appear there, but they don't appear in a very satisfactory way. It totally makes sense also for you to note, like, yeah, maybe this is just the most common story.
28:57We don't necessarily have to invoke any exotic theories of why it keeps talking about slides. But I wonder if the non-RLHF version, if you had access to kind of the base GPT-4 model, if that might have been different. I had the opportunity to red team the GPT-4 early before it had all the safety measures, but it did already have the RLHF and kind of the instruction following capability. I never saw the like totally base model. I don't think that many people did. In the technical report, they said it was not, people weren't sure what to do with it, right? So I think they maybe put that one out to pasture.
29:37Did you guys try any different versions of GPT-4? Obviously being at Microsoft, you might have some privileged access to different versions that the rest of us wouldn't have. Yeah. So we did have access to an earlier version that, I mean, you know, we're not sure about the exact technical, like what the difference is from the model that's now available to the public. But that model had less safety features on it. But as you said, it may be the same model you had access to. So it did have a certain extent of RLHF, right? Now, I think if you take, you know, just the language model trained on the pile without any RLHF and you say, you know, create a story such that blah, blah, blah, maybe the most plausible completion in terms of a random, you know, entry from the distribution of web pages is I don't want to do it.
30:39like write me a story such that that's the question answer i don't feel like it or you know it could be that that you know the net instead of completing the story it's just gonna ask you know another question without any rlhf without any alignment that the model has to to i mean the model doesn't know that it's supposed to actually perform your instructions, right? You know, that's the most basic part of it. But I think that the RLHF they did on GPT-4 is just good enough so that it makes sure to, as accurately as it could, actually, you know, satisfy the constraints that you give it. So it almost always combines the words that you ask it to combine into the story.
31:36And also, it almost always actually writes a story that only uses simple words. Yeah, I think, as Ronan pointed out, the biggest difference of the base-like GPT-4 model compared to this RLHF version is just following instructions. I mean, for base model, you can give it as a beginning of the story. The completion is very good. but if you just say write me some story that combines certain elements then the model has a very hard time understanding the instruction because those things are very rare in the internet it probably needs some fine tune so the model understands what does it mean by an instruction it's not like a conversation it's an instruction when you ask write me a story yeah makes sense At best, you could maybe do a few-shot approach, but then you probably have – I've seen a lot of issues with over-indexing on the examples as well.
32:31So, yeah, I totally get it. A few-shot approach would probably prime the model into thinking of specific plots also. I'm doing some AI 101-type education at a friend's company right now called Athena. and I was just doing a webinar this morning where I was getting into that with folks saying, if you are going to do a few shot, probably don't do one example in your few shot, at least do two, because otherwise you tend to get this like over in context learning on just the one example you gave. So yeah, I'm with you on that for sure. So a little bit of math. So there's like 2000 words. there are, you know, let's say that's just to take a round number, um, a thousand each of kind of, you know, the verb, the noun, the adjective.
Read the full transcript
33:20So that in, you know, pure, you know, extended, expanded form is a billion possible bags of three words, roughly order of magnitude, right? So then you made a million stories. So I just wanted to establish that the space of possibility versus the actual data set that these models were trained on is about a thousand to one ratio. Do I have that roughly right? That's true. I think it's pretty accurate. Only in addition to those three words, we also have another way to add diversity, which is a bunch of features we ask GPT to add to the story, such as a plot twist, a bad ending, dialogue. So that adds a little bit more diversity.
34:10But one in a thousand is a pretty good ballpark estimate of the ratio. Cool. And then just cost of this, if we were going to pay retail price for GPT-4 to write all these million stories, that would be, if each one is say 300 tokens, I'll just take a nice round number because that maybe equates it to roughly one cent per, it would be like a$10 ,000 GBT4 retail price to generate the data set. That's pretty accurate. Okay, cool. So then this curriculum concept is, I think, super fascinating. And this is one of the areas that had me so intrigued by the paper. You're taking inspiration, obviously, as you said, from human development and starting with simple words, which definitely makes sense as an approach.
35:05I always kind of try to keep in mind as well that these things are very alien. And I'm very intrigued by this curriculum sort of approach, but I wonder, what about more weird curriculums? this is maybe outside of the scope of like this particular research, but I kind of keep waiting for somebody to show up with a, like we trained it first on like pure logic notation. You know, we've seen this a little bit. It's kind of been discussed a lot recently that the code pre-trained models seem to demonstrate better reasoning, you know, once the kind of language part gets added on, you know, and obviously who knows exactly what that baking recipe looks like.
35:50How do you guys think about that? Like, do you expect the same thing that somebody is going to pop up with a, hey, we did a pure logic or we did like, you know, just massive amounts of like abstract algebra first and kind of taught some sort of, you know, structure that we then were able to layer natural language onto? It could definitely help because based on our previous research on the attentions in language models, there are some simple attention mechanisms that the language model may have. The first one is just associating two tokens that are exactly the same. And the second one is after it associates the two tokens that are exactly the same, it also copies the tokens around the first token to the second token.
36:39So it's just like us when we read some word, we go back and see what's the previous time that this word appeared and what's the surrounding context. And I think just training this height is actually pretty expensive. It requires a lot of training data. And something like coding or logic is the perfect way to train those has. Because for coding, when we define a variable, we definitely need to look back, like, what's the previous definition? Or when we call a function, like, we check that function, we see what the function is doing. So it kind of may set the language model into learning those important concepts, like look back or checking the surroundings.
37:18And that may serve as a very good warm start for training on other things like simple natural languages. It makes a model learn much faster. Yeah, so maybe let me, I mean, another way to say what Yuanji said is, yes, I mean, we do observe this to a certain extent that, you know, maybe coding improves models reasoning. I think at this point, there is no overwhelming evidence that this is actually the case, but there are some observations. But we are not sure at all that the reason behind it is that when the model learned how to write code, it actually learned how to reason. It looks like rather the reasons that this works, the explanations are just much simpler.
38:12You just managed to calibrate the exact attention heads that you need. And those attention heads don't have any particular sophistication in them. They might just be able to very accurately look at some relative position to a given token or just compare two tokens in a very precise way. And so that means the reason is more like the types of components in your neural network that are required for coding are already there. But those components are pretty simple. It's not like the network has like very sophisticated neural paths that, you know, emerge after the training that actually know how to do reasoning.
39:06And for that, we actually have a paper we wrote at Microsoft Research about a synthetic task called Lego. So this is a very basic synthetic task that has the core elements of reasoning. And what we observe there is that we use a transformer based on a BERT architecture. And what we observe is that the pre-trained BERT transformer basically grasps this reasoning task much faster. The task is something simple. It's basically, you get a string that looks like A equals to one, B equals to minus A, C equals to minus B, D equals to C, et cetera, et cetera. And you have to resolve the values of all variables.
40:00And, you know, at first we thought, you know, maybe there's some, in some kind of profound sense, the pre-trained BERT model has learned how to reason. And this is why it grasps this task so well. But if you dig into it just a little bit, what you realize, the explanation is just much simpler, much more superficial than that. It's just the pre-training has given rise to some simple attention heads that if you just initialize the model with those attention heads, then it basically grasps reasoning much faster. This explanation is much closer to what actually happens when you train a model to code and then it exhibits better reasoning capabilities.
40:54So this is maybe a good time to talk about what we mean by reasoning. And I see a ton of confusion out there about this. Maybe you can help us get a little bit more clarity. I guess one thing that I kind of observe is, you know, of course, people are debating this capability. And it seems like, you know, you've got kind of different standards of evidence or like, you know, people put the burden of proof in different places. to put my cards on the table. You know, I kind of conceive of myself. I call myself these days an AI scout. And I'm really interested in like, what is possible? What can be done?
41:38Not necessarily holding the systems to the standard of that they can do it every time or that they can do it in all cases, you know, or it certainly matters like how adversarially robust they are. But, you know, I wouldn't say, oh, well, we found an example that it failed on. Therefore, like it, you know, can't do X if it could do, you know, X nine out of 10 times before it got to that kind of crazy example. But how are you guys thinking about reasoning as, you know, something more than a binary, obviously, in the context of this research? Initially, we think reasoning as something that is just a subset of consistency.
42:19when we generate sentences or when we say things we need to make sure that they are consistent with what we said before and there's like the first level of consistency which is just a nearby words they need to follow some grammar rules and follow some basic semantics and those are not really reasoning it's just some simple more like stochastic parrot where you just do simple pattern matching just look at the previous couple words and just generate one that is consistent. Well, I think what goes to reasoning is when the consistency goes to the next level, which is you really need to be consistent with something very far away from the current token, like something consistent with a general plot of the story.
43:04For example, there's a word but, and then you need to say something in the opposite order. And those levels of consistency, they are the primitives of reasoning. So we do think that anything beyond just a local consistency should be think of some ability that is reasoning. The first thing we have to say is the type of reasoning we're thinking about is a very, very basic. It's just some, you know, basic core capability that comes with speaking coherent English. Some people, I guess, still say that, you know, large language models will never be able to reason. I guess they have a very, very different definition of what reasoning means than what we have, right?
43:53I mean, what we mean by reasoning is really like, you know, the capacity to just apply some basic logics when you generate text, right? Right. And I think maybe to be concrete, we can look at one of the examples in the paper. So if we look at the sentence like Lily likes cats and dogs, she asked her mom for a dog and her mom said no. So instead she asked and then you do autocomplete. We kind of see it as a hierarchy of capabilities. So some words in this sentence, in order to complete them, to know what the next word is, you just need some very, very basic grammatic rules. For example, she asked her mom for, the next word is a.
44:53For this, you only need to know a little bit of grammar, and that's it. Now, the next word after A, she asked her mom for a dog. Well, you know that if you just know grammar, you know that the next word should be some noun. But here you already need to have some, you know, contextual tracking of, you know, what's going on in the text. The relevant nouns here could be, are probably dogs and cats, right? Those are the two objects that were mentioned in the sentence before. Now we go to the next sentence and, you know, you say like her mom didn't let her have a dog. So she asked for a, and when you try to autocomplete this, now, you know, the most common noun that you've seen so far, also the most, the proximate one is dog, not cat.
46:01Dog already appears twice in the sentence. You know, she likes dogs and cats. She asked for a dog. Her mom said no. So our smaller models actually complete this by saying dog. And even GPT-2 XL, which has 1.5 billion parameters, its most likely completion is still dog because it's still at that level where it did resolve that there should be some noun there. and it does know how to look back in the sentence and see that, okay, there are two nouns, dog and cat, but, you know, dog appears more. So it's more likely that, you know, if I just had a dog in the previous sentence or in the just five words before, it's going to be dog again.
46:58But on top of that, if you have a very, very basic reasoning capability, then you're supposed to be able to apply elimination and realize that, okay, she can't have a dog. We had the set containing the two objects, dog and cat, but now dog is not allowed. So what's left is cat. But we thought this is one of the most basic examples of a completion that would require some extent of reasoning. There's always this notion, I mean, this intertwine between reasoning and planning. For example, when we say reasoning, many people would think about mathematical reasoning, like I prove a mathematical theorem.
47:51And that's not only reasoning, it's also planning. I need to come up with the correct method. I have the intuition, like what's the next step should be. And for us, the reasoning that we are more interested in is just consistency. Like you should say something that are consistent with what you previously said. And the consistency is not only local, it's global. And that's where kind of we think as the reasoning for natural languages. The only thing it needs to do is generate text that's consistent with the prompt. Like that's the only objective a language model has to, you know, fulfill. The next word it generates should be as consistent as possible with all of the prompt.
48:38In order to achieve this consistency, there are several different levels. So for most words that it generates, the only capability that's actually needed is grammar. or maybe not for most words, but for many words. Just by knowing some grammatical rules, you know that if you have a sentence, Amy wanted, the next word is probably to. She wanted to something. And you don't need to know anything beyond that, right? Now, the next level after that, and this is again, very vaguely speaking, it's not like there is a very strict hierarchy. but the next level is to have some semantic understanding of what's going on or just to understand what are the relevant nouns actions stuff like that or maybe which action is could be related to which object etc etc and you know if if and i i guess if you look at uh models of size say, around 1 billion, it's very, very good up to that level.
49:56It almost always gives you a word that is grammatically correct and also has the, I mean, this word is kind of well-related. It works well with, you know, the previous few words that you saw, or it fits well with the previous few words in the prompt. But the next level after that is sometimes already requires, you know, first order logic or second order logic, etc. So you're kind of breaking it down into like micro skills. You know, this is a ridiculous analogy, but I'm kind of thinking I follow this guy on TikTok who coaches basketball micro skills. And it's amazing how many micro skills there are, you know, involved in being a good basketball player.
50:52And, you know, mostly the untrained eye, even among basketball fans, like can't really enumerate them. But this guy has enumerated them and now he's teaching them, you know, one by very small one. And so it's, you know, maybe similarly where you wouldn't say like this person can like fully play basketball. That probably doesn't even like make a lot of sense or sounds already sounds strange. But you wouldn't say if there's any missing micro skill that they can't play basketball. You have some sort of continuum there that's like, you know, people can be better and worse at playing basketball. People can certainly be better and worse at reasoning.
51:30And language models, too, can be better and worse at reasoning. And that probably seems to map on to some sort of hierarchy of micro skills that it either has or it doesn't have or it's like, you know, maybe in the process of grokking, you know, at any given point in training. So that leads me to the other, you know, big, bold vocabulary word that I want to dig in on a little bit, which is emergence. Again, tons of confusion, tons of different meanings out there. I think some people mean things that surprise us that we didn't necessarily predict. Some people mean things that happen suddenly. I guess what I kind of think is like, it seems like there is some process where, but I always think back to the, the grokking paper and the Neil Nanda exploration of that, which I'm sure you guys are at least, you know, somewhat familiar with where there is a phase change from initial memorization to a circuit, which, you know, what's so amazing about their work is they actually show this circuit in very like concrete terms.
52:41And it's like, and this is the circuit that does this algorithm that allows it to generalize, you know, to the full set, you know, from just the sample data that it was originally trained on. You know, I don't know that we have any circuits here that we can kind of elucidate, but does that feel right to you? Do you feel like there's this sort of process of like memorization sort of being gradually replaced by like concrete circuits that solve particular micro skill challenges? Is that your model of what's going on under the hood here? Okay, let's take a step back, okay? Like people talk about emergence.
53:16I guess we can both agree that, you know, this is a very, this is not a well-defined notion at all. You know, it's not like it's also it's not like you see a sudden phase transition from the model not being able to do something to slightly increase the size and then suddenly it has this. It's really good at some capability. It didn't have it all before. Rather, you know, it's a vaguely defined term saying that there are some qualitative capabilities that, you know, the model at certain sizes has, whether at smaller sizes, there's almost no trace of these capabilities at all. Like, you know, GPT-2 did not know how to summarize text, and suddenly at GPT-3, you have this summarization capability.
54:12But the notion is not well-defined at all. On the other hand, we do see that as we increase the size of the model, suddenly you have some certain capabilities you didn't have before. And, you know, in a sense, I think a good analogy is if you compare dogs, monkeys, and humans, right? You increase the size of the brain, suddenly, you know, humans can do math, whereas monkeys cannot. So, you know, it's an ability that emerged when you made the neural network larger. Not that I'm trying to imply at all that the same mechanism explains both things. We have no idea about that. But what we say here is one of the most important abilities for generative models is to be able to speak coherent English.
55:10This is an ability that we see emerge also in much larger networks trained on those large language corpora. And I think tiny stories basically gives you a much smaller data set where you can observe this emergence at much smaller scales of models in the sense that if the model is 1 million parameters, then it can hardly generate coherent stories. And if you go to 10 million, then almost all stories will be coherent. Same with reasoning, like, you know, one to five million parameter model. All of our reasoning prompts fail, whereas for 30 million, almost all of them succeed. Now, as Yuan just said, all of this basically has to do with keeping coherence with the text.
56:20So you have the emergency of the ability to generate the next word in a coherent way on different levels of difficulty. So the easiest level of difficulty we could think of is just when you have something that follows from some easy grammatical rules. and then you know you have you you can think of like different level of levels of difficulty sometimes you need to know a certain fact in order to be able to complete the next word right jack was hungry so he went looking for some you to to complete this you have to know that to satisfy the desire for, to satisfy hunger, you need some food, right?
57:15So sometimes you need to know a fact. So there's all those core capabilities that are necessary in order to keep consistency along the text. And each one of them, we can actually witness its emergence as we increase the size of the model. So what then is the theory of what is happening? I have a theory, but I want to hear yours on. So you gave that example a minute ago, right? Girl wants either a cat or dog. Mom says no dog. So it's going to be a cat. But GPT-2, you know, much bigger model, whatever. That's like 30 times bigger than your biggest in this in this research, right? Like you're, you kind of max out around 30 million parameters.
58:04GPT two is like what, 1.5 or something. So it's a lot bigger, 50 times bigger, maybe it still says dog, which is, you know, every, we can all tell that's kind of obviously wrong. You've got these much smaller models that can get that right. You know, in some way there's this emergent, you know, observed phenomenon that It is able to get that, you know, exclusion concept. What do you think is happening there? Is that a micro skill that is like, you know, this sort of exclusion, you know, that that's like a little piece of reasoning that is grokked by the small model, but not by the big model? I mean, it seems like there's you could be a really good stochastic parrot, but it feels like there's something there that has like truly kind of settled in to the structure of the network.
58:55And maybe that didn't happen with GPT-2 because the data was just too noisy and it's kind of all over the place. and so it wasn't able to learn those same things. How am I doing here? Does that resonate as likely true or even plausible? Yeah, I think that's a valid conjecture. What we think is for GBD2, because it trained on Vipe data or just think of like Read Wikipedia, we will try to minimize the language model loss. The consistency is like the least concern. It's more about just getting the knowledge correct. Like, I mean, you're talking about some object or some person and you want to know his birthday or you want to know some specific aspect of that person.
59:39I mean, this has nothing to do with natural language. It's more about just the sheer amount of knowledge that we encounter in the web data or something like Wikipedia. So the model, I mean, because the model, I mean, I think both GPT-2 and our model, they are not large enough to minimize the loss to a full extent like GPT-4. So they have to select some part of the loss that they focus on. And if the data has too many knowledges or too many other nuances, then the model may just focus on other aspects compared to consistency. Well, here, our tiny story data really pinned down. Because the language is simple and the vocabularies are simple, the really difficult part is consistency.
1:00:26And that is where the model focuses to minimize the training loss. And that's why I think our model, although it's much smaller, it gets better consistency compared to the larger ones. Yeah, you can think about it as when you train a model on the entire pile, on those large language corpora, they have much more incentive to learn the preferred clothing styles of celebrities much before they learn how to complete this sentence with the dog and the cat. They definitely learned that Joe Biden is the president of the United States way, way before they learn how to reason. This is a conjecture. Of course, I didn't actually test it, but I'm pretty sure that's what happens.
1:01:23You know, you just overload the model with so many facts that appear in so many places. Only once in many words that you generate, this capability of reasoning becomes relevant. So I think it's a really good exercise is just to open a random Wikipedia article. It's one of my favorite activities since I started thinking about language models. And just go over the Wikipedia article word by word and just think for every word, what core abilities do you need to use in order to guess what the next word is? And what you'll see is, I think, definitely for a random Wikipedia article or a random example in the pile, in some web training set, you will see that only once in every 20 to 30 words to predict the next word, you need to use some reasoning capabilities.
1:02:38For one in every three or four words, you need grammar. For most words, you can just kind of guess the next word using, if I just tell you what were the nouns and verbs that appeared in the previous sentences are, without telling you anything about the context or, you know, things which are farther away in the text, you will still be able to guess them. So the reasoning is kind of a rare capability that you need. It only becomes relevant pretty rarely. And therefore, the capacity of the model will be dedicated to other things much, much before you will get any reasoning capabilities. That's also why we always kind of see there's an emergence behavior.
1:03:36It's really because the reasoning or those very rare consistency events, they actually happen very rarely. So only if you minimize the loss to a certain extent, you start to learn those rare events. And your model feels different. For example, just a simple example, say Bob feels hungry, but he doesn't like sweet food. So he went to eat. I mean, if the model says went to eat some candy, then we think that this model knows nothing about what he's talking about. But this is only just one word difference out of the 10 to 20 words. And only if the model gets to that extent, it starts to learn this consistency and we feel like the model starts to know what it's talking about.
1:04:24But in terms of loss, it's probably we only see like less than 10 % difference. I mean, if you turn into image classification, when I tell you my model got like 90 % correct and you have a model that gets 95, I wouldn't consider this as an emergent capability that you get five more percent. But for language model, this 5 % may actually be the emergent behavior. And especially for math, if I really want to solve a math problem, there's only the connection between the two sentences where I need to make sure that my proofs are extremely coherent. Well, most of the part, I'm just completing some formulas, write down the results.
1:05:07but it's only this very tiny amount that defines and that gives the model the emergence capability so i think the two things are connected that the models and learning reasoning is like a hard task that it only gets at the very end and also we see the emergence capability for if you grow the size of the model or if you grow the number of training data yeah maybe let me just expand on your example, you know, so, okay, we have this sentence, Bob could have either candy or pizza, Bob doesn't like sweet food, so Bob got some, and autocomplete, okay? Now, the model, usually, if the previous sentence says something about sweet food, if we don't read the entire sentence, the most likely completion is actually candy, not pizza.
1:06:02You have to read it in a pretty nuanced way in order to realize that Bob actually doesn't like sweet food, right? Sweet food and candy come together so many times in the training data, right? So the neural network has to, in quotation marks, choose between using Gapit's capacity in order to be able to resolve this nuance or to use its capacity in order to know that Joe Biden is the president of the United States, right? You can't have everything together. The model has a finite size and there is some theoretical limit to the amount of things the model can learn. The model will definitely prefer to learn that Joe Biden is president and many, many other facts because they are relevant much, much more frequently.
1:07:05The curriculum development space is likely to be a huge unlock over the not too distant future, right? I mean, you're kind of, it seems like you're probably just kind of scratching the surface here because we've got web scale data that, you know, is not built for this purpose, obviously, you know, where you're saying reasoning doesn't even, it's not even required that often. And so no wonder it kind of emerges late in the game that, you know, maybe pre-training on code, is kind of changing that in some interesting ways. But man, intentional design around what a gradually upstepping curriculum might look like, especially with the ability to create the training data synthetically to really kind of isolate and bring those key skills forward.
1:08:05It seems like you could rebalance training data and probably shrink it like a ton and get to a lot of the, you know, just by kind of shifting the balance, right, to get these kind of emergent things to be more important relative to just kind of, you know, mind numbing repetition of, you know, who's the president or whatever. It sounds like, I mean, you're nodding that, it sounds like that aligns with your expectations too. It's important when we design like a new version of the data, for example, to extend the degree of our thing to maybe elementary school or middle school, I think it's very important to balance the amount of knowledge in the data versus the kind of ability that we want the data set to teach the model.
1:08:50For example, if you go to elementary school, there's ability to do simple math or mathematical reasoning or some physics reasoning or do comparison of historical events. I mean, those will take some capacity of the model, maybe you could actually take a very big portion. And the remaining ones, I mean, if your data set has too many knowledges, then the model may just prefer to use its capability or use its capacity to actually memorize the knowledge instead of really learning those abilities. So we want to balance that there's some amount of knowledge that the model must have in order to do basic stuff like math or do basic physics reasoning.
1:09:32But more importantly, there should be a lot of data that's only emphasized on the ability side. There's no new knowledge. It's just a bunch of math training samples or a bunch of simple comparison of some basic historical events or some simple physical rules and their explanation or different varieties so that the model can actually focus on the ability part instead of just being screwed by the vast amount of knowledge. So I think right now the web data, they don't really balance between knowledge and ability training. So that's why training on them is not good for a small model because they need to allocate their capability or capacity to just memorization.
1:10:16So I think that's basically the criteria for maybe design like better version of synthetic data. I guess we can just kind of maybe relevant notions here would be the breadth and the depth of the data set or the capabilities of the model. So, you know, the entire web is very, very broad. It's also so by breadth, I mean, you know, it has a lot of facts. The vocabulary is very large. You need to have a lot of knowledge to capture the data set. And by depth, what I mean is, you know, it has first and second and third order logic that you can infer from learning this data set. And this is not well established.
1:11:10I mean, there is no I don't think there is any research that that that really establishes that there is a trade off between the two. But, you know, it's very, very reasonable to assume that there should be a trade-off between breadth and depth when you train the model. Yeah, I think there's some optimal ratio because without the knowledge, you can't really do reasoning. Like you have to have some basic knowledge, let's say candy sweet, in order to do reasoning. and when you go to elementary school, I mean, you have to know some basic events, like some basic rules for math, like one plus one is equal to two in order to do mathematical reasoning.
1:11:55So there's a balance. You cannot have no knowledge, but you cannot have all the knowledge. So maybe there's some optimal ratio between like breadth and depth. My head keeps coming back to this same kind of thing where, yeah, we should see so much gain from kind of rebalancing the data set and maybe even, you know, starting with some more abstract things, like I could see sort of A or B, not A, therefore, you know, and then do that with like, just, you know, all the letters and then start to, you know, introduce kind of these associations and layer that on. Like it, it sure seems like there is a lot of, of opportunity there.
1:12:34Yeah. You're just mentioning the extreme case of only teaching reasoning because they are, everything is just symbols. There is no knowledge. And it's all about just reasoning. And maybe it's good to combine this with just something that is pure knowledge. And maybe we can get something good by just adjusting the ratio. Yeah, maybe it's not clear at this point. If, you know, a human being, you can, when you teach a human being, you can kind of almost separate the the knowledge and the reasoning, right? You can just say fact A, blah, blah, blah, fact B, blah, blah, blah. Now here's how to reason.
1:13:13I mean, it has to be a pretty smart human being, but in general, you know, we are able to take those two things and then combine them so that in the end, we will have those core reasoning capabilities that we learned using, you know exercises that only involve if not a then if a a then b means not b then not a right when we learned when we study to the sats we have you know those rules and and uh and separately we have uh the knowledge of facts and we're able to combine them And I think it's an important question whether it's even feasible in language models to just take those things, separate them into two different modules.
1:14:06And, you know, will the model actually be able to combine those abilities? My conjecture, by the way, is no. As long as you don't combine them enough in the data set, the model is not going to be able to infer the way that a human consciously infers the connection. But if we could do that, then this would be, of course, a very, very powerful technique to train models. I just did another interview with a couple of guys from Mosaic ML, and they talked about training at times on like massive client data sets. Even when they do that, they typically still mix in, you know, kind of the general pre-training, you know, pile or whatever, because otherwise they see catastrophic forgetting.
1:15:01so they have to you know kind of keep some mix at at all kind of stages of training to avoid that so i think that would be very consistent i think with what you said like you you probably can't do it in strict phases um there's got to be some sort of like mixing strategy throughout the process right and that's going to be very non-trivial to do um it might be feasible but it's definitely not and it's not going to happen on its own just because, you know, what the model cares about is just being able to efficiently autocomplete samples in the data set. If the data set is either knowledge or reasoning, then it has zero incentive to combine.
1:15:49And even if you give it a few examples where this is actually combined, no one is, you know, no one assures you that it'll actually be able to take those two modalities and really use them for the combination. So I feel like this is a science that's really, I mean, we're only beginning to understand how these things work. Hopefully, we'll figure out a way how to actually do it. But I'm not sure, Yuanju, maybe, I don't know any concrete evidence of an example where this seems to work. Right now, we like the kind of concrete evidence that curriculum learning really works. But we believe this should be helpful.
1:16:39But I think overall, as Ronan said, the model has no incentive to connect the different phrases where you train things. they just i mean when you start a new phase it's just greedily minimize everything that related to this phase and it can't just forget everything that he has learned so it's definitely a very non-trivial task to get this curricular learning work first you just had a number of kind of interesting empirical observations and you can you know address this kind of however you want But you noted that grammar emerges before consistency and creativity. And consistency is, you know, kind of related there to reasoning in our, you know, per your telling earlier.
1:17:24Is that the same thing that I've observed in my kids? I'm, you know, my memory, maybe I'm sleep deprived, but I feel like grammar maybe came last for them. Like they definitely have a, you know, a certain consistency. Like they want what they want. you know they like there's um if they want ice cream like you know ice cream they keep the next token is ice cream and it's they're pretty consistent on that and you know creativity i don't know that's a little trickier but does this feel like it echoes human development in your mind like to me i'm not sure if it feels that way i love this example um so yeah i mean I think the learning process for children is very, very different that way, right?
1:18:06Children don't get, like, their incentive is not to say the correct next word. Children, they want ice cream, their incentive, the outcome needs to be, I get ice cream. So if they just say ice cream, ice cream, ice cream many, many times, Maybe they don't get the best grade for creativity, but they will likely get the ice cream, right? Depends on how basically on the self-discipline of parents and whether they are... Limited in this household. Yeah. I have a very close experience with this exact scenario. But, you know, like more seriously, like, you know, when children produce language, I mean, I'm only basing this on my observation.
1:18:59They have constant contact with the physical world. They know which entities are involved in the current conversation, right? We are talking about a book that I just read. So, you know, it's very unlikely that the next sentence will be about like a car because, you know, they have in their heads this entity like, no, we're not talking about a car. We're talking about this book, right? Whereas the language model, you know, if it just makes one mistake in one word, saying car instead of book, the loss that it incurs is actually not as big as the loss it would incur for incorrect grammar, which is relevant in almost every word it produces.
1:19:56And not only that, you know, the language model, it's actually much cheaper. It's much easier to have consistent grammar. You only have to kind of be consistent inside the same sentence for the grammar to be correct. You don't even have to look past five or six words back. Whereas for children, children have a few entities which are kind of active in their working memory. So this entity could be, I want to go to the park and have some ice cream. So the two relevant entities are park and ice cream. you're not going to replace those two entities in the middle of the conversation. That's, you know, what the child cares about at this point.
1:20:48So it's really like the incentive system is different, but also the whole kind of cognitive perception around that's basically underlying the conversation is pretty different. Yeah, I don't know if it makes sense. Like maybe a good way to think about it for a language model, definitely in its first phase of training, it doesn't need to look way back like three, four, five sentences before what it produces. So I think a good example is to look at one of the GPT-2XL generations that we have in the paper. So we have a story about like a snowman, two kids are building a snowman and a dog comes and destroys the snowman.
1:21:51And then like you autocomplete to see what happens. and maybe maybe i'll actually read it because i think it emphasized i think it emphasizes pretty well uh what i'm uh trying to say here so there's those two kids saran band they built a snowman a dog comes and uh destroys the snowman and the and now uh gpt2xl's completion looks something like that. Sarah and Ben cry and look at the snowman in the snow. Ben asks to do something. Sarah tells him to go away. Her hand tightens on the little red button. The snowman is covered in fur. It is red. It is real hair. It is a real snowman. It looks just like the ones they used to make.
1:22:46Ben has a big smile on his face. He looks at the face he has made. He walks over and kisses the eyes. He's happy to be alive. He has done what he was going to do. He has saved the world. So, okay, what am I actually trying to say here? Except that, you know, GPT-2 XL has a good sense of humor. What I'm trying to say is almost every sentence here, all you need to do in order, or GPT-2 XL does in order to generate it is look at the previous sentence. And maybe it knows that there is one entity, which is snowman, which is relevant here. Other than that, there is no consistency at all in the completion that goes beyond, you know, looking at two consecutive sentences.
1:23:39Language models are, if they're too small or in the beginning of training, what they do is they don't actually have enough incentive to know the whole context of what's going on. Because to complete most words correctly, you just need to have the context of the current sentence and maybe the one before, and maybe also one or two important entities. Whereas for humans, this is completely different. We have agency, we know what we want when we form the next sentence. And yeah, we care much less about grammar than about the ice cream. Yeah, I would make an analogy like human child is learning with our RLHF algorithm, just doing reinforcement learning with the parents' feedback.
1:24:36And obviously the parents are very robust to grandma mistakes, so they don't really want to optimize that in order to maximize their reward. They'll probably care more about consistency, the topics, and they want to get the topic correct so they can get the reward. Well, for language model, it's just next word prediction, and the grandma is going to be penalized much more severely compared to the global consistency. Okay. Yeah, I like that. I'm always a little wary of analogies and I always want to come back and think, what is sneaking into that analogy that I don't want to allow? So I'll bookmark that one.
1:25:13But certainly the surface level intuition there makes a lot of sense. And it's a very clippable analogy as well. How about on the depth versus hidden dimension size of the, which I often just call width, the number of layers versus width of a layer? You note some interesting trade-offs there. And I didn't really have an intuition necessarily for why that would be. But per the paper, you report that the fewer layer models will do better on grammar compared to consistency slash reasoning. And from that, it seems that more layers are important for this kind of reasoning consistency. How should I think about that?
1:26:03Is there a story that crystallizes why that would be? None of this is well established. It's all our conjectures that need to be studied much more. But I think a good way to look at it is, so depth tells you about how many times information can percolate between the tokens. So every time you have a global attention layer, like a transformer attention, certain tokens, the information inside certain tokens can percolate into other tokens. So, for example, if you have some instruction, create a story with this and these words, and I want a bad ending as well, and then I type in the beginning of the story and it needs to autocomplete it.
1:27:02these instructions can, in every attention layer, they only have one chance of percolating into the tokens of the story. And sometimes these instructions by themselves are nuanced. So maybe there is an instruction saying the story has a bad ending. It's not enough to know So in order to complete the next word, to know the current sentence and that the story needs to have a bad ending, you also need to have like the context, like the wider context of what happens in the story. So in order to fulfill these kind of instructions, you have to basically have the information percolate several times between the tokens.
1:27:55and that's also the case for reasoning. If you have first order logic, so if we think about this example of cat and dog, like Alice wanted a cat or a dog or mother didn't let her have a dog so she got a cat, how many times does the information have to percolate between the tokens here? So first you want to understand that there was a cat and a dog involved. that's one layer of global attention then you want to know that she couldn't have a dog and you really want to you and and so you have this set cat and dog and you want to do cat plus dog minus dog equals to cat so so after you know that you have either cat or dog you need the fact that she couldn't have a dog to percolate into my token in order to know that the only available option is cat.
1:29:05So there are actually two layers of percolation here. You need to know that it's not a dog. So the token not has to go inside the token dog to know that dog is not allowed. And then those two tokens, not plus dog, have to percolate into the generation in order to know that, okay, I had cat plus dog, but I have to do now minus dog to have a cat. So it's just several. So if you think about it as coding, you have several conditionals you have to do and several times that you need to have pointers to information that appears in other places in the text. Okay, this is very, very, I'm putting it in a very vague and non-formal way, but I think it's a good initial intuition probably to what happens.
1:30:11Whereas when we talk about facts, so if you have a completion, you know, I don't know, okay, like if I have a completion in a language model that's like the capital of France is, then all I have to do is I just need to have one kind of lookup table with, you know, all the countries and their capitals. I don't need to have many layers of global attention. It's enough to just take the two tokens, capital and France, put them together, and then just have one lookup table saying that capital plus France equals to Paris. Now, here, the dimension seems to play a much more important role because the bigger the dimension of the space is the more entities I can kind of squeeze into this vector space.
1:31:14And, you know, also the more neurons I have inside my lookup table to tell me, you know, that to have this list of all those possible facts. So is another way to say that, that within a single attention block, the attention relationships are not immediately transitive. And so they need multiple iterations of attention in order to create that transitivity. Like if the current token is looking back at a certain token, but then that token is looking back at a previous token, we need two rounds of this to move the two hops. Yeah, exactly. You first need the knot to go into dog to know that it's not dog.
1:32:04And then the knot dog needs to go together into this set that has both cat and dog in it. So that's already two leaps. Does that also suggest then that sort of for as many kind of logical leaps as you might need, you need like maybe that many layers? like you can't you're sort of bounded by if you have if you have two layers you can make maybe two logical leaps is that a general heuristic that seems sensible i think there's a depth and width trade off for example you can simulate two layer leaps just using one layer but you have to enumerate like all the two possible combinations which makes your size go from like say n to n square.
1:32:49So if you want to be kind of the most size efficient, then you will definitely have to go as deep as the number of logical loops. But if you are not that deep, then you can actually use a wider network to concatenate the two steps into one and just make the intermediate layer much bigger. So here also, I mean, your question kind of bursts into an open door in the sense that the paper we wrote at Microsoft Research about this synthetic reasoning task we call Lego, that's exactly a task where you have multiple leaps of reasoning. And we see a very direct connection between the number of layers that you need and the number of reasoning steps required to complete the task.
1:33:46But maybe somewhat surprisingly, we see that the model finds very interesting and sophisticated ways. This is actually not in the paper. This is kind of like a follow-up work. It finds very sophisticated ways to do multiple leaps of reasoning within a single layer. So definitely more layers help, but it's not like it's a strict upper bound for the number of leaps you can do. I could only wish it could be quite that simple. But that's really, really interesting information. I'm learning a lot from this. The interpretability part of this paper is also really interesting. You kind of break it down into the attention portion and then obviously the MLP neurons portion.
1:34:43And a couple of things jumped out at me. One was in the attention portion, you seem to observe that there's kind of two sorts of attention heads. One really just focuses on the distance relationship between the tokens. And then the other is more semantic. And the distance one, I was like, holy moly, does that look like the alibi scheme that has recently come to popularity with these super long context windows? So I don't know if you guys have had a chance to study that, but quite an uncanny resemblance, right? I mean, you're showing all these attention heads where it's like this one, you know, it's just a very tight attention range.
1:35:26And then, you know, there's different kind of lengths. And that's almost exactly what they cook up as the, you know, kind of substitute for positional embeddings in the Alibi research, at least as far as I understand it. Do you see that same parallel? Yeah, I think they are definitely doing the same thing, which could explain like why Alibi is very, I mean, it's very helpful because the positional embedding in Alibi is already initialized to do this multi-scale distance-based attention. Well, for absolute positional encoding, like what we use here, the model has to learn to discover this optimal kind of positional-based attention.
1:36:10So just hard code that positional-based or distance-based attention, I think is a really good choice based on our observation. And also we can talk like the shorter ones are really responsible to just learn the grandmas and the longer ones they may just make sure that your content are consistent globally or maybe just to grab the associated words for example you have an Alice in one sentence and then you have an Alice like five sentence ago you want to make sure that these two words you have a chance to put them together so so yeah let me just point out that, you know, clearly to complete the next word, you need two things.
1:36:55Usually you want to know what the proximate words are, what are the most kind of recent words you saw, and you want to know what are the most important entities in the story are. So these are going to be you know words with the relevant semantic meaning uh to what you want to complete now if i remember correctly what happens in alibi is you there is kind of a little bit of a mix of both of them so you just take every attention head and you make it decay uh for every attention head you just prescribe some scale and the the the strength of attention decays with the distance uh inside the text is that is that correct i if if i'm not mistaken that's that's what happens there and we actually see i mean one surprising aspect is we actually see a dichotomy there are heads that only care about distance and other heads that only care about semantics.
1:38:07And there is hardly a mix between the two. But we have to say that this is only for one attention block. We haven't really checked for multiple attention layers, what the transformer will do together. But if you just train a network with one attention block, it seems that the network learn to separate the distance-based attention versus the semantic-based attention, where some heads are just looking at tokens based on its distance, some other heads are just looking at tokens based on the semantic similarity. So is there anything else that you can say? I mean, that's pretty profound in and of itself, that that dichotomy emerges, because you didn't initialize it.
1:38:52I mean, in Alibi, they've engineered it that way through, you know, some probably trial and error and heuristics and guesses and whatever. But this is totally just happening on its own. Is there anything else we can say about what you see in the semantic one? When I looked at those, you know, I didn't know light bulbs went off in my head to kind of interpret those visualizations of the semantic blocks. But anything you would highlight from studying those? I think the most interesting one we see is just a symmetrical attention to the main character names. So for example, there's some hat where every token just attend to Tom and Lucy, which are the two main characters.
1:39:34I think this is pretty important. I mean, it's just try to identify what are the persons that are involved in the story. So the next time when it generates a new thing, it's not going to say, like, wow, it's going to say something consistent. So I think the semantic has, at least what we see in the OneTransformer block, is more about this type of attention, where it identifies what are the main objects in the sentence and just try to make sure that most of the tokens or the relevant tokens attend to those objects. Like when you have the or a, you will attend to like banana. So you know that you want to complete the next word as a banana instead of something completely made up.
1:40:17So I think those symmetrical attention has, they are really useful to just speak consistent English inside the transformers. Yeah, but let me add that, you know, it's very natural to complete the next word. You want to know what are the relevant characters, what are the relevant entities in the story. But no one expects a priori that you will have such clean attention heads, an attention head that exactly attends to the character and a different attention head that exactly, you know, attends to, you know, the objects. In the example we gave, it's like a banana and park. A priori, we might expect that it'll all be just a big mess, right?
1:41:04Every attention head attends to a little bit of everything. And why would it be interpretable at all? But it's quite surprising that when the model is small enough, it seems that we can actually give meaning to both attention heads and neurons. So does that kind of fall apart if we add a second layer? Does then it just become more messy again? Or what is that, as you start to stack layers, what does that start to look like? Yeah, I think when your transformers are getting higher or getting deeper or getting larger, it definitely becomes more messy because the transformer can simulate. I mean, if the transformer is small, you really need to learn those separate modules in order to minimize the loss.
1:41:52But if the transformer has a larger degree of freedom, it has a luxury to, for example, use five attention heads to simulate one or use three layers to do what could be done in one layer. It has no incentive to be like as precise or as kind of conservative as the smaller ones. So it's actually less interpretable. And we also observe that in when we try to interpret the neurons as well. Perfect bridge then to talk a little bit about the neurons. Maybe just give us a little bit of understanding of like the technique that you used to, you know, I'll try to summarize it real quick. You tell me where I'm wrong.
1:42:30You run a ton of stories through and you look for what tokens specifically are maximizing the activation of a certain neuron. And then you can kind of print out like, here are the snippets and the individual tokens that maximize the activation for this particular neuron. And then And holy moly, like it really looks like there's a pretty coherent concept, you know, as you just kind of scan down that list of things that, you know, corresponded to high activation on any given neuron. That's pretty accurate. So you have this these middle layers in the MLP, which we can think about as neurons. Those are really the coordinates that can either be activated or not.
1:43:17You have the only basically non-linearity in there. And yeah, like, again, just like the attention heads, a priori, it's not clear at all that they would have any meaning. Those are just, you know, different coordinates of a certain vector space. like no one promises you that you know the neural network is going to use one particular coordinate for one particular kind of task and um indeed so so maybe let me mention that this basically uh follows an idea suggested in a 2015 paper by lee et al called visualizing and understanding in neural models in NLP, which is like their idea is just to look at the tokens which induce the highest activations for every neuron in a certain text and try to see whether, you know, those tokens have a common role.
1:44:26And when we look at larger models like GPT-2 XL, what we, and we try to look at those tokens, at least the two of us could not find any common meaning. You know, the same neuron is activated sometimes on nouns, sometimes on verbs, sometimes on like, there's just no clear pattern whatsoever. Whereas when we take a small model, For example, there is one particular neuron that seems to always be activated when the main character of the story is introduced. And, you know, that kind of makes a lot of sense if you think about what the neural network needs to do. I guess, like, if there was a programmer writing code that tries to autocomplete, you know, stories, there would probably be a function that tries to locate the name of the main character because it's useful in many, many places when you autocomplete.
1:45:41In fact, whenever you know that the name of some character should appear, it's a pretty good guess to think that this is going to be the main character. So you have a neuron exactly doing that, and we haven't checked enough to be sure. But there's probably an attention head that then attends to what this neuron outputs whenever you know that the name of some character should appear. And when you connect those two together, what you will get is this mechanism that is able to copy the main character's name to different places along the generation. So, you know, this is a very basic mechanism that you can actually observe inside the neural network.
1:46:40And this doesn't happen in bigger models, at least not in a way that is, you know, so easy to trace. I guess my the simple version of it was I thought maybe it was just like maybe there are sort of a bunch of concepts that are easy to identify. Where you can just see like, OK, at a glance, I know what that is. And maybe there's just only so many, you know, like maybe we only have. So when you have 30 million neurons, you know, are 30 million parameters, you have however many neurons, you know, maybe that's kind of enough and you can kind of capture those. And then you go 50x and it's like, well, if you start fishing at random, you know, points in the network there, maybe you just miss a lot of those.
1:47:21You know, they may exist, but they're just kind of hard to spot because maybe they're sort of sparse, if you will. And then the things that are in between, I mean, I'm really getting out on a limb here, but I was kind of thinking maybe those are sort of analogous a little bit to like the subconscious processing that goes on in our brains, where I kind of know on some level that like processing is happening even for many concepts, you know, that I don't have like a clear label for. It's just, you know, there's some sort of churn happening in the brain, but then like only a certain, you know, small set of that kind of rises up to this level of like, you know, what I've called it a conscious concept that I can sort of say, like, I have a label for that and it's like a tidy enough thing.
1:48:05So I guess the two ideas there are, maybe they're just a lot more easy to find in the small network because, you know, you have to have them and, you know, they get packed densely versus a big network, you know, maybe they just are packed more loosely. And then those other networks or those other, you know, neurons maybe are just kind of analogous to some stuff that we don't understand very well in our own cognition. Yeah, it's definitely possible. I think that's an advantage of small language models. They may be more interpretable compared to larger ones because the smaller models can only do basic stuff, and only the basic stuff are probably interpretable.
1:48:42Like the very complicated stuff, for example, how GPT-4 write a code that are 1 ,000 lines. I mean, those things, it's almost impossible to interpret. But how could a small language model keep the main character consistent? In those basic questions, we can probably understand, and there are probably some neurons associated with that. And in GPT-4, like out of the, I don't know, maybe 10 ,000 neurons or even more neurons, there may also be some neurons that are dedicated to keep the main character consistent. But it's just so hard to find it because it may be in the 25th layer neuron, like 9 ,700, whatever.
1:49:25It's just so hard to locate. Well, for smaller models, because it's so small, every neuron must be doing some basic stuff because the complicated one, as we said, in the consistency hierarchy or in the loss hierarchy, they only consist of very tiny fraction of the loss. So the main fraction of the loss that contribute to basic consistency, grammar, and things, those are probably the things that are learned by the neurons in the smaller models, and they are more basic and more interpretive. So there's many ways for the neural network to solve a problem. Like given a problem and an architecture of the neural network, there are many different configurations of the weights that would solve the same problem.
1:50:13Some configurations might be more interpretable to a human, and some are just, you know, one big mess. Like every neuron is doing a little something of every possible task. And they are combined in very, very complicated ways. And the network has no incentive in the loss function not to be a one big mess. Like most solutions to the same problem are one big mess. This is where the entropy is located. Right. And when the model is small, it kind of has no choice. The neurons have no choice but to align with meaningful tasks because, you know, the neurons are where you have the nonlinearities and you just don't have enough of them for the one big mess type of solution.
1:51:14Somehow, the most efficient solution is the one that is not completely messy. And if you have a large network, then, you know, it'll just kind of find a way to do it that does not align with the coordinate structure of the neurons. Whereas when the model is small, you just have no choice. So interpretability appears as kind of a side effect. So anything else that we didn't cover? Yeah, so one thing, maybe going to the initial motivation in creating the data set, which is, you know, to have like some basically like a small data set, which is a testing ground for ideas in LLMs. Like an open question here is, do we even have a reason to expect that behaviors we witness in this compact setting will translate to LLMs, right?
1:52:18We don't know the answer to that. Like, suppose we find an architecture that works much better for the tiny stories dataset. Do we actually have a reason to expect that, you know, this architecture will also be better for LLMs, right? So I'm just saying this as a question. I think it's one of the most relevant questions that, you know, stem from this paper. And, you know, it kind of connects to a more general question, which is, you know, there's all those papers like the Google Chinchilla paper and the OpenAI Scaling Gloss paper, which try to suggest that there might be universal phenomena in LLMs.
1:53:06There is some, for example, a trade-off between width and depth that is perhaps, they don't suggest it explicitly, but a natural question that arises, is this universal in the sense that it does not depend on the exact mix you take in the data set and the exact architecture and the exact range of sizes you take. So the question here is, are there universal phenomena which will be common to the tiny stories data set and to LLMs being trained on these large corpora? And, yeah, I mean, we maybe let me just say we have just a few indications of some sorts of universality. But at this point, it's completely open.
1:54:03And, you know, we really hope for the sake of, you know, saving energy and also having, you know, just opening the door for PhD students to actually do LLM research. we hope there is some universality going on so that, you know, you could gain insights, not necessarily on tiny stories, but on any small data set, which would actually be of relevance to LLMs. Yeah, our future work is mainly just planning to extend the capability of tiny story. If we can create a story that captures like elementary school knowledges, I think this is already a really good data set. If we train a language model, for example, less than 1 billion, maybe 300 million parameter.
1:54:56And it's just cool that everything for elementary school or maybe even third grade of elementary school, I think that's already a very good model. I think people will love to interact with it. It knows how to talk. It knows the basic knowledges. and maybe it's for, I mean, the data set would be diverse enough to capture everything in real language, capture every aspect of real language, but just at like a downscale level. And once we have that data set, I think it really opens the door for everyone to do natural language research, not the ones that has like 100 A100 in their hands, but the ones with just a laptop GPU, they can train the model in like one or two days and they can get some interesting operation.
1:55:46I think what we witness in LLMs is kind of a mathematical miracle going on. And what do I mean by that? You know, you take this algorithm, which is pretty simple. It's, you know, it's gradient descent plus plus. I don't want to belittle, you know, all the basically really smart technical contributions that are inside that algorithm. But all in all, it's basically gradient descent with an architecture that's very clever, but still it's quite simple. And the miracle is you take all this huge training corpus, you fit it to the algorithm, and you don't just get a network that has memorized some text.
1:56:34You get a network that can actually genuinely, you know, create, like synthesize new content, show signs of reasoning, understanding, and so on. And we think tiny stories is just kind of a compact example where you observe the same type of miracle. Of course, it's not nearly as exciting of what happens in LLMs, but already there at this size, you see that there is some very interesting generalization and emergence going on. And even if it doesn't give us a lot of insights about large language models, This is still kind of a nice playground to try to develop maybe the mathematical foundations necessary to understand why neural networks are able to generalize so well.
1:57:44So maybe then just one final question, you know, and I'll encourage people to get in there and try it out. What other interpretability type work have you guys seen that has inspired you that you would recommend that folks in the audience go take a look at as well? Yeah, I think one of the works that inspired our research, I mean, it's the work from our group previously, which is called Lego. It's a synthetic reasoning task, which tries to understand what the model is. I mean, try to understand what the attention mechanism of the model is. We identify several types of different attention has in the transformer.
1:58:24some of them are, as I said, they are just looking at the tokens that exactly appeared before. Like Alice appeared before and Alice is associated with Alice. And there are some other more advanced mechanisms such as doing reduction or some other stuff. So I think this work is pretty inspiring. And it tells us like Transformer is at least doing something that is reasonable instead of a pure mess. So that's why we also have the interpretability section we want to look at the attention has. And we do see some very good behavior that's corresponding to some aspects of the natural language. Maybe there's one work I want to mention also.
1:59:11There is a paper called Transformer Feed-Forward Layers or Key Value Memories. That's another paper I like. that's a paper that tries to interpret what neurons are doing in, I think, basically BERT size model transformers. And I mean, they are basically able to show that at least some of the neurons have meaningful roles. In general, the theory behind the interpretability of neural networks is at its very, very beginning right now. So there are plenty of very, very clever works, but... it seems just very difficult. So in spite of really nice works in the literature, I think we are still light years away of being able to actually understand what's going on inside the model.
2:00:21And a priori, there's no reason to assume that we'll ever be able to really understand, right? I mean, we have a very limited understanding of how the human brain works. Like, it's not like we can point to a neuron and say, you know, this neuron has this and this role and, you know, that thought process. And there's just no reason that we'll be able to ever do it in neural networks. And like, there's also no reason to assume that the solution that the neural network finds, that solution that gradient descent finds to the problem is not a very, very messy and not interpretable solution. So we'll probably be able to find to come up with some, you know, kind of basic or small examples where which are partially interpretable.
2:01:22And we might have some insights about big networks. But yeah, I personally am not very optimistic about, you know, being able to interpret what's happening inside those models to a satisfactory extent that, you know, might lead us to being able to like control them and, you know, manipulate them. them make sure they have better alignment and so on and so forth yeah large-scale models interpret we may need to take a different approach so i think it's impossible to look inside the neural networks and just pin down like the attention is doing something or the neuron is doing something but maybe we have to do an approach that's more like our sparks of agi paper where we just talk to the model and we kind of try to, it's more about interpreting other humans' intention.
2:02:22When we talk to them and we throw a sequence of conversations, maybe we can understand what the model likes to do and what the model doesn't like to do, or what the model is good at, or what's the typical cases of the model's failure. And it's more towards like psychology study, but really for robotics, that maybe we need to take that approach for interpretability. You know, humanity has taken advantage of horseback riding for quite a long time. Now, we have no idea what every neuron inside the horse's brain is doing. We can't really interpret how, you know, we give some, I don't know how to call it, like command to the horse, like physical cue.
2:03:12and the horse obeys, and it's very, very, very useful. And we can actually rely on, like, you know, horseback riding is very reliable. There are very few cases where, you know, the horse has acted unexpectedly in a way that, you know, caused accidents. Like, humanity has profited from that vastly, maybe, you know, leaving animal rights aside here. and you know it works perfectly even without interpretability and you know we we just figured out figured out ways to align the behavior of the horse with our needs by you know taming the horse we can tame it without understanding the exact process right that's going on on there and you know, this is a big success.
2:04:10And I think, you know, I think it's just a good analogy, right? I mean, it's kind of horseback riding for the brain, those LLMs, right? They are just, they give us, suddenly we can go much faster to much longer distances, even if we don't exactly understand, you know what the horse is doing like definitely you know the mongols didn't understand much about the biology of the horse they could still use the horse like in a very reliable way so you know i i just think this is even though i'm pessimistic about uh actually understanding the inner workings of the neural network. I'm very optimistic about the usefulness and the fact that we will be able to align it efficiently.
2:05:10Ronan Eldon and Yuanjuli, thank you for being part of the Cognitive Revolution. Thank you very much. Thank you for the invitation. It's really my great pleasure.
From the publisher
Thanks to the over 1m people that have checked out the Rise of the AI Engineer. It’s a long July 4 weekend in the US, and we’re celebrating with a podcast feed swap!
We’ve been big fans of Nathan Labenz and Erik Torenberg’s work at the Cognitive Revolution podcast for a while, which started around the same time as we did and has done an incredible job of hosting discussions with top researchers and thinkers in the field, with a wide range of topics across computer vision (a special focus thanks to Nathan’s work at Waymark), GPT-4 (with exceptional insight due to Nathan’s time on the GPT-4 “red team”), healthcare/medicine/biotech (Harvard Medical School, Med-PaLM, Tanishq Abraham, Neal Khosla), investing and tech strategy (Sarah Guo, Elad Gil, Emad Mostaque, Sam Lessin), safety and policy, curators and influencers and exceptional AI founders (Josh Browder, Eugenia Kuyda, Flo Crivello, Suhail Doshi, Jungwon Byun, Raza Habib, Mahmoud Felfel, Andrew Feldman, Matt Welsh, Anton Troynikov, Aravind Srinivas).
If Latent Space is for AI Engineers, then Cognitive Revolution covers the much broader field of AI in tech, business and society at large, with a longer runtime to go deep on research papers like TinyStories. We hope you love this episode as much as we do, and check out CogRev wherever fine podcasts are sold!
Subscribe to the Cognitive Revolution on:
* Website
* Spotify
* Youtube
Good Data is All You Need
The work of Ronen and Yuanzhi echoes a broader theme emerging in the midgame of 2023:
* Falcon-40B (trained on 1T tokens) outperformed LLaMA-65B (trained on 1.4T tokens), primarily due to the RefinedWeb Dataset that runs CommonCrawl through extensive preprocessing and cleaning in their MacroData Refinement pipeline.
* UC Berkeley LMSYS’s Vicuna-13B is near GPT-3.5/Bard quality at a tenth of their size, thanks to fine-tuning from 70k user-highlighted ChatGPT conversations (indicating some amount of quality).
* Replit’s finetuned 2.7B model outperforms the 12B OpenAI Codex model based on HumanEval, thanks to high quality data from Replit users
The path to smaller models leans on better data (and tokenization!), whether from cleaning, from user feedback, or from synthetic data generation, i.e. finetuning high quality on outputs from larger models. TinyStories and Phi-1 are the strongest new entries in that line of work, and we hope you’ll pick through the show notes to read up further.
Show Notes
* TinyStories (Apr 2023)
* Paper: TinyStories: How Small Can Language Models Be and Still Speak Coherent English?
* Internal presentation with Sebastien Bubeck at MSR
* Twitter thread from Ronen Eldan
* Will future LLMs be based almost entirely on synthetic training data? In a new paper, we introduce TinyStories, a dataset of short stories generated by GPT-3.5&4. We use it to train tiny LMs (< 10M params) that produce fluent stories and exhibit reasoning.
* Phi-1 (Jun 2023)
* Paper: Textbooks are all you need (HN discussion)
* Twitter announcement from Sebastien Bubeck:
* phi-1 achieves 51% on HumanEval w. only 1.3B parameters & 7B tokens training dataset and 8 A100s x 4 days = 800 A100-hours. Any other >50% HumanEval model is >1000x bigger (e.g., WizardCoder from last week is 10x in model size and 100x in dataset size).
Get full access to Latent.Space at www.latent.space/subscribe




