In short
Summary of the TWIML AI Podcast Episode #674: OLMo with Akshita Bhagia
Podcast Overview Title: The TWIML AI Podcast Host: Sam Charrington Guest: Akshita Bhagia, Senior Research Engineer at the Allen Institute for AI Episode Focus: Discussion about OLMo, an open-source language model, and its surrounding ecosystem.
---
Key Themes
Introduction to OLMo
- Definition: OLMo stands for "Open Language Model."
- Variants: 7 billion and 1 billion parameter models released, with a 65 billion model in training.
- Motivation: Address the lack of transparency in the development of proprietary language models.
Importance of Open Science
- Access to Data: OLMo includes not just the model weights, but also datasets, training tools, and logs.
- Scientific Collaboration: By making the research accessible, the aim is to reduce redundant efforts across research groups and improve the quality of scientific inquiry.
Unique Features of OLMo
- Ecosystem: OLMo is distinguished by its additional tools and datasets, such as:
- DOLMA: A large training dataset (3 trillion tokens) created with publicly accessible sources like Common Crawl, Wikipedia, etc.
- Paloma: A benchmarking tool for evaluating language model performance across various domains.
---
Discussion Highlights
Development Insights
- Curation of Datasets: Emphasized the importance of careful curation to avoid toxic content and ensure quality.
- Training Challenges:
- Issues encountered with weight tying and layer norms during the training of the 7B model.
- Importance of transparency throughout the training process.
Evaluation Methods
- Paloma Benchmark: Provides a nuanced evaluation of model performance across 600 domains, offering insights into how well the model understands different types of content.
- In-Training and Offline Evaluations: Utilized to refine model architecture and ensure effectiveness.
Open vs. Closed Models
- Security Concerns: Open models can still pose risks, but transparency allows for better guardrails and ethical practices.
- Collaborative Improvement: Encourages community engagement and innovation, with the goal of building better models collectively.
---
Key Takeaways
- OLMo's Contribution: Represents a significant step towards open sourcing language models and their training data, facilitating research and industry application.
- Ecosystem Approach: Focuses on providing comprehensive tools that enhance the usability of the language model beyond just the model itself.
- Future Directions: Plans for ongoing development, including more advanced models and additional modalities.
Conclusion This episode of the TWIML AI Podcast features a deep dive into the OLMo project with insights from Akshita Bhagia on the motivations, features, and implications of developing open-source language models. The episode highlights the importance of transparency in AI development and the collaborative potential of the research community.
For complete show notes, visit [TWIML AI Podcast Episode #674](https://twimlai.com/go/674).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:05All right, everyone. Welcome to another episode of the Twinmill AI podcast. I am, of course, your host, Sam Sherrington. And today I'm joined by Akshita Bagia. Akshita is a senior research engineer at Allen Institute for AI. Before we get going, be sure to take a moment to hit that subscribe or follow button wherever you're listening to today's show. Akshita, welcome to the podcast. Thank you for having me. I'm excited to dig into our conversation. We'll be talking about your work with Ulmo, which is a language model and ecosystem really released by the Allen Institute. But before we dig into that, I'd love to have you share a little bit about your background and how you came to work in the field.
0:47So I graduated in 2015 with a degree in information and communication technology. And at my first job, I was building a platform for constructing financial data sets and models. This is software development at the periphery of machine learning. The company that I worked for, they also did a bunch of other cool ML projects. And that's kind of what got me interested in machine learning. And so then when I decided to pursue a master's degree, I applied to universities that had professors working in NLP. During my master's at the University of Massachusetts in Amherst, I worked on a bunch of different projects within the field of NLP, in biomedical NLP and digital humanities.
1:31And after graduating, I joined AI2 as a research engineer. Since then, I've worked on a lot of different things. I've worked on the original Allen and NLP library, a bunch of other general purpose libraries, and then a lot of research projects within various subfields of NLP. And then most recently, I've been working on the Ulmo project, or rather various projects under the umbrella of the Ulmo project. Talk about the project and the key motivations for Ulmo. So the Olmo project began at the beginning of last year. And one of the primary motivations really was the fact that a lot of language models were being developed behind closed doors.
2:14So these are really powerful models, but they were either hidden behind APIs. Or if the model weights were released, there was no information on the training data or any other pre-training details. And the thing is, these details are really important when you want to study these models scientifically. And so we felt it was really important for the research community to have access to truly open language models. And that was kind of the original motivation behind it. There's also the fact that a lot of this type of research is really expensive to conduct. And so if we are able to share as many details and findings as possible, that's just good for good science.
2:54like you don't want n number of different groups working on language models running experiments that take months to run and they are really expensive with gpus and so on and then you arrive at the same conclusions or you rediscover the same knowledge that should really be shared and that's also an important motivation for us so almost stands for open language model and it's a set of models released alongside the pre-training data. The toolkit used to create that pre-training data, the training code, the training logs, including weights and biases logs, the evaluation code, and the instruction tuning code.
3:34So currently the set of models that have been released are the 1B and 7B versions of the models. I believe we should also have the instruction tuned models out probably by the time this episode is out. And we also have a 65B model currently in training. With all of the models out there, it's easy to look at this as, you know, if you look at it just from the perspective of a model, oh, it's another 7B model. You know, it's like Lama 2, but late. But really, the differentiators are in this ecosystem of componentry and surrounding tooling that you're releasing. Before we get into all that, though, I'm wondering if you would even push back on comparing the models one to another.
4:21How does OMO as a model standalone compare to other models that are out there that might be similar? There are two parts to this answer. One is that on a lot of standard benchmarks, OMO performs as well or slightly better than a lot of models that are slightly comparable. I would also like to emphasize that all of the models that have been released prior to Olmo at varying degrees of openness, we have learned a lot from them. There's the Pythia set of models and there's even Lama, regardless of how open it was. There's the Bloom model. And so we have learned a lot from them. And if tomorrow some other group uses our findings or uses our pre-training data or uses our code base and comes up with a better model, that's really still good for science.
5:08The goal ultimately is to build better models more collaboratively rather than, you know, be on top of a leaderboard for like two weeks. The point is not necessarily to push this as the best model in and of itself and independent of the rest of the ecosystem, but really to point out all of the things that the research community would want to have from an open access perspective and to, for the first time, provide all of that in addition to the model itself. Yeah, absolutely. I think that having that kind of shared knowledge is really important for the community. One of the most important things that you released and a big differentiator, again, relative to a LAMA2 is that in addition to the weights for the model, you produce or publish the data set, which is called DLMA.
6:01Can you talk a little bit about that data set and how would you expect a researcher to use the data set now that you've released it and kind of what all they should expect when they go looking at the project? Yeah. So I think open training data sets are really crucial, especially for certain kinds of research. So for example, understanding the relationship between the inputs and the outputs of the model, understanding the model capabilities, how well it does on certain types of inputs, how well it deals with toxicity, whether it has seen any toxic content. And it was really important for us to release the training data alongside our models to enable really these kinds of research directions.
6:46And not just the training data, but also the way we create and curate this training data. So DOLMA is really a two-part release. DOLMA stands for data to feed almost appetite. We like fun names. Yeah. So there is DOLMA, the pre-training data set that we released. It's around Sheetal and Drogans. And there's Joel Mother Toolkit that we released, which really is a set of curation filters for quality and personally identifiable information and so on and so forth. It was very important that we were able to release this data completely in full. And so we purposefully chose data sources that were already accessible to the general public.
7:31And that was kind of one of the key motivations behind choosing the sources. So the data contains mostly common crawl, C4, code data, Reddit, academic papers, books, Wikipedia. You point out that this is all data that's accessible to the public. Is it data that was previously collected in existing datasets and you're collecting them in this umbrella dataset? Or is it public data that you've crawled yourself? So some of it was already public, like the code data set. We used Stack for this. But then some of the other data, like CommonTroll, we spent a lot of time cleaning that up and curating that.
8:12We also spent quite a bit of time curating the Reddit data source. Can you talk a little bit about the curation process and how you went about that? There's like an initial filter for the language. Dolma is primarily an English data set. And the only other languages are code languages. So there was at the first step, the language filtering. And the second step, there was like a basic quality filtering. Again, we don't really have a good definition for what we mean by quality. So we went with prior work and some of the general heuristics. Like, for example, we don't want a document that just contains a single vowel, for instance.
8:53So stuff like that. And then there was some degree of content filtering following that. So I think the two things that we focused on was toxicity and then removing personally identifiable information. The other interesting thing that we also took advantage of was the What's in My Big Data project, which is also something that AI2 has worked on. And it's really a way to analyze really large scale data sets. And so we use that as one of the steps to check for contamination with common evaluation benchmarks. Meaning you're checking to see if the evaluation benchmarks have contaminated the data sets such that you no longer feel confident running those benchmarks on the data set because the results are already in there?
9:39Yeah, exactly. It would be good for leaderboard numbers, but it's not really good for science. Right, right. Maybe taking a step back, I'm wondering how you thought about kind of task specificity for the model. And how do you think about the whole universe of possible tasks when you're building a language model like this? So it's, first of all, really impossible to be completely comprehensive and solve for every single task. I think as a first step for Olmo, we wanted it to be something that was comparable to other models in the same space. So that really gives us a way to talk about models in context of the field more broadly.
10:26So we focused on some of the standard benchmarks that were already being used. And then we also focused on code for primarily that reason. You mentioned the importance of this for researchers. Do you see it as being important for industry practitioners as well who might want to use the model or incorporate it into their products or fine tune it? Does having access to the training data set help in that perspective? So there are two ways in which that could be useful. First is for industry use cases, it's really helpful to know what kind of data the model has been trained on. So say, for example, you are using Olmo and you have a very specific use case in mind.
11:08It's helpful to know if the model can actually perform well on that use case. If it has seen, for example, political news, for instance, and if you know that that's available in your training data, then that's a good thing to know. So there's the knowing what goes into your model aspect of that, that can be useful. And then there is also this aspect of wanting to build your own models or maybe fine-tune on your own models. And because this data set is open and because this data set is directly usable, that's another aspect where you can directly use your data set, use this data set and apply it for your own models.
11:47One of the approaches that's evolved for creating a degree of transparency in training data sets without actually providing the data are things like data sheets for data sets and model cards and those types of artifacts that try to statistically describe the data set. You mentioned that you are providing a set of tools that it sounds like attempts to provide some of that same kind of information. Or do those tools produce a profile that is kind of as robust at helping a user or a potential user understand the data set as what they might expect from a data sheet? Yeah, I think it's definitely a bit more informative than a data sheet might be.
12:38It really allows you to look at individual instances in your data. And, you know, in the old times in LLP of like five years ago, we were building these really small data sets that were very highly curated. And that's no longer the case now that we are building these really large scale data sets. And the larger the data is, it's really difficult to actually know what's going into it. It's really difficult to point to a particular document and say, yeah, I know exactly what's in there. And so having automated tools that can describe this data, that can describe what the distribution looks like, that's really helpful.
13:13And so the What's in My Big Data toolkit really provides a way to do that. It's funny to hear you say the old times of NLP as five years ago. It's almost like two years ago as the old times of NLP. Yeah, I think last week was old by this point. Yeah, yeah. Yeah. I guess I'm wanting you to talk a little bit about some of the things that you learned as you tried to kind of train this model. You know, certainly Allen Institute has a lot of experience with, you know, building and training large scale models. At the same time, one of the challenges that you're trying to overcome is that people present these large scale models and some of their results, but don't give you quite enough information to reproduce it.
13:59So you're having to, you know, learn and bang your head against things along the way. So I'm assuming that, you know, there are some things that you had to bang your head against as a team along the way and producing this model. What were some of those things and kind of what did you learn about the process? Yes, we do have a training specific paper coming soon. And that's really going to get into the weeds of some of these findings that we had. I think the general picture is that there are so many confounding factors and so many different variables. and it's really impossible to ablate in all possible directions.
14:33I think a couple of the anecdotal findings that we had was that at the 1B scale, things are different than at the 7B scale. So one of the things that we found anecdotally was that wait time seems to really cause a lot of instability in your loss curves. So that was one thing that we found specifically for the larger model. You said wait time? Yeah. Weight tying. So this is when your embedding layer and your output layer are sharing the weights. So this works well for the smaller models, but not quite as well for the 7B model. I think the other thing that we found, which is a little bit different than what other prior work has found, which is that we really had a lot of issues with parametric layer norms.
15:25and so we decided to use non-parametric layer norm. This is different than what the other compatible models have done and I think what really drives the point home is that you really don't have a way of concretely saying right now why something works as well because you can't apply it very thoroughly so you end up with a lot of anecdotal data. I like to think that the Olmo findings are one data point in that research direction. and then there's a lot more to learn there really. I also wanted to share like one somewhat of a fun anecdote and this is like I really love this finding. So early on when we were starting to train the models we were observing a lot of irregularities in our training curves and we narrowed it down to being a function it being a function of the order of the training sequences.
16:20So this order was supposed to be randomly shuffled And yet somehow we are observing this weird pattern and we were not sure where that was. And ultimately, what we found was that Torch's random number generator, it's not quite as random for generating permutations. And I just love this finding because we are stuck on an experiment for two weeks and we are having all these discussions about model architecture choices. And it turns out to be something as invisible as the random number generator. and usually like knowledge like this is not something that usually makes it into a training paper like this is one of those things that you maybe share on twitter and and it gets lost and it's good to good to record that somewhere you would have thought that with all of the models that have been trained you know someone would have done a pull request in torch and it would have been fixed yeah exactly you mentioned weight tying and non-parametric layer norms i'm wonder if you can go into those in more detail and, you know, kind of talk about, you know, what they are and how they come up in training LLMs and, you know, some more specifics around what you found.
17:28Yeah. So I think the thing that I want to emphasize is that we took a lot of architecture decisions based on prior work because we wanted our model to be comparable. So a lot of those choices came from prior work's hyperparameters. And then once we started training is when we started having these loss spikes. And that's kind of what started us on this path of figuring out what actually is working for our case. Have you tested the thesis and given everything to someone and tried a clean room, reproduce everything that you've created? so in theory we would have always like clean ablations for every single hyper parameter and every single design choice in practice because these experiments are really expensive to run uh it's it's like a trade-off between the competition budget you have and and the time you have and what are some of like the best practices that you have seen before um and so it's really hard, which is why I think sharing this and having a shared research knowledge source, that's really important because we ran some set of experiments.
18:40The next group who's working on this doesn't need to run the same set of experiments. They can update on other things. And so collectively, that gives us better models in the future. How complete do you think what you're providing is? Could someone take it off the shelf and go and with sufficient time and budget, just kind of do what you say, follow all the steps, use all the tools, use all the data and pop a model out of the end? A, has that been done yet? And B, how, it's almost like from a coverage, how much coverage have you given of all the corner cases? Or do you think there are probably still things that they would have to bang their heads against to get everything working?
19:23I think we have tried to give as high a coverage as possible. We have shared the configurations that we used for all of our released models. We have also shared the training code, not just the inference code. So you can actually run the same scripts that we ran. And in theory, you should be able to get the model that we got. In theory, so it hasn't been done just yet. It depends on what you mean by hasn't been done yet. We ran experiments ourselves using the same code. We did run more than one experiment. In that sense, I would say it's been fairly well tested. But of course, it's not a perfect world.
20:01Talk a little bit about some of the work that you've done and published around evaluation. So for the Olmo project, we used both in-loop evaluation to validate some of our design choices, as well as offline evaluations for evaluating checkpoints at the end. And so for both of these, we used a mixture of commonly used downstream tasks. So these are mostly formatted as rank classification tasks. And we also use tasks to measure the intrinsic language modeling fit using a public city benchmark. And you published some of this as the Paloma project? So Paloma is the part that's the public city benchmark.
20:46Paloma stands for public city analysis for language model assessment. And the motivation really is that evaluating on downstream tasks gives you some notion of how the model is doing. But it really gives you a sense of how that model is doing on a task that's framed in a very particular way. So for example, for rank classification tasks, you are essentially giving model a set of possible sequences and you're picking the one that the model assigns a higher probability to. And there is some work now that suggests that based on how you frame the task, the performance of your model still varies, whether you frame it as a Q &A task, whether you frame it as a generation task, whether you frame it as a classification task, and the performance kind of varies.
21:35So downstream tasks are a useful metric, but they are not a complete metric. They are not giving you a complete picture of how well your model does. And so to address that, perplexity is a really common metric that is used to measure distributional fit. And so we kind of wanted to have that. But public city measurement on a large monolithic data like C4 can be a very crude measurement of language modeling. But public city really is a way to measure how well a model models a particular distribution. And when you have a distribution as diverse and varied as common troll, for instance, what does that really mean when you say that public city improved with scale?
22:22is it improving as well on every single domain in common troll or is it something that only a few domains are carrying forward the performance so for example if your model is trained on web text and it has seen a lot of political content it might be doing well on data of that nature but maybe it doesn't do quite as well on medical literature and so having that kind of fine-grained understanding of how well the model knows about different distributions. That's really useful information when you're deciding whether or not to use a model for your particular use case. How are you presenting that? Is it in terms of benchmark performance on domain-specific benchmarks, or is it formulated in a different way?
23:12So Paloma is a benchmark, and it's a set of around 600 domains and these are collected from 18 different data sources so by source I mean something like c4 so where the documents they belong to it because that's how the data was curated and then there would be something like a domain within that source so say for example political text or say books so those would be counted as like more fine-grained domains And so Paloma gives you these different fine-grained domains. And these are constructed by metadata that already exists in the source. For example, URLs or tags for academic field. And then for evaluating the language model fit, we compute perplexity on these fine-grained domains and use that as a proxy for model familiarity with a particular domain.
24:10So what the benchmark gives you is the data for these fine-drained domains, as well as a set of guidelines that you can use to actually compare your models or compare your training data. Are those evaluations for the final trained model or are you also providing evaluation during the training process? We also use it to provide some signal during the training process. I was wondering if you can give an example of some of the data that you're providing to that effect and how folks might use it to help ensure that their own training efforts are staying on track. Yeah, so I would say with Paloma specifically, I think the idea is to provide a more nuanced understanding of how well your model is doing on particular domains.
25:05and so it really depends on what your particular use case is. So for example, if you are a digital humanities researcher and say you're analyzing murder mysteries written in the 19th century, it's helpful to know that your model does well not only on books written in the 19th century, but maybe also is a little bit familiar with some toxic content because there might be violence. So it's really about picking and choosing the kind of domains that are useful for you. you're meant to get creative with this really like it's not meant to be a prescriptive set of like these are the 600 domains that you should be using and that tells you everything i think one of the key advantages of having a benchmark like this is that researchers who are not building these models even they are maybe even just using these models and they want to take more informed decisions they don't really have to create a very curated task and then evaluate the models on that task, what they can just do is provide essentially text for their particular domain, say medical literature, and then use that as a way to assess whether a particular model is suitable for their needs.
26:15How does Paloma compare to other approaches to evaluation that are out there like Helm and others? So Helm and Luther and so on frame their tasks in a particular way. So you say, for example, you have a Q &A task or you have like a generation task. Mostly it's rank classification. Paloma is really meant to be a bit broader than that. It's really telling you how the model is representing a particular distribution. It's never going to be perfect, but it is giving you an additional data point. Is the idea that a researcher who's using Ulmo and using the datasets would use those, what you've provided in conjunction with other tools like Helm?
27:02Yeah, absolutely. Again, it really depends on what your use cases. I guess I was curious to shift gears a little bit and talk about your broader perspective on kind of open and open sourcing models and data sets versus keeping models and data sets closed and protected from bad things happening. How do you think about that balance and the importance or danger of a project like Alma? I think that closed does not necessarily mean safe. There is always going to be a risk of malicious use with anything that you release. And even with closed systems, that's still possible. and the additional problem with closed systems is that even if you want to make them safer because you don't have any information about how these models are constructed it's really hard to create even guardrails around that i think for even developing better guardrails or developing better safer more ethical models it's still important to have that discourse in public and really ensure that the research community as a whole knows what's going on and so for that reason alone I think open sourcing models is really important.
28:17And not just the models, as I said, but releasing the entire process. I think that's important. Do you see the benefits in that regard of Ulmo and what researchers can learn from Ulmo as being limited to Ulmo? Or do you think that they extend to other models? With Ulmo, I've got the benefit of the training data. I've got these evaluation tools. You know, are they only going to, you know, help me explore, you know, the model that's created as a result of all these tools? Or do you think because you've built the model architecture on some of these other prior work, do you think that the things that they would learn could potentially extend to, you know, models of a similar class or type?
29:00Yeah, absolutely. I don't think that this is really meant to be an Olmo-specific work. All of the things that we have released alongside Olmo, I think they are useful resources in their own right. So for example, the pre-training data we talked about, it really can be used for conducting all kinds of analysis and also for building newer models with that data or by curating that data a bit more. All of the training code that we released, the evaluation code that we released, the All of these are meant to be more general purpose. So we use them for Olmo because that is what we were building, but really they're not really tied to Olmo in any particular way.
29:39So I think there's still advantage to having all of these resources. Can you talk a little bit about the instruction tuning part? I think that's not been published yet, but it's coming soon. Can you go into a little bit of detail into what you're doing there? Yeah. Yeah, so this is primarily based off the work that was released by AI2 in the Dulu project. So we are building off of that and the instruction tuned versions will be similar to how we do Dulu. Where do you see Omo and this project as all going? So the first set of Omo models was really the first step. We have a lot of interesting research directions, a lot of things that are cooking.
30:19We already have newer versions of the model in the pipeline. We have the 65B. We also plan to add new modalities. So yeah, there's a lot of interesting research directions that we want to take and we're excited about all of them. Should we expect to see OMO kind of track other popular, less open models? Like, should we expect to see an MOE style model come out of OMO, you know, 8x7B OMO or something like that or as i mentioned there are always there's always too many research directions and not enough time or not enough resources so what we really want to do is enable that kind of research so that we build certain things and then other people build even more things on top of that no expectation that ai2 is going to do everything it's uh we're providing some foundational steps and hopefully folks will take it and innovate on it it's really not possible to do everything the goal really is to build better models collaboratively, we are carrying the baton forward in some way.
31:22And then there will be other folks who will carry it forward a bit more. Well, Akshita, thanks so much for joining us to share a little bit about OMA. Yeah, it is great being here.
From the publisher
Today we’re joined by Akshita Bhagia, a senior research engineer at the Allen Institute for AI. Akshita joins us to discuss OLMo, a new open source language model with 7 billion and 1 billion variants, but with a key difference compared to similar models offered by Meta, Mistral, and others. Namely, the fact that AI2 has also published the dataset and key tools used to train the model. In our chat with Akshita, we dig into the OLMo models and the various projects falling under the OLMo umbrella, including Dolma, an open three-trillion-token corpus for language model pretraining, and Paloma, a benchmark and tooling for evaluating language model performance across a variety of domains.
The complete show notes for this episode can be found at twimlai.com/go/674.




