AI Fundamentals: Datasets 101

17 Jul 2023 · 1 h 1 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Latent Space: The AI Engineer Podcast

Episode Title

AI Fundamentals: Datasets 101

Overview In this episode of *Latent Space*, hosts Alessio and Swix delve into the fundamental aspects of datasets in AI. This is the second episode in their "101 Track," following the previous episode on benchmarks. The discussion centers on the importance of datasets, their construction, and how they influence the quality of AI models.

Key Topics Discussed

  • Common Misconceptions About Datasets
  • The popular claim that "GPT-3 was trained on the entire internet" is refuted, with clarification that GPT-3 was trained on about 600GB of data, primarily sourced from Wikipedia, books, WebText, and CommonCrawl.
  • The Significance of Quality Data
  • Emphasizes the principle of "Garbage in, garbage out," indicating that the quality of AI models heavily depends on the quality of the datasets used for training.
  • The episode touches on the challenges of acquiring high-quality data, particularly with restrictions from UGC platforms like Reddit and StackOverflow.
  • Tokenization and Its Importance
  • Explanation of what tokens are and how they are utilized in AI models, including the significance of token representation in language processing.
  • The hosts explain the differences in token efficiency across languages and the implications for training costs.

Critical Issues Addressed

  • Data Scarcity vs. Data Quality
  • Discussion on the notion of a "token crisis" and debates among practitioners and academics regarding data availability.
  • Scaling Laws
  • Scaling laws for models (e.g., Kaplan and Chinchilla papers) are discussed, showing how the size of datasets directly affects model performance.
  • Dataset Construction
  • Overview of major datasets used in AI, including:
  • Common Crawl: A major web dataset that is not exhaustive and has quality issues.
  • C4: A clean subset of Common Crawl, curated for language model training.
  • WebText: Sourced from Reddit submissions, with a focus on filtering to enhance quality.

Notable Datasets and Their Characteristics

  • Common Crawl: A fundamental dataset with 3.1 billion web pages but known for its quality limitations.
  • C4: A 10% sample of Common Crawl, filtered for quality, containing substantial chunks of high-quality text.
  • WebText: Curated from Reddit to ensure higher engagement levels, but lacks transparency in its cleaning process.
  • Books and Code Datasets: Highlights various sources for training data, including open-source books and permissively licensed code repositories.

Insights on Data Processing

  • Deduplication and Contamination:
  • Discusses the implications of duplicate data in training datasets and the potential for models to overfit on frequently repeated information.
  • Copyright and Privacy Issues:
  • Addresses the legal implications of using copyrighted material in training datasets and the challenges faced by organizations in navigating these waters.

Conclusion The episode concludes with a strong emphasis on the need for more dataset creators and the significance of understanding datasets for building effective AI models. The hosts encourage listeners to appreciate the foundational work done by dataset builders and to engage thoughtfully with the materials available for training AI.

Resources Mentioned

  • Various academic papers and model training references, including:
  • [Token Crisis Paper](https://arxiv.org/abs/2305.13230)
  • [OpenAI Tokenizer Tool](https://platform.openai.com/tokenizer)
  • [Chinchilla Paper](https://www.deepmind.com/publications/an-empirical-analysis-of-compute-optimal-large-language-model-training)
  • Links to dataset repositories such as Hugging Face and Common Crawl.

Additional Notes

  • The episode reflects a growing awareness of the challenges and considerations necessary for effective AI training.
  • The hosts maintain a commitment to keeping the content relevant and educational, suggesting a continuous evolution in the landscape of AI development.

For further details, listeners are encouraged to refer to the full show notes available at [Latent Space](https://latent.space).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:09Hey everyone, welcome to the Latent Space podcast. This is Alessio, partner and CTO of Residence at Decibel Partners. I'm joined by my co-host, Swix, writer and editor of Latent Space. Today, we finally, finally, we have the second episode of 101 Track. This has been a long time coming. So last time we did Benchmarks 101, I think we got a lot of mileage out of that. We understood a lot about the benchmarks, and we talked with a lot of our guests over the previous episodes about how they evaluate their models. And today we wanted to dive into datasets, what they are, how they're constructed, and why they matter.

0:44I guess I should go into why we wanted to do this episode. It's a little bit weird to separate datasets and benchmarks. So we did benchmarks first, but a lot of the benchmarks were datasets. So pretty much they're one and the same thing, right? And I think where they start to diverge is a cause of significant interest. But mostly, actually, I wanted to focus on data sets for one primary reason, which is that many people say that GPT is trained on all the internet. So first of all, this is actually not true. And second of all, it actually causes some potential misperceptions. I say potential because there is some legitimate debate about this.

1:26There are misperceptions about us running out of data. And we can discuss the pros and cons of whether or not we are running out of data. It's been named the token crisis by academics and quite a lot of commentators on AI. So in the show notes, we're going to link to a paper on to repeat or not to repeat insights from scaling LLMs under the token crisis. And then I'm also going to link to an opposite view from OpenAI with Ilya Suskever talking about how they're not anywhere close to running on the data that they want to train on. So whenever there's such an interesting divergence between practitioners and academics, I think it's a worthwhile thing to dive into.

2:06And just in general, I think there's a lot of foundational knowledge that people skip over when they assume that everyone knows what data sets we're talking about. Yeah, I was going to say, I think also in terms of like the knowledge that the models have, if you say it's been trained on the whole Internet, you would assume it knows everything on the Internet, but it obviously doesn't. you can go in there and ask about people that have online presences and are not actually in the knowledge base of the model. And this also helps when thinking about what data to then use to fine tune. So if you understand what's in the model, if you're trying to build a verticalized model for a specific use case, you can better figure out what's actually going to be meaningful versus what was already present in the first training run.

2:49Yeah, just for some comparison, in, let's say, the total size of the internet, some people have estimated at around 5 billion gigabytes. Most of the data sets that we are going to talk about today are in the hundreds of gigabytes range, and it's growing every single day. There's always new data being created every single day, and there's always new modalities to claim that data from. So a lot of the whisper behind OpenAI's whisper is that they're actually transcribing YouTube, which is a source of extra tokens. And we'll have to explain what tokens are. The first thing is about whether or not we're running out of data.

3:28The second issue is this divergence between data sets and benchmarks. And I wanted to dive into this specifically because they used to be essentially one and the same thing. In a very standard machine learning tutorial, you would do something like the IRIS data sets. And then you would do train test validation splits. and you would basically evaluate data based on samples from the data itself. But more recently, we actually have decoupled benchmarking from the datasets they're trained on, except for the calculation of loss. I think in our Discord, in Lanespace Discord, we've actually been doing a small paper club where we've gone over some of the foundational papers.

4:10And we actually recently went over the BERT paper, which is the bidirectional transformers paper that was a predecessor to T5 and is a predecessor to all the large language models today. And in BERT, they actually invented this concept of masking, which meant that datasets could create their own training objectives, which I think is super interesting. So basically what you have is, for example, a sentence. And out of the sentence, you mask one word and you ask the model to predict that word and you grade the model based on whether or not it's able to predict that word. This basically starts to go from supervised learning, where you have a data set that you're trying to train on to sort of self-supervised learning where you can just kind of let loose on unlimited set of data.

4:55And so this basically lets you scale as much as you want on the data side, as much data as you have, which I think is just really interesting and foundational. Like you don't have deep learning without self-supervised learning. You don't have a good training objective until you have the concept of masking. And once you have really, really good masking, then you start to find algorithms for that and it turns out that deep learning is is the way to do it to uh to achieve lost unseen in by any other algorithm and then once you have masking you can predict and then you can generate and and then you know everything kind of follows from there but that is the data set and not a benchmark right because you can have all of wikipedia as a data set but nobody bothers the benchmark on wikipedia because that's not a reasonable benchmark to derive any score on.

5:43But it turns out that training on a dataset leads to higher evaluation on benchmarks that we covered in the last episode, like common, what was it? Hella Swag, Big Bench, MMLU, Helm. Yeah, exactly. So I just think that's fascinating. We've just summed up, now it seems so retroactively obvious, but we've just summed up maybe about 10 years of progress on deep learning. Yeah, and finally, the But the most important thing about data sets is, one, understanding how big they are and then using scaling laws to work yourself back into what size model you can train with them. So we talked about the chinchilla scaling laws and we'll cover that later.

6:23But if you want to train a model that is 100 billion parameters, you cannot just pick any amount of data. It needs to be a lot of it. So understanding common crawl, how many tokens is that? c4 how many tokens is that helps you understand okay this is what i can get off the shelf this is what i need to provide and we'll dive into more of the data sets later but first i just wanted to do a quick explanation of what tokens actually are so the thing you read on the open ai docs it's like one token is like three quarters of a word so when i first read it i was like oh you're just doing character splitting but that's not really how it works so basically one token is a integer that can be up to, I actually don't know what the highest number will be, but it's an integer representation of words.

7:11And the same word can also have different representations based on where it is in a sentence, for example. One of the big things that Transformers did to be more efficient is the space is included in the token. So if you're doing a long sentence, instead of having one token for each space in it, the space is inside the token of the word. So it really cuts down on how many you actually need to go through it. But the funny thing is that then you have different representation for the same word. So if you take the word red, like the color, there's one token for lowercase red with a space in front.

7:48That's token 2266. Then you have red with a space in front with a capital R. That's 2297. Then you have red with a capital R, no space. That's token 7738. So you can see that when you say a trillion tokens, for example, it's not one trillion different English words. So that's also one of the main things that you got to be careful of. You can have a data set that is a lot of tokens, but has potentially a lot of repetition that is not as helpful. And also when you're doing things like a logget bias in OpenAI, where you can deprioritize certain tokens, you have to find all the tokens for that. So the example that they use is if you want to do a recipe for a cake with no eggs, you have to set the logit bias of both egg and like space egg, egg space tokens, not just one.

8:41So that's one thing to keep in mind. If you're coming from a non-ML background, that's probably one of the first gotchas that you have when talking about data sets. And just for more examples, OpenAI has a tokenizer tool that we're going to link in the show notes that you can use to see what any particular phrase translates to. So I just plugged in, for example, latent space. Latent space is three tokens. It's L-A-T, that's the first token, LAT. And then second token is E-N-T, E-U-N-T. And then the last one is space and then S-P-A-C-E. And that's the last token. So latent space is three tokens, and some combination of that will form other words as well.

9:21And it's an interesting tokenization scheme. As far as I understand, by the way, Alessio, I think the upper limit on tokenizers is between 50 ,000 to 80 ,000 tokens, which is amazing. It means that you can basically make up any sentence or language from these individual tokens. It's a small number. It's actually like five digits of numbers, not millions and millions of tokens. And keep in mind that they speak multiple languages. So you have to actually tokenize other languages as well as numbers, as well as symbols, emojis. I actually am pretty amazed at the depth of tokenization. And honestly, there was actually a well-known flaw with GPT-3 where you could actually do a quick test to see if you're talking to a bot or not.

10:07One of the reasons that GPT-3 is just not good at math is because they don't tokenize numbers individually. and they don't represent numbers the same way that humans do. Humans represent numbers digit by digit. But if I type in into this tokenizer, GPT-3, like I type in one to this, that's one token. If I type in one, two, which is 12, that's also a different token. And that's a single token still. And if I type in one, two, three, so it's 123, that's also another single token. But if I type in one, two, three, four, that is now two tokens that is made up of 12 and 34. So it's breaking up one, two, three, four, like an English word rather than a mathematical representation of 1 ,200, 3 ,10s, and 4 ,1s.

10:50And that's one of the reasons it doesn't do math very well, because it looks at things as though it's a word rather than numbers. Yeah. And you mentioned the language thing. That's another good point. Some languages are actually less token efficient than others. For example, Spanish actually requires more tokens for the same sentence. And they also have this issue where the lower the number of the token, the more common the token is. And if you tokenize a Spanish phrase and you have syllables like men in it, the token value of the token is very low because the English language is used very often.

11:27So you can have predictions that are a little weird when you use different languages. And we'll kind of go into language-specific data sets later as well. Yeah, there's a famous article about this called, why is GPT-3 15.7 times more expensive for certain languages? And so the English bias is real. So I want a pizza, it's four tokens in English. And then in French, I don't speak French, but je vous un pizza. That's seven tokens. And in Chinese, it's 15 tokens. I want a pizza. I don't actually know what these characters are. But anyway, so it's interesting, right? Like all of this is represented within the token space that GPC has trained.

12:11And it's actually pre-trained. This is one of the first examples of pre-training being useful. Because ultimately, language models or the transformers that power the language models are transforming a set of numbers to a different set of numbers or predicting the next token ID in a sequence of numbers. The transformers themselves don't actually know what word they're predicting. All they know is they're given some data sets of number after number after number, in fact, like trillions of them. And then their job is to predict the next number given a sequence of numbers. And then we take the tokenizers to convert those numbers into words or images or audio.

12:49We reference scaling laws a little bit. And basically this comes out of research around what the optimal sizing of a model should be for a given data set. I think in 2020, OpenAI published the first scaling law, which was the Kaplan paper. And that was estimated to be about 1.7 times tokens per parameter. So what that means is the reason that GPT-3 is a 175 billion parameter model is because they had around about 300 billion tokens to train with. So they worked backwards from the data set size that they had of 300 billion. They said, OK, based on this, the largest possible model that we can train is 175 billion.

13:29Let's go for that. That was the state of the art at that time. And we as far as OpenAI is concerned, larger models were always going to be better because of understanding around emergence and capabilities and honestly just research around what AGI could be. And just for reference, the GPT-2 was about 100 times smaller than that. So that's super interesting. And then the year after that, Chinchilla came out of Google DeepMind, where they actually optimized for a different metric, which is compute optimal training. So a given compute budget, a given amount of days and hours running a certain number of GPUs, so that gives you a certain compute budget in terms of flops.

14:10If you hold that budget constant, so let's say, you know that's roughly a few days or a few months of compute that is actually very material for a research team or anyone because that translates directly into dollars right like how much time are you renting on the shared gpus so for given compute budget what is the best model that you can train given some kind of compute budget and the number that he came up with was 20 times tokens per parameter so that's actually 10x what the captain laws were doing what deep mine did was they had sort of replicas of GPT-3 that they called Gopher. And then they created another replica called Chinchilla and showed that despite being about 10x smaller, they were able to match or beat Gopher.

14:56And the assertion they would also beat GPT-3 despite being 10x smaller. Most people call this is basically that GPT-3 was over-perimiturized. There was way too many parameters. 175 billion was way too much. And in fact, to train GPT-3 to a full 175 billion parameter model, you need 3.5 trillion tokens, not 300 billion, which I think is just fascinating and a little bit depressing. It means we just need a lot more data. Yeah. And it just shows you how early this whole foundation model space is. These papers are not coming every 10 years. They're coming every 18, 24 months. and the other thing you mentioned that with the cost of compute you know if you think that 1.7x is good you probably don't want to burn gpus for like a much larger training scale you know and now the the next the next thing is the llama optimal which is 200 times tokens per parameter so now to train gpd3 you need 35 trillion tokens which is like if you take all the books published each year, like all of Kindle Unlimited, all of that stuff, it's only 100 billion tokens.

16:06So to get to like 35 trillion, you need a lot of data. But again, is this the new optimal? We also don't know because it's not easy to find ways to train models at this size. Like it's not easy to find the computer. It's not easy to find the data. So once we record data sets, you know, 201 in a couple of years, we're probably going to say it was crazy that we were using like 20x. You know, it's actually like this iteration now. One thing I want to basically mention is that as far as I can tell, the researchers that I'm following talking about this stuff, because we've asked this question on the podcast every opportunity that I've had.

16:44And essentially, we're going from compute optimal, which is like a training time to inference optimal, which is at inference time. Right. LAMA was designed for an inference optimal situation where we start caring about the latency of the inferences. And basically, that just means the size of the models have to be smaller. They cannot be hundreds of billions of parameters. And they're all round about this double digit parameter size now, maybe even single digit with the 7B and 3B type models from MPT that we talked about with Jonathan Frankel. This is a nice evolution. It basically balances practical requirements.

17:22You know, it's very funny coming from a software background studying AI, because in software, performance means speed. Whereas in machine learning, performance does not mean speed. Performance means capabilities and evaluations on benchmarks, where inference time doesn't matter. But now inference time does matter because AI is crossing over from research into software, into practical applications. So now we do care about things like inference costs and memory that you need to run all these systems. So I just think it's just fascinating that Lama Optimal is purposely overtraining Chinchilla. DeepMind showed that Chinchilla is compute optimal.

18:00And now we're saying we just don't care. We will purposely overtrain and be suboptimal there in order to be more optimal in inference because inference is more important now. Yeah, exactly. The question is optimal for what? If you're writing a paper, you're optimizing for training and building the proof. If you're training a model for production use case, the training cost is actually just a small part of overall lifetime cost of the model. So the pendulum is going to keep swinging depending on the application. Maybe I'll go on to the next bit, which is LLMs is databases. Okay, this is interesting.

18:34So we have this concept of tokens, and then we have the concept of training in a compute budget. Basically, what the training process is, it can be abstractly viewed as a way to compress the data sets. For example, we have actually a nice conversion ratio between billions of parameters and the amount of data that they generate. So just as a rule of thumb, right? Let's say each parameter is 8 bits. Usually it's 16, right? We always talk about FP16 in our podcast with Jortats. But let's just say 8 bits is one parameter. then 175 billion parameters is 175 gigabytes, right? One billion parameters is one gigabyte.

19:14That is actually the definition of a gigabyte, which is one billion bytes. So that's super intuitive. And so therefore a full point, falling point precision 16 bits means 175 billion parameters uses 350 gigabytes to store parameters. That's how much memory that you technically need to do that inference and just to load the model itself. And, you know, most graphics cards will not even have that. Even the professional grade, like A100 cards, would be like 80 gigabytes for a single one of them. So to fit 350 gigabytes, you need to network all these things together with very high bandwidth. But they're trained on 3 ,000 gigabytes of data.

19:54So that is 3 ,000 gigabytes of data being compressed into 350 gigabytes of data. And that is a form of compression because from there you can sort of, it is lossy compression, but it's a lossy compression in a way that learns how to decompress itself such that when you ask it to spit out some facts, it spits out something in the approximate neighborhood of what you started with. So I think that that's an interesting analogy. Some people don't like this idea that LLMs are databases, but it's super cool. Another very prominent description of this I remember from last year, which was stable diffusion, which compresses all the images that it creates and all the images that it trains on into, I think, something like two to four gigabytes.

20:36Like it has knowledge of flying saucers and horses and humans and beaches and, you know, computers in images, all in like a downloadable file that you can host on your laptop, which is crazy. Yeah. I think people get very surprised by how many things the model knows, you know? And if you were putting the raw, again, going back to how much memory you need, if you put the raw data in memory, it would be like impossible. to actually run. But if you take, especially if you think about 16-bit versus eight versus four and like all the quantization work, it's like compressing the compression, you know, now all of a sudden you can put data sets and like knowledge that like was impossible to like put on a certain device.

21:19Like I can run the Red Pajama, like 2 billion model on my phone and my phone is like 64 gigabyte of storage. If I were to put all the data that got into the training of it, would be like running my phone way out of surge. Instead, it's only like a few gigabytes of model. So that's another interesting thing, especially as we think about using these models at the edge and how much stuff we can get there on time. But that's for another episode of Quantization 101. And for those who are, again, coming from our George Hatz episode, around the one hour mark, he talks about compressing humanity, a person's consciousness or knowledge or all your life experiences, How much information, how many bits of information is that?

22:02He thinks it's two gigabytes. That's probably too much compression. You'll probably be a really bad copy of yourself. But there is a point at which you should probably be able to replicate yourself as a digital twin, which is the whole mind-uploading phenomenon that we discussed. Cool. I think we sort of maybe knocked on that a little bit too much, in fact. I want to basically verbally go through this chart of tokens and scale because when we talk about, when we say things like billions of tokens, trillions of tokens, billions of parameters, it's just very, very big numbers that we don't know how to estimate.

22:35And this chart that we had from this person, S Rush on Twitter, actually was super helpful to me. I was actually planning to make something like this, but this guy already made it. So I'm just going to quote him. So maybe we'll just kind of go through the token chart. So this is just memorized tokens in terms of orders of magnitude. Being able to do token math, I think is super important. And order of magnitude math is super important when it comes to deep learning, right? Because everything is like so big that like the significant numbers don't really matter. It's just the order of magnitude really matters.

23:09Okay, so 10 to the power of zero, which is one, that's on the order of like a hello world, like individual word tokens, right? And then 10 to the order of one, which is like 10 tokens would be like, you know, one sentence, one phrase, or whatever like that. 10 to the power of two, that's 100. that would be blank space chorus apparently which by taylor swift yeah this is definitely a swifty because he talks about taylor swift later 10 to the power of three that's 1000 tokens uh that's the wikipedia article on fermi estimation 10 to the power of four again that's the taylor swift article on wikipedia 10 to the power of five that is the gpt3 paper itself including the appendices 10 to the power of six is one year of the new yorker so that's a big jump right 10 to the power five to 10 to the power six.

23:54That's a one order of magnitude jump, but we've gone from a single paper to one year's worth of magazine. 10 to the power of seven is the whole of Encyclopedia Britannica. 10 to the power of eight is number of Reddit posts per month. 10 to the power of nine is English Wikipedia, all of Wikipedia. 10 to the power of 10 is the number of WhatsApp messages per hour. 10 to the power of 11 is the number of published books per year. Not the number, the amount of tokens inside of published books per year. And then finally, we get to 10 to the power of 12, which is the order of magnitude that large language model data sets operate in.

24:27And so that's intended to power 12 is 1 trillion. I was mostly surprised by the fact that Taylor Swift's Wikipedia page is longer than the Fermi estimation one. But the amount of data that you need for these models is great. And yeah, I think the 1 trillion tokens is kind of like the MPT 30B that was released yesterday. So when we publish this episode, it's going to be pretty new. But that's the amount of tokens that they used for there. Yes, we're trying to keep this episode current, even though it's an evergreen episode. Yeah, that's how much was used there. But it just gives you, again, a way to reason about this.

25:03So when you read online next time, it's like a trillion tokens. You understand this like 100 times all of English Wikipedia, which is a lot of text to collect. It took us years to get that online. There's a question about whether we're running out, right? Is there more orders of magnitude to this? And arguably there are, but that is one of the issues that we'll discuss at the end. Just to recap, because I'm so excited about this, because this to me is an evolution in terms of smaller models. TPT3, a breakthrough that got a lot of us interested, was 300 billion tokens for 175 billion parameters.

25:38That's the 1.7 ratio from Kaplan. But then LAMA is 1.2 trillion tokens for a 7 billion per model. model. So one order magnitude higher tokens for one order magnitude lower parameters. That's crazy. It really is. And I think like, again, we already made this point, but it's like, we're just really early in terms of like how much data we need, what model size we need, that I wouldn't preclude use cases today based on the size of these models. And also for an enterprise, the other thing, the analogy that I like with databases is that when you get a database like, you know, Postgres, MongoDB, there's nothing in it.

26:16There's kind of like a cold start problem. Like you install the database, you need to start putting stuff in it. Now, machine learning in the past used to be similar, where before you even start using it, you need to collect a lot of proprietary data, you train your own model, and then you start to do the production. The way it works now is that the foundation model labs and researchers like OpenAI, Anthropic, and the likes, they've already done all the work for you. So you already get these models and meta, of course, we cannot miss one of our open source paladins. Once you get that at a company level, what you need to understand is, okay, I have this model that knows a lot.

26:54How do I prompt it and how do I fine tune it to make it good for my use case and then run my own inference on that fine tune plus prompted model? so the scale of data that you need as an enterprise is like so much lower like you don't need to collect billions and billions of like examples and tokens the model for example for code the model is already really good at code you just need to give it you know dozens of examples you know maybe a hundred if you want to be really thorough and you're gonna get really good performance and i know shan you were a fan of carpave's presentation at build conference yeah it was a really insightful and authoritative, I guess, coming from Andre.

Read the full transcript

27:37And recently at the Microsoft Build Conference, Karpathy had a State of GPT talk that I think was very well received. And he is just really good at outlining the important things in a lot of the mess that we have to wade through in order to understand language models. And so he basically outlined this famous slide, which is the GPT-assisted training pipeline that outlined essentially four kinds of data sets that go into making something like ChatGPT. And obviously he's extremely authoritative on that. And I just think it's useful to have this in your head, this slide in your head. Again, refer to the show notes if you want to see the image.

28:15But the four datasets are raw internet, which is just the raw datasets that we pull from Common Crawl and Wikipedia and books and all that. And the second one, which is the demonstrations dataset, where you're demonstrating ideal assistant responses. So basically prompt and response pairs. Comparisons. So basically comparing outputs between output A, output B and seeing which one's better and just sort of reward modeling that on the language model. And then finally doing reinforcement learning with props. And so those are different kinds of data that have to be collected in different ways. There's different orders of magnitude of them as well.

28:51So there's trillions of low quality, large quantity data from the raw Internet. and then there's tens of thousands of demonstrations, hundreds of thousands of comparisons because that's easier to just choose A and B, and then tens of thousands of prompts. So that's basically the kinds of datasets that they think is representative of the ChatGPT, like goes into building something like a ChatGPT. Do we want to at last get to the datasets part? I know we had a little bit also on instruction tuning, but we have a bunch of things to get through. So maybe you want to start there. I'll just quickly mention that instruction tuning was another paper coming out of OpenAI and obviously a very important part.

29:32That is under a subset of the demonstration stuff that we're talking about. There is some debate from the Lima paper, LIMA, about how much data we actually need. And so we interviewed Databricks. And Mike Conover is a very good friend of the pod. And they collected like 15 ,000 pieces of data to instruction tune themselves. Open Assistant from Yannick Kilcher also collected tens of thousands of pieces of responses to instruction tune-on. And there's always something on some research that indicates that maybe it's not so much as we need. So Lima is something that we'll call out as interesting there.

30:07But yeah, we can move on to the major datasets because this is Datasets 101. It doesn't have to be today in datasets, which can be a whole different podcast. So first we start with Common Crawl. That is the OG. That is the bread and butter of every single data set, including the image ones that we'll talk about later. So it was founded by this guy, Gil Elbaz. And I actually did some research on him. And then you fact check my research. So this guy, Gil, is actually, he has his own Wikipedia page so you can look him up. He basically started the predecessor to AdSense and sold it to Google and worked at Google for a while.

30:40And obviously, Google, and this is like in the late 1990s, early 2000s. and he saw firsthand how important, like obviously that Google's crawling was important, was for Google and basically quit Google and started Common Crawl. Like basically quit Google and started empowering competitors to Google, which is kind of interesting and scandalous. You did some research. What did you find? That's kind of like the beginning of the open web. So mostly the data was used for like surfacing pages and then the whole big data thing kind of came to be. And one of Gil's ideas was like, okay, this data is not only good for like Google search, like indexing.

31:22There's a lot more work that you can do with it. We're like, I think 20 years into it almost. Yeah, in 2008 was like when they first published the first data set. And it's one of the biggest ones out there. So there's 3.1 billion web pages, 400 terabytes of content, 43 million hosts. Only 46 % of the content is in English. Going back to our discussion before. It's run as a nonprofit. So there's a lot of unique things about it. I don't know if today you will see a nonprofit getting started just to provide massive amounts of data to the public, especially in this world of AI where everybody's hiding the data that they have.

32:02So maybe we do need it, but I'm not sure if we're going to get it. I think it's actually super interesting how it got started itself. This is obviously a very significant effort that all research, which all NLP research and all language modeling is downstream of. They were started in 2008, but I think they were stealth for four years. Because the earliest example I can find of them releasing any of the data was in 2013. Sorry, they called it 2012, but the press release is 2013. And it looked like they used to crawl once a year, crawl as far as they know, all the internet. And now they crawl once every two months, right?

32:37And it's just an interesting example of a nonprofit-driven approach that people don't really question or look into, but it's actually secretly driving all of the LLMs that we had today. There's some issues that are very well known in Common Crawl. So Common Crawl actually says that they only cover a fraction of the web. It's a nonprofit, works on nonprofit resources, doesn't cover all of the web. And all language models are trained on Common Crawl. Therefore, our language models are not trained on all of the web. It is also a biased sampling. It's definitely biased towards the United States. There's a lot of data quality issues.

33:14I think we talked about in some of our other episodes where the labels for some of the languages that Common Crawl has might be completely off. I think someone mentioned about the Arabic issues being tagged. If you actually look at all those pages that are tagged as Arabic, they're not Arabic at all. So just really basic coding errors or nobody checks these trillions of pages that are being crawled. So it's just really, really difficult. If you have robots.txt that blocks Google, you will also block Common Crawl. If the page is too big because of just the sheer amount of data that's on it or the images on it, the pages are deleted or if they're duplicated on multiple sites because of spammers, it's just really, really difficult.

34:00Or if your pages are written in SPAs as JavaScript because Common Crawl doesn't render to JavaScript. It only executes limited JavaScript. So, for example, much of Facebook is not under Common Crawl, right? And this is increasingly a problem with the closed gardens or the walled gardens of the internet, right? As information migrates from the open web into apps like Discord and Slack and Facebook, it's just not available to Common Crawl. No, no, no. Then Common Crawl has its own clean subset, so to speak, which is Google's C4, which was created during the training of their T5 model. Jonathan Frank on our podcast called it Weirdly Good.

34:40I think that kind of explains a lot about data sets. Sometimes they're good and we don't know why. And C4 is made by using a few heuristics. So it's about 10%, I think, of Common Crawl. Like it's much smaller. And it tries to filter Common Crawl by different ways. So one, it's using this open source thing called list of dirty, naughty, obscene, and otherwise bad words, which is 402 terms written in English and one emoji, which you can guess which emoji it is. The list was created by Shutterstock, actually. They basically wanted to avoid bad words to be autofilled in their search. So they created this blacklist of things that they wouldn't autofill for the user, which I think is a funny way to end up being one of the foundation pieces of modern large language models.

35:33The other thing is there's a lot of stuff that didn't get filtered out, like certain pieces of 4chan, like Kiwi Farms, things like that. Again, the episode is not about giving our judgment on the datasets. It's just about what's actually in them. and there was a Washington Post article that kind of went through the whole list that we're going to link in the show notes. The other thing that I found fascinating is that if you look at what domains are in the C4 dataset, the patents.google.com website is like twice as large as the second one, which is Wikipedia. And then Wikipedia is like three times as larger as the third one, which is script.com.

36:14So there's, you know, It's obviously like less than half of a percentage point, but it's still interesting to see very formal kind of like text as the largest represented one. But yeah, C4 is another obviously core data set. So if you're looking to train your foundation models, that's one you should check out. Yeah. And in fact, a lot of models will list both Common Crawl and C4 as part of their data sets. And it will be a very, very heavy weight. It will be something like 30 % to 60 % of the amount of the token budget that they have, which is super important. I mean, this is the starting point of all of our language models.

36:55It's really, really fascinating. Going off of a list, basically trying to reintroduce where GPT-3 gets its data sets from. So if you pull up the GPT-3 paper, we're basically going in order of explaining which of the data sets and telling a little bit of the story behind each of the data sets. Wikipedia is the next data set. And obviously, it's a very high quality data set because a lot of people have spent a lot of hours editing them. But except for the fact that Wikipedia itself has its own bias. I always have this fun fact that I pulled from Google here. 77 % of Wikipedia articles are written by 1 % of Wikipedia's editors, meaning there's just an extreme, extreme skew in terms of the representation of the kind of people that write Wikipedia articles and the decisions that are being made, right?

37:49In particular, this one guy, Stephen Pruitt, because he's constantly made the rounds as the highest ranking Wikipedia editor. He's made over 5 million edits and has made one edit to one third of all English Wikipedia articles. So if you want to seriously affect machine learning datasets or large language models, you should edit Wikipedia. That's funny. It's not just Wikipedia. Yeah, Reddit. Yeah, exactly. WebText is another major dataset that was also used in GPT-2. It's about 45 million links in the text of those webpages. the way they collect it is basically scrape every url from every reddit submission up to december 2017 that add at least three upvotes just to make it somewhat i guess like less spammy and they've removed all of wikipedia from it so there's no wikipedia in this it's all reddit links that are not wikipedia between 2017 and i forget when the release date of this was and then they did another a round of heuristic-based cleaning, which again is like, who knows what that means, which makes it kind of complicated to then scale these models, these data sets, right?

39:00Because if we knew the cleaning process, then we could say, oh, let's take all submissions from like 2010 and clean them up. But we don't actually have the step-by-step rules for some of them. That said, it's been replicated by the Luther organization, right? Luther is something, is one of the organizations that was consistently shouted out by our guests as doing really good work in LLMs. And so they've replicated web text from OpenAI. So OpenAI, as far as I know, did not release web text, did not really release the rules around them. But the Eleuther organization created OpenWebText, which is an open source reproduction of web text.

39:35And so we're going to leave in the show notes the link to OpenWebText 2, which is the latest reproduction of web text. We also mentioned the issue with Reddit, which is also a very hot topic right now with them shutting down their APIs. But let's keep moving on in terms of data sets. Next, we'll go on to books. So this is just basically, as far as I can tell, open source books, quote unquote open source books. It's a data set of 196 ,000 books in plain text for training large language models such as GPT. And it's also included in the pile, which we'll talk about later. But one of the interesting factors in the books data set is that apparently the copyright for some of them is not clear.

40:17So, for example, when Jonathan Frankel and Abhi trained MPT-7B, they actually got attacked by this person who was basically saying, how can MPT-7B or MPT-30B be commercially available or commercially licensed if the dataset it was trained on had some issues with copyright? And they went back and forth on it. We linked to that discussion on our episode page. But it's an interesting thing. if any of these are compromised, if they're not copyright-free books, then that's an issue. So we need to actually make sure that our data sets are clear such that our models are clear. Yeah, and then copyright expires.

40:56So hopefully the more time goes by, the more data we get. Yes, yes, yes. But unless Disney says they want more time and then Disney will extend the copyright. But as far as I understand, I think Sherlock Holmes copyright recently expired. So a lot of people are making Sherlock Holmes fanfic now because you can start using that. And I think maybe like Winnie the Pooh, maybe. I don't know. I forget what the recent expiries was. But every year there's an expiry day and a bunch of stuff becomes public domain, which is useful for, I guess, training and children's book storytelling. Okay, then we go from, everything we've covered so far is natural language datasets.

41:34But then we go from there to code datasets. And there's a lot of different code datasets. Salesforce CodeGen, I think I would point to as a predecessor to what I'm going to talk about. The most significant one in my mind is the stack from Eleuther. It's basically a scrape of a GitHub archive, six terabytes of permissively licensed code data. So the beautiful thing about code data sets is that most of the code on GitHub will include a license file telling you what license they have. And so you can just kind of pick the set of licenses that you think are permissively licensed and you can kind of just get them out.

42:07So the raw data was 102 terabytes from 153 million repos. and that's 320 times larger than Wikipedia. Then you clean them for file extensions. So for example, you can take out the images and take out blob files for whatever reason. And then you're left with 69 terabytes. Then you clean out the licenses and that's 90 % of code that's thrown away because they're not permissively licensed or they have no license at all. Please, please, please, by the way, if you open source any code on GitHub, please add a license file so that issues like this don't come up. And so we end up with six terabytes of permissive code data that anybody can train on.

42:43It's great. Yeah, and they didn't stop there with datasets. There's also the pile, which it's like my favorite dataset name. It's sort of like a pile of data. It's 825 gigabytes from 22, basically like 22 smaller datasets combined. So in there you'll find PubMed data, archive papers, GitHub data, the FreeLaw project, the Ubuntu IRC channel, Anchor News, YouTube. There's a bunch of things in there. The thing that we know about this from Stella and Luther from one of her tweets is that only about one third of the contents are duplicated. And we'll chat a little bit later about deduplication and whatnot.

43:27But they actually deliberately upsampled the original dataset to include some of the duplicate data. So even without that, the data set is quite large, but at 825 gigabyte of kind of mixed data is another one of the core data sets people use. Yeah, so obviously we can keep going and going and going. There's no end to these data sets. Alessa just mentioned a whole bunch of interesting data sets that we don't have time to go into, right? But you can always research them if you want to. We're just trying to give a one-on-one. But no one-on-one would be complete without mention of other modalities of data sets, So I'll just mention two of them, and then hopefully we can just kind of go on to issues.

44:08Otherwise, this will be a 10-hour lecture. LION is the Eleuther for images. LION stands for Large-Scale Artificial Intelligence Open Network. That is the sister organization that came out of Eleuther that Stability and Iman Mostak worked with to create stable diffusion. So this is all the predecessors of stable diffusion. And actually, it came out of COVID. Right. Everyone was kind of bored sitting at home and looking at Dali and going, why don't we have an open source replication of Dali? And so the first thing that you do when trying to replicate a model is you go collect data. Right. And so Lion in 2021 collected 400 million images from where?

44:46From Common Crawl. Right. Like Common Crawl keeps coming up. It's the OG. It's so goaded. And then in 2022, they released Lion 5B, which is five billion images for data sets. They've also released an aesthetic subset of the Lion 5B dataset. And basically, there's a lot of filtering when it comes to images. Obviously, there's a lot of porn that you need to filter out and not save for work. And also copyrighted images, right? Like you have to make some decisions around, do you want images of Spider-Man? And I was just at the Figma conference yesterday where Adobe was very proudly showing off Adobe Firefly, which has no idea what Spider-Man is.

45:25And for them, it's a feature because for the kind of companies that are using Adobe to create images, they need to not run into trouble with Disney. So it's fantastic. It's very hard to get images of Spider-Man out of your image dataset. And then Whisper is the other one. So instead of images, we need transcription, like ASR. So Whisper, if you look at the paper from OpenAI, again, is Liberty Speech, Common Voice, VoxForge, Switchboard, and Fisher Corpus. That's all I know about them. there's a lot. Whisper is so good. I feel like nobody's saying, let's do another Whisper. You know, I think the NLP and text part is the most active one right now.

46:04Yeah. So for those who want to like research more data sets, I think the best place to go is probably Hugging Face Hub. A lot of people for training data sets, they also go to Kaggle. I maintain a list of useful big data sets on my repo of useful resources. So I'll link that in the show notes. But yeah, I think that's going to be the high speed tour over all data sets. So again, we'll come back to the key question, why are datasets important? First of all, we have to know the fundamentals about datasets to figure out whether or not we're running out of it or how we're using it. But also, fundamentally, I'm a little bit concerned about the number of dataset producers to the number of the ratio of dataset consumers.

46:43Everyone wants the glory of training language models and saying, I train this model that does X. But not that many people are interested in cleaning data. right like it was just a common meme in the enterprise sort of data science world machine learning world that everyone wants to be data scientists no one wants to be a data janitor right so i don't know if you've like run into any of these uh conversations in your line of work yeah that that makes sense i think like especially now there's a lot of pressure on companies to use ai and everybody wants to use ai nobody wants to do the work of getting the data ready to make ai useful for their company so it definitely resonates yeah and And companies that are sitting on top of a lot of data are actually realizing that they are sitting a lot of gold.

47:27I call this data is the new oil part two, right? Bloomberg recently came out with Bloomberg GBT because guess what? They have the license on a few decades worth of Bloomberg financial news reporting. So if you ever need to generate or do any reasoning around financial data, Bloomberg has the best data set by far. And it's closed source and you have to pay Bloomberg to do it, right? Like Notion, if you think about it, Notion has 22 million users and all their users use their Notion as a knowledge store, right? So Notion has a tricky issue because Notion doesn't own their customers' data. They just hold it for them.

48:04And so they don't actually have the right to train on them. And so it's like this tricky dance between them. But if you enable, for example, customers to fine tune on their own company data within Notion, just with a single click, that becomes extremely valuable as well. So individual companies are realizing their moats. Just yesterday, the Stack Overflow CEO was saying they're joining Reddit in terms of shutting down their API access because they realize they have good data as well and they want to be paid. And that's a part of the issues that we'll go into. But finally, we should also cover the counterculture movement for open data sets, open replication of data sets.

48:39And Yannick Kildscher, I think, is one of the leaders here, as well as Eleuther, for reproducing the instruction tuning datasets that people will need to train their own chat GPT. It's pretty interesting that YouTube influencers are coming to the rescue of open source because there's no other source of influence that is powerful enough to compete with OpenAI except for YouTube. Yeah, I think that was true, Sean. Maybe we want to run through some of the issues and kind of things to keep in mind for the datasets. Yes. Okay, so I put this first because it's also fairly current. There's always this issue of dataset quality, right?

49:16So when I ask researchers, like, hey, why do you think we're running out of data? There's obviously hundreds of petabytes of data produced every single day. Why don't we just use that as data? And the typical answer to me is, well, that data is low quality, right? Which is true. And you need diverse sets of data as well. So for example, one of Stella's tweets about how we're not running out of data says, oh, there's like, you know, hundreds of terabytes of legal filings that are generated every single year. But all those legal filings have the same format. We don't learn very much by going over, you know, 1 ,000 pages of parking fines or whatever, right?

49:52Like they have to be unique and actually useful and high information. And so curating those data sets and making them useful is emerging, is becoming more and more important. And the way that we know this is actually from something that happened recently, which is Microsoft started training this small language model, Phi1, that is 1.3 billion parameters. So not even a large model anymore. It's just 1 billion parameters. So this is the size of GPT-2 on a dataset size of 7 billion tokens. So again, way smaller. Like this is a 7 to 1 ratio, not 200 to 1. Way, way, way smaller. And it's basically comparable in terms of the benchmarks to state-of-the-art models that are something like 10 times bigger, right?

50:39Like it's comparable on human eval because it's a co-gen model. It's comparable on human eval and MBPP, which is their own benchmarks. And so it's just like very interesting. I think that's something that's an emerging area of research, like how to improve the quality of data such that you train smaller models with fewer tokens, right? That is the final step of this evolution. But also Falcon 40B, the model that came out of the UAE that is now the top open source model, that was also on a proprietary new data set that the UAE government collected as well. And so that's just super interesting that you are not competing on size anymore.

51:14You're competing on quality. Yeah. The other issue that we mentioned, obviously, is copyright and privacy. We're not going to go over the same issues again, but there's a couple of interesting things going on. So there's the stable diffusion litigation, which is from the same counsel that is doing the get up co-pilot litigation. Basically, the whole idea is that, hey, this is not really fair use. You cannot really use my data to do this. So they're basically suing the model trainers on whether or not they should be able to use their work, even though they didn't specifically license it to not be used.

51:50There's kind of like an ethical question there. The interesting thing here is that now some of the AI providers are siding on the user's behalf. So, for example, if you use GitHub Copilot and Copilot generates GPL code, which is licensed and in theory should not. It's basically like a contamination license. If you use GPL code in your code base, all of your code base becomes GPL and you need to open source it. I actually wrote a very long post on the history of open source licensing, which I'll link in the show notes. But basically, GitHub is saying, we'll literally pay for your lawyers and we'll send lawyers to fight on your behalf.

52:28So it's really interesting how the risk piece is not being clear by saying, hey, we definitely 100 % not use copyright data. They're basically saying, maybe we do. And if we do, we mess up and we're going to pay for it. But ideally, we're not doing it. And there's also different articles out there that I mentioned before by the Washington Post and companies like that, where a lot of the training data comes from newspapers, comes from magazines, and people have not always opted in to having that as a training of the model. So some of the work then comes out in the inference. So yeah, that's kind of like another interesting piece of development.

53:09And again, if you're training your own model, you should be really careful. If you're using an off-the-shelf model, you should also be somewhat careful, but it seems like there's a lot more insurance on that. Yeah. Licensing issues also come in the form of terms of use, which is not an official license, but it's something you agree to when you use services. So OpenAI has this very famous clause in their terms of service, which basically states that you can't use OpenAI output as input for your training models, which is exactly what the Alpaca Vicunia students did in Stanford to train their models that now compete with ChatGPT.

53:48so this is why in our conversation with Mike Conover he talked about he was very excited about Red Pajama which is an open source replication of Lama because Lama also has similar licensing issues like Lama doesn't allow you to use it commercially right all these licensing issues copyright issues permissions issues are emerging areas that are being litigated people are coming up with different ways to license this stuff so for example Hugging Face has this rail license, responsible AI license that is different from MIT, different from Apache 2. And that's the license that Stable Diffusion is under.

54:22But it has never been litigated in a court of law. It's not accepted by the OSI Institute as open source. So it's just unclear. Like, can you use it? You have to consult your lawyer, quote unquote, which is a real cop out to basically say nobody knows until some judge rules when a case is brought up. So that's the licensing thing. So there's a lot of work there too, like Hugging Face has built a PI removal pipeline for their development. You can also go on Hugging Face and check if your information is in the stack and whatnot. But again, we're going to put all of this in the show notes. We're not going to bore you live on all the details.

54:57Yeah, just so people know, we have 17 pages of show notes that we've been collecting for two months. So this is a little bit of a crazy thing to compress into one hour, but let's try to do that. All right, the next few issues are to do with quality as well. So we'll talk about deduplication and filtering. And basically, like the amount of duplicates that... There have been studies done on the amount of duplicates in open web text in C4, and there's still quite a significant amount. It's pretty interesting because every time you duplicate something, you're exposing the model to that set of raw data again without knowing it.

55:33And therefore, it's basically going to try to memorize that text way more because it's just been exposed to it a lot more. and you're just basically wasting compute because that's not the kind of training that you want. So there's a bunch of research here that is reflective of studying these data sets and identifying those duplicates and removing them. But just the impact of this is really interesting. So we have this paper here that basically states that a sequence that is present 10 times in the training data is on average generated 1 ,000 times more often than a sequence that is present only once.

56:03So basically there is a disproportionate amount of weight that is placed on repeated information, which makes sense, right? Like if you show a model repeated information, it's going to overfit to that set of information. But it's on the order of 1 ,000 times more frequent. And that is a concern when trying to train general purpose models. The other thing is contamination. It's especially related to benchmarks. So for example, one of the things that GPT-4 showed is like, oh, they do very well on code forces, code puzzles. and but you'll see that the models for example does 10 out of 10 on the pre-2021 problems and then on the more recent ones post training cutoff date it does zero out of 10 it's basically i think i looked at the scoring it's like worse than a person doing it wrong on purpose which is you know it's actually pretty impressive so when you're training a model understanding what goes into it also helps you understand how to benchmark it like if you're benchmarking it against things that the model has learned is not super helpful.

57:06So that's another thing to keep in mind. Yeah. And this is why also, you know, releasing models and showing how you evaluate your models or releasing datasets, right? That's the topic of this episode. It's very important. So when Falcon 40B came out, they went right to the top of the Hugging Face leaderboard on the benchmarks that Hugging Face leaderboard provides. But we actually don't know if Falcon 40B's datasets were contaminated with the things that they're evaluated on, right? If you just copy-paste the exact results of the test that you're going to be tested on, of course you're going to do extremely well on it.

57:39There's actually a current thread of people, now that the model at least has been released, the dataset has not been released, but the model, as far as I can tell, has been released. People are replicating those evals and finding that it actually is falling shorter than claimed. So I haven't done it myself, so I can't really claim to know one way or the other. But I think it's just one of those things where you have to have some healthy skepticism because there's a lot of people trying to gain benchmarks. And the easiest way to gain benchmarks is to conveniently forget to remove the benchmarks from your data sets.

58:08Final point, because we're running out of time, data set imbalance. Obviously, we're all talking in English. The world is very English-centric, but there's other languages in the world. And we already talked about the tokenization issues, which will cost more, right? Because all these APIs are charged based on the number of tokens generated. And if your tokenizer uses more tokens per language, then you will cost more to generate those languages. Some amount of that is honestly not the fault of anybody because the language is just more complex, right? As a Chinese speaker, I'm well aware that the average Chinese person is required to learn 10 ,000 Chinese words to be considered literate, which is absurd.

58:48Right. But obviously China is making a lot of progress on language modeling as well. So there's actually a lot of papers coming out of China for Chinese datasets and English to Chinese conversion datasets, right? Which is a lot of the original translation benchmarks. So there's some Chinese datasets that we've outlined here, the CMRC, DoReader, and CHID, all of which we'll link in the show notes. Is there anything else that you wanted to comment on in particular? No, I think that's a lot of it. Actually, one of my friends and former co-founder, Andrea, he's working on an Italian language model.

59:20So I'm curious to see more of them come online. And I think like in the episode we're going to release with the practical AI crossover one, we talked about how language is also used differently in different countries. So some are very oral driven. So like a language model that is only text is not as important. So the voice data is also crucial. So, well, again, data sets 201. We'll get back to it. But I think we're already at one hour 10. So I think we covered a lot. Yeah, yeah. Hopefully that was a good overview, especially for people who keep hearing about things like Common Crawl, keep hearing things about contamination, keep hearing things about tokens even.

1:00:00And this is a ground up reintroduction to these concepts. We are recording these one-on-one episodes in order to be evergreen, right? That you can listen to this a year from now and hopefully it'll still not be out of date. Who knows? Hopefully. We can keep going at this pace. but you know i really want to like emphasize yeah data sets are great let's spend more time applauding data set creators because we're downstream of them yeah if you want to train on the latent space podcast please go for it we got all the descriptions in the show notes so we're doing our part we're doing our part all right everyone thanks for listening

1:00:53All right. Bye-bye.

From the publisher

In April, we released our first AI Fundamentals episode: Benchmarks 101. We covered the history of benchmarks, why they exist, how they are structured, and how they influence the development of artificial intelligence.

Today we are (finally!) releasing Datasets 101! We’re really enjoying doing this series despite the work it takes - please let us know what else you want us to cover!

Stop me if you’ve heard this before: “GPT3 was trained on the entire Internet”.

Blatantly, demonstrably untrue: the GPT3 dataset is a little over 600GB, primarily on Wikipedia, Books corpuses, WebText and 2016-2019 CommonCrawl. The Macbook Air I am typing this on has more free disk space than that. In contrast, the “entire internet” is estimated to be 64 zetabytes, or 64 trillion GB. So it’s more accurate to say that GPT3 is trained on 0.0000000001% of the Internet.

Why spend $5m on GPU time training on $50 worth of data?

Simple: Garbage in, garbage out. No matter how good your algorithms, no matter how much money/compute you have, your model quality is strongly determined by the data you train it on and research scientists think we just don’t need or have that much high quality data. We spend an enormous amount of effort throwing out data to keep the quality high, and recently Web 2.0-era UGC platforms like StackOverflow, Reddit, and Twitter clamped down on APIs as they realize the goldmines they sit on.

Data is the new new oil. Time for a primer!

Show Notes

* Our 2 months worth of podcast prep notes!

* The Token Crisis paper

* Ilya Sutskever on datasets

* OpenAI Tokenizer

* Kaplan Scaling Laws Lecture

* Chinchilla Paper

* Sasha Rush’s Tweet

* Karpathy’s Build Conference Presentation

* LIMA Paper

* Phi-1 by Microsoft

* Washington Post Article on datasets

* Our episode with Jonathan Frankle

* Our episode with Mike Conover

* BloombergGPT

* Datasets

* HuggingFace Hub

* CommonCrawl, Overview

* C4

* List of Dirty, Naughty, Obscene, and Otherwise Bad Words

* OpenWebText

* books3

* OpenAssistant

* The Stack

* The Pile

* LAION

* Audio:

* LibriSpeech: A dataset of audio recordings of audiobooks

* CommonVoice: A dataset of audio recordings of people speaking different languages

* Voxforge: A dataset of audio recordings of people speaking different languages​

* Switchboard: A dataset of audio recordings of telephone conversations​

* Fisher Corpus: A dataset of audio recordings of news broadcasts​

* Chinese:

* CMRC (Chinese Machine Reading Comprehension 2018)

* DuReader

* ChID

* Copyright & Privacy:

* https://stablediffusionlitigation.com/

* https://haveibeentrained.com/

* https://githubcopilotlitigation.com/

* https://twitter.com/moyix/status/1662131770463072257

* OpenAI Opt Out Process

* Check if you’re in The Stack

* Deduplication

* Deduplicating Training Data Makes Language Models Better

* Deduplicating Training Data Mitigates Privacy Risks in Language Models

* Contamination

* CodeForces example



Get full access to Latent.Space at www.latent.space/subscribe

More from Latent Space: The AI Engineer Podcast

All 247 episodes
AI Fundamentals: Datasets 101Latent Space: The AI Engineer Podcast · 1 h 1 min
Listen in VO