In short
The episode explains how AI training data shapes what generative models can produce, why companies keep training sources secret, and how large datasets are assembled and used—often via scraping and “data laundering” through nonprofits/universities. It argues that training data selection is a core determinant of model capabilities and that commercial AI has shifted from research norms to profit-driven extraction.
Guest backgrounds
Alex Reisner is a staff writer at The Atlantic; he has spent years investigating training data, reverse-engineering sources, and studying industry research papers and developer forums.
Key claims
Training data is more fundamental than model architecture; secrecy protects competitive advantage and avoids backlash over acquisition methods; early models trained on Common Crawl were initially poor due to “mostly junk” internet content; synthetic data is overstated and can cause “model collapse.”
Notable examples
Common Crawl (crawling since ~2009; monthly hundreds of millions of pages; widely used for early LLMs); Lion (Europe) dataset of ~12 million YouTube songs; YouTube as a common, easily downloadable source (e.g., Whisper); “data laundering” via universities downloading millions of images/articles for AI training.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Importance of Training Data
0:45 to 3:52
Discussion on the significance of training data in AI models.
“particularly generative AI as a creative expression.”
Data Acquisition and Company Secrets
3:52 to 9:41
Exploration of how AI companies acquire training data and the secrets surrounding it.
“I think we have talked a lot on this show about why people feel the way that they feel about AI and the sort of instinctive reactions to the whole idea of AI.”
The Role of Databases in AI Training
9:41 to 14:00
Analysis of the role and existence of databases used for AI training.
“The research papers read very differently now than they did a few years ago.”
The Origins of Data Sets in AI
14:00 to 17:10
Explore the various purposes behind the collection of datasets for AI research.
“And so that's why I'm particularly fascinated by where these data sets even come from.”
YouTube's Role in AI Training
18:26 to 19:14
Discuss the significance of YouTube as a widespread source for AI training data.
“This is kind of a diversion, but I'm so struck in reading all of your work by how often YouTube appears as just it is everybody's favorite source for everything.”
The Ethics of AI Data Usage
19:14 to 21:35
Examine the ethical implications and the notion of data ownership in AI training.
“It's what there's a lot of stuff based on videos.”
Synthetic Data: Myth or Reality?
21:35 to 24:39
Analyze the viability of synthetic data in AI training and its potential pitfalls.
“because the thing that we're doing is so important that we must have it no matter what.”
The Future of Data for AI
24:39 to 26:01
Predict the future landscapes of data sourcing for AI and the role of creators.
“And it's not hard to see why that could be the case, right?”
The Future of Data for AI
28:04 to 28:25
Predict the future landscapes of data sourcing for AI and the role of creators.
“You think you know a browser, but Gemini and Chrome, that's new.”
Transcript
Automatic transcript. May contain errors.0:02Hello and welcome to the Vergecast, the flagship podcast of music that sounds eerily but not exactly like other music. I'm your friend David Pierce and today on the show we're talking about training data. It's the raw materials of everything that is AI and we probably don't understand it or talk about it enough. I'm talking to Alex Reisner who's a staff writer at The Atlantic and over the last couple of years he has spent a lot of time investigating training data and how all of these books and all of these articles and all of these YouTube videos and all of these songs are compiled into these gigantic data sets that AI companies then use to form the basis of their models.
0:40I think the way that these models are created and the sources of these data has a lot to do with the way that we feel about AI, particularly generative AI as a creative expression. And understanding how this data works, where it comes from and how it gets used, I think is really important. So we're gonna have Alex on, we're gonna dig into it. I'm very excited about it. But first, here's everything else happening on The Verge today. This is 90 Seconds on The Verge for Thursday, June 25th, 2026. Apple just raised the prices of a huge number of its most important products. iPads and MacBooks in particular are now up anywhere from$100 to$500.
1:15And even like the Apple TV is way up. It's now$200, roughly the price of 900 Amazon Firesticks. We knew this was coming, but this is still a big moment. Apple is as well-managed a supply chain company as anybody and exists at really high margins. If it can't keep prices down in this era of AI-driven shortages of memory and storage, nobody can. And I suspect this is going to get even worse. My advice from yesterday still very much holds, which is Prime Day is not over. Go get your deals while you still can. And here's a way to help you pay for all of those more expensive gadgets. Go get some money from Disney.
1:49If you have YouTube TV or DirecTV Stream, you might be eligible for a cash payout from a new$50 million settlement. The case started in 2022, and the argument was essentially that because Disney owns ESPN and Hulu, which are both powerful streaming services in their own right, Disney was able to drive up the prices of all of its rivals to unnecessary heights. Live TV is lucrative and competitive and is going to keep being messy in this way. This will not end things. But hey, go get some of that$50 million. And finally, a quick PSA to advertisers everywhere. Yes, I get it. The whole idea of using AI creative and AI targeting to put your AI ads in front of AI people is exciting, I guess.
2:30But maybe check your ads to make sure that you're not, I don't know, advertising a bike with two sets of handlebars like REI did this week. That may not go over so well with people. Just a little free advice from me to you. You can read more about all of this at TheVerge.com. That's 90 seconds on The Verge for Thursday, June 25th. Support for the show comes from ServiceNow. AI is moving fast across the enterprise. But without visibility, it's just chaos. Different tools, different models, different teams using AI in completely different ways. ServiceNow turns that chaos into control. With the AI control tower, you see all your AI across the business in one place.
3:07What it's doing, what it's done, and what it's about to do. So you stay in control. To put AI to work for people, visit servicenow.com. We all do it. You have a night for yourself but don't like the sound of the silence, so you turn on the TV just for the ambience. It's a little trick that helps you feel like you've got company and aren't alone. And other insurers, well, they may make you feel alone. But when you switch to GEICO, you've got claims reps available around the clock. So whenever you need, you'll have people around to help. And let's turn on the washing machine, just for good measure.
3:44Wow, isn't that soothing? It feels good to have support. It feels good to Geico. All right, let's talk training data. Alex Reisner from The Atlantic is here. Hi, Alex. Hey, David. Thank you for having me. Very excited to talk to you. I think we have talked a lot on this show about why people feel the way that they feel about AI and the sort of instinctive reactions to the whole idea of AI. And a theory that I have is that a lot of it is about training data. And I just want to talk about that. You've done a lot of work on this stuff, and you've been investigating how AI models get trained for a long time now.
4:20But I want to just lay a little bit of groundwork here. And just to ask the very obvious question up front, why does it matter what data is used to train these products? What difference does it make what's inside of these models? I think it is potentially the most important aspect of a model is what it's trained on. I mean, if you take a model and you train it on, let's say it generates music and you train it on 1950s jazz, that model will be very good at generating music that sounds a lot like 1950s jazz. If you train it on recent hip hop, it's going to generate music that sounds like recent hip hop.
5:00you know these models have names like chat gpt and claude but i think you could make an argument that the the right name for a model is actually the description of the data it was trained on because that is a description of its capabilities that's what it can output and so i think that the training data is really uh it's really fundamental to the model uh maybe more than the architecture to some degree. Interesting. And I think, I mean, that kind of answers to some extent my next question, which is why is training data such a closely guarded secret from these companies? And it seems to me that there is both a straightforward business answer and maybe a slightly more nefarious answer as to why these companies would be so secretive.
5:48But what is your read on why this is such a closely guarded secret, what data is being used to train these models? Yeah. I mean, the companies have argued that they need to keep this secret because the data that they have selected to train on is their competitive advantage, right? Like, Anthropic has done a better job at selecting data than Google and OpenAI. And if they were to let that come out in a court case or be public in some way, they would lose their competitive advantage. you know there's another pretty obvious reason which is that they have gone about acquiring a lot of this data in ways that the people who've created the data the authors of the books and the creators of the videos and the music would not be happy about and in a lot of cases they just don't know that the their their work is is being used and when they find out they're not happy about it and i think it's a conversation that the ai companies have tried to just avoid having So a big part of what you've been up to recently has been sort of reverse engineering this process to like peel the model apart to figure out what data is inside of it.
6:55And it seems to me that you've had to do sort of the reverse of what all of these companies have done, which is go figure out where these mass sources of books or web articles or videos or songs are. So tell me a little bit about your process. How do you go about finding these things that are kind of otherwise closely held secrets? Yeah, my process is actually maybe not that different from the company's process. I am a programmer by trade. I worked in tech for 20 years. I built websites and apps and did some statistics work. And so I've been aware of these models for a long time. And at a certain point, I started hanging out in the forums where AI developers were hanging out and talking about the work.
7:42And, you know, they're talking about what data they're using to train these things. And so that was a useful source. And I've also been reading a lot of their research papers. You know, selecting the data is really challenging. There's a lot of effort that they put into it. They have to come up with some notion of what is high-quality data. They're like, what do and don't we want in the model? And they write papers about that. I think they want to be involved in this conversation that's happening about training data within the industry. One thing that's been really helpful to me is that there's an open source AI development community.
8:17And they believe that the work should be done more out in the open. And I think a lot of them are doing really good and important work. And they have really interesting things to say about AI just philosophically and socially. and so they you know they're pretty transparent places like allen ai and luther about what they're using to train and so that's been helpful but then even at the companies you know ai the ai world is a little bit like academia in that it is good for your career to publish papers and the companies don't like that but they also acknowledge that they have to let the employees publish something and so the lawyers will go over it and you know tell them what they can and can't say and over time they've clamped down more.
8:59And so the companies are revealing less through the research papers. But yeah, a lot of my research is just from reading a Google paper, you know, for example, for this last article, they said we trained on, you know, tens of millions of songs. Yeah, I remember not that long ago, Apple going through a big cultural issue with this, with its AI team, because Apple is so secretive and so reluctant to let any of its work be public that all of its researchers were like, well, if you're not going to let us publish, we don't want to work here. And that push and pull seems like it has kind of morphed in a bunch of different directions over time.
9:33But interesting to hear that it is definitely, everybody's retrenching a little bit as this space gets hotter and hotter. Yeah, it's not 2021 anymore. The research papers read very differently now than they did a few years ago. That's really interesting. But tell me a little bit about these databases. One thing I think I had not given enough thought to until I started really reading your work is is that there's a real business in making and maintaining these databases. Like one company you wrote about, Common Crawl, I think is like a thoroughly fascinating player in this. And they essentially, as far as I can tell, just crawl the Internet and make it available to whoever wants it, which is weird and complicated.
10:15And we should probably come back and talk about Common Crawl sometime. But it seems like if you look around a little, these gigantic databases of books and articles and videos exist. Like, are people making these and selling them? Why do these databases exist? Mainly, well, I mean, for AI training, why else would they exist? It's a huge, you know, it's an extremely labor-intensive process to find, you know. In one sense, you just go and download all of Library Genesis or all of Anna's archive. That's one sort of naive thing you can do, but the companies realized pretty early on that they need to filter the stuff pretty carefully.
11:01So the organization you just mentioned, Common Crawl, yeah, they've been crawling the web since the late 2000s, maybe 2009 or something like that. and they just make the whole thing available. Every month there's a new, they've scraped a few more hundred million web pages and it's available to anyone who wants to do any kind of research with it. In fact, it's mostly AI researchers who are using it. But all the early large language models were trained on common crawl. If you go back and read those early open AI papers, everyone was training on common crawl and at first, and the models were terrible.
11:40Because if you train a model on the whole internet, it just, you know, you get all the, it says all the junk that people say on the internet. Right. Along with the intelligent things. Yeah. Mostly junk. Yeah. I would say statistically mostly junk. Statistically, yeah, mostly junk. And I think the early large records models were proof of that. Sure. But, yeah, I think they're, you know, Common Crawl is a nonprofit. So, you know, they would argue it's not a big business. they do get a lot of money from AI companies and AI investors. But yeah, I think the topic of training data selection, the challenge of selecting the right data for a model is still really hard.
12:23The AI companies, I would say, still have a very primitive understanding of what data will make their model better. It's an area of research that I think even they at this stage are not very good at. They do it mainly by trial and error, as far as I can tell. Interesting. Again, so the reason you're asking why do these data sets exist? I think it's people trying to share what they've learned from curating data sets in different ways and training models with them. Yeah, I mean, part of the reason I ask is I think one of the things I have come to believe about the AI industry is that this shift that went from AI research being fundamentally a research thing.
13:11Like if you go way back, OpenAI was basically a research organization, right? And you talk about Common Crawl and I think it's early users were largely researchers and these things were academic things for academic purposes. And from what I understand, And the kind of rules of the road are different, right? That like, if you want to make copies of a bunch of things for academic purposes, these things are generally considered less problematic, right? But if you then do it and become a trillion dollar company on the back of it, people are going to rightly feel differently about the way that you went about getting that information.
13:46And just the speed with which AI commercialized, all of these companies just moved so fast from we are essentially an academic thing to, oh my God, we're making so much money, everyone's filthy rich, that it feels like they just hoped everybody would ignore the ways in which they got this information. And so that's why I'm particularly fascinated by where these data sets even come from. Because it does seem like, in many cases, like you're saying, it's not that there is some gigantic business in being the one to sell the songs to somebody. It's that stuff is being compiled for other purposes. It's just that now the main purpose, because of the sheer volume of work, is AI research.
14:31All this stuff has been sort of co-opted from every other purpose to AI. And it just all feels so concentrated now. Does that feel right to you? Does that make sense? Yeah, I think that I agree with most of that. I do think that I'm not sure how much these datasets were really collected for other purposes. Common Crawl likes to talk about, I think they're one case, they probably have the strongest argument that their data could be used for other purposes. But when you go back, you know, they've been cited by over 10 ,000 papers. I didn't read all 10 ,000, but I read a lot of them. And they are mostly AI, right?
15:11And it's early. A lot of it is stuff that people wouldn't mind as much as with generative AI, right? Like Common Crawl, I think without Common Crawl, you know, AI translation tools might not be as good as they are. I think it was really a huge help because they scraped web pages, the same page in multiple languages, and people were able to train translation models based on that. So that was helpful. But the thing that, you know, there is still a, what I would call a data laundering network where the AI companies are still relying on, they'll do a collaboration with the university and they'll have the university download, you know, millions of images to train a model or download millions of articles to train a model.
16:02and the AI company can say like, well, we didn't do it. This was like an academic thing. You know, the same goes. Common Crawl is not the only nonprofit that's like doing a lot of this scraping for the AI industry. One of the datasets I reported on in the music, the article about music training data is this organization based in Europe called Lion. They have a dataset of 12 million songs from YouTube. So anyway, this is like, is it academic? Like, not really. Like, this is, you know, technically, yeah, there's universities and nonprofits, but they're all receiving money from the AI industry. When I got a new car, I thought my insurance premium would increase and empty my bank account.
16:44Like, if Fetween won the lottery. I've invested most of my winnings in chicken tenders because they're bomb. But, bro, I bought a house and it's sick, bro. I'm thinking the floor is going to be all trampoline, bro, with the helipad on the roof. The contractor said it's structurally unsound. They're just being babies. But switching to Geico saved me hundreds, so my bank account is safe. It feels good to save some hard-earned cash. It feels good to Geico. Support for this show comes from Fetch Pet Insurance. Do you have a pet? Every six seconds, a pet owner in the U.S. gets hit with a vet bill of over$1 ,000.
17:20And it's almost always an unwelcome surprise. That's where Fetch Pet Insurance comes in. Fetch is the most complete pet insurance. Get paid back up to 90 % of vet bills. You can use any vet in the U.S. and Canada. All vets are in network. Go to FetchPet.com slash save right now for your free quote. That's FetchPet.com slash save. I'm pretty confident talking into a mic. Hey, I'm doing it right now. But Home Projects, I second guess everything. is that noise normal is that water damage and who should i even call that's where thumbtack comes in upload a photo or voice note and their ai powered search helps diagnose the issue and match you with the right top rated local pro instead of second guessing or searching for hours you get clarity and can hire the right pro with confidence for your next home project try thumbtack they know homes hire the right pro today
18:26I'm not giving up. I am selling the building. The final season of FX is the Bear. The restaurant is flooded. Everything's either gonna be okay. No, stop. Or not. We are outgunned and we are outmanned. We have each other. FX is the Bear, the final season. All episodes now streaming on Disney+. This is kind of a diversion, but I'm so struck in reading all of your work by how often YouTube appears as just it is everybody's favorite source for everything. Like, it's what, you know, it's what OpenAI used allegedly to create Whisper. It's what a lot of the music stuff is using. It's what there's a lot of stuff based on videos.
19:17Like, what is your sense of YouTube's role as an AI training force? because it seems to be everywhere. Yeah, that's accurate. YouTube is an extremely common source. I think one reason is there are tools for downloading from YouTube that work really well. They're really easy to use. And it's pretty common for AI developers to just use those tools. And it's just kind of a, it's become a custom. But also, you know, and that includes, stuff on YouTube is just less protected, I think is one way of saying it. Like there's music, you know, if you're a musician, your song might be on Spotify, but Spotify's website has digital rights management protections.
20:00It's really hard to download from Spotify. It's much easier to get the same song from YouTube. And so many songs are also on YouTube. So I think it's just ease of downloading. YouTube obviously would say out loud that this is not allowed, right? That it violates the terms of service. and yet it seems to have either it can't or it just hasn't done anything to stop this really. Yeah, that's a question that's been in the back of my head for a long time and which I've asked YouTube and which they don't really answer. They have said that they consider it a violation of their terms of service to be downloading their videos.
20:39But yeah, they haven't, you know, years have passed and it's just as easy to download from YouTube now as it was a few years ago. Yeah, I downloaded a YouTube video this morning. It is just a thing you can do. It's really true. You mentioned the sort of data laundering stuff, but it also seems to me that more and more the people doing the AI training are just completely unapologetic about it. like you quoted Rich Screnta, the CEO of Common Crawl, who literally said to you, like, if you don't let AI robots crawl your data, you essentially don't exist. Like you're going to miss out on the future of the internet.
21:21And I think about, you know, Marc Andreessen, even a couple of years ago, being like, none of this would work if we couldn't just take the training data that we needed. there is this almost like manifest destiny sense of the AI industry that we can have this data because the thing that we're doing is so important that we must have it no matter what. What is your sense of the trend there? Because it occurs to me that that's happening even as the backlash from people who hate the experience and feeling of AI in part because of the way that this stuff is trained just keeps getting worse. Like these things are just running away from each other at like record speeds?
21:58Yeah, I, that's, that's a huge question. Um, I think it, you know, I think it has something to do with the fact that we just have not done a great job in this country with establishing the value of data and who should be able to have data, right? This, this is something that privacy advocates have talked about. Jaron Lanier, I think was one of the earliest people to be talking about this. I think he wrote in 2007 that you should get paid, like you were being surveilled basically, and companies are monitoring everything you do. They're generating data from your online activity. That data is extremely valuable to them.
22:43It might seem like nothing to you, but you should be getting paid for that. People thought that was crazy back then. And I think a lot of people still think that's crazy now, but taking people's music to build models that generate songs that compete with them is just the next version of that. And it's just going to keep going. If we don't acknowledge that this data is incredibly valuable and figure out a way to write laws around that or just have better business practices or something, this is just going to get worse. This is going to be more exploitation and more mining. the next step i think a lot of people have perceived for a while to be synthetic data right that eventually we're going to get ai models that are so good at making new things that then we can use those things to train new models and eventually they don't they don't need existing recorded music that synthetic data is the future and that's how we get to everything um and especially now i mean you look at spotify and there are tons of ai generated songs on Spotify that some people are listening to.
23:49There are a billion AI generated podcasts out there. The content is being made. Is that the next phase of training data? Is synthetic data coming into its own in such a way that we're going to start to see the next generations of these models built on the things made by the last generations of these models? Absolutely not. Really? I don't think that there's any evidence that that actually works. I think when AI companies talk about training on synthetic data. They choose their words very carefully and they always exaggerate the extent to which it's happening. There's a lot of research out there on a phenomenon called model collapse, which is what happens when you train a model on its own outputs.
24:34It very quickly, it doesn't get better. It very quickly degrades. And it's not hard to see why that could be the case, right? Like AI is kind of an averaging machine statistical average between different types of content and putting that into some new kind of more average type of content. And there's just not enough weirdness or interestingness or something like that. There's some quality in the work that humans do that's not in the work that AI does. And I think that's actually proved by the model collapse phenomenon. So I'm amazed that AI companies are going around talking about synthetic data still There's so much evidence that it doesn't work.
25:17So is there a next untapped place filled with data? I mean, these models keep getting bigger. They keep needing more data. They got to go somewhere, right? Is there a next sort of unturned place to go for these companies? I think they just pay people to make it. I think there's already a gigantic industry of writers who are writing for AI, musicians who are making music for AI. I, you know, after I published this story, I got an email from a company that is doing this. They claim to have paid creators over$10 million just to make things for AI training. It's very strange, but this is the next, like AI as your audience is the next frontier for creators.
26:07It is a deeply strange thing to think about. But also like, you know, we're going to put this on YouTube and it's going in there. anyway. So who knows? Maybe this is all of our destinies, no matter what. I certainly hope not. I don't think so. I think we can have a conversation and arrive a more reasonable future for ourselves and the culture. I'm with you on that. All right. Well, Alex, thank you so much for being here. I really appreciate it. Thank you, David. That's great. All right. That's it for the show. Thank you to Alex for being here. And thank you, as always, for watching and listening.
26:40If you have thoughts, questions, feedback of any kind, if you have a favorite AI song you want to send me, just not the Puerto Rico song. I already know that one. Send me all your favorite AI songs. No judgment. I just want to hear all of them. You can always hit up the hotline, 866-VERGE-11. You can send us an email, vergecastatheverge.com. We love hearing from you. And as a reminder, the best thing you can do to support all of this is to subscribe to The Verge, theverge.com slash subscribe. It gets you all of our podcasts ad-free, including this one. It gets you all of our exclusive newsletters.
27:09It gets you all of our coverage. Terrence O 'Brien on our team has been covering Suno a lot and doing a really terrific job. Lots to come on that. Subscribe to The Verge. I think it's a pretty good website. The Verge Cast is a Verge production and part of the Vox Media Podcast Network. This show is produced by Josh Cajas, Eric Gomez, Brandon Kiefer, Travis Larchuk, and Aaron Lacasio. We'll see you tomorrow. Rock and roll. We all do it. You have a night for yourself, but don't like the sound of the silence. So you turn on the TV just for the ambiance. It's a little trick that helps you feel like you've got company and aren't alone.
27:41And other insurers, well, they may make you feel alone. But when you switch to Geico, you've got claims reps available around the clock. So whenever you need, you'll have people around to help. And let's turn on the washing machine, just for good measure. Isn't that soothing? It feels good to have support. It feels good to Geico. This episode is brought to you by Google Chrome. You think you know a browser, but Gemini and Chrome, that's new. It can help you with practically anything on the web, like restoring a vintage motorcycle from a 50-page restoration block, or finally break down that long article you've had open for weeks.
Read the full transcript
28:16Gemini and Chrome is here for it. Ready to make anything online make sense? There's no place like Chrome. Check responses set up required. Compatibility and availability varies 18+.
From the publisher
Training data is the raw material of the AI industry. Claude, ChatGPT, Gemini, and the rest are built on top of oceans of stuff. What is that stuff? Books. Blog posts. YouTube videos. Reddit comments. All of it and more, in virtually incomprehensible quantities. Alex Reisner, a staff writer at The Atlantic who has been investigating training data, explains how AI companies get all this data, why they'd really prefer you not know what's in it, and whether training data could ever be a fair trade.
Further reading:
Apple raises prices on Macs, iPads, and more by hundreds of dollars | The Verge
Disney agrees to pay $50 million to YouTube TV and DirecTV subscribers | The Verge
Two handlebars are better than one, right? | The Verge
At Least 15 Million YouTube Videos Have Been Snatched by AI Companies
The Hypocrisy at the Heart of the AI Industry
The Millions of Songs Mashed Into AI-Generated Music
Common Crawl Is Doing the AI Industry’s Dirty Work
Subscribe to The Verge for unlimited access to theverge.com, subscriber-exclusive newsletters, and our ad-free podcast feed.
We love hearing from you! Email your questions and thoughts to vergecast@theverge.com or call us at 866-VERGE11.
Learn more about your ad choices. Visit podcastchoices.com/adchoices
