In short
AI Today Podcast - Episode Summary
Podcast Title: AI Today Episode Title: ChatGPT's Training Journey: Unveiling the Unlikely Websites
Episode Overview In this episode, the host dives into the specific websites that have contributed to the training of ChatGPT, uncovering unexpected sources that have shaped its conversational abilities. A significant focus is given to the types of data included in training datasets and the implications for the model’s performance and biases.
Key Discussions
Understanding Training Data
- Corpus of Text Data: Mention of common misconceptions about the amount of data used (e.g., one-third of the internet).
- C4 Dataset:
- Description: This dataset is a massive snapshot of content from 15 million websites.
- Use: Utilized for training major AI models like Google's T5 and Facebook's LLaMA.
Composition of the Data
- Data Size & Scope:
- The C4 dataset is large but only a fraction compared to models like GPT-3 and GPT-4.
- 30% of the content in the dataset is no longer available online, making it exclusive to those who have access to it.
- Content Breakdown:
- Business Content: Approximately 16% of the dataset, with notable sites like Fool.com (investment advice), Kickstarter.com (crowdfunding), and Patreon.
- News and Media: Dominates the dataset with well-known outlets like The New York Times, LA Times, The Guardian, and others reflecting diverse perspectives.
- Personal Blogs: Significant presence, indicating a wealth of individual opinions and thoughts.
- Religious Websites: Approximately 5% of the dataset, offering insights into diverse beliefs and practices.
Ethical Considerations
- Scraping Concerns:
- Many news organizations criticize AI companies for using their content without consent.
- The ethical debate surrounding the use of potentially untrustworthy sources in training datasets.
- Handling Controversial Content:
- Discussion on how certain words (e.g., "swastika") are treated in datasets, including the challenges of censorship versus necessary inclusivity for historical context.
- Commentary on the balance between safeguarding against harmful content and allowing for comprehensive historical understanding.
Future Implications
- Changing Landscape:
- Companies like Reddit planning to charge for data access could reshape the availability of training data for AI models.
- The implications of exclusive datasets (e.g., those no longer online) may lead to competitive advantages for companies like Google.
- Emerging Trends:
- Increasing interest in how AI models can derive insights from various data sources (e.g., patents and educational content from Coursera) that could expand their capabilities in teaching and idea generation.
Conclusion The episode wraps up with reflections on the future of AI training data and the discussions surrounding ethical considerations in its use. The implications of content diversity and the accessibility of datasets for training models remain central to the ongoing evolution in AI technologies.
Key Takeaways
- Diverse Data Sources: The training of AI models like ChatGPT relies on a wide range of sources, including business sites, news outlets, personal blogs, and more.
- Ethical Challenges: The ethical implications of using scraped content and managing potentially harmful or controversial data are critical discussions in AI development.
- Future of AI Training Data: As companies begin to monetize access to data, the landscape for training AI models is likely to shift significantly.
For those interested in exploring the intersections of AI training, data ethics, and the role of diverse content, this episode provides an insightful perspective on the nuances involved in shaping advanced conversational models.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00We hear all the time that ChatGPT and all of these different AI models are trained on an enormous corpus of text data. But what exactly is this data, right? A lot of people have, you know, come up with different things like, oh, it's like a third of the internet, or it's 90 % of the internet or 80%. Today on the podcast, we're going to talk about what exactly is inside the black box known as ChatGPT, Google, Facebook, all of these different AI companies that are training their AI models on data. We're going to be talking about what exactly this data is, what's inside it. I'm like, what are the actual websites?
0:35This is an incredibly interesting podcast in my opinion and you're going to want to listen close because other than just talking about the specific websites, it's going to give you a really good idea of why ChatGPT and all these different AI models are really good at some things and not really good at others because you're going to be able to, we're talking about what goes into them and it's going to kind of give us some ideas about why it says some of the things it says. So first off, a lot of this that we're going to be AI. And they actually went and they decided to dig in a little bit into what's called Google's C4 data set.
1:14So it's essentially just a massive snapshot of content from 15 million different websites that were used to construct some of the most high profile English language AIs. And so that would be Google's T5 and Facebook's Lama. OpenAI doesn't necessarily disclose what data set they're using to train the models backing chat GPT. But we can assume it's some pretty similar things. And you'll see why in a little bit. But you know, suffice it to say this is going to be the data that's in Google Bard. And that is coming out of a lot of the Facebook products as well. And inevitably, this is probably what's in chat GPT as well.
1:53Now, this being said, I would say it's important to know that of all of the content we're going to talk about today, this is is still only a small amount of data. This is what was essentially scraped in April 2019 by the company, it's a nonprofit called Common Crawl. That's for Google's 4C dataset, which is using a lot of different AI, just use this giant dataset. And so I think that while it's really huge, you know, 15 million websites and all of the content on them, and some people call it like a gargantuan dataset, um it's still you know it's probably about 40 times smaller than what gpt3 was trained on so this while it might seem big is actually still a lot smaller than gpt3 and the assumption is that gpt4 is even much bigger than this although they haven't actually released uh how many parameters are in gpt4 so um what what's important to know on all of this is one of the biggest content sections in this entire data set is just in business in general.
3:05So business and industrial websites, they made up about 16 % of the entire data set. And the number one website in that was fool.com, which if you don't know, it's just like a financial investment advice kind of website. So that is kind of interesting as a lot of people are experimenting with different use cases for investments and other areas like that out of chat gpt so not very far behind fool.com was kickstarter.com that is the crowdsourcing you know raising money for different businesses and that kind of stuff um below that a little bit they had patreon which was a pretty big chunk of that which if you don't know it just helps creators collect monthly fees from subscribers for exclusive content so that is actually pretty interesting that they were able to get the patreon list what people are essentially selling.
3:57And I feel like between Fool, Kickstarter, and Patreon, that's kind of like three different industries online that process a lot of money. And so I'd be really curious to see what kind of data it gathered from that that would be useful in generating business ideas in general. Kickstarter and Patreon might give the AI, I think a lot of access to different ideas for marketing and just like technology and just a lot of really different interesting ideas. The next biggest area behind finance that these models were trained on would appear to be the news. So news and media accounted for, I think about half of the top 10 sites overall in the entire data thing were news outlets.
4:44So New York Times was number four, LA Times was number six, The Guardian number seven, Forbes, Huffington Post, Washington Post. And so I think like artists and creators, a lot of these news organizations have criticized tech companies for using their content without authorization or compensation, right? Like essentially, they're just scraping all of this kind of news data. And a lot of these news companies are complaining about that. So it's going to be interesting that all of those that is getting sucked into this as well. So I think they found that several different media outlets that rank low on NewsGuard's independent scale for trustworthiness showed up in there.
5:31And the Washington Post did a news article about this whole thing. And they specifically kind of commented on this. You know, this is also interesting as a lot of these different like trustworthiness or fact checking organizations have come under scrutiny and have also themselves had a lot of flack. I think overall online, a lot of people don't really love the whole fact checking going on beneath posts. I noticed on Twitter that's mostly been replaced with community notes. So just allowing anyone to go. And, you know, if a community creator is range high enough, then they'll be able to post like a high quality link, um, disputing a specific claim, uh, which is kind of interesting because that in a sense democratizes it away from, uh, I think in the past, a lot of fact checking websites, um, were selected essentially by routers or other organizations and, uh, people complained about, um, who got to pick them, yada, yada.
6:33So I think, uh, uh, democratizing a little bit was good. In any case, uh, Washington Post wasn't happy that Russian state-backed news site RT.com was included in the list of media outlets. And also they were complaining that Breitbart.com, which is a right-wing news opinion website, was on there. It's kind of interesting. regardless of anyone's political opinions right or left or whatever I think it's really important to have news from different perspectives and all this kind of content it's kind of a big debate in AI you know people I a lot of opinion pieces on the Washington Post specifically talk about you know like why would we have untrustworthy training data put into this you know that's going to propagate bias and propaganda misinformation all that kind of sort of all that sort of stuff i actually think it's pretty important to have a wide variety of opinions you know false and real um that this is trained off of because these represent the opinions um of a wide range of people in the world and anyone that says their opinions uh their biases because everyone has biases um that theirs are the exclusive right ones uh lacks a lot of perspective obviously because people have a lot of different opinions and a lot of different perspectives and i think it's important to kind of encapsulate all of that and you can pick what you believe or what you don't believe but um you know i i saw a you know the washington post in that same article was complaining that there's like a list essentially of words that get blacklisted so if one of these words in an article it doesn't get added to that it's not supposed to get added to the training data one of those words was swastika obviously swastikas reference nazis and hitler and all that kind of bad stuff but they were complaining that the word swastika still showed up in this giant training data set over 75 000 times even though technically it was a blacklisted word and that kind of got me thinking that it's i feel like it's not a very good idea to have completely blacklisted words even though obviously swastikas represent a political party that did a lot of horrible things in the world why would we want to remove that word right like why wouldn't we just want to say you know swastikas and germany are bad but obviously it's part of something that happened in history right we can't just uh erase that and hope that ai never talks about it and i think it actually is it could cause more harm than good because i think it's important to um have these AI models ingest words or whatever that might be deemed bad because it's important.
9:18And then, you know, you could train and tell them obviously swastikas and Nazism and killing people is horrible. But I think it's important that that's in there because, you know, it needs to, it needs to, it needs to know all of those different concepts. And I think it's important that we're not just, you know, I get really nervous looking at how AI models are trained. When people are trying to put any sort of bias on the model or remove different segments or remove different blacklisted words, it just doesn't feel very good. It feels a lot like censorship. And I think it's fine to have all the content on there and then people can choose what they believe.
9:58And, you know, you can put safeguards in there and say X, Y, and Z topics are bad. And I think that's probably what Google and OpenAI are doing here against the, you know, to the complaints of the journalists. They're including words like swastika in there that obviously could be a red flag, but I'm assuming they're using, they're referencing, I would be, I would hope, or they have worked this in, referencing that, you know, obviously swastika is a symbol that represents something not good, right? But I don't think you want to just like pull that out completely because then it doesn't know anything about a very important, you know, atrocity that happened in the world.
10:33So I don't see why you'd want to remove that completely. In any case, my swastika tangent there, I think one of the other really big areas that is in this is religious sites. I think about 5 % of all the content were religious sites. Obviously, this makes sense, right? There's over a billion Muslims in the world. There's over a billion Christians in the world. These are, you know very large percentage of the population so I don't think that's shocking um despite some commentators I don't know having different opinions on that one in any case uh yeah and I also think that that's uh really useful for anyone that would like to learn about different cultures or religions or people to have all that different uh content on there so it's just seems it would seem like it'd be a it's a big part of the world it would be a good thing to have incorporated into these AI models.
11:25Another area that seems to have made a pretty big chunk up of this 15 % is personal blogs. So it's actually the second largest category. Sorry if I said it was, what did I say it was before I said news, news section number three, second biggest is personal blogs. And oh, the one other really interesting thing about this that I forgot to mention at the beginning is that 30 % of all the content that is in this giant data set that's used for so many different AI models is currently not available online anymore, meaning it was websites that expired, they got taken down, people removed them, changed the URL, whatever it is.
12:03And the reason why I bring this up and why I think this is so important is because if you think about it, Google that collected this giant data set, 30 % of all the content they collected is now essentially exclusive to them, right? They have like 30 % of it is gone now. So they're the only ones with it. And it's going to be interesting because people are talking about, you know, claiming rights to their data. Reddit recently said they're going to start charging companies to train models off of Reddit, which was a really big part of OpenAI's data set for ChatGPT. But if companies like Reddit are going to start charging, this data obviously is getting more valuable.
12:40And now Google has like this massive like uh i call it like a black hole of data like 30 of all their data is exclusively to them because it's gone off the internet but they have access to it so that's really interesting i wonder how valuable that is because no one else will have access to that data and if it's off the internet probably no one else will ever claim copyright or uh access or like ownership of it down the line into the future another really interesting thing is that a lot of these ai models it has been said are actually when they categorize all of the data that they're collecting they're not really categorizing a lot of it is not categorizing the authors and different things like that because they're kind of worried about personal data that's getting sucked in which we'll talk about in a second because there is a lot of personal data that has been sucked into these models and I'm curious how that's being used or pulled out or scraped or integrated but in any case, as we were talking about, number two biggest chunk of this entire thing is personal blogs.
13:41And that includes a lot of different platforms like sites.google.com, which can be, you know, anything from, you know, like a Catholic preschool in New Jersey to a, you know, a judo club in New York, whatever. So these are all just random things. And that's a really big chunk of this. So I think that's more than half a million personal blogs were pulled into that, which is, you know, representing about 4 % of the total categorized tokens. So of the actual text length, 15 % of all of the sites, but 4 % of the actual text that was inputted to train things on. So it's pretty interesting. A lot of this is WordPress, Tumblr, Blogspot, and LiveJournal.
14:27And I think this is why there's just a lot of different, I think this is a really important chunk because this really just gives a lot of perspective to people's feelings and thoughts around a vast array of different topics. I think that's a really valuable part to have. Like I was talking about before, a lot of companies like Google really heavily filtered the data before feeding it to AI. So C4, which this data set is called, it stands for Colossal Clean Crawled Corpus. So in addition to removing, you know, all of the duplicate text out of the whole of the whole thing, so it's not training on the same text twice, like I mentioned earlier, Google also uses a list of what they call, you know, dirty, naughty, obscene, or otherwise bad words, which is about 402 words in English and one emoji.
15:20And the company typically uses, you know, high quality data sets to fine tune the models and essentially they're just trying to shield users from unwanted content right you don't want to ask chad gpt something and have it cuss you out so i think it pulls out a lot of that kind of stuff um and like we mentioned earlier there's a lot of um controversy that goes around that i would say from both sides of the political spectrum um you know i saw you know, uh, the Washington post recently was criticizing the fact that it, um, that it, they, you know, obviously we're happy. It pulls out racial slurs and obscenities, but, um, they said that they were, they're disappointed that it eliminated some non-sexual LGBTQ content, um, by pulling out obscenities.
16:09So that's their complaint about it. Um, and then their other gripe is that it includes the word swastika so it would seem like they would like it to exclude less things from obscenities and more things from related similar to swastika and I just think that at the end of the day I don't think we're I think the less the probably the less bias or the less fine-tuning the better we'll let the users do that in my opinion I think that's going to be the the companies that win in the end um and it would appear that uh google and open air are trying to stay true to that sam altman has discussed that you know really letting people use this the way they want um obviously not making this like a horrible uh racist hellscape but um making it so it's safe but you know has a lot of variety of opinions and uh thoughts etc on there so yeah i'm sure we've talked about that enough but in any case uh it's really interesting if you go look at the Washington Post, check out their article that they just came out with today.
17:11You can look and see if your website was trained on some of this AI data. And it shows you the top sites, which are patents.google.com is the number one website that was used here. So all of the patents that have ever been written about any content, which is really, really interesting. And that is not just in America, but all over the world patent documents. So it's pretty interesting all of that has been pulled in there so a lot of really cutting-edge stuff as well which makes me think you could probably ask it if i was apple and wanted to write a patent for xyz what would i do that's a really interesting topic i think patents that might be a seek that might be like a treasure trove that might be a gold nugget from this episode um and from from some of this research is the fact that patents is the biggest data set in there and you could get a lot of content and ideas out of that the second is wikipedia.org obviously all of wikipedia is a ton of information open source it's kind of the perfect data set to be honest for training the third is called Scribd which is essentially audiobooks and digital books so just a lot of all of the content of books that have ever been written number four is the New York Times then they have journals.plos.org which is science and health they got the LA Times Guardian Forbes Huffington Post number 10 was patents.com so just more patents in there number 12 is Coursera, which is interesting.
18:35I'd be curious to see exactly what that entails. But Coursera has a lot of courses teaching you about a lot of different topics. So that might be why ChatGPT is actually pretty good at teaching things. Fool.com for business and industrial. And there is a handful of others. So really, really interesting to see what's going to go, what's going to happen in the future. Like I said, Reddit recently said that they are going to start charging to train on their data. And I'm expecting to see a lot of other websites do that. So I believe that the people that got in early may actually have a good advantage in a sense that they were able to get all of this data for free.
19:15You know, chat GPT open AI trained off a lot of Twitter data before Twitter shut off its API. And now we see that Elon is going to make an AI play, probably training off of a lot of Twitter data as well. So it's gonna be really interesting to see what happens in the future in this whole industry. This was a long podcast today for me. You know, usually I try to do 10 minute bites. We're at over 20 minutes. So I'll leave you guys here. But I hope you have a wonderful rest of your day.
From the publisher
In this episode, we explore the diverse array of websites that have played a role in training ChatGPT, shedding light on unexpected sources that have influenced its conversational prowess.
-
Invest in AI Box: https://Republic.com/ai-box
-
Get on the AI Box Waitlist: https://AIBox.ai/
See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
