In short
The 404 Media Podcast - Episode Summary
Podcast Title
The 404 Media Podcast Description The 404 Media Podcast presents a weekly recap by co-founders Joseph, Sam, Emanuel, and Jason, discussing stories from their journalist-owned digital media company focused on technology and its societal impacts.
Episode Title
Google Is Exposing Peoples’ ChatGPT Secrets Episode Description The episode discusses the exposure of nearly 100,000 ChatGPT conversations indexed by Google and explores Wikipedia's new policy on AI-generated content. The episode includes insights on the history of gaming platforms Steam and Itch.io in the subscribers-only section.
---
Key Topics Discussed
- ChatGPT Conversations Exposed
- Incident Overview: Nearly 100,000 past ChatGPT conversations were inadvertently indexed by Google, exposing sensitive user data.
- Exposure Mechanism:
- Users could share ChatGPT conversations, which Google indexed, making them searchable.
- OpenAI's failure to prevent Google from scraping these pages led to public exposure.
- Types of Sensitive Information Found:
- Private queries such as "write my essay," "plagiarism," and personal information like Social Security Numbers.
- Sensitive corporate data, including financial settlements and non-public information about mergers.
- User Education: Many users are unaware of the implications of sharing their conversations and the potential for information leaks.
- OpenAI's Response: They have since removed the sharing feature and are working to take down indexed conversations.
- Wikipedia's AI Slop Policy
- AI Slop Definition: AI-generated articles flooding Wikipedia, often containing inaccuracies or fabricated information, are dubbed "AI slop."
- New Policy Implementation:
- Wikipedia has adopted a speedy deletion policy for certain AI-generated articles to manage the influx.
- Two conditions for speedy deletion:
- Usage of language indicating AI generation (e.g., phrases directed at the user).
- References that do not exist (fake citations).
- Community-Led Change: The policy change reflects a community-driven response to the challenges posed by generative AI, contrasting other platforms' more accommodating stances on AI content.
- Importance of the Policy: The move aims to uphold the reliability of Wikipedia while minimizing the workload on editors.
- Future Implications and Reflections
- User Takeaways:
- Users should recognize that anything shared online is potentially leakable.
- Greater caution is necessary when using AI tools to prevent unintentional information exposure.
- Broader Lessons for Platforms:
- Other platforms can learn from Wikipedia's proactive approach to filtering AI-generated content.
- There's a need for policies that respect user privacy and provide clear guidelines for content moderation.
---
Additional Information
- Subscriber Content: The episode concludes with a teaser for the next episode, which will delve into the history of gaming censorship, particularly focusing on Steam and Itch.io.
- Call to Action: Encourage listeners to subscribe for an ad-free experience and access to bonus content.
---
Closing Thoughts The discussions highlight the urgent need for awareness around privacy and data security in the age of AI. Both the ChatGPT and Wikipedia cases illustrate the complexities and responsibilities that come with using and moderating AI-generated content.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00YubiKeys are the original passkeys. small sturdy and easy to use physical security keys that prevent phishing attacks and account takeovers. YubiKeys are manufactured by Yubico which is a company with headquarters and manufacturing centers in both Sweden and the United States. Unlike basic multi-factor authentication methods such as SMS one-time passcodes or mobile authenticator apps YubiKeys provide modern MFA and are a proven security solution that cannot be hacked or bypassed by malicious actors, stopping AI-powered cyber attacks, online identity scams, fraud, and account takeovers. YubiKeys help businesses of all sizes, from large banks and tech companies to critical manufacturers, energy concerns, and government agencies stay ahead of evolving cyber threats and regulatory requirements.
0:47They also protect individuals and everyday users by securing email, banking, and social media accounts, password managers, productivity tools, developer tools, and more. For more information on how YubiKey secure applications, services, and accounts for both individuals and businesses, visit yubico.com slash 404media. And for a limited time, get$10 off your order of exactly two keys from the YubiKey 5 series or Security Key series using the code 404media10 at checkout. That's yubico.com slash 404media and use code 404media10 at checkout. That's ubico.com slash 404media and use code 404media10 at checkout.
1:52where we respond to their best comments. Gain access to that content at 404media.co. I'm your host, Joseph, and with me are 404 Media co-founders, Sam Cole. Hello. And Emmanuel Mayberg. Hello. So, first of all, we had our party in Los Angeles, party slash live podcast recording, the audio and the video, I guess, because we did it on YouTube as well. We will put that into your feeds soon. Sam, what did you make of the party? Did you have a good time? Yeah, it was awesome. So many people came. We packed the house. It was a little bit crazy at times how many people were in there. There were concerns about code and capacity, but it worked out.
2:45It was great. Yeah, we did a little test live stream. We'd never done live streams before. Um, so Jason and his friend Raul, uh, came through and helped us set that up. Um, there was a really good, if you go to the YouTube and you go to the live section, there's a really great, um, conversation that happens just organically while we're setting up the live stream with the folks at RIP space, which is the hacker space where we had the event. Um, they just like sat down with Dexter Thomas, who, um, does the kill switch podcast and was also helping us do this. They just sat down with him on the couch and started streaming and did an impromptu panel.
3:27It was awesome. It was very sick. So that's the first hour of the YouTube Live. And then you'll see us come back on and do our panel, which you and Jason talked about your reporting on ICE and surveillance and Flock and all that good stuff. I was in bed watching the live stream, like the Wolverine meme. Yeah, I assumed you were calling us in the chat. You were probably a non in the chat. But yeah, it was really good. What did you think, Joe? You were there too. Yeah, I really, really enjoyed it. And mentioning the live stream, it was just a very, very low stakes. Hey, we're testing out this tech.
4:08We've never done it before. I think I literally posted that to Blue Sky. Like, we're testing out this technology. Join if you want. And a fair few people jumped into the YouTube live chat and we're really, really supportive. So it shows, oh, actually we can do this sort of thing and we will do that in the future. As Sam said, if you want to see the panel discussion, you go to the YouTube link and all that sort of stuff. It's just on our YouTube channel. If you don't want to fuss with any of that, as I said, don't worry. We will get just the stuff from the second half, like the live podcast, from the video and the audio, and we'll put that into the feeds as normal.
4:49So just keep an eye out for that at some point if you don't want to faff of all of the timestamps and everything. But speaking of events as well, we did that. We did that live podcast recording. Now we might be doing another one. Well, we're definitely doing another event. We'll see exactly what we do there. But Sam, do you just want to tell people about this New York party that we have on the horizon? Do you just want to tell them about that? Yeah, I got back from LA on Monday and then realized that it's August. And our second anniversary is on August 22nd. So we're having a second annual anniversary party in Brooklyn on the 21st.
5:30It's going to be at Farm One, which is a vertical farm slash microbrewery in Brooklyn. We're working on getting the ticket sales page and all that good stuff set up. We're hoping to maybe do some kind of similar panel and maybe even a live stream where we can figure it out. But TBD on the details, just watch for that ticket information and get yours when it comes out. Last year's anniversary party was also a huge hit. And those tickets sold out really fast. So, yeah, not to leave people on a cliffhanger on that, but it's coming soon. Yeah, keep an eye out in the weekly newsletter. We definitely announce stuff there.
6:14And of course, the next episode of this podcast, there'll be details of where and how to get tickets for subscribers and non-subscribers. Okay, with that housekeeping out of the way, Sam, do you want to take the lead on this story I wrote? And you can grill me about it because you edited it. Yeah, I thought the story was super interesting. And as with many ChatGBT leak slash exposure type stories is horrifying. And it blows my mind that people put private information into ChatGPT every day. So the headline is, Nearly 100 ,000 ChatGPT conversations were searchable on Google. So this saga begins before today.
7:00Do you want to kind of walk us through where this story started? Yeah, it begins while we were doing the LA party or something around there. So I kind of missed this when it broke. But there was a Fast Company story last week, July 30th, I think, and Fast Company said it found that people were, it seems inadvertently, exposing the contents of some of their conversations with ChatGPT. and the way that works is that ordinarily when you're speaking to chat gpt those conversations are basically private i say basically because of course open ai can go in and probably review stuff with that with that caveat but it's not like you can just stumble across the the contents of someone's conversation um it is you're logged into the service right and to view any to view any of those conversations, you'll also need to be logged in.
7:56But there's this interesting feature built into ChatGPT that allows people to share their conversations. So they scroll down on the page or whatever, they select, I would like to share this, and ChatGPT gives them a little warning saying that, well, you know, this is going to be accessible by more people, probably. Of course, that's what people want to do because then they share a link to the chat gpt conversation um and they provide that to somebody now they don't need to log in and they can just read the content the problem is google sees that because of course it's indexing the public web and what this feature has done is essentially create a public web page of that conversation so the result that fast company found was that oh a bunch of people it seems may not realize they are basically exposing their chat gpt conversations not just to the person they shared the link with but to the wider internet because open ai hasn't configured it in a way where hey google please don't scrape this webpage?
9:09Yeah. I mean, ChatGPT being private by default is more than Venmo used to do for memos or even that Meta used to do for the chatbots that we found were just blasting conversations out into the open. So there's that, but it being so easy to find these cache links was really interesting to me. So people being people on the internet what did they do with this revelation from the fast company article yeah so people obviously started going to find these indexed web pages these chat gpt web pages on google one example was the sort of osint trainer and um practitioner henk ven s um they they ironically used Claude, they say, to generate some Google Docs to figure out, well, how could I best search this for sensitive information?
10:15So they do that. And then the Google Doc returns. And sorry, when I say Google Doc, that just means a very specific Google search. Like a really basic one would be file type colon PDF, then your search term. And that would only return things which are inside a PDF for whatever you're looking for, your keywords. So when he did this, he was doing searches like, I want to search ChatGPT for phrases like, Write my essay, or plagiarism, or my assignment due. Another one is, Without getting caught, avoid detection. Then some corporate stuff, such as my company, strategy, or revenue, or acquisition. all of that sort of thing.
11:02And then even my SSN, and you can see where he and others are going with this. And he says they found pretty sensitive stuff in there, such as confidential financial data about an upcoming settlement, non-public revenue projections, intelligence about companies that may be merging together, and some NDA stuff as well. I will say I haven't seen any of these search results so I can't verify or vet for those specifically but yeah when Fast Company first reported this researchers then jump on it and they start finding all of their own stuff as well. Yeah so you mentioned some of what was in the chats like the plagiarism stuff what else was in there was there anything like super sensitive like what are we talking about when we're talking about what was specifically in the chat.
11:56Well, that's the thing. Hank FSS doesn't really provide a ton of specifics, but what they... They say that they're sensitive corporate stuff in there. They didn't want to quote them directly. And I think that makes sense because, of course, at the time this is going on, again, we were away. So I'm kind of coming to the story a few days later with our own story with new information. But this is all out there. And if you are being too specific, you can inadvertently direct readers to go dig this up themselves. And if there's anything truly sensitive, you don't really want to do that. And I could say this later, but I might as well say it now.
12:41OpenAI is dealing with it. They are removing certain indexes from Google. and OpenAI has removed this opt-in share feature as well because it appears they realize that, oh, people are doing this and they don't fully understand the consequences of what's going on here. Yeah, because they think that this stuff is private. So they're just going for it. I thought that this is so low stakes out of everything else that could be in Chattapiti messages, but you said people were asking Chattapiti to write their LinkedIn posts, which I would guess is not only a lot of LinkedIn at this point, but also a lot of what TrashGPT gets asked to do on top of term papers and stuff like that.
13:32I don't know. It's just depressing to me. Well, yeah. And that leads to our story, which is that when people were digging through it, they were dealing with maybe sort of hundreds of queries, that sort of thing. and just seeing what they could pull out of the indexed Google pages, I then get this tip that a researcher who I granted anonymity to had found many, many more. They had found nearly 100 ,000 chat GPT conversations, whereas earlier people had just probed, I think, 500, something like that. they scraped all of those pages they also scraped um so not just the google results but the actual content of the chat gpt messages themselves as well and then gave me access to it and yes that's when i i was i log in and i start probing around to see what is interesting and sensitive and that could come up and it really goes from the sensitive and the delicate to the really benign as you say with like the LinkedIn stuff.
14:39I think to find that I literally typed in, write my LinkedIn, my LinkedIn, or something like that. And so many results come up, like write me a LinkedIn post in this vibe to touch on these points or whatever. There was one in there that again, I didn't want to quote particularly directly, but the vibe is that clearly somebody, it appears a man, is thinking about their ex-girlfriend and he's asking chat gpt why is she not looking at my stories and like clearly having some emotional distress over this relationship and chat gpt is walking them through about how you shouldn't break up you should talk to your current girlfriend about these feelings like stuff you really you don't want on the internet you know i would also probably say you probably shouldn't talk to chat gpt about this stuff either ideally i don't know how good the advice is but if people get closure or benefit out of that hey sure go wild but you don't want that being um on google right you don't want people to people to be able to come across it and then dig through it or anything else and it looks like many of the chat logs were generated by people anonymously.
16:00It doesn't have a username for all of them. That being said, I mean, there are clearly names in there. There was one I saw where someone had obviously built a rapport with this chatbot and sort of their version of ChatGPT because, of course, you build a history and a dialogue with these tools. And I can't remember the person's name, but he said something like, yo, yo, I'm back. It's uncle name. And then ChatGPT greeted them once again. Maybe I should have quoted that one a bit more directly because it didn't seem to be particularly sensitive. But yeah, there's stuff in there that you do not want to be online.
16:37And I also saw corporate stuff. I saw someone uploading what they said was a copy of OpenAI's own non-disclosure agreement for visitors to the company's headquarters. I emailed OpenAI asking, is this really your NDA? And they didn't get back to me. Somebody thought it was though and they were pulling it into chat gpt and now ironically that has now been posted online because of chat gpt's privacy settings it's so crazy it's so i mean it's like i think that's a big benefit a lot of people find with chat gpt is that you can be totally cringe and like open with this thing in a way that you might not want to be publicly so it's not surprising that this stuff is contained in these chats, but, oh man, it is.
17:26It's, I don't know. It's scary, it's funny, and that I think is usually our wheelhouse. It's like, it's intersection of like, I'm laughing because I'm terrified of the future. So knowing all of this, knowing that this was kind of almost like a built-in feature that ended up leaking these chats, what do you take away as far as like the privacy lessons? like are there any how can people avoid something like this happening in the future what should companies do like what do you kind of see as the lesson here yeah um it feels kind of similar to the early days of app development on smartphones and and all of that sort of thing which is not to say that chat gpt is like a super basic app or anything like that i'm sure they have pretty damn good security over there or i hope so considering you know how much money they have and I know some of the people they have hired are very, very competent.
18:22It's much more on a user level in the same sort of way where you grant location data permissions to an app and you may not fully understand what's going on there because it's very opaque to you. And then lo and behold, your location data has now been sold to a company that then sells it to DHS or something, which is not going to be a concern for a lot of people, but it definitely could for some others as well. It's similar to that where as a consumer or a user of this tool, there are these second or third degree events that you may not necessarily understand. And it even just reminds me of people sharing like a Google Drive or a Google Doc link without fully understanding.
19:06Anybody who clicks on this link is going to be able to read your terrible article draft. I mean, that would be in my case or anything else, right? Or your calendar can sometimes be accessible. And when people are using these tools, they may be very focused on, well, I'm not going to put sensitive information into the chat dialogue. I'm not going to give my banking information, although it seems some people did that as well. And they think that's sort of the be all and end all of privacy when it comes to these tools. But there's like an app development sort of issue as well where the user can really mess up if they do that without fully understanding.
19:49But it's definitely not all on the individual consumer because OpenAI, it seems, didn't communicate this fully and didn't take steps to stop Google scraping it. You know, a Google Drive link is not typically going to appear in Google search results unless it's been pasted on the web page or something like that. So kind of everybody's at fault. And I think as more people use AI for more and more sensitive stuff, you've just got to be really, really careful about how that might trickle out if you're not entirely sure what you're doing, you know? Yeah, at this point, I treat everything, almost everything that I type onto a screen as potentially leakable.
20:37Or I think about it a lot more than I used to. And I used to think about it a ton. So I'm thinking about it all the time now. But if you're typing something out, it could go anywhere. If you're typing it into a platform that you don't own, that doesn't have deleting messages or something like that, it could go somewhere that you don't want. Screenshots exist. um you know it's people are just i don't know people are putting wild things into these tools that they ultimately have no control over which is wild to me um yeah if you don't have any closing thoughts i will play us to the next one what do you usually say i'll play us out in the next story.
21:20Usually what I do is, I mean, I will do it now, but I do a little preview of the next one and then we go to it. But I will say, as a closing thought, there are more exposures here. I don't really want to say right now because we're pretty busy. So maybe we don't get this story out in time. But there's other stuff going on here. So we probably won't talk about on the next episode because we try not to repeat it too much. But there might be another article coming. I'll say that. And we'll leave that there. And we'll leave that there. That's what I was trying to think of. Sorry, I didn't write that down.
21:56I should put that in the Google Doc. And we'll leave that there. And when we come back, we're going to talk about one of Emmanuel's stories about a policy change on Wikipedia that might protect the site and the platform from AI slob. We'll be right back after this.
22:26Why drop a fortune on basics when you don't have to? Quince has the good stuff. High quality fabrics, classic fits, and lightweight layers for warm weather, all at prices that make sense. I'm always on the lookout for new basics companies, and everything I've ordered from Quince has been nothing but solid. Quince has closet staples you'll want to reach for over and over, like cozy cashmere and cotton sweaters from just$50, breathable flow knit polos, and comfortable lightweight pants that somehow work for both weekend hangs and dressed up dinners. The best part? Everything with Quince is half the cost of similar brands.
23:03By working directly with top artisans and cutting out the middlemen, Quince gives you luxury pieces without the markup. And Quince only works with factories that use safe, ethical, and responsible manufacturing practices and premium fabrics and finishes. is. In the last few weeks, I picked up a 100 % European linen shirt that has entered my regular rotation. That's this right here. And I've also got a few 100 % cotton tees, which have gotten compliments for how well they fit. And they come in a lot of varieties, which is good because I really like a sturdy, thick look and Quince has those as well.
23:37We also picked up a basket weave quilt that has quickly become our favorite blanket in the house. Keep it classic and cool with long lasting staples from Quince. Go to quince.com slash 404media for free shipping on your order and 365 day returns. That's q-u-i-n-c-e.com slash 404media to get free shipping and 365 day returns. quince.com slash 404media. One of the scariest parts about building 404media was figuring out the logistics of, well, how to do business. There's a handful of tools that make 404media run, but it's been a real pleasure to use Shopify, which has given our company a footprint in the real world with our merch store.
24:22Without Shopify, I don't know how we would have done it. So if you're thinking of starting a business, start with Shopify. Shopify is the commerce platform behind millions of businesses around the world and 10 % of all e-commerce in the US. From household names like Mattel and Gymshark to brands that want to be household names like 404 Media. Shopify's got you from the get-go with beautiful ready-to-go templates to match your brand style. Their easy-to-use backend helps you manage your store's inventory and makes creating an attractive shop for your customers really easy. They also help you find new customers with easy-to-run email and social media campaigns.
24:58And if you get stuck, Shopify is always around to share advice with their award-winning 24-7 customer support. So turn those dreams into and give them the best shot at success with Shopify. Sign up for your$1 per month trial and start selling today at shopify.com slash media. Go to shopify.com slash media. Shopify.com slash media. This is an ad by BetterHelp. These days, it feels like there's all kinds of ways to treat your mental health. Cold plunges, gratitude journals, screen detoxes, but it's hard to know which one of these will work for you and what is just noise or fads on the internet. One thing that's long tested and long trusted is talking to live therapists to help get you personalized recommendations and help you break through the noise.
25:51As a therapist gets to know you, they can provide personalized suggestions for positive coping skills, stress reduction techniques, and strategies that will help you become the best version of yourself. BetterHelp is easy to use and easy to plan around. It has more than 30 ,000 therapists and has served more than 5 million people globally, meaning you can fit therapy into your busy life. Join a session with a click of a button and switch therapists at any time. As the largest online therapy provider in the world, BetterHelp can provide access to mental health professionals with a diverse variety of expertise.
26:27Talk it out with BetterHelp. Our listeners get 10 % off their first month at betterhelp.com slash 404 media. That's betterhelp.com slash 404 media.
26:52All right, and we are back with a manual story. The headline is Wikipedia editors adopt speedy deletion policy for AI slop articles. First of all, Emmanuel, what is the AI Wikipedia problem? Does it have a big AI slop problem? We've spoken about it a little bit before, but what's the problem there for Wikipedia? I would say Wikipedia has a bigger problem than you or any other average user of the site realizes. and that is because of the incredible effort that the wikipedia editors the contributors the volunteers the people who maintain wikipedia um put into the site to kind of protect you from that problem so it exists some of it is visible in the sense that sometimes ai generated articles that are wrong, that are filled with hallucinations and fabricated information, do make it to the live version of the site.
28:03But behind the scenes, the editors who approve articles and discuss articles, they're dealing with a huge flood of AI-generated articles in the same way that all the platforms that we talk about here every week are dealing with that on Facebook, on Instagram, on Twitter, YouTube, just people are flooding Wikipedia with their generated articles. You're not seeing it because the editors are filtering that stuff out. Yeah. Maybe we don't know this, although I feel it came up in conversations before. What does some of that slop look like? Obviously, Wikipedia, it is articles about specific subjects.
28:45Is it people trying to fuck with Wikipedia? Is it people who think they're really smart and they've discovered something through ChatGPT and they're like, I have to now tell the world this on Wikipedia. Do we know what the slop is exactly? I presume it's varied. Yeah, it's varied. It's funny you mention it. I talked to you about this months ago. I doubt that you remember. But after I wrote this article about a group of Wikipedians that have this initiative to protect the platform from AI-generated content, They put together this document showing examples of AI-generated content on Wikipedia. And one of them is an article about the Rule Aloum du Bond, an Islamic seminary in India.
29:35And it's like, there's an old painting showing some of the people in that article. and it's just AI generated and they have six toes and stuff like that. There was another article about... It was a castle in Turkey, I believe. Long article, thousands of words, completely made up. Just total fabrication of ChatGPT. And I got in touch with the guy who made that article. And I was like, what's up? Why did you do that? And I got a very confusing answer. It sounds like maybe he's Armenian and there is some sort of attempt to show... No, he's Turkish who takes issue with the articles on Wikipedia about the Armenian genocide.
30:27And he was kind of trying to show that anyone could get anything up on Wikipedia. So it was like some sort of attempt to undermine the validity of Wikipedia as a reliable source of information was his reasoning. but it was kind of nonsensical. It was somebody fucking with the platform to make some kind of point. It's like the AI generated article that I saw. Right. And it does remind me now you mentioned it of that American who ended up writing half of Scott's Wikipedia. Like obviously a specific language, right? And this teenager did not speak the language but somehow blagged their way through filling up massive swaths of Wikipedia, which, I don't know, obviously that's a massive time commitment because that didn't involve AI.
31:16And maybe ironically, if you were trying to do something like that, you know, to be a bit of a dick or for your own ego or whatever, like if you used AI, it might actually be harder because maybe you'd be detected by Wikipedia more, you know what I mean? Rather than the artisan hand crime control. Yeah, I mean, it's interesting. you bring that up because one of the editors I talked to for this article brought up this point. The AI generated articles are a bigger problem for them for the same reason it's a problem for moderation on other platforms, which is that AI content is being generated at the rate of a machine, whereas the verification of the articles is being generated at the rate of like a team of human beings reading every word of the article so there's like an inherent imbalance there so even though the stuff can be very easy to detect sometimes there's just so much of it that they had to change their policy about how they review articles which is kind of like the the policy at the heart of this article yeah so So what is this new policy exactly?
Read the full transcript
32:35Because it sounds like they're already pretty well equipped. What's the new policy? Yeah, so Wikipedia is run by a community. Everything is done by open discussion and consensus or voting. And articles get deleted all the time for a variety of reasons. Sometimes articles are just plainly planted advertisements for companies or people promoting something. Sometimes they're just pure gibberish. Sometimes people will argue about whether an article deserves its own... Whether something deserves its own article or it should be a segment of another bigger article. And most of those deletions happen after a seven-day period of discussion.
33:26It's like somebody says, hey, I think this should be deleted because this article is not notable, meaning it should be either deleted or integrated into another existing article. And then editors have seven days to discuss this and hopefully reach a consensus or at least a vote about whether that is true or not.
33:48for some of the more obvious cases like these gibberish articles that i've mentioned they have a speedy deletion policy which means one person flags the article an editor or an administrator sees it they confirm it's just like it's nonsense it's not even english and then they can just delete it without this like week-long discussion and the proposal that was eventually adopted yesterday is a speedy deletion process for certain types of AI articles because there are so many of them now. So changing the policy to include another type of article for speedy deletion is a big deal on Wikipedia. They can't do it for the majority of AI-generated articles because the majority of AI-generated articles, there's some doubt.
34:44People can argue like, oh, there's bullet points and bolded chapter heads and em dashes and all these signs. Right. I don't need to get those on Wikipedia necessarily, but that sort of thing. Yeah, like all these signs are clearly the style of a chat GPT. But people can argue like, A, is this really AI-generated? And B, I don't know, maybe there's a value to it. So those still have to be debated. But there's another category of AI-generated articles where there's two conditions that they have to... One of two conditions that they have to meet in order to be eligible for one of these speedy deletions that don't require a discussion.
35:27One is if they include what is clearly language directed at the user. So we have used this method to identify AI-generated, LLM-generated content many times. And that's when you're talking to ChatGPT, you say, ChatGPT, tell me about the history of aviation or something, right? And it will say something, as per my last knowledge update, which means according to all the information I've ingested as an LLM last time in January 2024, this is the history of aviation. And that language is clearly an LLM talking to a user that prompted it. And if it includes that, A, it means the article is clearly AI-generated.
36:23And B, more importantly, it means that the person who submitted the article didn't even bother to read it. because if the person read it, he would notice that and remove it and then submit it. And that's the worst bit. That's just rude. It is. No, that's actually the policy. They're like, if the person who submitted the article didn't even read it, we're not going to sit here and debate it for seven days whether it's worth deleting. That's just indication nobody cares about this article. It's just white noise. Let's get rid of it. The other condition is if the article includes references that don't actually exist.
37:02So I think Sam reported on this a few times. A bunch of lawyers have been busted doing this where they'll submit a complaint and they'll cite cases like Cox versus Cole. And in 1999, you look it up and it's like, it never happened. It's not a case that exists. It's a classic. What are you talking about? Precedent setting case. So when that happens, Wikipedia takes sourcing very, very seriously, I would say it's one of the primary uses of the site and why it exists. It's not just about reading the summary, but being able to follow the sources to justify why stuff is included in the article. And when you include a fake source, you really violate a basic principle of Wikipedia.
37:48So if those citations are clearly fake, if they don't resolve, then lead to a non-existent page, a 4-4 page on a scientific journal or something like that that is also a reason to just remove the article without discussion yeah that makes sense i guess the last thing well no i'll ask two more things you touched on this and maybe it's a little bit unclear but those examples are very obvious going back a little bit to the harder ones where they have to have this longer discussion any idea how they're gonna determine it or is it more case by case and you figure it out because I mean, it's getting harder and harder.
38:29Some are very obvious, but it is getting trickier to figure out if something's AI. Have they given any specifics on how they might do that? Or maybe they find out now this policy is out there, you know? So for the story, I talked to Ilyas LeBlou. I hope I'm pronouncing that correctly. They are one of the founding editors of the Wiki project AI Cleanup, which is this group of editors that are actively trying to protect Wikipedia from AI-generated content. And the way they put it to me is that this is definitely an improvement, and it puts Wikipedia at a better place than it was. But the AI issue is definitely not resolved.
39:18The issue will persist and is still a big issue. It is helpful in two ways. One, there's a category of AI-generated content now which can more easily be deleted. And that's a big deal. That just reduces their workload, which is the main purpose of this policy. It frees them up to review all this other content. The other reason that it is useful is AI-generated content and how it is being used on Wikipedia has been pretty controversial. I wrote a story a couple of months ago about the Wikimedia Foundation, which is the nonprofit that owns Wikipedia. They introduced a measure or a feature. They were piloting a feature that would include an AI summary of the article at the top of the article.
40:18And the Wikipedia editors basically revolted and got so mad that Wikimedia decided to pull it back. And this is kind of the first editor-led policy change that has a clear anti-AI policy. And LeBlou, this editor, thinks that it's useful in that sense because it's like black and white policy anti-AI, which will signal to Wikimedia and will signal to other editors and will signal to users that it's like, they're very vigilant. They're taking this issue seriously. They're not going to roll over like some other platforms have for AI-generated content. Speaking of other platforms, the last thing I wanted to ask is what can other platforms learn from this?
41:06Wikipedia is obviously a pretty unique thing, but everybody is dealing with AI slop. What do you think other platforms dealing with that problem can learn from Wikipedia here? So I think 10 years ago, 5 years ago, we wouldn't cover policy changes, obscure policy changes that Wikipedia editors are putting in place. The reason that I think this is newsworthy and the reason I've been paying very close attention to how Wikipedia is responding to AI is because, again, it's a community-led project. And when the community is at the wheel, we're seeing them actually take a position that I think, you know, it is anti-AI when compared to Instagram or Meta generally, which is like, we love AI, we make AI products, flood the zone with all the AI-generated content you can possibly imagine.
42:14Like, LeBlue was very measured. They were like, it's possible that in the future, generative AI tools will make Wikipedia better. We already use some AI tools to do certain things. but to speculate on how it might be useful in the future or not is blinding us to the problems that exist currently and the problems that exist currently is that we're getting hallucinated facts and false citations and articles so we need to deal with that now and i think that just like a good model a it reminds me a lot of like our approach to technology reporting. It's just like, let's talk about what's happening now, what's happening to users.
42:58And yeah, I just, I think it shows that there's a different way. You don't have to roll over and just allow it to flood your platform because it's the newest, coolest thing. You can respect the users, respect their time, and set some sort of boundary and have some really practical policy. The thing I discussed about the language that's clearly an LLM talking to a user, we see that prompt in digital libraries and scientific journals and social media posts all over the place. You can imagine having a filter in place at all those places that at least flags that content for review or something. And nobody really does that yet.
43:47So I think to see Wikipedia do it is encouraging and hopefully other people will take note and adopt a similar policy. Yeah, I think it's really interesting that a relatively small policy change could actually be an indication of how platforms, people, companies, organizations can resist AI if they wish to do so. I think it's a really interesting model and policy update for that. Okay, we will leave that there. If you're listening to the free version of the podcast, I'll now play us out. But if you are a paying 404 media subscriber, we're going to go into a really deep dive into one of Sam's stories about how we got here with Steam and Itch.io and gamer censorship and all of this stuff.
44:38It absolutely did not come out of nowhere. So if you want to understand how we got there to this point, you can subscribe and gain access to that content at 404media.co. As a reminder, 404 Media is journalist-founded and supported by subscribers. If you do wish to subscribe to 404 Media and directly support our work, please go to 404media.co. You'll get unlimited access to our articles and an ad-free version of this podcast. You'll also get to listen to the subscribers only section where we talk about a bonus story each week. This podcast is made in partnership with Kaleidoscope. Another way to support us is by leaving a five-star rating and review for the podcast.
45:21That stuff really, really does help us out. This has been 404 Media. We'll see you again next week.
From the publisher
We start this week with Joseph’s story about nearly 100,000 ChatGPT conversations being indexed by Google. There’s some sensitive stuff in there. After the break, Emanuel tells us about Wikipedia’s new way of dealing with AI slop. In the subscribers-only section, Sam explains how we got to where we are with Steam and Itch.io; that history goes way back.
YouTube version: https://youtu.be/mQJvOTHu61I
Nearly 100,000 ChatGPT Conversations Were Searchable on Google
Wikipedia Editors Adopt ‘Speedy Deletion’ Policy for AI Slop Articles
The Anti-Porn Crusade That Censored Steam and Itch.io Started 30 Years Ago
Subscribe at 404media.co for bonus content.
Learn more about your ad choices. Visit megaphone.fm/adchoices
