We Spoke to an Amazon Worker Destroying Books for AI

2 Sep 2026 · 41 min · 14 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Amazon worker describes how books are bulk-shipped to a Las Vegas warehouse (VGT3), de-spined, scanned (barcode images), and then destroyed/recycled for AI training data; episode also covers “AI ghosts” (LLMs reusing the same fake author names in academic publishing).

Guests/backgrounds

Joseph (host) with 404 Media co-founders Sam Cole and Emmanuel Mayberg. Emmanuel is the reporter who previously investigated the Amazon book-scanning warehouse and now interviews a worker there. Sam discusses related AI “ghost author” research.

Key claims

Books are cut at the spine, pages scanned via ~20–25 “cash-counter”-style scanners, then loose pages are tossed into large “shuttles” (Amazon “Gaylord” containers) with no salvage; at least some shipments end in a Mexico paper mill producing toilet paper/paper towels. LLMs generate recurring fake names (e.g., “Alina Vasquez” and “Marcus Chen”), which appear in ~1,655–2,000 DOI/Zenodo records, spreading to Google Scholar/ResearchGate.

Notable examples

A bookseller’s tracking device ends at a Mexico paper mill; library-tagged books and UK/university reports appear in shipments; “Elias Thorne” lighthouse-keeper persona repeats in AI stories and became sellable.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Anniversary Event Details

1:16 to 3:00

Discussion about the upcoming anniversary events and ticket information.

“gain access to that content at 404media.co.”

Event Perks and Merchandise

3:00 to 5:30

Details on subscriber perks for the events, including free access and exclusive merchandise.

“But definitely get in there if you're already a supporter.”

Amazon Warehouse Book Destruction

5:30 to 9:12

An update on the Amazon warehouse that destroys books for AI training.

“for doing anti-AI whiteboard, like pen on paper adverts.”

Insights from Warehouse Workers

9:12 to 11:10

Discussion on the types of books being destroyed and the workers' experiences.

“We always want to speak to essentially as many people as possible.”

Operational Insights of Scanning Process

11:10 to 14:00

Details on the scanning operation within the warehouse and the technology used.

“And that could be from a library liquidation sale.”

Inside Amazon's Book Destruction Process

14:00 to 22:21

Learn about the workflow and machinery involved in Amazon's book destruction practices.

“that lawyers use agents all the time or chatbots all the time.”

Ad: Protect Your Privacy with Incogni

22:21 to 23:27

Discover how Incogni helps remove your information from data broker sites.

“So for fun, I googled myself last week and found my home address, my phone number, and a relative I'm fairly sure I never met.”

Ad: Protect Your Privacy with Incogni

23:46 to 24:57

Discover how Incogni helps remove your information from data broker sites.

“If I'm walking outside, working around the house, or getting a workout in, I still want to hear the person talking to me, the car coming up behind me, or whatever else is happening around me.”

Ad: Protect Your Privacy with Incogni

25:02 to 26:36

Discover how Incogni helps remove your information from data broker sites.

“I think traveling gets a lot more fun when you can actually participate in what's happening around you.”

Ad: Protect Your Privacy with Incogni

26:38 to 26:51

Discover how Incogni helps remove your information from data broker sites.

“That's Babbel, D-A-B-B-E-L.com slash 404 for up to 60 % off.”
Show all 14 chapters

AI Ghosts in Academic Publishing

26:51 to 28:00

Explore the implications of AI-generated personas in academic documents.

“This is one Emanuel wrote, but I definitely have some questions for Sam as well because it strongly relates to an article she recently wrote.”

Examining AI-Generated Content Detection

28:00 to 31:16

Learn about the statistical patterns in AI-generated content and how researchers are identifying AI-generated academic papers using specific names.

“So there is this phenomenon with LLMs where they keep producing the same names in certain contexts.”

The Case of Elias Thorne and Model Collapse

31:16 to 36:37

Discover the phenomenon of 'model collapse' in AI storytelling, focusing on the recurrent character Elias Thorne and its implications.

“And this is in order for you to then submit the paper to an actual journal.”

Statistical Reinforcement in AI Behavior

36:37 to 39:22

Explore how AI models reinforce certain names and behaviors through statistical patterns, affecting content generation and detection.

“oh, this is a story that people are interested in.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Emanuel:Hey Chicago, class it up with Crocs. You know back to school is coming in fast. So why wait to find your new fave footwear? Step into a local Crocs store and step into your new look. Try it. Style it. Make it yours. Because the right pair doesn't just show up, it shows off. First day fits, handled. Walk out ready for whatever's next. Visit your nearest Crocs store today. A bookseller independently decided to put a tracking device in one of their books. The final ping from their tracking device was in a giant paper mill in Mexico, where they receive recycling materials and then they turn it into like toilet paper, paper towels and stuff like that.

0:47Wow. The circle of life is a beautiful thing. You might be wiping your ass with like a copy of whatever. Which just means scammed by Amazon. Yeah.

1:01Hello, and welcome to the 404 Media podcast, where we bring you unparalleled access to hidden worlds, both online and IRL. 404 Media is a journalist found in the company and needs your support. To subscribe, go to 404media.co. As well as bonus content every single week, subscribers also get access to additional episodes where we respond to their best comments. gain access to that content at 404media.co. Also, do remember to subscribe to our YouTube channel where you can watch all of our episodes, including this one. Subscribe at youtube.com slash at 404media.co. I'm your host, Joseph. And with me are two of the other 404 Media co-founders, the first being Sam Cole.

1:47Emanuel:Hey. And Emmanuel Mayberg. Hello. Well, Sam, first of all, this Thursday and Friday, those are the days of our three-year anniversary events, the live podcast recording on Thursday, the party on Friday. What's the ticket situation looking like for people if they still want to get a ticket? Yeah, tickets are very limited at this point. I would say they're like edging selling out. So if you haven't, get your tickets. Thursday, we're doing, like you said, the live panel at the Green Space. It's a WNYC event space that's down... I guess it's considered like Hudson. It's like Chelsea-Soho intersection in Manhattan.

2:37Emanuel:So we're going to have Matthew Gault joining us. Becky Ferreira is coming out to talk about cool science stuff. So definitely get your tickets for that if you haven't. That at this point is very much almost sold out. And then on Friday, there is an open bar party at Three's Brewing in Gowanus. And that is free for paid subscribers. So if you're a subscriber at the supporter level, get your ticket. You still need a ticket. You still need to sign up. But definitely get in there if you're already a supporter. It's free for you to come out. We're going to have snacks. We're going to have good snacks like fries and veggies and stuff like that.

3:11Emanuel:I don't know. Bar snacks, I guess. And good beer. And good beer. Good beer. wine cocktails all that good we'll have the photo booth again we'll all be there most importantly to party with you and if you're not a supporter you can actually get 50 % off of that event if you are coming to the green space live panel so lots of ways you can get into this and we'll put the link in the show notes for all the details here is a little secret in case it wasn't obvious in us promoting this event over and over again you can buy a ticket to the open bar for about$30 or you become a subscriber at$10 and then you get access to an open bar.

3:51So you're just paying$10 for like somewhat unlimited beer.

3:55Emanuel:It's a good deal. So how many beers do I need to drink to make up the cost of$10? Like three quarters of one. Okay. And how many beers do I need to drink to cover the cost of an annual subscription? An annual? Well, I wouldn't do that. Yeah, like 10. I mean, buy an annual subscription. I don't think you should probably drink all of those beers in one go. Listen, here's the thing. If you come to the bar and you show me that you're a super fan, which is a$1 ,000 subscription, I will drink the cost of a super fan subscription with you. Good lord. Don't do that because Emmanuel will die and then we will be down.

4:36Emanuel:I will die. That is true. I mean, that's the offer. That's the offer, yeah. The sacrificial lamb. If you would like to kill me. Yeah. It will be easy. Yes. As Sam said, check the link in the show notes. We'll have merch, auto discounts. So if you're coming in person, we'll have merch that's cheaper and easier to get, obviously, because we'll just hand it to you than it is online right now. And then we'll also have exclusive posters, which just came in the mail today. They're like holographic, really cool event posters that we're going to have. They look so good. Yeah, they look very sick. And then we'll have special stickers that I just slapped together and basically in paint and put it into the printer and go.

5:18Emanuel:So yeah, lots of things to not miss this week. Hell yeah. I like the paint stickers. It is kind of like a story we did where those businesses are going viral for doing anti-AI whiteboard, like pen on paper adverts. And we're doing that as well. All right. As for this week's stories, let's start with a really interesting update to one of Emmanuel's about the Amazon AI training warehouse that you'll be recapped on in a second. The headline of this one is Inside the Warehouse where Amazon Scans and Destroys Books for AI Training. This is obviously a huge story that we did a couple of weeks ago at this point, Emmanuel.

6:04Just recap our listeners super briefly. What is this warehouse and how did we find it? Yeah, so the warehouse is in Las Vegas. It is called VGT3. And it is at least one location that we know of where Amazon, they purchase huge bulks of books. They're shipped there. People who work at the warehouse cut us the spine and scan the loose pages in order to create AI training data. And we know this because we had suspicion that many AI companies were doing this. We weren't able to say that for sure because the buyers on these marketplaces are anonymous. But we put a tracking device in one of the shipments and saw that the shipment ended up at this warehouse where we were able to confirm that this is in fact what was happening.

7:02so that reporting obviously as you say was based mostly on the apple air tag and then conversations with the bookseller source and you also did find discussion online from people who worked at this warehouse saying oh yeah you know we take the spines off the books blah blah all of that but you managed to get in touch with somebody who actually worked inside this facility we'll get to what they said in a minute. And the article is sort of formatted as a Q &A. And frankly, I basically recreate that Q &A by asking you the same questions and you repeat their answers. But before we get to what they actually said, why did you want to speak to somebody inside?

7:46What were you hoping to learn or get by speaking to somebody who'd actually been inside this facility? Well, initially, when I was reporting the previous article, I wanted additional confirmation that they're destroying the books. But what I thought was useful about this conversation is a few things. One, there is a better idea of the scale of the operation, just in terms of how many people are there and what does the space look like? What does the equipment look like? How much equipment there is? How many books are coming in? And then also, what type of books are coming in. And then as always, as a publication, we're always interested in the perspective of workers.

8:38It's like, this is a big operation. Presumably, this is happening in many warehouses across the world. And there are hundreds, maybe thousands of people involved in this. And just like, what are their working conditions? What do they feel about the objective of what they're doing? And we're interested in that always. If we're reporting on Uber, we want to know how drivers are doing. If we're reporting about Amazon warehouse, we want to know how the warehouse workers are doing. And they often provide the best insight, which I mean, there's definitely great insight in this conversation, I think.

9:13Yeah, it totally makes sense. We always want to speak to essentially as many people as possible. And I think for us especially, we do focus on people who are actually doing the work. Like we're never going to get interviews with Mark Zuckerberg or Jeff Bezos. Or we're never going to get interviews with the executive class, basically, right? And frankly, I think when you have conversations with those sorts of people, you actually really don't get that much out of it. They are trying to push a certain message. But speaking to an actual worker, you're going to get more specifics on what is actually going on.

9:50So first of all, what did this person say about the sort of books Amazon is scanning? Because the article focused on rare books because that is what was being ordered. But what else did this person say about the sort of books that Amazon is buying to then de-spine and scan and destroy ultimately? So first of all, there's confirmation that it's any and all books, right? It's like you can see some logic in terms of there being a lot of textbooks and stuff like that. But there's a huge variety of books. The worker did, however, see a few things that I thought were interesting. One is they would get big shipments of specific languages, right?

10:41So it's like they come into work one day and there's a palette of Japanese books. They'll come into work one day and it'll be a palette of Russian books. And that doesn't mean that this isn't happening at other AI companies around the world. In fact, there was some reporting in Europe that was able to identify a Chinese AI company that was buying books. but we've seen reports of booksellers reporting the same thing uh i heard uh but in in the netherlands in spain and germany and i was wondering like okay are those local companies or are all these books being shipped to the u.s and at least in this case we could we could say at least some of those books are are in fact being shipped to the u.s uh another thing that was interesting is some bulk orders clearly included library books they had all the the the tags and the markings of books that you would lend from the library like the tags still on them or like yeah like you know on the spine you have like the code and and all that so it's like they They would get big shipments like that.

11:55And that could be from a library liquidation sale. It could also be from a bookseller who bought it from a library liquidation sale. It's hard to say for sure. But we do know that library books are ending up in there. The last thing I thought was interesting in terms of what books they were seeing is there was another shipment that clearly came from the UK. and seemed like it came from a university. And there were a lot of textbooks and stuff like that. But there were also books that didn't actually have spines. They were stapled together. And they looked like government reports. I think as the worker speculated, and I agree, they're most likely public reports.

12:40It's not like it's super secret information or anything. But it's just interesting to note that it doesn't have to be like a bound book. They would get other reports and guides and stuff like that that weren't necessarily published books. Yeah, because obviously with a lot of US reports, you can just get those online. And AI companies would have grabbed those in their mass scraping of the web. But in some cases, and I'll say this is definitely the case for the UK, where information just isn't as available on the internet in a lot of UK stuff. So I can see that, oh, maybe somebody had to print it out or something to scan over.

13:29We don't have any evidence for this, but it just kind of made me think. I wonder if the AI companies are scraping PASA, the US court record system as well, and ingesting that. I can't say for certain, but I guarantee it. I'm linked to that money on it. I'm just trying to think of the value because there's a lot of... There definitely is value in that to be like, well, here's the understanding of the law or something, and this is how lawyers are using it. Or just formatting stuff in the style of complaints or motions and stuff like that. I mean, obviously, we know from other reporting we do that lawyers use agents all the time or chatbots all the time.

14:06Yeah. It's just interesting because there's also a lot of wrong information in court records as well, where it's all allegations hasn't been proved. to and that's what we... Anyway, maybe we'll put a pin on that because that is interesting. Maybe we should look into that. But back to this Q &A, what did this Amazon worker say about the scanners? Like, how many are there in this warehouse as they know of? Like, how big are they? That sort of thing. Yeah. So a little bit about the operation. It's pretty much what I imagine and what I gathered from discussion online. But basically, trucks come in, they take everything off the truck in an area they call staging then this person you know saw that there's a team that's in charge of like taking stuff from like this staging area and then putting stuff into bins that then another team will take over to an area called receiving which is where the books are scanned in as in they're not scanning the content of the book They're scanning the barcode to see what are they getting.

15:14From there, they're taken to a bunch of stations where they cut the spines off the book. This is a big machine with a blade. You put the book into this secure area. You remove your hands from the area. You press a button. A big blade comes down and cuts the spine off the book. Then the loose pages are put onto these carts. and they're separated by little pieces of cardboard and they're taking over to the scanners where the worker said there's about 20 to 25 scanners there. And I haven't been able to nail down exactly what the machine is. There are a bunch of machines, a bunch of scanners that are designed for this process.

15:59The worker described these as, you know, those cash counting machines You always see it in crime movies. You feed it a big wad of cash and it makes that satisfying flipping noise. So it's like... And very neatly churns it out. Yeah, yeah, yeah. Yeah. So it's one of those, but for scanning. And there's a monitor next to the scanner that shows you what you captured. There's an image of the page. from there everything all the loose pages are tossed into these uh giant cardboard uh open boxes uh people in the business of of book book selling call these gay lords amazon calls them shuttles the workers speculated i didn't put this in the article but the workers speculated it's because people were like being juvenile about the term gay lord obviously the term has taken on a slang meaning and a pejorative meaning in modern context.

16:59So having a corporation like Amazon officially referred to something as Gaylord may not go with the company's HR policies, I guess. Right. So they call them shuttles. And then everything is kind of like loosely tossed in there. And we'll get to what happens to the books after. But the worker observed, and I think it's inarguable that there is no way to salvage what goes into that that giant box like it just it just piles of loose paper and uh they also noted that not all the books that they get end up getting scanned they don't know why maybe it's a duplicate maybe it's not a book that they're interested in maybe it's something that they like they didn't actually order but ended up in the shipment.

17:49And those books end up in that pile as well. So these are books that are functioning, salvageable books, right? They're not destroyed. But almost certainly, they're ending up in a pile of books that will get recycled. Damn. So yeah, they are. I mean, on one side, they're obviously destroying the books and then scanning them for those AI purposes. But in some cases, they are just destroying some books that don't even get scanned at all. Just a couple more brief questions. What did they say about the job itself? Because in your original article, based on what people were saying online, it seems like quite a desirable job, at least inside Amazon itself.

18:32What did they say about that? Yeah, so I didn't get the feeling that this worker found this job particularly desirable. like I said in the previous story the same structure the same building where VGT3 is also has a separate operation for print on demand it sounds like maybe that is a harder job by comparison and maybe that's what makes people say that VGT3 is desirable but it sounds like a normal Amazon job you're at one of these stations so you're either cutting the books you're lifting a bunch of books or you're tossing them everything else they have to say is normal complaints that people have about their workplace.

19:16And typical of Amazon, which is it's a big operation. There's a lot of turnover. And this worker got the feeling that they would come in and... Amazon was still figuring out how this works. Amazon was still figuring out the ins and outs of the operation and how to make it efficient. So they would feel the order of things would change daily or who was doing what would change daily. But I think it's hard to say whether that says anything in particular about this operation or that's just what it's like working on Amazon or that's just what it's like working at a big warehouse. Yeah. And just to round it off, since you published this piece, you've done a ton of media interviews and appearances on TV and for other outlets and that sort of thing.

20:06And we've also had more mainstream outlets like the Wall Street Journal follow up on our reporting and publish their own stories. Have we learned anything else either about the Amazon story specifically or the broader trend of AI companies scanning and destroying books since we published? There has been a lot of pickup on it. I think what other reporters were able to get is interviews with booksellers and bookstores that are confirming the same activity, which is to say giant purchases of books. there was one interesting thing about that wall street journal story which i have some corroborating evidence for but basically a bookseller independently decided to put a tracking device in one of their books just because they were curious about what was happening to the to the book and they they saw it kind of make a similar route across the country as what we saw but they ended up like the final ping from their tracking device was in a giant paper mill in mexico in mexico and i looked up this location uh i was able to find it and yeah i mean it's like without a doubt that is what is happening like the the tracker that was in this book eventually ended up at this giant paper mill where they receive recycling materials and then they turn it into like toilet paper paper towels and stuff like that so um we can say that at least in some cases that that's the ultimate location for the books wow the circle of life is a beautiful thing dude i really i was thinking about it it's like you might be wiping your ass with like a copy of whatever you know which just means scanned by amazon yeah all right uh with that we'll leave that there.

22:07When we come back, we're going to talk about AI ghosts. No, we're not personifying them. Don't worry. It's another sort of ghost. We'll be right back after this.

22:21Emanuel:So for fun, I googled myself last week and found my home address, my phone number, and a relative I'm fairly sure I never met. All on directory sites I never signed up for. That's the data broker economy. They scrape your info, package you up like a trading card, and sell you off to whoever's paying, which is exactly how identity thieves go shopping. Bad luck when your entire job is not getting doxxed. So we use Incogni. Instead of spending hundreds of hours filling out opt-out forms across 40 browser tabs, Incogni gets to work by sending legal removal requests to over 420 harmful data broker sites.

22:55Emanuel:It's independently verified by Deloitte, resubmits requests every 30 to 60 days as brokers try to quietly relist you, and you can track every request right from the app. On the unlimited plan, the exposure scanner digs up the weirder stuff across the wider web, like old forums, random blogs, or obscure directories you might be listed. It surfaces all of your exposed profile links right in your dashboard, where you simply select which links you want taken down, and Incogni's team of human privacy experts will handle the manual corporate takedowns for you. Take back control today. Go to incogni.com slash 404media and use code 404media for 60 % off an annual plan.

23:32Emanuel:That's incogni.com slash 404media, promo code 404media. Risk-free with a 30-day money-back guarantee. They can't harm you if they can't find you. This message is sponsored by Raycon. Sometimes I want headphones that completely block everything out. But most of the time, I actually don't. If I'm walking outside, working around the house, or getting a workout in, I still want to hear the person talking to me, the car coming up behind me, or whatever else is happening around me. That's where Raycon's essential open earbuds have really been useful. They sit just outside your ear canal so you get clear audio without completely disconnecting yourself from the world.

24:09Emanuel:I can have a podcast or music going and still hear what's happening around me at the same time. And because they're super lightweight with a flexible rotating earhook, they stay in place without feeling like I've got something jammed in my ears all day. They also have multipoint connection, which is great because I can have them connected to multiple devices and switch between them without doing the whole Bluetooth settings dance every single time. And the battery is kind of ridiculous. Eight hours of playtime and up to 36 hours total with the charging case. I've been using various Raycon headphones for months now and I really like them.

24:39Emanuel:I wear them when I'm walking my dog at the gym. I use them to take calls. I use them to listen to podcasts and music. So I highly recommend them. Raycon has more than 3 million customers. And there's also a 30-day guarantee if you want to try them out for yourself risk-free. The Essential Open Earbuds are the perfect companion for your fall reset. Go to buyraycon.com slash 404 open to get 20 % off. That's buyraycon.com slash 404 open to get 20 % off. Thanks to Raycon for sponsoring. I think traveling gets a lot more fun when you can actually participate in what's happening around you. You don't need to know every word, but being able to order dinner, ask for directions, or have a simple conversation with someone when you're having a beer or over dinner or coffee, instead of just pointing at things and hoping for the best?

25:24Emanuel:Well, that changes your entire experience. That's why I've been into Babbel. Instead of having you memorize a bunch of random vocabulary or stare at verb charts, Babbel focuses on real conversations you might actually have. Even just 10 minutes a day can help you start having real conversations in as little as three weeks. Babbel's lessons are quick and practical, and they're built by more than 200 language experts. You also get interactive dialogues, personalized reviews, and even podcasts, So you're practicing the language in a bunch of different ways. That means you won't get bored of any one way of learning.

25:57Emanuel:Babbel teaches you in all sorts of different ways, which is what I really like about it. If you have a lot of time, you can spend all day on the app. But you can also fit a lesson into a coffee break, your commute, or 10 minutes before bed. Babbel's award-winning app has sold more than 25 million subscriptions, and it's backed by a 14-day money-back guarantee. So if you're traveling somewhere this summer, Don't make the first time you try speaking a language the moment you're standing in front of someone trying to order dinner. Start now and you'll be ready before you go. If you've got summer travel coming up, now's the time to start so you can actually use what you learn on the trip.

26:32Emanuel:Right now, Babbel is offering listeners up to 60 % off. Go to babbel.com slash 404. That's Babbel, D-A-B-B-E-L.com slash 404 for up to 60 % off. Rules and restrictions may apply.

26:50All right, and we are back. This is one Emanuel wrote, but I definitely have some questions for Sam as well because it strongly relates to an article she recently wrote. But the headline for this one is The AI Ghosts Contaminating Academic Publishing. So I'll just read the lead and then I'll ask you about Emanuel because I think it sums up really, really well. Or rather, I think it's a quote from the study. Alina Vasquez and Marcus Chen have appeared as volcano experts, astronauts, thriller protagonists, podcast hosts, and academic co-authors across hundreds of independently produced AI-generated documents never having lived.

27:34That's a pretty compelling sentence. As people can probably get from the headline, these names, and just for the sake of this conversation, these people, these personas, these characters, whatever, they are AI generated. And a new paper is looking at that. So what's the deal? These names keep appearing in AI generated text. What did this paper find? So there is this phenomenon with LLMs where they keep producing the same names in certain contexts. Obviously, it's a lot more complicated than this. I do not mean to minimize the technology of large language models, but it is a statistical model. So statistically, the same data will often produce the same concepts and also the same names.

28:41and this is known uh there's been other reporting on this uh we'll talk about sam's story in a bit but it just so happens that they could appear in other contexts but mostly when you're looking for if you tell cloud like hey like uh tell me a story about an expert in chemistry it will often come up with the same names. And essentially, the researchers here were trying to use this as a way to find AI-generated content on the internet. And it proved to be, I would say, rather successful. You can't rely on it solely as a way to determine whether something is AI generated or not. But because Elena Vasquez and Marcus Chen are so closely associated with a specific Claude version, if you see those names together, that is a pretty strong signal that you're looking at something that was AI generated.

Read the full transcript

29:47Yeah. So there's this phenomenon, as you say, just to sum it up, where LLMs keep spitting out the same names for some reason, right? You know, the article also mentions if you ask something to do with software development or something like that, Marcus Chen will come up sometimes that way. So this phenomenon is noticed where these LLMs keep using these same names for whatever reason. These researchers then took those names, searched for them across academia, and found, whoa, there is a ton of academic publishing which is using those names, which would, I think, strongly indicate that they were AI-generated papers or AI-assisted papers in some way.

30:31Obviously, apologies to somebody who's actually called Marcus Chen and you work in academia. Sucks to be you. Sorry. But what were they finding that these names were appearing in what exactly? All sorts of academia? What's the sort of scale we're talking about? So there is a repository that allows you to create DOIs. DOIs are the ID numbers for academic publishing. And there's a website, for example, called Zenodo, where anyone can go in and submit a paper and automatically, like without any human verification, generate a DOI. And this is in order for you to then submit the paper to an actual journal.

31:25But they searched this database and found 1 ,655 ghost-authored records, as they call them. So it's like papers that include one of these names or combinations of these names indicating that they were AI-generated. I should clarify something. The paper determines that it is much easier to identify AI-generated names if they come in specific groups. It is statistically more likely that the paper is AI-generated if it includes Elena Vasquez and Marcus Chen together as opposed to one of them alone. So they use this to kind of see that there's almost 2 ,000 AI-generated papers on this database, which is maybe not as catastrophic as it sounds.

32:23Because like I said, it's an open access thing. I can go there and make up a bullshit paper. You think it'd be more, potentially. Yeah. Yeah. But the thing is that there are other services that scan Zenodo and kind of pull out all the DOIs and pull them into other databases like Google Scholar, if you've ever used one of those. Like Google Scholar is, if you search for an academic paper, you're probably going to end up there. You can click on the name of the author to see other papers they've written. ResearchGate is another similar service like this. so they're going to this free service but that is feeding other databases so it's kind of like infecting the internet from there on yeah and that's where the contamination like analogy yeah i should also say like the the the researcher joke that it's like they made the extremely scientific experiment of like googling the names uh which then turned up a bunch of air-generated books on Amazon and Google and so on.

33:26Yeah. Yeah. So it is spreading. And I guess in academic study, it's like in this case, well, we just focus on this database because those are the parameters of the study, obviously. And it's like, yeah, you could also do the same thing on Google Scholar and you could search that. But this is like a really, really interesting baseline of where these names are appearing and how frequently. Sam, this is obviously very similar to a story of yours I think from a few months ago. Could you just remind us, what was that? And what was this character, persona, AI-generated thing that kept coming up there?

34:07Emanuel:Yeah, so people started noticing that if you asked an LLM like ChatGPT to tell you a story without giving any kind of details about what kind of story, a lot of the time it would tell a story about a clockmaker, a lighthouse keeper, or a librarian. And a lot of the time it was a lighthouse keeper. So it would be kind of like a, it was like a mystery or like a cozy story about these specific tropes from literature. And a lot of the time they were naming, the LLMs were naming the character in these stories, Elias Thorne specifically, which was weird. It was like Elias Thorne is all of these different characters but doesn't actually exist really strongly in mainstream media, I guess.

35:03Emanuel:So people were noticing this. I got tipped off to this by a software engineer named Daniel May. And then researchers were writing about this already and found that it was all coming down to model collapse, basically. So all the models are trained off of these open models, specifically OpenAI's first chat GPT model, which was GPT 3.5, is kind of the root of this family tree. And it was used to make wild chat, which is a training set that then was used to make other training sets. And I mean, they were all using kind of these same tropes, but also these specific stories were very safe for work. And when the models and people making LMs were trying to create kind of like safety guardrails and weights that would make safe stories when anybody asks for a story, a lot of the times it would just default to these specific types of stories about clockmakers and lighthouse keepers.

36:06Emanuel:So, I mean, it's not even that like there's a ton of stories about Elias Thorne in literature. It's that the models landed on this as a very safe and sanitized output that they could repeat over and over. and then they did. And then people started making like AI-generated Elias Thorne Lighthouse Keeper mystery books for Amazon and selling them. And so it kind of like, it perpetuated it more and more as like it would kind of like deepen that connection to, oh, this is a story that people are interested in. So I'm going to repeat it again and then. Yeah, it was, it's kind of like a weird, there's a ton of them on YouTube, which is kind of odd.

36:49Emanuel:You know, if you search Elias Thorne on YouTube, that you find a lot of like AI generated dramas about. It's always an old man, but yeah, weird stuff. So Emmanuel, Sam there mentioned model collapse and sort of the origins of that there when it came to producing this character. These researchers that you looked at with this academic stuff and these sort of new characters like Chen and the grouping of those together, do they think it's the same reason? Like it's model collapse? Or have we learned anything else about why the LLMs are doing this from the study you looked at? Yeah, I mean, it's to begin with, it is, again, statistical, right?

37:42So it's going to end up landing on the same names because it's all coming out of the same model. And then they did note that there is like a reinforcing cycle where it's more likely to produce these names. Those names end up on the internet. And then presumably the LLMs are fed additional training data from the internet. So it will kind of repeat and only become stronger. I think he also said that that ultimately might complicate the goal of using this as a way to detect LLMs because the models will correct themselves, right? It's like presumably OpenAI will go in there and be like, okay, well, don't keep coming up with Marcus Chen as a name.

38:42But I think we need to live with these models a little bit more in the wild before we see if that's the case. Yeah, it might be like right now there is a window of time where the LLMs are doing that and it acts as a convenient and useful indicator of a potentially AI generated article. but in the not so far future that avenue may not exist anymore because either potentially in the training data itself as it maybe gets more varied but also just as you say from direct intervention by the AI companies who are making this which seems more likely. It doesn't look good for them if it's constantly churning out the same names.

39:23You know? Yeah. All right. We will leave that there. If you're listening to the free version of the podcast I'll now play us out. But if you are a paying 404 Media subscriber, we're going to talk about a bunch of recent immigration and customs enforcement purchases, including those Boston Dynamic robot dogs. Remember the cute ones that do all the cool dances and stuff? Yeah, they work for ICE now. You can subscribe and gain access to that content at 404media.co. As a reminder, 404 Media is journalist-founded and supported by subscribers. If you do wish to subscribe to 404 Media and directly support our work, please go to 404media.co.

40:06You'll get unlimited access to our articles and an ad-free version of this podcast. You'll also get to listen to the subscribers only section where we talk about a bonus story each week. This podcast is produced by Alyssa Midcalf. Another way to support us is by leaving a five-star rating and review for the podcast. That stuff really, really does help us out. This has been 404 Media. We'll see you again next week. Thank you.

From the publisher

Take control of your data footprint risk-free with a 30-day money-back guarantee. Go to https://incogni.com/404Mediaand use code 404MEDIA for 60% off an annual plan. That's code 404MEDIA at https://incogni.com/404Media

We start this week with Emanuel’s follow-up to his Amazon book scanning story, in which he spoke to someone who worked in the Amazon warehouse which destroys books to train Amazon’s AI products. After the break, Emanuel and Sam tell us about the same few names appearing in LLM output over and over again. In the subscribers-only section, Joseph tells us why ICE is buying loads of data about ‘voter fraud’.

You're Invited: 404 Media's Third Anniversary Live Podcast and Party!

Inside the Warehouse Where Amazon Scans and Destroys Books for AI Training

The AI ‘Ghosts’ Contaminating Academic Publishing

ICE Plans to Spend Millions on Boston Dynamics Dog Robots

ICE Wants the Country’s Voter Data

Youtube Version: https://youtu.be/PcNdLnLF7mo

Subscribe at 404media.co
Learn more about your ad choices. Visit megaphone.fm/adchoices

More from The 404 Media Podcast

All 164 episodes
We Spoke to an Amazon Worker Destroying Books for AIThe 404 Media Podcast · 41 min
Listen in VO