In short
Post Reports Episode Summary: The Quest to ‘Destructively Scan’ All the World’s Books
Episode Overview
- Podcast Title: Post Reports
- Hosts: Martine Powers and Elahe Izadi
- Episode Title: The quest to ‘destructively scan’ all the world’s books
- Release Date: January 29, 2024
- Description: This episode discusses Anthropic’s Project Panama, an ambitious and controversial effort to destructively scan millions of books to train AI models, raising significant ethical and legal questions.
Key Themes and Concepts
- Project Panama
- Objective: To destructively scan all the world's books, converting them into digital formats for AI training.
- Methodology:
- Purchase books in bulk from used book warehouses.
- Physically slice off spines to facilitate rapid scanning of pages.
- Resulting digital library intended to enhance the capabilities of the chatbot Claude, developed by Anthropic.
- Legal and Ethical Implications
- Copyright Lawsuit:
- Anthropic faced a class-action lawsuit from authors alleging copyright violation.
- The case was settled for $1.5 billion, surfacing more details about Project Panama through unsealed court documents.
- Fair Use Argument:
- A judge determined that scanning the books for training AI models could be considered fair use, as it was a transformative use, differing from the original works.
- Comparison to Other AI Companies
- Industry Practices:
- Other tech giants like Meta, Google, and OpenAI have also faced similar lawsuits for copyright violations.
- Allegations against these companies include using unauthorized copies from shadow libraries (pirated content).
- Cultural Significance of Books
- Value of Books:
- Books are viewed as high-quality content, curated and edited, unlike much of the lower-quality material available online.
- The belief that understanding and utilizing books can significantly enhance AI's knowledge and abilities.
Key Arguments and Discussions
- Destructive Scanning Debate:
- The physical destruction of books for scanning raises ethical concerns about preserving literature and the rights of authors.
- Legal Precedents:
- Current legal frameworks struggle to address the complexities introduced by AI and mass data acquisition, leading to varied judicial interpretations in copyright cases.
- Impact on Authors and Creatives:
- The ongoing legal battles could set precedents that affect how creators are compensated and how their works are utilized by AI companies.
Conclusion This episode of Post Reports highlights the tension between technological advancement in AI and the ethical considerations surrounding intellectual property rights. It raises important questions about the future of literature, how knowledge is consumed and transformed in the digital age, and the implications for authors as AI becomes an increasingly dominant force in content creation and distribution.
Additional Notes
- Host and Guest: Martine Powers interviews technology reporter Will Oremus, who provides insights into Project Panama's implications and the broader context of AI and copyright law.
- Production Credits: Produced by Rennie Svirnovskiy, edited by Dennis Funk, mixed by Sam Bair.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnveiling Project Panama
0:45 to 2:25
Exploring Anthropic's controversial initiative to destructively scan books.
“It thought if it could get all the pros from all the books in the world, that it would be able to build the best AI chatbot of all.”
The Copyright Lawsuit
2:25 to 4:05
Discussion of the legal ramifications of Anthropic's actions.
“Today, Will tells the story of how one Silicon Valley company built its AI by buying, scanning and destroying millions of books.”
Destructive Scanning Explained
4:05 to 6:00
Insights into how Anthropic scanned books and the ethical implications.
“And just quickly, can you tell me about the lawsuit against Anthropic that these court filings came out of?”
Anthropic's Methodology
6:00 to 7:45
Details on the approach Anthropic took in acquiring books for scanning.
“So tell me a little bit more about Project Panama.”
Industry Reactions and Concerns
7:45 to 10:00
Exploration of the tech industry's take on Anthropic's project and its implications.
“So they're just basically like ripping these books apart.”
Legality of Scanning Practices
10:00 to 12:20
Analysis of the legality of Anthropic's scanning practices versus traditional ownership rights.
“They would say, look, in a way we are saving this.”
The Fair Use Debate in AI and Books
14:00 to 19:00
Explore the nuances of copyright law and fair use in AI models that utilize books.
“Like if I buy a book, I can do what I want with it.”
The Value of Books in AI Training
19:00 to 20:30
Understand why AI companies prioritize digitized books despite abundant online content.
“After the break, why these AI companies value books so highly, and how AI lawsuits might impact artists, writers, and other creatives in the future.”
Tech Companies' Legal Battles Over AI Training
20:40 to 27:40
Discover the ramifications of ongoing lawsuits against tech companies regarding copyright violations.
“Like, there's a lot of written word on the Internet that is, you know, free to access and easier to access and they could just suck up.”
Transcript
Automatic transcript. May contain errors.0:01So you might have heard of the AI startup called Anthropic. This is the company behind the AI chatbot, Clod. Well, in 2024, executives at Anthropic ramped up an ambitious project that they'd hoped to keep quiet. It was codenamed Project Panama, and internal planning documents recently unsealed in legal filings described it as their, quote, effort to destructively scan all the books in the world. Project Panama was Anthropics' ambitious project to buy as many books as it possibly could. It would take them to a scanning center. It would slice off the spines, scan every page one by one, feed it into this digital library.
0:47It thought if it could get all the pros from all the books in the world, that it would be able to build the best AI chatbot of all. This week, technology reporter Will Oremus first reported on Project Panama's details exposed in these legal filings. It was really evocative that this company was literally destroying hundreds of thousands, maybe millions of books. Because that's sort of what creatives are worried about, right? Like the destruction and the aggregation of their work into these gargantuan AI systems that are going to hoover up all the knowledge in the world. It was just sort of like a physical manifestation of that concern.
1:28Initial details about Anthropik's hunger for books emerged in documents filed in a copyright lawsuit. That lawsuit was brought by book authors against the AI startup in 2023. It was settled in August for more than a billion dollars. But documents from the case also show something else. that the company may have crossed legal lines while pursuing all the world's pros. Even though it sounds really bad to try to destructively scan all the books in the world, this was actually the company's effort to do this in a more legal or more ethical way than everybody else was doing it. The alternative was to download them from these really shady pirate sites.
2:11Which Anthropic had also done and is what ultimately landed them and other AI companies in legal trouble. From the newsroom of The Washington Post, this is Post Reports. I'm Martine Powers. It's Thursday, January 29th. Today, Will tells the story of how one Silicon Valley company built its AI by buying, scanning and destroying millions of books.
2:47Will, thank you so much for joining us. Thanks for having me. All right. So tell me how you first heard about Project Panama. The name Project Panama actually had not been reported before. Our amazing researcher, Aaron Schaefer, keeps all these alerts out for court cases that he's tracking and that various teams are tracking. And one of those sets of cases were these AI copyright cases against these big tech and AI firms. And by the way, it's not just Anthropic. It's pretty much all the tech giants, or many of them, getting sued with allegations of violating people's copyright in various ways.
3:23But he noticed that this case that had already settled, and we reported on it when it settled, Anthropic settled with the authors for$1.5 billion in August, had some new alerts on it. And so he went and he was like, oh, what's that? Why would there be new files in a case that's already settled? And it turned out that they were files that were part of the evidence in the case, but now they were significantly less redacted. So a bunch of stuff that was blacked out before was no longer blacked out. And part of that were the details of Project Panama. We knew some of the vague outlines of it. We didn't know what it was called or everything that was involved in it.
3:56And by the way, when we reached out to Anthropic for comment, they emphasized that the case has settled, that the judge in the case found that much of what they were doing was, in fact, legal. And just quickly, can you tell me about the lawsuit against Anthropic that these court filings came out of? So these were court filings in a lawsuit that was filed against Anthropic in 2023. And it was part of this big wave of lawsuits by book authors suing pretty much all the tech giants, alleging that they had violated their copyright. It wasn't just book authors either. It was online publishers. It was news outlets.
4:35It was videographers, writers of screenplays, photographers, visual artists. As they started to realize that their life's work had been vacuumed up by tech giants and fed into ChatGPT and other AI models like that, they were like, hey, nobody asked us. Like, we didn't say you could do this. Is that legal? And we're still in the process of finding out the answer to that question. And when you were looking through these files, what stuck out to you here? Like, why did this kind of part about destructive scanning catch your attention? Well, first of all, that was just a really evocative quote, right?
5:16Like, destructively scan all the books in the world sounds like something that a James Bond villain would be trying to do. It's interesting because Anthropic is actually, they sort of have this ethos of wanting to be seen as the good guys in the AI industry. This company was founded by people who split off of OpenAI, the maker of ChatGPT, because they were worried that OpenAI was straying from its original mission of saving humanity from runaway AI. And here they were, they're not literally destroying all the books in the world. I think they only needed one copy of each, not all the copies. But it was just a really evocative phrasing.
5:52And the name Project Panama, we still don't know why they called it that, and they didn't say when I asked. But we just wanted to dig in and figure out what this was and why they were doing it. So tell me a little bit more about Project Panama. Like, how did it work? It might help to start with who they hired at the outset of this project. They hired a guy named Tom Turvey. Now, most people don't know that name. But if you are familiar with the history of scanning books, you know that name. because when he was at Google, he helped oversee the massive project called Google Books, where Google Books wanted to scan all the books in all the libraries and digitize them and make them available online.
6:36So Anthropic brought this guy on to help lead Project Panama. And what he did was not to just go out to publishers and authors, which is what I think the plaintiffs in the case would have preferred. There is some evidence from his deposition that he did a little bit of outreach to publishers and authors, But he quickly determined that this wasn't a viable strategy. It just wasn't practical to pay to license books on such a massive scale. So instead, he settled on a different approach that included going to these massive used book warehouses with names like Better World Books, where you could buy hundreds of thousands of books at a time in bulk for the cheapest possible price.
7:17and then the cheapest and fastest way to scan them turns out to be, you know, it's kind of hard if you ever tried to spread out a book on a, you know, a Xerox machine, right? Yeah, yeah. It's like the part of them in the kind of near the spine always gets cut off and you're just trying to like press down on the book as hard as possible. You miss some words. And it's also just really slow. So instead they did this practice that they invented, but it's a practice that's out there where you slice off the spine so that you don't have to worry about that anymore and you can just quickly feed all the pages in to a machine that will scan them all.
7:47So they're just basically like ripping these books apart. Yeah, I mean, it was actually very neat. It was more like a paper shredder than a kid carrying up pieces of paper. But yes, they were destroying the books. And then there was a proposal from a vendor in the court files that indicates they recycled all the materials afterwards. Because I guess who wants to store giant warehouses full of destroyed books once you've scanned them all? And the result was that they had the books in digital form And now they had this massive, massive digital library. They didn't get all the books in the world, but they got a lot.
8:21And now they could draw from that library to train their AI models. And we should say that Better World Books did not respond to our request for comment. How many are we talking about in terms of the number of books that they were able to scan? So the exact number that they scanned is still redacted in the files. And keep in mind, these filings are from a year or two ago, so they wouldn't have the updated numbers anyway. But there's evidence that they had at least one project proposal to acquire hundreds of thousands of books from one outlet. And a judge last year said they spent many millions of dollars to purchase millions of print books.
8:57And Anthropic didn't just go to these massive used book warehouses to acquire books. They were trying everything. There's evidence in the files that they approached The Strand, you know, the famous bookstore in New York that promises 18 miles of books. A spokesperson at The Strand said that didn't end up happening. They also thought about approaching libraries to non-destructively, in this case, to scan their books, more like the Google Books project where they keep the spines intact. But they were looking all over at all kinds of different sources, whatever they could get their hands on. Wow.
9:29What you're describing here almost strikes me as like a reverse Noah's Ark for books. Like you take one of each book and you take it to this warehouse to be scanned, but instead of saving the book, it's like the book is sort of sacrificed to the recycling gods in the pursuit of AI. I mean, it's just pretty bizarre. I love that metaphor. You know, there are people in the tech industry who would say that it's actually more apt than we're giving it credit for. They would say, look, in a way we are saving this. I mean, the people who run Anthropic and a lot of the people in the AI world seem to truly believe that AI is going to be like everything, you know, in the future.
10:12You know, in a few decades, AI will be doing all kinds of intellectual work. It'll be what we turn to for everything. And so in their minds, I'm just hypothesizing here to be clear. We don't have evidence of this in the court documents. But, you know, I think they might say we are saving these books. We're making sure that the content of these books, many of which, by the way, are probably, you know, somewhat obscure. Obviously, they're not super rare and valuable books or they wouldn't be in these bulk warehouses for the most part. They might say we are saving these books. We're saving it, you know, for a world where AI does everything.
10:45Now we made sure that the AI will know what was written over all these centuries of human authorship. Hmm. Well, I want to talk more about why this came out in court and why people had concerns about this. I mean, obviously, I too could go to a bookstore, probably a used bookstore because I don't have enough money to buy all these books new. But I could go to a bookstore, buy a bunch of books, like do what I want with them and throw them out later. Why is this, I mean, I guess, is this illegal? And why do people have concerns about what happened here? So the potentially illegal part was actually what they were doing before this used book buying.
11:25One of the things that was really interesting to me about this story is that as bad as it sounds to try to go out there and destructively scan all the books in the world, this was actually seen by a lot of people in the industry as a more legal or more ethical way of doing it than what a lot of the other companies were doing at the time. Authors have also sued Meta, which declined to comment for our story. And by the way, there are pending lawsuits still against OpenAI and Microsoft. There's a lawsuit against Google. They're all making broadly similar claims that these companies violated authors' and publishers' copyright.
12:05What some of these companies were apparently doing, and what Anthropic apparently did before Project Panama, was going to these vast repositories online that are known as shadow libraries. And these are unauthorized copies of millions of works that have been put together into these giant data sets. And they're traded around for free on the internet via the software called Torrent Software. And there is evidence that Anthropic torrented the entirety of two huge shadow libraries of books and other copyrighted works. There is evidence that Meta did the same thing. There are allegations that other tech companies did this too.
12:49And that's seen by a lot of people as even worse because they're, A, nobody's getting paid, right? B, a lot of these sites are in legal trouble. They're like, you know, the FBI is not happy about their existence and is investigating them for all sorts of things. And sometimes they go dark because they've gotten busted by some authorities or another. And then there's a third thing involved, which is that when you use this torrent software, you sometimes end up making the pirated copies available for other pirates to take for free. And so that's what Meta is still accused of doing. They're accused of while they were trying to download all those books secretly and for free from the pirate sites, they were also making stuff available for other pirates, which might be a more straightforward copyright violation, at least that's what's alleged, than using them to train an AI.
13:38And then in terms of the buying of the books and scanning the books, was there anything illegal about that? Like, how is that different from me just buying a book and scanning a book, obviously on a probably smaller scale? Right. So with the caveat that I'm not an expert in copyright law, I want to amend something you just said, which is, yeah, we have this intuition. Like if I buy a book, I can do what I want with it. Right. But that's not quite true. You actually can't buy a book, copy it, make your own version, and then sell that. That would be pretty clear cut violation of copyright law. You can't buy a book and then digitize it and then sell it online or make it available for free online.
14:22What the judge said in this case was that Anthropic was actually doing something different. They were taking the books and transforming them into something else. They're transforming them into these AI models, including Anthropik's popular AI chatbot, which is called Claude. And Claude is a fundamentally different product from a book. And so they're not competing directly with the books that they're acquiring. The judge found that this falls under a doctrine in copyright law known as fair use. This is where you can make use of copyrighted works without permission if you're using them for certain purposes.
15:00Often you can do it for teaching purposes, right, in schools. One of those purposes is you can copy somebody else's stuff without permission if you're then transforming it into some other innovative thing that isn't the same as the thing you copied and doesn't compete directly with the thing you copied. So the argument here from these AI companies is like, yes, we wanted to train our AI models, but what we're putting out isn't just copying these books, that it's a new thing, that it's this innovative thing that is not a copyright infringement, but instead like a thing onto its own. Yeah, exactly.
15:37So Anthropic's not out there saying, hey, like, you know, you want to buy a John Grisham book? We've got him here for 10 cents apiece, right? And in fact, in the court cases, there is some of the evidence turns out to be like, can you get Claude to spit out a whole verbatim copy of one of these books that the company acquired? And so far, the answer is usually no. the chatbots don't do that. And I think the companies probably try to train them not to do that. And so that's one of the factors that the judge is weighing. But I should note here that this is really unsettled law. The comparison that I'm making in my head, and maybe I will age myself with this, is Napster, right?
16:14That like in the age of Napster, where everyone was like, just downloading random stuff off the internet, you know, music and uploading it again. And a lot of it wasn't legal. And you didn't know where any of this music came from. You certainly weren't paying for it, that it's basically that was happening with books. But it sounds like the allegation here is that huge billion dollar companies were engaging in this, not just like a bunch of teenagers in their basements. Yeah, you're exactly right. I mean, Napster is a great point of comparison. Yeah, I'm old enough to remember like the stories of federal agents showing a knocking on the door to some teenager's basement where he's like downloading some Nirvana tunes or whatever.
16:51But in fact, there is a way in which this is sort of like the teenagers in their basements. I mean, there's evidence from the court filings, especially in the Meta case, that some of the engineers who were working on this project were really worried that what they were doing was illegal. There's chat logs between employees at Meta. There was one where an engineer who was working on downloading the shadow libraries said, quote, torrenting from a corporate laptop doesn't feel right. And then later on, he shared a concern with the company's legal team that using these torrent sites could mean uploading pirated works to other people.
17:27And he said that it, quote, could be legally not okay. And then you started to get the sense that they were told, don't worry, this has all been approved. There was a quote from the filings that is going to refer to MZ. That's apparently Mark Zuckerberg, the CEO of Meta. Quote, after a prior escalation to MZ, Gen AI has been approved to use LibGen for Llama 3 with a number of agreed-upon mitigations. So that's a lot of jargon there that's internal to the company, but basically it seems to be saying, yeah, go ahead. Go ahead and do this stuff. You know, I know you're worried that it's illegal, but it's been approved.
18:03We're going to do it. But then you've also got people who defend this type of file sharing, people who say that information wants to be free. One of these famous shadow libraries, it's called LibGen, which is short for Library Genesis, actually arose out of this culture in Russia when there was censorship and Russian academics couldn't get access to published research from around the world. They started finding ways to attain academic literature and putting it together. And it was this big cooperative underground project to be able to share with each other the accumulated knowledge of their academic peers around the world so that they could advance research.
18:46And this was seen by a lot of people as a really noble thing, trying to break down the walls to the collected knowledge that humanity has put out.
19:00After the break, why these AI companies value books so highly, and how AI lawsuits might impact artists, writers, and other creatives in the future. We'll be right back.
19:21This is the year you stop overthinking and start building. The year your side idea becomes something real. Founder, creator, business owner, it all starts with one decision. And that decision is launching with Shopify. Maybe it's a product your friends already asked to buy. A service you know you're great at or a brand that's been living rent-free in your mind. January is your window. Before another year slips by, 2026 is when you turn the idea into income, and Shopify is how you begin. Millions of entrepreneurs have already made the leap, from household names like Heinz and Mattel to first-time business owners just getting started.
19:59Choose from hundreds of beautiful templates that you can customize to match your brand. Create email and social campaigns that reach customers wherever they scroll. In 2026, stop waiting and start selling with Shopify. Sign up for your$1 per month trial and start selling today at shopify.com slash reports. Go to shopify.com slash reports. That's shopify.com slash reports. Hear your first this new year with Shopify by your side. Why have all of these AI companies been so desperate to attain this massive amount of digitized books? Like, there's a lot of written word on the Internet that is, you know, free to access and easier to access and they could just suck up.
20:47What is it about books that they were, you know, going to such extraordinary lengths, including buying physical copies from a used bookstore and scanning them very quickly one by one? Like, what is it about books that is so valuable here? Oh, don't worry. They were definitely also vacuuming up, like, the entire Internet and using that to train their models, too. They were, I mean, you know, more or less the companies were trying to get their hands on all the created works that they possibly could, whether that's a blog post, whether that's the Library of Congress, whether that's copyright filings over the decades, or whether that's books or videos, movies, screenplays, you name it, right?
21:32Books, though, were seen as particularly valuable because the internet has a lot of detritus. You know, like there's a lot of cruddy writing across the internet. There's a lot of stuff that's not true. No kidding. Right. There's stuff that's written by bots for bots. There's stuff that's fake news and propaganda. And so the quality of any given internet content is not very reliable. And you might end up with your AI model spitting out fake news or spitting out conspiracy theories or spitting out just pure nonsense because that was in its training data. And so books were seen as especially valuable because by and large, I mean, there's bad books out there too, but by and large, books are carefully curated.
22:09They're edited. Somebody went to a lot of trouble to gather facts or to create a work of fiction and to craft the prose. And so Anthropik's executives are shown in the court documents talking about how valuable books can be. You know, this is the way that even though, so Anthropik, by the way, is an underdog. in this AI race. OpenAI, Microsoft, Google, Meta, these are much, much larger companies with a lot more money. And Therapeutics thought a way that they can kind of bootstrap their way to competitiveness with these bigger companies was to focus on books and to focus on better quality training data.
22:45Wow. You know, what you're saying there, I feel like there's a little bit, like, I feel a little bit of, I don't know if it's like patriotism or pride in hearing you say that, this idea that even in this moment where, you know, it feels like nobody reads books anymore and obviously AI is going to take over all of our kind of literary habits and in the future no one will even be writing books because AI can do it for us but that in this moment still like the the crowning achievement of of human linguistics is still a book and in some cases a physical book and that that like still has a lot of value and that these companies recognize that being able to inspect and understand a book and how it works and how it's written is something that is valuable to do.
23:33Yeah, I think you're exactly right. And, you know, authors should be proud of what they've produced. I think their objection to this practice is that they're not seeing any of that value, right? The tech companies clearly value the work that these authors poured their years of their life into, in some cases, their heart and soul into, but they still didn't want to pay the authors for it, right? Like they didn't want to pay the publishers. I think what a lot of the publishers and authors would prefer is for these giant tech companies to come and say, hey, look, we want permission to use your work.
24:04We think it's really valuable and we're willing to pay you some for that, right? We're going to pay you to license your work. That's what they would have liked to see. And that's in most cases, not what happened with books. Now we have seen some types of licensing agreements with news outlets. In fact, we had to disclose in the piece that there's a content arrangement of some sort between OpenAI and the Washington Post. And in these cases, the tech companies are paying outlets in order to use their material in training. Because news is another thing that you can't just, you know, you can't just find it in a reliable way for free on the internet always.
24:40But again, that didn't happen in most cases with the books. And that's why you're seeing all these lawsuits from the book authors. I should also mention that there are other copyright lawsuits out there from photography wires, from video makers, from illustrators and visual artists. Everybody's suing the tech companies, and we just don't know yet how it's all going to shake out. Yeah. Are these lawsuits having an effect in terms of changing the expectations around whether these AI companies are paying artists, writers, authors, musicians, filmmakers in the future that because they can't get away with just like stealing people's stuff for free, that they are, you know, standardizing some sort of process going forward to make sure that people are paid for the work that's consumed by these AI models?
Read the full transcript
25:30Both the books industry and the tech industry are watching extremely closely to see how the judges rule in these cases, because the details of their decisions will shape whether the tech companies have to pay for the stuff that they train their models on, how much they have to pay, what they can get away with as fair use, and what they can't. And so absolutely, everybody's trying to see how this will shake out. In the Anthropic case in particular, the judge issued kind of a nuanced ruling. He dismissed the parts of the case where the authors were alleging that Anthropic broke the law by training its AI models on their books without permission.
26:09He also dismissed the part where they scanned all those books as part of Project Panama. He said that was probably actually okay. That was fair use. What Anthropic got in trouble for were the books that it pirated that it didn't use to train its AI models. As weird as that sounds, his decision was, as long as you're using this to train AI models, that's fair use. You're transforming the books into something else. But all those books you downloaded and then didn't use, that's just, you just created a library. You just created like your own library of books and that actually could be illegal copying.
26:47And that's what Anthropix settled with the authors to avoid going to trial over. They ended up settling for$1.5 billion. But in the Meta case, the ruling was a little different. The judge ruled that the author's lawyers, he basically was like, you guys, you missed the point here. You're like, you have to show me that this is hurting your ability to sell books. I'm sure there's a million ways you could have showed me how an AI could hurt your ability to sell books. But you guys just never did that. And so I can't find that they violated the law here. But that judge thought other plaintiffs in other cases should be able to show that it's hurting their book sales or their ability to sell a screenplay or to be a graphic artist in Hollywood or whatever it may be.
27:33So, again, a lot of this is still unsettled, and different judges are going to be arriving at different conclusions because there's never been a case exactly like this before. The idea of hoovering up all the world's knowledge and putting it into AI models is just not something that the original copyright laws had explicitly anticipated. Will, this is so fascinating. Thank you so much for explaining all this. Thanks again for having me.
28:04Will Oremus is a tech reporter for The Post. That's it for Post Reports. Thanks for listening. Today's episode was produced by Rennie Svernovsky. It was edited by Dennis Funk and mixed by Sam Baer. Thanks also to Aaron Schaefer and Tom Simonite. We would love to hear what you think about this episode. If you've got thoughts, questions, or ideas for future shows, give us a shout. Send an email or a voice memo to postreports at washpost.com. I'm Martine Powers. We'll be back tomorrow with more stories from The Washington Post.
From the publisher
In early 2024, executives at artificial intelligence start-up Anthropic ramped up an ambitious project they sought to keep quiet. It was code-named Project Panama, and internal documents filed in court described it as an “effort to destructively scan all the books in the world.”
According to the filings, the company had spent tens of millions of dollars to acquire and slice the spines off potentially millions of books, before scanning their pages to feed knowledge into the AI models behind products such as Claude, its popular chatbot. A judge ruled this fair use.
Details of Project Panama emerged in more than 4,000 pages of documents in a copyright lawsuit brought by book authors against Anthropic. The company agreed to pay $1.5 billion to settle the case in August – but a district judge’s decision last week to unseal a slew of documents in the case more fully revealed Anthropic’s zealous pursuit of books.
Today on “Post Reports,” technology reporter Will Oremus explains the lengths to which AI firms such as Anthropic, Meta, Google and OpenAI went to obtain colossal troves of data with which to “train” their software – a frantic and sometimes clandestine race to acquire the collected works of humanity.
He and host Martine Powers discuss how AI companies’ efforts sometimes might have crossed over into the illegal, and how authors and artists might fare in an AI-centered future.
Today’s show was produced by Rennie Svirnovskiy. It was edited by Dennis Funk and mixed by Sam Bair.
Subscribe to The Washington Post here.



