In short
AI Today: Episode Summary - AI Companies Take Note
Episode Overview In this episode of AI Today, the host discusses the implications of Anthropic's historic $1.5 billion copyright settlement with writers, highlighting the immediate changes in the AI industry and compliance strategies that may become central to AI companies moving forward.
Key Points
- Anthropic Settlement: Anthropic has settled a lawsuit with writers, compensating them for copyright infringement related to their training data sources.
- Industry Reactions: The reaction is mixed; while some see it as a positive step for the industry, others, particularly writers, have expressed dissatisfaction with the settlement terms.
- Significance of the Settlement: This is noted to be the largest payout in U.S. copyright law history, with approximately 500,000 writers eligible for about $3,000 each.
Detailed Breakdown
- The Settlement Context
- Settlement Amount: $1.5 billion, aimed at compensating writers whose works were included in AI training data.
- Legal Precedents: A federal judge ruled that using copyrighted material for AI training could be considered transformative, falling under fair use provisions.
- Controversy Surrounding the Settlement
- Criticism from Writers: Many writers feel that the settlement does not provide adequate compensation or ongoing rights for the use of their work in AI training.
- Industry Perspective: Some industry leaders view the settlement positively, believing it sets a precedent for future AI training practices.
- Training Practices of AI Companies
- Data Sources: Anthropic reportedly used a combination of licensed books and pirated sources (shadow libraries) for training their AI models.
- Legal Justifications: The court's ruling indicated that purchasing a book and using it to gain knowledge for creating AI outputs is permissible.
- Future Implications for AI Companies
- Compliance Strategies: The settlement is likely to shift compliance strategies within AI companies, as they will need to ensure that their data sourcing practices adhere to legal standards.
- Ongoing Lawsuits: The episode highlights that many other companies face similar lawsuits, indicating a larger trend in the AI industry's legal landscape.
- The Bigger Picture
- Precedents for Future Cases: The decision is seen as setting a key precedent for other ongoing litigation involving copyright law and AI training.
- Challenges in Compensation: The episode raises questions about how to adequately compensate original content creators in a landscape where AI models are increasingly using vast amounts of data.
Conclusion The episode wraps up with reflections on the broader implications of the Anthropic settlement for the AI industry, emphasizing that while it may be seen as a win for tech companies, the sentiments among creators suggest a need for ongoing dialogue and reform in copyright and compensation practices for AI-generated content.
Additional Resources
- AI Box: The host encourages listeners to check out their startup AI Box, which offers access to multiple AI models for a monthly fee.
- AI Community Engagement: The episode invites listeners to explore the AI Hustle community for further discussion and networking.
---
This summary captures the essential discussions from the AI Today podcast episode, emphasizing the complex interplay between copyright law, AI development, and the rights of content creators.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Anthropik has just reached a historic$1.5 billion settlement with writers for copyright. Now, this is basically a lot of people are excited about this. And there's also a lot of people that are not excited about this. Personally, I'll fall in the camp that is happy with the direction of the settlement and basically the verdict of it. I'll break down both sides. But there's definitely people that are unhappy. If you go over to TechCrunch, they have a whole article that says, screw the money. Anthropics$1.5 billion copyright settlement sucks for writers. So I'll put it out there. Not everyone is happy about this.
0:33But for AI companies, this is, you know, positive for the industry. So let's dive into all of it. Before we do, I want to say if you want to try any of these models I talk about on the show, including all the cloud models from Anthropik, I'd love for you to go check out my own startup, which is AIbox.ai. We have the top 40 different AI models all in one place. So you pay 20 bucks a month and you get access to 40 different AI models. So you don't have to have subscriptions to everything. But in addition, we just launched our AI app builder, which is basically a little box like ChatGPT and you type in a tool that you want to create and it will chain together multiple AI models.
1:07You will put prompts in and it will basically build you a tool. You can go and customize it. We're really excited about this and this is what we've really been working on for the last two and a half years. So if you want to go try out the no code AI app builder on AI Box, there's a link in the description and I would love to hear your thoughts as we're actively fixing things, adding things and it's an exciting time for us over here. All right, let's get into the episode. So basically, if you go over to TechCrunch, like I mentioned, they're not super excited about this. But how this is rolling out is that about 500 ,000 writers are going to be eligible for about$3 ,000 in this$1.5 billion settlement.
1:46A group of writers brought this lawsuit against Anthropic. It's kind of interesting because it's not just that Anthropic trained off of their data. And that's kind of what a lot of people are complaining about with this lawsuit is how that was, how that is kind of the shakeout on that. But basically, this is the largest payout in the history of US copyright law. And it is, I think, really exciting. So for me, anyways, but some people do not think this is a win for authors. It is just a win for it is by TechCrunch. It is quote, yet another win for tech companies. Everyone is trying to get as much data as you possibly can to train the models.
2:22I think we all know this. Everyone basically scraped the internet at the very beginning. OpenAI scraped the whole internet at the beginning. Everyone did. And then people complained about that. Oh, you scraped the blog post. So that was kind of like a thing. What actually ended up happening is these AI model companies ran out of data. They wanted more data. They ran out, like they scraped the internet. And so a really interesting untapped source was books, right? Because books a lot of times are not actually online. Google has kind of their Google Scholar, I believe, a kind of project that has like photocopied books and put all the pages on.
2:52I don't think they allow people access to basically use that. And there's not all books on there. There's a ton that are not. So what Anthropic ended up doing was, and this is what got them in trouble, is they went to a bunch of pirated sources. They're called, quote unquote, shadow libraries. So there's millions of books in there and they're pirated sources. But it's not just a photocopy. People have uploaded the full book. There's the transcripts you copy and paste, right so it's very easy data for these ai models to ingest and it's virtually impossible to have gotten that that data set as fast as it did so they were able to grab all of these pirated libraries throw them into the model and claude got way better i think this is one of the reasons why even compared to open ai from the early days claude has always had a much better tone and how it talks and writes it's been way better for writing basically it's kind of ironic but all the writers i know use claude because like yeah the tone's way better and that's because they grabbed a copy, a pirated copy of every single book.
3:45Now, I think they kind of knew they were in hot water with this. They were in trouble. Maybe they're trying to cover their tracks or cover their back. And what they ended up doing was going and buying one of like every book in the world, like something crazy, right? And when you have billions of dollars, it's just a cost of doing business. Then they basically had a robot that would take each of these books, would flip through the pages, scan the pages, and then transcribe the pages. And then like basically take a picture and they can read the picture and then include that into the model data training.
4:12They did both of those things. And when the lawsuit came, they got in trouble for obviously the millions of books on a pirated library. And this is actually what the billion dollar settlement is coming from. Now, the judge actually ruled in this case that what they were doing where they were scanning all of the books and uploading them. And basically they purchased the book and then they included it in the data set, they said that is allowed because it's the same as if a person goes and buys the book and reads the book, has the knowledge, and they go write like some sort of paper or some sort of essay on it.
4:44And that's monetized, like you're allowed to do that because you gained knowledge. And so this is kind of like what they did, they paid for the book, and they gained knowledge. So you're not allowed to use pirated books. But if you buy the book, you can include it in your data set. And so a lot of people are upset because they're like, you know, those authors should have reoccurring compensation forever. If you want if they want to be included, they should be able to be pulled in and out. I think the cat's out of the bag. It's kind of too late. Honestly, with the shadow libraries, it's too late now anyways, because basically if you have the pirated copies in the model, you could just use the old model to train a new model.
5:15And so even if you're like, okay, we're not using the old model anymore, the data is already in there, the tone's already in there. It's kind of too late at this point. And so it's now just like, what's their fine? So the fine was$1.5 billion. What's interesting is there is dozens of lawsuits filed against companies like meta google open ai and mid journey over basically all the legalities of training ai on copyrighted work so this isn't the first i don't think this is going to be the last copyright one that we see come out i think anthropic is going to come out ahead for this and i think a lot of people are happy with the precedent because now they know like the right way they can do this i think everyone's kind of solved this problem at this point but it's nice to know that for them anyways that the way they've solved it is something that they can continue to do into the future and they're not going to get in trouble for so right now writers are basically getting a settlement if their work was included in all of the in all of the pirated stuff so what's interesting is anthropic actually just raised 13 billion dollars i did a podcast on that if you're interested on the dynamics of that but they've just raised 13 billion dollars so paying out 1.5 is not going to kill them right their last raise was 3.5 and i imagine if they had to pay out 1.5 billion of 3.5 just for that one lawsuit that would be hurting them quite a lot i think with this fresh round of funding they can move forward and they'll be fine but all this happened because in June, federal judge William Alsup sided with Anthropic and ruled that it is legal to train AI on copyrighted material.
6:37He argued that this use case is transformative enough to be protected by the fair use doctrine that is carved out of copyright law that was set back in 1976. So he said, quote, like any reader aspiring to be a writer, Anthropics LLMs train on open work upon works, not to race ahead and replicate or supplement them, but to turn a hard corner and create something different. The piracy obviously was a completely different problem. And that's why he let the case go to trial was because of that. And this is what Anthropic said about this whole thing. They said, quote, today's settlement, if approved, will resolve the plaintiff's remaining legacy claims.
7:12This is Aparna Sridhar, who is their deputy general counsel at Anthropic. And then they also said, we remain committed to developing safe AI systems that help people and organizations extend their capabilities, advance scientific discovery, and solve complex problems. So there are tons more cases that are currently being litigated right now between AI and copyrighted works. But because of this, I think this, you know, BARTs versus Anthropic is going to be basically a precedent in all of these things. And so I think that there's some people think that because of the ramifications of it, maybe judges are going to arrive at a different conclusion but i think basically this precedent is going to hold and we're going to start to see that a lot of these cases i won't move forward if you're using pirated stuff obviously you're going to get in trouble but if you purchased whatever the original work was and like you kind of think of like music generators which i i you know this would be funny and they'll probably all get slaps on the wrist too but like you'd have to go buy one copy of like every top song ever recorded for the last hundred years and then i don't know whatever a billion dollars you spend on that, you know, all the music in the world, then you can feed that into your AI model.
8:18So to make music. So I think basically we have a precedent for how it should go. In my opinion, I think this might be the way you have to go because one of the big problems that like Adobe tried to solve with image generation was they were like, they're like, look, we'll pay people if you include your images in our data set, but they pay the original photographers for Adobe Firefly images. But it's impossible to know, like when I say, you know, generate a picture of a green plant on a stand with a flag in the background, like what data was used to create that image, like what was needed. So it's not like you could do it like Spotify, where if you listen to a song, they get, you know, you, they just listened to that song.
8:54So now you give them a couple cents, you give them a penny in streaming revenue. It's impossible to know like what the original source was. So basically Adobe did it where they just took in a huge data set of images and they're like, look, you know, we took in a million images. So let's say there's a million photographers each put in one image they all get like one one millionth of the of the royalty or revenue or whatever and so it's it's basically like the same thing with a lot of these where it's impossible to i think it's i don't think it's realistic to set up systems where like once you're included in a data set now all of a sudden you can use that model to spit out more outputs that other models can use to train on it it's just really it's it's lost so i think it's impossible to track everyone's copyrighted data forever.
9:36And we probably should just move forward if we all agree that these AI models are more useful for us than harmful. Let's just move forward. And yeah, that's my opinion. But I know everyone has different opinions on this. In any case, thank you so much for tuning into the podcast today. Make sure to go check out AI Box. There is an amazing new no-code AI app builder that we just integrated and launched. And I'm super excited about it. I'd love to hear your thoughts on it. Thank you so much. And hope you have a fantastic rest of your day.
From the publisher
Every AI company is now on alert after Anthropic’s deal. We discuss the immediate changes happening behind closed doors. Compliance strategies may now become a central business focus.
Get the top 40+ AI Models for $20 at AI Box: https://aibox.ai
AI Chat YouTube Channel: https://www.youtube.com/@JaedenSchafer
Join my AI Hustle Community: https://www.skool.com/aihustle
To recommend a guest email: guests(@)podcaststudio.com

