AI on Trial: Inside the NY Times vs. OpenAI Lawsuit with Cecilia Ziniti | E1874

4 Jan 2024 · 1 h 16 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: This Week in Startups - Episode E1874

Episode Title

AI on Trial

Inside the NY Times vs. OpenAI Lawsuit with Cecilia Ziniti

Episode Description In this episode, host Jason Calacanis interviews attorney Cecilia Ziniti about the landmark lawsuit involving The New York Times and OpenAI. The discussion centers around important themes such as fair use, data scraping, and the potential ramifications of this case.

---

Key Themes and Highlights

Introduction

  • Host: Jason Calacanis
  • Guest: Cecilia Ziniti, a lawyer with extensive experience in technology law.
  • Context: The episode explores the significant lawsuit between The New York Times and OpenAI, which has implications for the future of copyright law, particularly in regard to AI and digital content.

Overview of the Lawsuit

  • Plaintiff: The New York Times
  • Defendant: OpenAI
  • Allegations:
  • The Times claims OpenAI used a significant portion of its content to train its AI models without consent.
  • Two theories of infringement are presented:
  • Training AI on NYT articles.
  • Generating outputs based on NYT content.

Discussion on Fair Use

  • Fair Use Test: Key legal framework governing the use of copyrighted material.
  • Four Factors of Fair Use:
  • Purpose and Character of Use: Commercial vs. transformative use.
  • Nature of the Copyrighted Work: Creative works get more protection.
  • Amount and Substantiality: Amount used in relation to the whole work.
  • Effect on the Market: How the new work affects the original market.
  • Implications for AI: Current legal precedents do not explicitly address the training of AI models (e.g., LLMs) on copyrighted material.

Historical Comparisons

  • Roy Orbison vs. Two Live Crew: A landmark case reinforcing that parody and commentary can constitute fair use.
  • Google vs. Oracle: Illustrates the complexities involved in cases of copyright and technology.

Potential Outcomes

  • Settlement: Likely outcome where both parties come to a compromise.
  • Legal Ruling: The case could lead to significant legal precedents regarding AI and copyright.
  • Legislative Action: Possible new laws or amendments to existing laws governing AI usage and copyright.

Market-Based Solutions

  • Licensing Schemes: Transition towards a marketplace where content providers are compensated for their material used in AI training.
  • Subscription Models: Potential implementation of authentication systems where users access AI-generated content by verifying subscriptions (e.g., to NYT).

Conclusion

  • Future of AI and Copyright: The ongoing debate about how technology companies will navigate copyright laws as AI continues to evolve and integrate into daily life.
  • Cecilia Ziniti's Startup: The formation of her AI-focused startup for legal applications.

---

Key Takeaways

  • Legal Ambiguity: The intersection of AI technology and copyright law remains largely uncharted territory.
  • Societal Impact: The case reflects broader implications for technology, innovation, and the rights of content creators in the digital age.
  • Expectations for Resolution: A resolution to the lawsuit could shape the future landscape of AI and how it interacts with existing intellectual property laws.

---

Resources Mentioned

  • [Cecilia Ziniti's Twitter](https://twitter.com/CeciliaZin)
  • [Cecilia Ziniti's LinkedIn](https://www.linkedin.com/in/ceciliaziniti/)
  • [This Week in Startups](https://www.youtube.com/thisweekin)

---

Final Thoughts This episode provides deep insights into crucial ongoing legal discussions surrounding AI and copyright, emphasizing the need for clarity and adaptability in laws as technology advances.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00The music industry is a great example of really the market wins. And like, that's one of the points I made in the tweet. And I think is important to think about when you think about this case that I'm not a doomer in the sense of like, this isn't going to end AI. Like there's no universe where this case would end AI. And so the result is, do we end up with a licensing scheme? Like, is this Napster to iTunes, right? But to your point that this is like, it's going to be a fight and it's going to be a lot of discovery. I would predict that. This Week in Startups is brought to you by Mev. Tired of the dev shop rollercoaster?

0:35Mev is your reliable technical partner offering a well-established software development process designed to consistently deliver unparalleled value to their clients. Get$30 ,000 off your first three months at mev.com slash twist. Northwest Registered Agent. When starting your business, it's important to use a service that will actually help you. Northwest Registered Agent is that service. They'll form your company fast, give you the documents you need to open a business bank account, and even provide you with mail scanning and a business address to keep your personal privacy intact. Visit northwestregisteredagent.com slash twist to get a 60 % discount on your next LLC.

1:18And the paintbrush loan is the earliest startup financing on the internet. No pitch deck, no business plan, no minimum time in business, and no warm intros. Plus, you get to keep your equity. Visit GetPaintbrush.com to see if you qualify for a$50 ,000 startup loan in less than two minutes. All right, everybody. Welcome back to This Week in Startups. You probably heard about this major New York Times lawsuit against OpenAI, the makers of ChatGPT. this is really a groundbreaking lawsuit here i think this is going to be the most important lawsuit that we've seen in ai perhaps in technology ever and so i wrote a blog post about it some of you may have read it at my substack calicanus.substack.com one of the great things about the x platform and twitter uh formerly known as twitter is that you meet interesting new people well one of those new people i met was chichilia ziniti uh and she is an actual lawyer and she did an incredible breakdown on her Twitter while I was writing my sub stack.

2:23So I invited her to come here on This Week in Startups so that we can break down what is happening in this lawsuit. And this is an absolutely critical episode for all founders because you can get yourself in a lot of trouble if you don't follow the rules. And this is uncharted territory. I think you would agree. Welcome to the program, Cecilia. Thank you. Yeah, excited to be here. Thanks for having me. So just your bona fides, as it were, you wrote a great tweet storm, by the way, and you have a background in legal. So maybe just share with the audience who you are and why you're taking the time to comment on this issue.

2:58I'm a lawyer for tech companies, been in tech since I joined Yahoo in the early 2000s when they were still competing with Google and always been interested in the legal side. And over the years, that's taken me different places. I was at Morrison and Forrester, a big law firm, represented Apple and Apple Samsung, which is a huge case of the day. From there, I joined Amazon and they said, you have all this mobile phone experience. I thought, surely I'll be working on the Fire Phone. I get there. They're like, no, we're going to have the more experienced attorneys on that. You're going to work on this device.

3:30It doesn't really work. It's called Doppler. And that turned out to be Alexa. And it was a great career move. So I was the first lawyer on Alexa, had a great experience there. and then went on to be a GC of different tech companies. You might have heard Anki was Andreessen Horowitz. It was an early robotics company. Spent some time at Cruise. And then most recently, I was the general counsel for Applet. Oh, wow. So what an incredible career thus far. Let's get into this case, because this is a very unique case in the history, I think, of copyright. And correct me if I'm wrong, having been in content my whole career as a journalist, publisher, silicon reporter blogs at weblogs inc i've dealt with a lot of these fair use claims and i've dealt with a lot of copyright claims i've dealt with cell phone manufacturers you know emailing us oh my god you have a leak that's our copyrighted information all this stuff and so there's um a lot to unpack here but when you saw this lawsuit drop and you you started unpacking it how important is this lawsuit and what is the nature of the lawsuit for people you know who uh you know, maybe are new to this, just briefly, what is the nature of this lawsuit?

4:42What is the New York Times claiming here? Yeah, so New York Times has a content library, one of the few content holders more prolific than you, Jason, perhaps, going back to 1851, right? So they reported on literally the Civil War, right? So that amount of content, millions of articles, the allegation is that those articles were used in a couple of ways by open ai without consent so one way is training right so in the complaint new york times actually breaks down that it was a decent percentage of the articles used to train open ai i think you know in the like one or two percent something where it's actually measurable you know one random blog post that i wrote you know not going to move the needle but the entire new york times archive, you know, maybe it does.

5:30And that's the allegation. So that's one. The second theory is more on the output side. So when you go to chat GPT, and you ask for an article, they've got this exhibit, New York Times made this exhibit, exhibit J, you can look it up, it's great. But essentially, it has 100 instances of somebody putting the first paragraph of an article in, and chat GPT gives you the rest verbatim, you know, like almost, you know, one or two word changes. is. But that is kind of a different theory, and it triggers the law differently. I can get into that if it's of interest, but that's really the core of this.

6:06Yeah. And so the nature of fair use, I am very familiar with because I've had many people claim that we use their content, let's say in a blog post or in this very podcast, where we might use a short snippet of a song where I'm doing commentary on it or a clip of a news event that occurs. And so I'm pretty familiar with the four-part test, but maybe you could run our audience through the four-part test because OpenAI, I think, believes that what they're doing is fair use. And then as part of that, I don't know that training as a concept has existed in the copyright law. This idea of training something, I believe, is novel to copyright law.

6:48Am I correct in that one? That's right. There hasn't been at least an adjudicated case on training yet. There have been a lot of fair use cases that I think OpenAI and New York Times will each point to ones that go their way in technology, but there hasn't been one on training that I'm aware of that's gotten to that point yet. But in terms of the fair use test, it's a super fun one. It's four factors, as you said, but they are non-exclusive and it's very squishy. It's literally courts are directed to judges apply and not juries. Courts are directed to balance the interest and they can consider other factors.

7:26And no one factor is fancy word dispositive. No one factor decides. so it really is something where there's a lot of discretion and the optics of it and how like whether the judge wants to rule your way tends to matter so the the four-part test is not um you can take five percent it's not you can take 12 it's not your you can monetize it a little bit over here it's open for interpretation exactly and you have to as a judge when you make these decisions look at the totality of those four parts so let maybe we get into those four parts and then go into some examples yeah let's dive in i have a i actually have a slide what i should i pull that up all right awesome yeah let's do it i mean wow i love a guest with a slide deck i love it i wanted to be a law professor and then decided other things were more lucrative so this is my my law professor dying to be free but essentially um fun thing about fair use it was uh the original fair use case was in the 1800s and it was a about writings that george washington had and another biographer copied 353 pages of washington's original writing and lost it was not fair use 350 pages was too much and then that opinion um from the 1800s got codified into the copyright act so here they are let's go through the four factors so i used emoji because this is the new generation, the Zoomers will do that.

8:57But essentially, the first one is the purpose and character of the use. And this is really where all the play is in technology cases. So I've got here, I've got the emoji for theater for like, how are you using it? The emoji for a video game controller, because video game cases are actually pretty instructive here. And then the emoji for the web, right? So this is where what the court considers here is how is the infringer using it? Are they making a commentary? Are they making a joke? Are they making a parody? Are they famous case, Perfect 10 versus Google? Perfect 10 was a pornographer and said to Google, hey, your thumbnails are infringing our content because they're literal copies that people see.

9:42Google defended saying this, we're using this for a different purpose. You're not trying to be pornographically entertained when you're doing a thumbnail search. Maybe you are, but it's not a good substitute. Google won that case on this fact. So that was for Google search. Now, let's go through some of those cases. If you were doing commentary, there have been many cases where people will take a movie or there'll be a documentary film about a movie or might use movie clips. And if you're doing commentary on that, even if it's commercial, there's some leeway allowed for that. And then there's parody.

10:16So if you made a parody movie like Spaceballs is the famous Mel Brooks parody of Star Wars, you can make a parody, you can make a joke of something. And the test, I believe, like the subtest here is the confusion of the audience. Does the audience know who the original author is or not? so if saturday night live does a parody of for two or three minutes of harry potter nobody is confused that that's actually harry potter i mean there it's pretty obvious right so this is part of it whereas if i did i wrote my own fan fiction of harry potter and it was really good and it was a full book you might be like wait a second i can't tell if jk rowling did this or not so there's something about the audience that matters here in this purpose as well correct yeah so it basically um fanfic is a great example it actually the examples that you gave implicate not just the audience's view but really the full factor test and the factors kind of like it's like a like an inverse scaled you know that one goes up another one goes down but in the case of harry potter great example so jk rowling sues fan sites and she wins because her stuff is so creative that, you know, if you have a fan site that says, okay, this is Hagrid and has big chunks of paragraphs and they're getting all this revenue, lots of clicks, you know, SEO optimized website, that's a fan site.

11:42JK Rowling testified and she said, I mean, so creative to even testify this way. She's like, it's as if someone came into my plum pie. I had cooked and picked the great plums out. And so it was like the creative aspects. Those were kind of what triggered the case. A lot of founders are great at going from zero to one. This takes vision, creativity, hustle, all that great stuff. But those same people often struggle with going from one to a hundred. If you want to scale and you want to do it efficiently, you're going to need process and you need structure. And that starts with your product. So if your startup needs a more structured engineering approach, you need to check out MEV.

12:19MEV helps businesses build and maintain their products faster and more effectively. They'll make your product more stable, scalable, and secure. They'll build custom infrastructure that scales, and they can help build additional features for your product and more. For each of your needs, Mev organizes an entire tech team comprised of senior engineers, delivery managers, DevOps, Q &A, and designers. And they've been in business for 17 years. And they've helped the following companies build complex tech products, Cartier, Tuit, and Ozempic maker, Novo Nordisk, my favorite. So let Mev help you increase product velocity and make product engineering more sustainable mev is going to give you thirty thousand dollars off your first three months that's right get ten thousand off per month right now at mev.com slash twist that's mev.com slash twist for thirty thousand dollars off your first three months interesting on the parody side you know the supreme court weighed in on it um i have a sound clip if you're oh let's do it let's do it okay because i remember when i was coming up in the industry i always found this one fascinating there was a game called mist it was a famous game and then somebody made a parody of mist and they sold it as packaged software and people got really uh they weren't confused by it but you know they had to make some concessions i think but and most of these lawsuits am i correct are settled out of court they don't go all the way people just say like hey this is not reasonable this feels unfair and then the other party says okay well if we put parody on it and we made these changes would that be okay with you and they kind of negotiate their way out of it yeah exactly they're pretty rare to go fully to the supreme court or even just to court in general because usually the parties work it out they're expensive unpredictable etc this one did go all the way commonly it's record companies or your sony's of the world your new york times that are pretty big and have you know kind of pockets to do it or a big financial reason to push it right they have a lot of stake exactly exactly so in this case they have to hold the line in some ways right if you don't defend yourself in an instance then the next instance it becomes harder to defend yourself is that correct that's right um also it's like uh commonly people who are trying to claim for use like they sort of know in advance and that's certainly the case with open ai like they knew the copyright issue was coming you know one of the things in the complaint says that their board member you know helen toner the one who departed one of her issues with sam altman was not addressing copyright properly so if you know that then you hire people like me to help you like okay where are the edges how can we wait on these different factors and so yeah so in this case this one is pretty woman so the classic roy orbison song the guitar riff is pretty recognizable and two live crew made um a version of it that i can i can play part of it sure let me see if you can get the audio let's see yeah is that coming through yeah it's coming through and now we're gonna get a copyright claim here exactly exactly no i'll just play just enough we will we will defend it as fair use because we're doing commentary on this exactly yeah right that is i mean it's literally it's literally the same right you know it's a cover song in a way it's like it's a cover yeah Yeah, so, but what they did, and this is kind of interesting, is, you know, instead of sort of the Oh, Pretty Woman lyric repeated, they made it Oh, Hairy Woman, Oh, Bold Woman.

15:49And like they made it, it's kind of a raunchy song, but the point is, is like, it's different enough that the argument was made that like, this is a commentary on, you know, sort of society and wealth. It talks about, you know, there's different, you know, you could argue that Two Live Crew was just in a different societal place than Roy Orbison in the 60s or whenever he wrote the song. And so it went to the Supreme Court, and this case stands for the proposition that a use can be fair, it can be parody or commentary, even if you're making money. Like, you know, this is a two live crew, they were not professors, they were not, you know, just like writing a blog no one would read, they were selling music.

16:29And so that case was important, and it got really into the four factors, so it's considered one of like the canonical cases. And how did that case work out? Did it wind up in a settlement, I would assume? Uh, you know, after the Supreme Court ruled in Tulip Cruz favor, I'm not sure what happened. I think, um, the Roy Orbison estate like lost rights to the work or something like that, but it turned out to be a sad thing for Roy Orbison in the end. And now when people do do samples, there is, and the music industry is the toughest. They're the most hardcore because it's a small group of people and they work together in unison.

17:04You know, let's be honest. They're just super sharp elbowed people. They've always been in terms of IP. So they have said like, hey, listen, you want to do a cover. Here's the mechanicals and the licensing for that. Hey, you want to use a sample. You have to have permission in advance. And then, hey, you can do a sample. And then Kanye just did a Backstreet Boys cover. when people were wondering how he got the rights to the sample it wasn't a sample it was a cover and so they they have their own little mechanics and traditions in the music industry that or standards right they've established exactly uh for this that yeah i mean that's the music industry is a great example of really the market wins and like that's one of the points i made in the tweet and i think is important to think about when you think about this case that i'm not a doomer in the sense like this isn't going to end ai like there's no universe where this case would end ai And so the result is, do we end up with a licensing scheme?

17:56Like, is this Napster to iTunes, right? So Napster comes out, you know, in the time I was in college, it was like, you could get the entire Beatles library from somebody in the dorm next door. Like, you know, it was clearly like, it felt sort of bad. Yeah, it felt like stealing. It felt like stealing, right? Well, because there was no difference between downloading on Napster or downloading on iTunes or buying a CD. It was this, you did one in place of the other. That's exactly right. And iTunes came up after, right? It was like, okay, this is a legitimate way to pay for digital music. And people like, you know, me or you or whomever, it didn't feel like stealing when you paid for 99 cents.

18:37And it wasn't when you paid for it on iTunes or Spotify or wherever now. And that industry has come out of that. So you can see, you know, with open AI, they could have a system where they figure out kind of the provenance of different outputs and pay in some way. Or, okay, you want me to do a verbatim Luigi? All right, there's your five cents to Nintendo or whatever. Like, this is not, the tech will find a way. I'm very confident of that. This argument by technologists is that this is too hard to do attribution is nonsense. I mean, if you can, you would agree. Yeah, I agree. I mean, if you can create this incredible AI that's able to make images, you should be able to figure out what was the source of those images.

19:22And if you can go find these libraries of content to train it on and then train these very sophisticated things and set up 10 ,000 computers or 100 ,000 computers and billions of dollars worth of computers with thousands of engineers, I think you figure out attribution. It's not that hard. And in fact, there are services that are already out that are in the chat GPT mode, which actually do do citation. So the market has already proven it's possible. Let's talk about this one piece of the four-part test You've got the purpose and the character of your use. Is it parity? Is it education in education?

19:51If you're not making money or in society, if you're doing commentary, you get a bit of protection We want that in society. We want mel brooks to be able to make jokes. Got it We want a professor to be able to show You know star wars and give commentary in a class in a non-commercial setting for people to learn It's not going to compete with people, right? So that's all really good stuff. That's good stuff that we want in society. We also want people to be able to make fun of things and do commentary. So if Jon Stewart or John Oliver want to take, I don't know, a talk that some, you know, President Trump or President Biden did, and they want to make fun of it and use parts of it, well, we want them to be able to be mocked in a free society.

20:33And that doesn't kind of conflict with anything. So we understand those. Yeah, kind of a fun, one fun little point on that is that it comes from the Constitution. So copyright law is actually, it's federal law. And the IP clause is Section 8 Clause, Article 1, Section 8 Clause 8. And it says, to promote the progress of science and the useful arts, Congress can secure limited monopolies for authors and inventors. And that initial to promote the progress of science and the arts, that's been used by courts to limit copyright. So copyright could go really, really far. Like you could allow copying never.

21:13But that idea that it really is about societal progress, that's also what helps tech companies, right? So that's also why Google won the thumbnails case because they're like, look, you know, it's super useful to have search. How else are you going to have image search if you don't know what image is actually in the results? Right. And they also had the argument, I think, in that case that they were doing very tiny images, smaller percentage of the original work, and that they weren't taking every image. I believe there was like, we're only taking a small amount of it. And then they also, I think, had the sort of ultimate rebuttal, which was you can also, I think they created robots.txt around that time where you could just say, you know what, I don't want my site index.

21:52And then Google was like, if you don't want to be in the index, you don't have to be. And so Perfect 10 then could just not be in the index and problem solved. So then they had to make the trade-off. Okay, give a little bit of my content, a thumbnail image of, you know, some photo of an adult nature. And then I, but I get some traffic. So maybe it's worth it. And then the copyright holder can make that decision. Just like I think Star Wars, Lucas was very cool with fan fiction as long as it, and fan movies even, as long as you didn't try to monetize it. Starting a business used to be a pain. You needed a lawyer.

22:26There were hidden fees. It was a mess. Now with Northwest Registered Agent, it only takes 10 clicks and 10 minutes. Northwest provides everything you need to start and maintain your business. Every LLC, corporation, or nonprofit at Northwest Forms comes equipped with registered agent service, a business address, a website and hosting, email, a phone number, and this is all covered by Northwest's privacy by default. Again, your full business identity will be live in 10 minutes and in 10 clicks. So here's your call to action for$39 plus state fees. They'll form your LLC corporation or nonprofit and launch your business in just minutes.

23:06Visit Northwest registered agent.com slash twist today. That's Northwest registered agent.com slash twist today. So if you go onto YouTube right now, you can watch all these really creative kids running around dressed as Jedi fighting each other and releasing episodes. They don't get cease and desist. But J.K. Rowling might say, hey, with my art, I want a different standard. Exactly. Exactly. And, you know, OpenAI to the point that, you know, you get lawyered up and you kind of realize what you have to fight about. OpenAI has been savvy about this. And they announced in the summertime that they'll respect robots.

23:41TXT go forward. And so these kinds of systems where you're giving the owner control, that's going to be the kind of thing that OpenAI will argue, you know, matters here. And they're doing that, you know, partly informed by precedent, but partly also because from an economic standpoint, it's the right thing to do. Okay. You own your content. You have this bundle of rights. You want to license your content to make a Harry Potter restaurant. Fine. I mean, that was another case, actually. Funny enough, somebody tried to do a, um, uh, it was called. I think it was a cafe. So it was, it was a restaurant actually of SpongeBob.

24:14Um, so SpongeBob, there's the Krusty Krab, which is a SpongeBob character. and there was a restaurant in houston called the rusty crab and they they tried to say oh this is social commentary but it really wasn't it was just a spongebob restaurant and so viacom went after them and won so now the percentage of the work matters when in in this fair use test as well correct that's right yeah so it's called um this the second factor is the amount and substantiality of the portion used um okay and sorry no actually no i can pull i can pull up the slide that's the third factor technically i got confused here i'm not i'm also not an ai one second we've proven it exactly exactly like a lot lots always um okay so nature of the copyrighted work that's factor two it doesn't get a ton of play because it's um you know most work is creative but in this case um new york times anticipated this issue and they've got a long thing in the complaint about how creative their journalism is and they're right like you know they spend a lot of times i mean obviously you know you've been a journalist it's not just pure facts there's a lot of ways and funny enough i i didn't expect this but in response to my tweet thread there was a lot of political things it was like oh my goodness i can't believe the new york times is such a chunk of opening eyes training data that's why you know gpt is so woke exactly yeah exactly but this is interesting facts and data points are very hard to copyright so if you there's like a website i use often which makes beautiful graphs i forgot the name of it but it comes up all the time it's like the world in data or something like that there's world in data and then there's another one that comes up in seo and all this company does and they charge like a subscription for it is take other people's data and make a very beautiful standardized chart you know i was looking for some market maps for one of my investments or some market sizing and it had you know it was like something super obscure it was like the world button market or something and had like all the countries yeah yeah yeah and then you look in the credit it says source you know this is uh you know pew research pew data this is from this data so you can literally make any chart you want on anybody else's data as long as and i think in terms of fairness you just put that that's the source of the data but data is not copyrightable is this correct like facts and data are not copyrightable.

26:37Yeah, so a couple of big cases on that one. One was Feist versus Rural Telephone, and it's kind of a little antique, but essentially one telephone maker took the phone numbers and names from another, made their own phone book, put their own ads in it, and essentially that case was pretty important because the Supreme Court said, look, copyright is not about labor. It's not about the work you put in. It's about the creativity, remember, to promote progress of science and the useful arts. Is this really about creative progress is copying of a telephone. Now there's other ways, maybe contract or other ways that you could go after.

27:15But in your scenario, Pew, you know, if it's reported as a fact, you know, percentage of Americans on the internet every day or whatever it is, then you would be able to use that compilation. Now there's some nuance around creativity in the compilation. So the other big, big case on facts is Oracle versus Google, Right. So that went to the Supreme Court. Google copied Oracle. I think it was declaring code. And so essentially, in order to be able to to have Java on on Chrome, they did that copying. And the Supreme Court that was heavily litigated over years, I think 10 or 11 years. But in any event, the Supreme Court found for Google in that case.

27:55so the nature of the work yes statista is the name of the website that we sometimes use and then there was another one e-marketer and they've gotten in all kinds of like legal letter kind of trouble i believe i remember seeing it i'm not sure which side had that but then like you know other kind of reblogging sites started doing the same thing so if you want to make a great business you can just take other people's facts and make beautiful graphs out of it you see people do that all the time but that makes sense and and then scraping data there was a israeli company that was scraping linkedin data and they were saying hey this is just facts that's another area scraping and fair use there i don't know if you've seen many cases there but they they yeah then you get into international jurisdictions like what people think in japan india you know the middle east and europe could be very different the jurisdiction could be very different in how you use data i think linkedin and microsoft sued this israeli company and lost um yeah yeah it's interesting i mean with scraping you know thinking about sort of the startup angle some of it is also contract law like i've seen scraping cases get on like you're literally trespassing and this is why also you know some of the technical means like you know if you're scraping in such a way that you're like ddosing the site or you're hitting it so much that of course there's other claims against you and you know every website terms of use has an anti-scraping that I've started to see in my practice.

29:17And then a lot of companies now are putting, it's against our terms of use for you to use our data for training. And you can imagine in vertical AI, people doing, I don't know, AI for doctors. There's a website called Doximity, which is like, it's the LinkedIn for doctors. Okay, are they, if you scrape that content, maybe you're individually doing it, you're breaking the terms of use, especially if you're doing it locked in. so there's kind of other there's other theories but yeah the linkedin case was a big one and scraping overall you know certainly in e-commerce it's everybody does it yeah but knowing the price of a product the price of a product across 10 different websites across 100 different days doesn't feel like the nature of that copyrighted work is not like some artists invested a lot of time in it now if you said 10 people to the front in the ukraine or in ukraine rather sorry um and And, you know, you spent a million dollars putting them there for six months.

30:15You know, this is a whole different ball of wax. There's a lot of work. And that's what the New York Times is claiming here. The amount and substantiality of the portion used. That's the third part of the test. What does that mean? Yeah. So this is getting at our particular phrases, copyrightable. So interesting one here, Taylor Swift with Shake It Off. She said something like, player is going to play. And there was a rap song called Player is Going to Play some years prior. And they sued Taylor. But she won, or at least the case went away. They agreed to drop it because there's not that many ways to say that concept.

30:49So there's this thing called the merger doctrine. And this is actually an issue where if there's only so many ways to do something, then you can't copyright that thing. But this is a fun one. And I have a visual on this one that I think is funny. And it's actually a doll. It was the Seventh Circuit. So that's the circuit over Chicago. And there was this company talking about e-commerce that made apparently very lucrative to the surprise of the court, which they say in the opinion, but basically farting dolls like you buy it at the mall or wherever and you get a doll and it makes a rude noise. The doll on the right was basically the makers of that doll had gone to like a toy show or something, seen it, and copied it.

31:36So copyright suit brought by the makers of the guy in the green chair. And the court said, look, the concept of a farting doll, that's not copyrightable. No. I can go make one. But the court has this amazing paragraph in the opinion that's like, they could have given him a mullet. They could have given him flannel. They could have done it. They could have put him standing up. They could have had him wearing boxer shorts, you know, whatever. Like the point is these little details, too substantial of a portion of the original was used. And it was not associated with the idea. Like had nothing like the idea itself.

32:09You can express it a bunch of ways. So people in this case, in the open AI case, you know, the art stuff is super fun because it's visual. So there's been a whole meme and I did another tweet on about it, about Super Mario and Luigi. And, you know, if you ask for an Italian plumber, that's what you're getting. Yes. You know, there's other ways to have an Italian plumber, right? Maybe he's really stylish and wears Prada, you know? You could make an anime version of it. But the fact is, the most iconic one that has had a lot of money invested in it was by Nintendo. Listen, not every business is venture scale.

32:43If you're not, you won't be able to raise money from VCs. We all know that. And not everybody has a rich family member to do their friends and family round. So if you want to jumpstart your business with$50 ,000, let me tell you about Painbrush Loans. Painbrush has created a new kind of loan product. They connect idea-stage startups with bank capital, so you don't need to give up any equity and there's no pitch deck or revenue required. And the Painbrush Loan is available at the idea stage. In fact, you can apply the moment you incorporate your company. Monthly repayment is a flat, predictable amount, which makes cash flow planning really simple.

33:19So here's your call to action. If you're a founder in the US, go to get paintbrush.com to see if you qualify for a$50 ,000 startup loan in less than two minutes. That's get paintbrush.com to see if you qualify in less than two minutes. One of the things I tell young founders or people in content is like, if it feels unfair, then perhaps it is. And you have to have empathy and take into account what the other party is going to think their opportunity would be. And I think this gets us to the fourth part of the test, which is if a new product or service is going to be made from J.K. Rowling's books or from the New York Times archive, who deserves that opportunity?

Read the full transcript

33:56Am I correct? That's the fourth part of this test? Exactly. And you're correct in two ways, both on the test and on the feeling. Like, you know, I've been practicing law a long time. And a lot of these cases really are, like you said, Napster felt kind of wrong and it kind of was, right? And so it does turn on that. But in terms of the factors, let me pull that back up and I can, I was very proud of my emoji. So I can show you the emojis here. But yeah, so basically like, you know, you see the flying money, but it's literally like, what is the market for the original and the value of that market and who gets to exploit that, right?

34:34So intellectual property is similar to regular property, right? If you have a piece of land, who gets to put a hotel on it, right? If you have this really juicy piece of land, right? So similarly here, New York Times got this factor by saying OpenAI has already made deals for this. They know how to license data. Like they've already done it with Politico. They've already done it with the AP. It's not like there's no market for this. And so, you know, others in the thumbnail case, that was harder to prove. There was evidence put forward, oh, you can use a thumbnail for, you know, at the time we had those Nokia flip phones with the tiny little lock screen.

35:07And it was like, okay, there's a market for that. But evidence came out in that case that those were fake licensing deals done just for the litigation. Here, it's clearly not. So, yeah, so that's another factor. And then, but it's not, like I said, it's very squishy, four factors, kind of, and they're not exhaustive either. So, you know, a court could say, you know, there's four factors, and I'm going to introduce a public good factor, you know, like basically make one up. For the effect of use on market and original value, this is where I'm going to do a follow-up post to my original post, which is I pay as a user 20 bucks a month or so, 20, 30 bucks a month for New York Times, and I pay 20, 30 bucks a month for ChatGPT4.

35:47I'm paying for both. And I recently was going to, and I'm a huge fan of the Wirecutter, and I am like a crazy product research guy. I just love researching products, restaurants, et cetera. I use Yelp. I use everything. And I love Wirecutter. In fact, I tried to buy Wirecutter or invest in it back in the day before they sold it to the New York Times. And so I did a search for coffee grinders and some other stuff. And I actually did it on, I believe, ChatGPT and Claude and a couple other ones. And I was just testing it. And it was pretty clear that they got their information from Wirecutter because, you know, it was kind of like the answers were very similar.

36:21I am like, I think the tip of the spear here. If I get my New York Times and my Wirecutter from ChatGPT before, I might cancel my New York Times over time. Is it not the case that the product OpenAI built, the ability to use a chatbot to talk to the archive of the New York Times, that is the New York Times' opportunity, not OpenAI's? Yeah. So, you know, the example I use, you know, Martha Stewart, great media conglomerate, and she's very, very savvy, very tech forward. She was talking a year ago, right after ChatGPT came out about creating a Martha AI that you could talk with because she's similar to the New York Times.

36:59She's got decades of really high quality content that in a particular voice. Right. And so, yeah, that's one way. Another way is New York Times made the case in the complaint that they calibrate very carefully what's free versus paid, right? You know, the amount of gift things that you have. You know, if you click from Instagram, sometimes you'll get the gift version. Basically, like, that's their, like, the rights holders, essentially property to exploit is how they would say. And how they subdivide it and where they put that line, how much admission they charge, all of those things they would say are within their rights.

37:34Now, OpenAI would probably make similar arguments to you that, okay, under that first or similar arguments to what they make out of the first factor, which is, you know, it's a different thing. It's a new, you know, having an LLM, specifically a very large language model trained on that number of parameters, the amount of investment that they've made, they've changed it into something different to where it has a different purpose. You're going to chat GPT to have generation as opposed to have, you know, pre-made research on a particular thing. their claim would be and i'm trying to take their claim seriously here in feral yeah their claim is hey we did this first we made a language model first therefore since the new york times they didn't get to it yet this is new because doesn't the new york times have an unlimited amount of time to exploit their own content like yeah and they could choose not to yeah or they could not to but so if disney said you know what we we bought marvel and you know we haven't made a marvel theme park ride yet that doesn't mean somebody else gets to make the marvel theme park ride exactly yeah no so the time aspect if i if i mentioned a time aspect um that would be i misspoke but no no you didn't mention it i was just building on your thoughts on it which is they're saying hey we spent all this money to build this thing it's like yeah you did we are planning on building it as well at some point therefore it's our opportunity so i think that That one fails miserably.

38:56They fail. Where I think they could say is like, we're not trying to replace the content. They could say our aim is to just merely have the best language model possible. And therefore, you know, the literally more, the higher volume of text that we use kind of the better. And, you know, it's going to be more like the job, the Oracle case that I mentioned where, you know, Google's argument was like, Like there's only, you know, Java programmers already know how to declare these variables. We're going to copy the declaring code to literally advance the progress of engineering. And so OpenAI could say something like, it's a different thing.

39:35You know, we use the New York Times content and other content for training to advance kind of the state of the art of the actual LLM, how it generates words, how it's better. And people can kind of see this. And what they would say is like, look, GPT 3.5 and GPT 4, very different. right gpt4 is much better because we did more training on more content ergo it's not the actual content itself or the creativity of the content it's just the fact of having content so that's another another way they could they could take it based on your gut let's say this goes to the mat and we went through this four-part test based on your gut percentage wise new york times wins their argument that you can't train on our data and they have to they get an injunction what are the chances as that happens.

40:20I mean, I'm really putting you on the spot here. The odds of an injunction are very slim. The standard for an injunction is that it causes irreparable harm to whoever it is, the plaintiff. And it's harm that cannot be fixed with money. And so there's very few harms, really, that can't be fixed with money. And so that, and then the test for an injunction is another four-factor test. And so when it's sort of close and when you have a technology that definitely has societal benefits. So, you know, opening, I'll say, look, we've got people, you know, diagnosing things with Chachi PT. We've, you know, saved marriages, you know, whatever, all the amazing stories about Chachi, which honestly, like it's an incredible productivity, but, you know, I use it every day.

41:08Like I'm a huge user of it. And so just on that societal benefit, I would be very, very surprised. So injunction unlikely. Injunction very unlikely. So let's work backwards from that injunction less than 10 % chance. But, uh, yeah, I would say less than 10%. I think that's not zero. Non-zero zero. I mean, you could have something very strange, like, you know, like in the Apple case, the patent case, it went sort of all the way to Biden and stuff. So, so you could maybe, but I think it's unlikely. So if they were to lose, then you would be in the damages, but then they would also have to remove it.

41:41Right. is that's a possibility, is that they have to retrain, that the settlement could be that they have to retrain things and take the New York Times out of it. Yeah, I mean, it could be, that being said, we, you know, 4.5, GPT 4.5 is rumored to be coming out. And so it could be that, you know, they sort of skip straight to that. And they've known about this case for a while, like the complaints as they've been negotiating since April. So my guess is OpenAI has probably already kind of firewalled off the New York Times content. Got it. New York Times, I think 535 other journalism publications, everything from down to like the St.

42:20Louis Post-Dispatch have put themselves on that do not train list. And so I don't know, but it's very tough for an LLM already developed, right? It's back to J.K. Walling's example. Like you can't put the plums out of the case. In this case, it's almost like, I don't know, it's baked a cake and it's like the vanilla. like how are you gonna get the vanilla out of a baked cake like it's not a thing if we know that they trained it on one or two percent in the first versions and they should be able to determine that because there's going to be discovery in this case and this case is going to keep going i don't think there's a set i don't believe there's going to be a settlement i think they're going to take this to the mat new york times because i think they regret not taking to the mat with google back in the day so this is i think existential for them where they view it as such therefore they're going to go to therefore there will be discovery and in discovery there will be slack messages or emails or conversations about what are we going to include and they're going to have that open crawl and that open crawl is going to be plain as day what they put in there is going to be in a hard drive somewhere and then there's people talking about it saying the new york times is really high quality we should move their weight up and we should make this like more important than say business insider which is a lot of like fakakha nonsense and then you know oh and then there's like 4chan or reddit like maybe we'll make those a little bit you know uh less valid or maybe make them more valid who knows uh for for in the case of reddit so that's all going to exist in discovery and that's going to be super damaging is it not and then the discovery part of this could be explosive yeah i mean it could be super damaging but it also could be helpful right so i mean open ai they went for the nonprofit model, in part because they saw, I mean, copyright is one flavor of issues, but they saw certainly societal issues.

44:05And so, you know, I've done, you know, interacted with OpenAI, Replit had a deal with OpenAI going back to 2020. So they're pretty thoughtful. So I mean, yes, you could get, but any discovery is always a wild card. You could get crazy emails. In the Google case that I mentioned, the Oracle Google case early on, there was like 150 000 worth of litigation over one email from an engineer saying hey i don't see any way how to get out of this without licensing from oracle he just literally put it in there yeah exactly and it was an email to you know like larry and sarah gay and it was like the lindholm email and it was like famous and this guy who's like you know a director of engineering was like had his moment in the sun from that email so yeah this is a reminder never put never discuss legal topics on electronic communication always discuss them on the phone or on the thread like i actually use it to teach privilege because he um he cc'd a lawyer but it wasn't to a lawyer he wasn't asking for legal advice he was declaring so it's a fun one but yeah it's it's um i actually could show the email if you want oh yeah that'd be great uh and so the the issue here though is open ai can't have their cake in it too they can't be selling billions of dollars in secondary and claiming the nonprofit for the good of the world when literally the same executives who i'm going to use the word liberated or took without permission the new york times took without permission are the ones cashing in their shares at 100 billion dollar valuation and pretty illogical if this thing is worth if the new york times was two percent of the training data and if it was let's say the best of the training data and they said this is five times better than anything else okay that's 10 percent of the good stuff okay 10 of 100 billion is 10 billion so we want 10 billion or if this thing's going to grow to a trillion we want 10 of the value of the company and when it becomes worth a trillion you know we're going to 100 billion yeah i mean and these kinds of cases like you know it's always and this is where you know i i love being a lawyer and like i i generally think yeah right i love you being a lawyer oh thanks no it's like i i think it's like where the advocacy really matters.

46:16So, you know, one of the things I worked on early in my career was the Apple Samsung case, and I was the associate on damages and figuring out, okay, what's the value of a rounded corner on a phone? Like that was like, how do you assess that? And so similarly here, there's a lot of unknowns. We don't know how valuable OpenAI is going to be. We all think it's going to be worth trillions, but we don't really know. There could be some meta could break out or one of the others could break out. Yeah, it could become worthless. It can become natural. Exactly. Or, you know, a lot of the research I've seen in the last maybe couple of months is that AI can generate its own training data.

46:49So there's people literally saying that we don't even need the New York Times anymore. We can use the AI we have to write the New York Times. Synthetic data, yeah. Synthetic. Exactly. And so that's like another thing. So it's really, you know, it definitely is not for the kind of faint of heart or stomach. There's a billion ways to heart and do anything. But here, let me show you this Lindholm email because it's so fun. One second. there's always somebody on the staff while you pull it up that thinks they're an attorney like me because i'm sitting here with my non-legal degree but i've got a lot of experience and uh i always tell my team members like you're not an attorney do not talk about any legal issues ever we can have a phone call and talk about when i'm an attorney but be careful okay here we go exactly so this one so this is from tim loneholm who was um i believe he was an engineering director and And he sends it to Andy Rubin.

47:36And then Ben Lee was a lawyer at Google. But he says, context for discussion, what we're trying to do. And he calls it attorney work product, which again, he's not an attorney. So he tried. He calls it confidential. And then he says, this is a short pre-read for her call. And then this is the famous line that got a lot of play in litigation here in San Francisco. What we've actually been asked to do by Larry and Sergey is to investigate what technical alternatives exist to Java for Android and Chrome. We've been over a bunch of these and think they all suck. We conclude that we need to negotiate a license for Java under the terms we need.

48:12And that was the key issue in the case. Funny enough, to your point of going to the mat, Google ended up losing on this at the trial level, but went up to the Supreme Court and ended up winning over the needed copying. But to your point that this is like, it's going to be a fight and it's going to be a lot of discovery, I would predict that. All right. So what else are we missing here? because you in your deck had some of the examples i think is super compelling and because technologists you know you work with technologists they tend to a portion of them think if i can technically figure out how to do something it's legal or it should be i don't know what to call this but like it's sort of might is right if i can technically figure out how to scrape your website and create this or create that and well then it should be legal which is how the napster folks felt like well we tech it's and there's also the technical inevitability argument well it's going to happen And so we might as well do it.

49:03Yeah. Yeah. I mean, there is something to that. So one of the cases is, was an emulator. So Sony is a very common either plaintiff or defendant in IP cases because they have a lot of valuable IP. And basically someone made an emulator of a PlayStation, an early PlayStation on a PC. And the graphics were actually technically better on the PC. And that case went to the night circuit and the emulator maker won. because and there was a lot of copy involved they had to have they had to basically reverse engineer the entire playstation to be able to do it and of course they copied it like the literal bits and bytes of of the code were were put onto into the emulator and so that there is something to that um i mean well this is also the great irony of this is that while open ai is an organization you can sue because it exists as an entity the open source community is a little bit harder to sue because they don't exist as an entity you have contributors so maybe you could speak to that because if let's say open ai does lose this case or settle which i believe is what it's going to be one of those two things a massive settlement nine figures minimum is my prediction and uh but it will not be disclosed but it'll be at least nine figures and with some kind of licensing going forward but even if you were to do that what's to stop as you know all these open source projects come out there and somebody's decides they're going to roll their own model as hardware gets better and better that they just rip the new york times and you could buy the new york times archive probably from somebody in india in manila in israel there are scraping companies that sell these things on the what i'll call the gray market maybe illegal here maybe legal there maybe there's no laws there so maybe you could speak to that do you think all this is for open source open source is an interesting angle i mean i think open source had its own you know one of the things that was that was interesting in the wave of ai regulation we've seen you know from the eu and others was you know open source had a lot of the same objections of like um you know people had t-shirts with algorithms printed on it they're like okay if we open source then you know all these bad guys will get will get the code but sort of the market worked out here i think it's tougher i think um you know it's going to involve calls by the right rights holders and then what i think will happen is what we talked about at the top of the hour which is like as the tech emerges a market like tech for the market for it will emerge you know we're going to get the itunes equivalent and i think there are some startups being funded in there they're still at this point I don't think it's a, it's not a, before this case, I don't think it was being talked about enough to be a problem with a big enough market, but now it is.

51:49Here's a possible solution. Let me see what you think of this. I buy ChatGPT for 20 bucks. And it says, if you authenticate with your New York Times subscription, so your ChatGPT, and I authenticate with my New York Times subscription, then it says, okay, you're going to use ChatGPT 4.5T for New York Times. Yeah. Yeah. And so but if you don't have the tea and you do it on 4.5, and you say, hey, wire cutter, what are the best things says, hey, you need to have a New York Times subscription. So authenticate with that. And then you say, hey, I want to make Star Wars carry says, oh, you know what, you have to use open AI.

52:25You have to use chat GPT with Disney plus so authenticate your Disney plus. And now you can start to have fun with the Disney characters in Dolly or whatever it is. And then they could license that to the highest bidder. because when i you know if you use hulu and you have hbo max or nba or use apple tv they just authenticate each other's subscriptions you have this sort of subscription death by a thousand subscriptions uh kind of concept what do you think of this concept i mean i think it's it's certainly that that shape of a solution it sounds right to me like i i think technically viable too right i mean technically doable the other interesting thing is there are a bunch of startups trying to do sort of like your digital life right where twitter search is like notoriously terrible you literally can't find anything on twitter and how often does it happen to me that i'm like oh i remember there was a tweet about that and then like i can't find it so you can imagine an llm that's actually trained on your entire everything you've ever consumed and then by the nature of if you've consumed it then presumably at some point along what you have the rights to it yeah there's like there's a million you know ways to do it and that's kind of the why i i characterize the lawsuit it is historic is that we're at this moment where we don't know what the what the tech and the market solution to this is yet and it'll emerge it's just you know maybe not in the exact way we did it so another fun thing um andreason horowitz there's an investing partner that she writes um i think it's connie chan she writes a lot about um china and media in china in china when you buy a kindle book or any kind of book digitally you pay by the page right so it's like not actually a thing so the fact that we happen to buy whole books here in the u.s that's the market that emerged not necessarily a foregone conclusion so in your example you could have you know that you're like do you want wire cutter and it could literally just be like the wire cutter slice of the new york times thing or it could be like by the query or it could be some kind of rev share like you know as you said music industry super sophisticated on this i think you know the words and kind of digital print publishers let's be honest publishers are kind of dopey they've they've been dopey historically they've never really been smart about their approach legally they've never held the line they let google run amok and you know rupert murdoch got it right he's like google is nothing without us if those publications had grouped together in that era and told google listen you know the top thousand publications are going to no index unless you pass a licensing fee and here's what we want.

54:55Google would have paid it, I'm sure. And they just never had the coordination or the chutzpah that they needed to. I think the New York Times today is so sophisticated because they're a subscription-based business. The move to subscription-based makes them understand the value of their content. And because it's subscription, doesn't that change everything on a legal and technical basis about this case? The fact that there's a firewall, maybe you could explain how the subscription wall changes this a bit. Yeah, so, um, strong plus one on New York Times having kind of jumped the digital divide or jump the digital, you know, evolution there.

55:34The New York Times food app would be, you know, it's all a startup in its own right, and in the hundreds of millions in terms of revenue. And, you know, recipes themselves are not copyrightable. Obviously, the rest of it is, but I pay for New York Times food and I have for since it came out, because it's so nicely compiled, and they do the, you know, 10 recipes to make for the New Year. and whatever and so beautiful it's worth it exactly it's a thousand it's a thousand percent worth it but publishers and another another example you could look to here is kindle right so they did one of the things that i i point out in the thread and i think is or you know in the responses to my thread was so amazon kindle had uh the guy who's now the chairman of co2 ventures dan rose was the head of business development for kindle and he basically oh good yeah he's done a bunch of um you know did a bunch of deals with the publishers and initially he's public about this bezos said don't tell them we're making an e-reader and he's like well how am i going to get them to do deals with me if i can't actually say and so eventually they did but you're right that that was a moment where you know the tech company kind of had this power but what was different about kindle was kindle was still kind of unproven at the time versus here we've got chat gpt clearly it's a runaway success it was the hundred million people using it yeah exactly you know they're at i think it was 160 or 1.6 billion arr now um and so they can't claim poverty or this isn't a real business this is not a student project so i mean even if you you know made the argument i don't i don't know enough about publishers to know if they're dopey but even if they were like you can see the money i knew them they were they were for 20 years when the digital but now they're super sophisticated the ones that survived it's kind of like a darwin thing like if you survived as a publisher you're sad yeah makes very full stop okay now in your deck you had some other examples is there anything else in the deck that's super compelling we should rip through here before as we wrap up let's take a look i love i love a guest showing up with a deck amazing you've turned out to be a great guest thank you for coming on the program yeah super fun happy any any time you have anything legal happy happy to dive in there was a fun part of the thread where people were making Luigi fan art and you could kind of tell when the model was getting was was being aware of copyright.

57:54So this is kind of a fun one. So I started getting errors that said, okay, put Luigi in the background of my chat GPT says I'm unable to create an image with Luigi as it doesn't align with the content policy for image generation. Okay, so this was like yesterday yeah and then okay but clearly they didn't care about the grinch and blues from blues clues coca-cola and then i threw in um in the background uh no it's the castle from downton abbey oh down now you're right yes that is yeah yeah so i'm sorry i've heard i didn't watch down and abbey oh it's amazing no i did so much it's so good i'm being cheeky i did watch it okay yeah no the movie itself if you if you just want the movie it's pretty good so for copyright violations.

58:38Exactly. I mean, you had trademark there with the Coke and everything. So yeah. And then, you know, with the Grinch, it tried to actually, at different points, this is kind of funny, it would try to actually do different things. I can see if I can find you the, let me see if I can find the Grinch that it did. It did a Grinch that was, let me pull up my GPT history, always scary on a live demo. Yeah, always scary to pull it up. You could have all kinds of interesting. Exactly. No, no, no. I will pull it up because it is funny. I think I asked for, a green character that hates Christmas or something like that.

59:09And what it did, this was really fun. So I was with my three-year-old and she wasn't fooled. Like this one, she did not think was the Grinch. Like she said Grinch, but he's kind of different. He's got like - Kind of looks like Sesame Street character. Yeah. You know, this is clearly not the Grinch, but later in the thread, I'm like, no, make it mean. That's like so clearly the Grinch, right? Jim Carrey. Yeah. And then, you know, also the Grinch, but then this one, it kind of goes back to - star grinch it goes back to a pixar grinch so it's like not really right and then this one is like like a wizard like what do you what even is this disney disney maybe a disney grinch i don't know yeah so i asked for nordic princess sisters obviously on an elsa you know so they're right there with the braids you know so well this is the the thing you can you can know the the keywords very easily of these are ip so if you just said hey give me all the disney characters all the marvel characters put their names in here people ask for that just tell them it's against the yeah moana it just says no if he's just saying make me moana we'll do it yeah and so i did this where i was trying to make a my bulldog into darth vader and then it says we can't not gonna can't can't do that i said make a sith lord bulldog and it's like yeah of course here you go so i think they're trying to get this copyright thing under control but the truth is especially for images there are there's a finite number of styles in the world and so it's very clear that they have a pixar style and they have a marvel style and they have stolen those styles those are not their style so maybe you could speak to the concept of a theme or a style the pixar style is unique to them is that defensible and if you if they if you say i want to make this in the style of pixar should a language model that makes images be able to make you a pixar character should they be able to do that yeah i think the idea of the pixar style should they be able to do in that style or inspired by i think so i mean this is like okay you know a round face and you know i actually represent this is fun this actually came out one of the cases i worked on at my firm was barbie versus brats so the founder of um mga which makes brats worked at mattel which is very active rights holder they sue people for a lot of barbie things and then obviously the barbie ip super valuable billion dollar movie this summer so he worked there during and one of the defenses was that you know it wasn't infringement because it there's only so many ways to make a doll so in the office really fun we had these like big doll heads everywhere and we looked at anime, we looked at whatever.

1:01:46And the case also 10 years of litigation. But should you be able to make a big headed doll like in the style of a Bratz doll or whatever? Probably. I mean, so I don't know. I think it'll be tough where you can ask GPT now to give you a Taylor Swift style song. And it does a pretty good job. So where it's something new, a little bit better. That's why the exhibit J with the hundred verbatim things is so important. So copyright law isn't gonna isn't gonna stop um you know make me a pixar style character of you know jar of mustard or whatever like you know like whatever you want to pick it does seem that some people are confusing non-commercial use with commercial use so they're like well i could draw a jedi bulldog is that illegal if my daughter makes one versus i'm charging 1995 for a product to do this and at scale with 1.6 billion in growing in revenue.

1:02:43So can you explain to people why, you know, these are two different things in the eyes of the law? Yeah. So that was actually a big issue in the Betamax case. So the Betamax case was VTRs or what is now VCRs. And the funny thing happened, which is Disney was one of the groups that sued Sony. And the Supreme Court held, you know, there is a substantial non-infringing use, that's the language, which is time-shifting. So, you want to watch the game, you use your VCR, you record it, and then you watch it later. And there was all kinds of evidence this is how people were using VCR. Nevertheless, Disney was one of the petitioners in that case.

1:03:24I think it was less than six years later, Disney was the single biggest seller of VCR tapes. and so literally like the tech finds a way right and the market finds a way and so in terms of commercial use you know there were a bunch of people in the comments and some beautiful article in i think it was the guardian about um i can't remember his name the the guy who came up with mario the game designer the famous guy yeah i know he's talking about shiguro i think yeah anyways saying that he was the architect of children's dreams for a generation which is like a beautiful quote but essentially like if people you know making actual marios okay i think that nintendo should be able to go after that but making mario style video games no i don't think they should be able to go after that try to inform the audience of where we think this is going the prediction for what happens in the long term here with this case and then how it affects the wider industry so it does seem the number one possibility in all cases of a copyright claim is settlement so i guess that would be one possibility settlement then there's go to the mat and take many years um and then get a judgment right that's a second possibility here so and then i guess there's the courts throwing this out or dismissing it right or something so are those the three buckets we should be looking at here like either new york times wins or loses or settlement happens those are the three possibilities broadly speaking yeah i mean winning and losing is like you know even in some of the famous cases there's a process where it gets remanded so sent back so some of these things for the infringement that or you know if it is infringement, the things that have already happened, your Exhibit J's type of examples, you know, New York Times will preserve claims on that.

1:05:24But as OpenAI makes changes, I think you're right. I think it's either the case gets settled. And one way that could happen is, you know, they announced some kind of copyright holder symposium or something. And New York Times is like the head of this consortium. And it's like some kind of opt-in system where, you know, So New York Times content is part of it and publishers can go there and maybe they get a little royalty. You know, something accelerated similar to what the music industry has, where there's a very sophisticated thing where you need to get the mechanicals and you need to get the performance rights.

1:05:55And it's it's it's there's like a known system and libraries for that. We'd call that a marketplace solution emerges. Exactly. So like one possibility is like a marketplace solution. Another possibility is, you know, the court comes out with a ruling that says LLMs, just the fact of developing an LLM, is not copyright infringement, provided you have some kind of substantial protections. And it could come out with a test that says, okay, if somebody asks for Luigi or Moana, you know, anything that's like very obvious, that should be fixed. And there should be measures taken to address that. another possibility which we wasn't on the table is congressional action that's very possible so we actually have that that has happened where courts have i'm sorry congress has codified things i mentioned fair use in the 1976 copyright act with the internet we have the dmca and it's a very robust system right somebody asks you to take something down you know they can contest it so an amendment to the dmca also possible there's a california senator who um he was a cs major super cool um so california i think senator um or u.s senator from california who um is proposing ai regulation and that could be a possibility so that would mean whoever gives the most money to be a bit cynical here whoever gives the most money to their senators congressmen whatever politicians and has the most influence in the deepest pockets for these old people uh the geriatric geritocracy or something that that run uh washington you know that would kind of feel like it would be in favor of the copyright holders because copyright holders in the united states we really do protect them uh in a major way so they could just say listen you got to get permission full stop you know that's possible i do think you know like i said open ai has been very savvy and thoughtful in a lot of ways.

1:07:51And part of Sam Altman's sort of charm tour last fall was on this. Like they know who he is and he's like, look, I'm not Zuck. And I think he was very successful in showing that he's not Zuck. And so that's another possibility. I think the Europe did regulate AI and they were very proud of that. So far it has not been regulated here, but I think there will be a situation And this is Google made a bunch of good law on sponsored search. So initially, if you search for a term on Google, if you search for, you know, Acme, Acme's competitor can buy that. Or if you search for Ford, you can get a Chevy ad.

1:08:32And that was actually it wasn't clear that that wasn't trademark infringement. Google spent about 10 years litigating that issue and just won. And I've seen legal theorists make the case that part of the reason Google won is that judges love Google because it's so useful. And so I think similarly for ChatGPT, it's so useful for us lawyers in particular because we are stock and trade as words that either I don't think that OpenAI will lose here. I actually disagree with you. I think that OpenAI will win some key points. You know, they'll probably have to make some concessions. They'll have to have a copyright.

1:09:10you're in center and they'll have to, you know, have DMCA style things, but I don't think they're going to straight lose on the fair use, or at least not without going all the way to the Supreme Court. That's fascinating. I'm taking the other side of that. I think it's going to be, we're going to come down in favor of, for people who have at scale copyright libraries, you're going to need their permission ahead of time. And if you've used it, I think you're going to have to unwind it, which is what I agree with you that they're probably in the process of doing that. Sam's pretty smart. And I think it's easier to just be like, you know what, we took it out, we redid it it's no big deal to retrain our model without new york times they could do that there's been some memes on that on the um you know the like square jaw meme and it's like take my content out and the guy's like fine yeah okay sure um that could be possible yeah i think that's a that's a distinct possibility i do think you bring up a really good point which is having seen up close and personal what happened with uber and also airbnb i wasn't an investor in that one unfortunately because people loved the service so much and became addicted to it by the time the lawsuit started to pile up when austin got rid of the city of austin got rid of uber and lyft at one point people went nuts and then when campaigns i mean it was amazing yeah and their public policy it was incredible like doing that the one distinction i would make so uber is a great example where like i even tell i advise my my clients like product market fit is an incredible drug right it really makes your lawsuits better and honestly uber every time they had a lawsuit their usage just went up so it's like this by the way also for uber yeah this one what's slightly different is uber was like in the trenches city by city versus this is federal so you run into some of the the gerontocracy or whatever issues that you mentioned airbnb the same thing you know people were like well i want to have choices of where to stay and i want to be able to monetize my home or my second home or my guest house and it just felt like those companies were on the right side of history vis-a-vis consumer choice lower prices etc and i think that's what chat gpt really has going for them which is we all want to be able to make luigi characters and make a birthday card for you know our family or make a party invite that has the silver surfer and marvel characters on it so if that's the case we're kind of like well that's kind of the world we want is where we get to use your copyrights without your permission it totally is but i mean this really gets like it goes back like kind of way back machine like i remember when i was at yahoo and yahoo was trying to launch like a subscription or a paid e-card service and i was like i'm not gonna pay for that i can get that free on blue mountain and it just happened to be that ads is what emerged.

1:11:55But one of the things that Europeans point out, and I think a lot of thinkers point out is that ad supported tech is not necessarily how it had to be. No, it could have been another way. And so, you know, similarly here, like it could be a subscription, it could be a licensing, like there's many different ways. And it'll be interesting to see, like, eventually the law catches up. It just like you said, takes time. I think now there's really very clarifying to have you on the program because the market-based solution is the likely case here so i think that's where i come to after an hour with you market-based solution always the best solution parties get around a table and and hash it out and then there's of course some liability for the mistakes that open ai made so they pay a speeding ticket they give them 100 million dollars as part of this new thing no harm no foul they can afford it they got 10 billion laying around it's all good but i like the market-based solution and i think there's something very interesting in how the cable tv system worked or how bundling and subscriptions work now and authentication because chat cpt knows how to do that sam waltman the team over there know how to do that like they already have api keys so the new york times subscription is like an api key to unlock some things in the new york times right it could be like a really cool feature like maybe the open ai markets to people hey you if you have a new york times subscription this is going to get a lot better for you because when you ask your queries it's going to give you a bunch of stuff and say and also from the new york times for further reading boom boom boom boom uh and would you like us to bookmark this is a world of possibilities of how open ai could work with you know new york times to make interesting stuff vis-a-vis recipes hey you i want to ask it about recipes here's the i took a picture of my refrigerator and then it went to the new york Times Food app and told me possibilities of what I can make based on my spice draw.

1:13:45You know, that's kind of interesting. It's amazing. And that was one of the things, you know, I mentioned that I spent, you know, three and a half years at Amazon, and that's kind of how Bezos thinks. And it was woven into everything where it was like, okay, how can we have this tech work together? How can we get paid for one piece of content multiple times? How can we turn something into self-serve? Like, you know, I loved Marc Andreessen's AI piece. Obviously, he's super in the AI optimism phase, but I'm very optimistic, too, for all these, like daily life fun use cases so yeah good stuff as career technologists you and i it's pretty clear that this is the one this is the chosen one this technology is the manifestation of everything that's come before it from the pc revolution to the internet to mobile and cloud and all this and then big data all of this is built up to this moment in time and so it's really important we get it right chuchilia you are amazing where people find more of you yeah so you can follow me on twitter it's chuchilia zinn um or i am pretty active on linkedin as well i'm launching my own startup which is ai for lawyers yeah oh i know and i know an angel investor i know an angel investor yeah he's really good at getting you your first hundred customers amazing yeah no it's excellent does that have a name yet or um yeah it's gonna be general counsel AI.

1:15:03So that's who I am. And I thought, okay, GCAI, but love it. Essentially still sort of, I guess, you know, stealth because we're developing the product, but I have an engineering co-founder and we're, we're pretty well, well there, but to the point that we talked about, like LLMs are wordsmiths, right? So what are lawyers also wordsmiths? So I've had, you know, when you said that this is the chosen technology, I had that feeling very strongly. I've never been excited about legal tech before it's like you know clm snooze but this was actually like management exactly you know but like this is something where to the cnd point i could write a cease and desist letter and i just say like here's the here's the infringement here's the whatever and make very light edits and a one-hour task becomes a five-minute task not even and i give training classes for lawyers you can find me on maven i know you you're friends with gagan and them too So I teach on Maven because I just have so much energy that I got to get out.

1:16:01Right. So fantastic. Everybody check out Maven and we'll put some links in the show notes. You are awesome. Please come back. We should do a check in when we. Yeah, let's do it. This is so fun. Okay. Happy to do it. Have a good one, Jason. Thanks a lot. And we'll see you all next time on this week in Star Wars. Bye bye.

From the publisher

This Week in Startups is brought to you by…

MEV. Tired of the dev shop rollercoaster? Mev is your reliable technical partner, offering a well-established software development process designed to consistently deliver unparalleled value to their clients. Get $30,000 off your first three months at http://www.mev.com/twist

Northwest Registered Agent. When starting your business, it's important to use a service that will actually help you. Northwest Registered Agent is that service. They'll form your company fast, give you the documents you need to open a business bank account, and even provide you with mail scanning and a business address to keep your personal privacy intact. Visit http://www.northwestregisteredagent.com/twist to get a 60% discount on your next LLC.

The Paintbrush Loan is the earliest startup financing on the internet. No pitch deck, no business plan, no minimum time in business, and no warm intros. Plus, you get to keep your equity. Visit http://www.getpaintbrush.com to see if you qualify for a $50K startup loan in less than 2 minutes.

*

Today’s show:

Cecilia joins Jason for an in-depth discussion about the New York Times versus OpenAI case, delving into the intricacies of fair use and analyzing the Fair Use Test (6:07), examining the legal complexities surrounding data scraping for Large Language Models (28:21), exploring possible ramifications of this legal confrontation between these titans (40:06), and more!

*

Timestamps:

(0:00) Cecilia Joins Jason

(2:44) Cecilia’s Background in Law

(3:44) Jumping into the case of NY Times vs OpenAI.

(6:07) Exploring Fair Use legal tests

(11:57) MEV - Get $30,000 off your first three months at http://www.mev.com/twist

(13:50) The case of Roy Orbison vs Two Live Crew and the music industry’s rules on fair use.

(19:04) Picking apart the defense of attribution.

(22:22) Northwest Registered Agent - Get a 60% discount on your next LLC at http://www.northwestregisteredagent.com/twist

(24:43) Fair Use Test: Factors two and three

(28:21) Legal challenges in data scraping for LLMs

(32:40) Paintbrush - Visit http://www.getpaintbrush.com to see if you qualify for a $50K startup loan in less than 2 minutes

(34:19) The fourth and final factor in the Fair Use Test.

(40:06) Potential outcomes of NY Times vs OpenAI case

(47:04) Google vs Java and legal discussions on digital platforms

(51:49) Jason shares a possible solution to this case and how the subscription wall could change things.

(57:54) Cecilia’s Grinch images on X

(1:02:22) Legal viewpoint regarding commercial vs non-commercial use.

(1:04:14) Summarizing where all this is going with the NY Times and OpenAI trial.

(1:12:18) Reviewing the market-based solution to this case.

* Check out GC AI: https://getgc.ai

Website: ceciliaziniti.com

Check out Ziniti Law: ⁠https://www.zinitilaw.com/

Check out Cecilia’s Maven Course here: https://maven.com/ceciliaz

*

Thanks to our partners:

(11:57) MEV - Get $30,000 off your first three months at http://www.mev.com/twist

(22:22) Northwest Registered Agent - Get a 60% discount on your next LLC athttp://www.northwestregisteredagent.com/twist

(32:40) Paintbrush - Visit http://www.getpaintbrush.com to see if you qualify for a $50K startup loan in less than 2 minutes

*

Follow Cecilia

X: https://twitter.com/CeciliaZin

LinkedIn: https://www.linkedin.com/in/ceciliaziniti/ *

Follow Jason:

X: https://twitter.com/jason

Instagram: https://www.instagram.com/jason

LinkedIn: https://www.linkedin.com/in/jasoncalacanis

*

Great 2023 interviews: Steve Huffman, Brian Chesky, Aaron Levie, Sophia Amoruso, Reid Hoffman, Frank Slootman, Billy McFarland

*

Check out Jason’s suite of newsletters: https://substack.com/@calacanis

*

Follow TWiST:

Substack: https://twistartups.substack.com

Twitter: https://twitter.com/TWiStartups

YouTube: https://www.youtube.com/thisweekin

*

Subscribe to the Founder University Podcast: https://www.founder.university/podcast

More from This Week in Startups

All 653 episodes
AI on Trial: Inside the NY Times vs. OpenAI Lawsuit with Cecilia ZinitiThis Week in Startups · 1 h 16 min
Listen in VO