In short
Podcast Notes: This Week in Startups - How AI Startups Can Navigate the Legal Landscape with Adam Shevell
Episode Overview In this episode of This Week in Startups, host Jason Calacanis is joined by Adam Shevell, a partner at Wilson Sonsini, to discuss the legal landscape pertaining to AI startups. The focus centers on gaining access to datasets, understanding copyright implications, fair use, and the intricacies of intellectual property in the context of AI.
Key Topics Discussed
- Gaining Access to Datasets (1:10)
- Importance of datasets as a lifeblood for AI models.
- Different sources of data: universities, open-source software (e.g., GitHub), and the internet.
- Legal risks associated with data scraping, particularly regarding copyright infringement.
- Copyright Infringement
- The necessity for startups to understand the implications of using copyrighted materials, especially well-protected assets like Disney's IP.
- Example: Getty Images suing Stability AI for unauthorized scraping of images.
- Fair Use Concept (6:29)
- Explanation of the multi-part test for fair use in the U.S.
- Four main factors considered:
- Purpose and character of the use (transformative nature).
- Nature of the copyrighted work (creative vs. factual).
- Amount and substantiality of the portion used.
- Effect on the market for the original work (market competition).
- Advice for New Founders (11:26)
- Importance of understanding legal risks and crafting a coherent strategy.
- Navigating the complexities of data usage and IP rights is crucial for securing venture investment.
- Key Legal Cases Discussed
- Getty vs. Stability AI: Scraping images without permission, focusing on copyright infringement.
- Andy Warhol Supreme Court Case: Importance of understanding how market impact alters fair use considerations.
- Strategies for Startups (17:42)
- Emphasizing the importance of getting permission when using third-party data.
- Highlighting the risks of scraping and the potential legal repercussions.
- Considering a pragmatic approach to risk assessment and data usage.
Key Takeaways
- Understanding Copyright: Most data available online is generally under copyright protection unless explicitly stated otherwise.
- Fair Use is Subjective: Determining fair use is context-dependent and often requires a court's interpretation.
- Legal Precautions are Essential: Startups must approach data usage with caution, balancing innovation with legal obligations to minimize risks.
- Market Impact Matters: The competitive nature of the use can significantly influence the legal standing concerning fair use.
- Consulting Legal Experts: Engaging with a qualified attorney is crucial for navigating the complexities of copyright and IP rights.
Conclusion This episode highlights the delicate balance AI startups must maintain between leveraging data for innovation and adhering to legal frameworks surrounding copyright and intellectual property. As technology evolves at a rapid pace, understanding these nuances is vital for sustainable business practices in the startup ecosystem.
For more information, visit [This Week in Startups](https://thisweekinstartups.com) and check out the legal basics series for additional insights.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Hey, everybody. Welcome back to one of my favorite series to do here at this week in startups. We call it our startup legal basic series. We do it with what's considered. the top firm one of the top firms uh in technology and startups in capital allocation wilson cincini now wilson cincini's got a lot of talents he got a deep bench i work with them they're uh my attorneys and we had a very niche discussion we wanted to have for this year's startup basics around ai and so when i was talking to becky degras who usually does this series with me and you could see the archive at thisweekinstartups.com slash basics she said you know i got an expert for this topic look can i bring in uh go deep on the bench and adam chavelle is a partner at wilson susini he works on ip what's ip that's intellectual property and ai is coming up big time you've heard me talk about this both on the all in podcast and this week in startups of hey what should the framework be if i'm training an ai model on somebody else's data set this is an emerging topic welcome to the program adam chevelle thanks for having me it's really great to be here all right so i got a lot of startups they're looking at data and they're saying i got a great idea look there's all this disney uh ip in the world can i take all the marvel scripts and comic books i find and then make a new character out of them hey i would like to uh do something in recipes can i go to condé nass website or some recipe database or uh some famous chef mark strousman's incredible recipes from mark's off madsen in new york can i go take his lasagna and um use it to iterate on lasagna recipes what's the answer here yeah well i know it's a moving target i'd be all over mark bitman's recipes but yeah no it's a great question, right?
1:58You know, getting data, access to data is like the lifeblood of anyone who's building a model, right? You need that huge pool of data to make it execute really effectively and do that magic that we see when we go on to the tools that we have access to, right? And there are different places you can get them, right? Like universities have databases that they've compiled over the years. The challenge with those is typically they're non-commercial. So, if you want to make a product that you're going to sell eventually, you can't really use those, right? It's against the rules. Open source software, right?
2:30GitHub repositories, there's tons of data there, it's available. And then obviously, the internet itself, massive treasure trove of data. But the challenge is, if you go out there and scrape data, there are a bunch of legal risks you're taking, right? And so you really should go in with your eyes open to understand, hey, here are the risks we're running, there might be ways we can mitigate these risks. but uh it's not a a risk-free opportunity right um the biggest one is copyright infringement so you mentioned disney marvel obviously they have fabulous uh ip assets that they've built over decades um they won't take it lightly if you go in there and grab their stuff you know they don't like you know they they are aggressive and rightfully so because they have a really valuable uh asset to protect right so go ahead yeah i mean and they have spent billions of dollars acquiring some of those marvel star wars uh and pixar were multi-billion dollar acquisitions that then had multiple billions of dollars put into them and my understanding is uh these are all protected ip there is a concept and and disney has actually just to pick them specifically they have worked really hard with ip law here in the united states to protect those characters and i think some of them like the original mickey mouses there's this concept in ip law of like 75 years since the creator or something and i'm not up on that exactly right now but if you're taking data off the internet by definition it's under 30 years old so it's going to be protected unless somebody explicitly didn't protect it am i generally correct in my framing here Absolutely.
4:11So for the purposes of what we find on the internet, almost all of it is copyright protected. Now, there are artists, let's think of like Vincent van Gogh or Mozart, who lived a long time ago, their copyright protection is gone. And so their work is in the public domain. But stuff like Disney, like that's still copyright protected. And copyright law gives the owner some exclusive rights, right? No one else can reproduce that IP, that work of authorship. No one else can prepare derivative works of that authorship. distribute them publicly before them so basically it gives them a monopoly over using that that art and that that all that that work of authorship right um so if you're scraping let's just use disney's art or or data as an example if you are scraping their site and pulling their art down you know you're reproducing right which is a violation of copyright in order to train your model right and actually there's there's a lawsuit currently getty is suing stability ai And the argument there is when Stability AI scraped Getty's website, they reproduced without authorization and they violated Getty's copyright.
5:16That's the argument they're making. It remains to be seen how that goes. Yeah. The fact is when the examples given in that suit and that claim literally showed the watermark from Getty. It did. It did. this was i would say not a very thoughtful execution on the part of stable diffusion and the best practice always in law because i'm a content producer for many years as a journalist is when in doubt get permission now there is a concept of fair use that's right i have had long i have a long history with this um just as a content creator again um if people want to understand fair use they need to understand this is a multi-part test here in the united states again you can correct me where i'm wrong but i try to explain this to my founders this is a multi-part test that is subject to the interpretation of the courts it very rarely gets to a decision because people tend to have a lawsuit or legal letters or a debate over do they feel that you're being fair in your interpretation of fair use so let's give a little primer to people listening of the multiple part tests of fair use and and where it's obviously applicable or where people maybe are selectively interpreting fair use incorrectly sure so um the the biggest challenge with fair use is that you don't know what fair use is until a court tells you what it is each time is different right you there's no uh there there are there's a test and we can go through the tests and tell you what different parts are but ultimately if uh the case of stable diffusion is or is not fair use is going to be cited by one judge maybe it gets appealed but but essentially the court's going to decide and if there's another case that getty sues let's say a different company in a different case that that would be a different a test and we'd reapply the same legal test to the different facts, right?
7:21So it's factually dependent. The four main factors for fair use that the courts weigh are the purpose and character of the use, right? This one gets a lot of attention, right? Is the new use new in some way? Are you adding a new transformative nature to that work of art or work of authorship, right? The second is the nature of the copyrighted work. A technical manual for your dishwasher gets less protection than, let's say, a novel by Dickens, right? So, the artistic merit of the actual work weighs in. The amount of and substantiality of the portion you've copied is also important. If you took a small snippet, as opposed to the entire work, that will weigh differently.
8:06Is that fair? Are you taking a very small piece? Or are you just kind of taking the whole thing and copying the work and doing something with it? lastly this yeah it was very interesting because you know i am obsessed with fair use as a content creator i remember there was like a famous documentary and it was at sundance and they were talking about you know the rating system and they came out i remember talking to some attorneys it would come to be the name of the film but it's not important they wanted to use certain scenes um to explain this point and they up front because usually when you do apply for distribution for a film they want to see that every single clip has been cleared they said nope this is a documentary it is fair use we're using it and i'm jumping the gun here a little bit sure for educational purposes etc we're using a tiny portion of these original films and we will not um get permission in advance and they did not get stopped because again with fair use you have to have somebody on the other side as you very clearly said a judge has to decide it's open for interpretation the other person has to feel it's worth going to bat to protect their ip and if a documentary film wants to use five seconds or 10 seconds to make a point that's educational and it's obviously uh you know in context uh you know the the other copyright holder might not even feel this is an existential threat to them in any way and they don't take action so yeah well actually that's a great story that's the pragmatic part of it so go ahead yeah yeah i know it leads into the fourth factor which is the effect on the copyrighted works market right?
9:38If you're taking the original work, making a different version of it and selling into the same market, you know, that first owner is going to lose market share, right? You're a competitor. So if you're competing with the original owner, that's also going to weigh against fair use. If you're taking it and putting it to a completely different use and it's transformational, then that's more likely going to be determined as fair use. But again, each time is fact specific, right? And what's interesting is, you know, you talked about the parties really having the desire and resources to go all the way and fight the fight for fair use you know a really high high profile case was the oracle v google case right with the java um apis and oracle sued google in 2011 okay or 2010 and it took almost 11 years to get to the supreme court ruling where the supreme court ruled that google's use of the apis in this instance didn't say all use of copying API, you know, command lines are fair use.
10:37But in the systems, it was fair use. It took 11 years. And you had two companies with some of the biggest resources, biggest pockets on the whole world. I think we know who won that case. We know who won. The attorneys won that case. The attorneys did fine, I'm sure. However, the, you know, but if you think about 11 years, okay, from start to finish, how does that match up with the pace and rate of change in generative AI? it's so it's so mismatched like how can the court systems be relevant with the rate of change that's happening in in the ai industry it's it's it's it's going to be really interesting to see because a lot of these questions like when i advise my startups that my clients like hey can we do this the answer is always like i i don't know yet we're going to find out here here are the factors that go into it yeah see this is the frustrating part for i think a lot of young individuals starting companies who haven't had to deal with this issue and an attorney cannot tell you here's the bright line and here's where you're good and then here's where you stepped over the line and what you have to then take into account is the totality of the situation and a framing i use and i'm curious uh what you think of this is it relates that fourth part of the test which is hey is this going to affect the original copyright holders ability essentially to uh monetize to exploit their their creation in the future now if you look at what stable diffusion did this will be my opinion i don't speak for you or wilson sancini but my opinion is it's obviously uh going to affect it because getty could create their own stable diffusion which is by the way based on an open source project, they could create their own product that builds off of that.
12:23And the IP holders of those original photos, if they are in a revenue sharing deal, would be able to monetize that. So of course, they're infringing on it. But a more simple step to take is how fair does the person on the other side of your innovation feel about it? And when you look at Google, their use of content with snippets, you know, they're only using a snippet of your website's information that felt fair to people and there was traffic being sent to you therefore you very rarely saw anybody saber rattle or threaten to sue google for putting a little snippet because there was a blue link directly to your website and it essentially was like a little amuse bouge that you know first 500 characters and by the way you could opt out of it too you can do robots.thc which facebook did which craigslist does so let's talk about the fair in fair use and how founders that you work with, you could kind of walk them through how to think about this and avoid problems in the future.
13:24Yeah, sure. So, you know, as you allude to, the test is technical and almost no one ever finds out if it actually is or is not fair use, right? You make a pragmatic decision balancing all the risks. And what you mentioned, how impactful is this going to be on the company or the individuals that we are reusing? their original material, are they going to be impacted negatively? Are they going to dislike that we're in the market doing what we're doing? Or actually, they might like that we're channel for them in the case of Google and the snippets, right? So that's a real sniff test as far as are we going to come down the line and in a couple of years have people throwing claims at us or not, right?
14:07And the more competitive your use is with the original data owners market, and the more competitive we are with them. I mean, it's just common sense. You're running a higher risk of running into issues. And the other thing is it's important to have a coherent strategy, whatever you do, right? Like there are lots of different ways to get there, but being able to articulate what the risks are, how you balance them and why this balance and approach is right for your company is not good for the company. But when you get into term sheet land and you have a venture firm coming to invest in your company, they're going to want to hear this strategy, right?
14:47They're not dumb. They have lawyers who are looking at this very intently. And so being able to articulate why you're taking this balanced approach to this risk is going to be really important for your company's success if you're looking for venture investment. What are the major cases right now? I know there's a GitHub. Open source folks are suing GitHub for their co-pilot. but explain that one maybe in broad strokes. This is really interesting, right? Because it has to do with open source software. And for those who know open source software, basically it's software that the owner has made available to the public to use, right?
15:23So on its face, you think, okay, well, here's some source code. So-and-so wrote it and they've released it to the public, right? So the arguments in this case, where you have some software developers suing GitHub, I think OpenAI and Microsoft is in there too, because Microsoft now owns GitHub. They're arguing that even though the open source software is made available and anyone could grab it, download it, fork it, develop it, whatever they want, there are still licensed terms that apply, right? And even the most permissive licensed terms have some requirements that weren't being followed by co-pilots and the developers.
16:00And the argument is, they allege that there is both a duty to notify future users that part of the software was copyright so-and-so, whoever wrote it, right? And so by ingesting all the software and then producing new software that could match, you know, word for word, line for line, the old software, they weren't providing that copyright attribution notice that they're required to. So there's a breach of contract, right? You had a license with me. I gave you a license to use my code. And one of the few requirements was just tell people when you distribute it attribution yeah it's giving credit it's a core tenet of open source if you want to use this you just gotta and i think they even say in their license link back and give credit and there's creative commons in open source and there's a very granular thing this would be de minimis for microsoft to actually put links in fact i was so vocal about this with open ai that i noticed with open ai's web crawler with bing they now will from time to time put a citation in and i was using bard the other day they now have images and thumbnails in there and they give credit to yelp and they linked to yelp again back to the fair in fair use it's nonsense that these ai models cannot point to where they got the information from if properly constructed you could say some of this information came from here some of it came from there um it's kind of nonsense to say you can't give a link and so i think this one will be solved with the links and credit being given to folks and also permission let's talk about permission the gold standard is to get written permission from people in advance of doing stuff nobody in technology likes to do that they like to beg for forgiveness so what is a proper strategy for folks is it to beg for forgiveness Is it to ask for permission to not kick the hornet's nest as it were and just put out a little experiment without monetization and see how the market responds to it?
18:09What's the pragmatist approach to this? Yeah, sure. So it really depends on the data source. There are certain data sources that are known to be challenging. Like for instance, Craigslist. You cannot take Craigslist's data. They will fight you tooth and nail. It's a matter of philosophy and principle for them. And so in that case, ask permission or don't do it because they will hunt you down if you take it without permission and they will get you as best they can. Right. So understand whose data you're using. Craigslist does not stop. They've been very clear from the beginning. Yes. And then also you're seeing a shift in the market from companies with huge pools of very valuable data.
18:49Like look what happened with Reddit, right? Yes. They went from a free model to a pay to use model. From their perspective, that makes a lot of sense because, hey, our data is now so much more valuable with the huge rush to make these models that like we're a for-profit company, we need to make some money and here's an asset we could leverage, right? So you also have to know like how is the market moving, right? And so there, they want to give permission. They want you to pay for it, right? Rather than coming and taking it without their understanding. understand. And so going back to your initial question though, you know, as long as what you're doing is sort of measured in steps, like the challenge is if you do a test and maybe you don't ask permission and maybe you take some data, you run the model, you see how it works.
19:33You're like, Hey, this is great. Let's keep doing it. If you get to a point down the road where someone sends a cease and desist letter, you don't want to be at a point with your product development where you can't put the genie back in the, in, in the lantern, right? You want to be able to stage how you're using the data in a way so that if you do run into a claim down the road, you can kind of, I'll say, hide your tracks a little bit if you can, right. And, and sort of, um, be able to, um, take away maybe the data that they're complaining about from your product without completely ruining how it works.
20:04Right. And so just, just thinking about contingencies, really, um, you know, I wouldn't necessarily go out and ask permission each and every time, even though that would be the legally correct thing. I think pragmatically for startups who are resource strapped and have to be a little bit scrappy, sometimes you need to take a little more risk. And all I say is just take that risk very calculatedly. Be thoughtful. You know, this is, I think, great advice. I can give really granular examples here. People want to do all kinds of things with the archive of this week in startups and all in. And one of those things a couple of years ago was, hey, I want to make clips on TikTok.
20:41And we had a couple of people who wanted to do clip shows. And I said, yeah, I tell you what. And they contacted me. Is it okay? I said, yeah. If you're a fan of the show and you want to do a couple of clips, you know, go for it. Just always link back to the original episode and just say that, you know, you're not the copyright holder, the copyright holder is whoever. And, you know, let's check back in a year or two and see where it's at. Now, some of these things have become big, and they're not monetizing. So I'm like, okay, fine. And then in some cases, if they want to monetize, I might be like, well, what is it making?
21:10Just keep me informed. if you're making ten thousand dollars a year or less and you're putting hundreds of hours into this i guess i'm okay with it you know it's promoting the show but then i had this one group that was taking the entire episodes um and then using ai to put them into sections clipping them and then putting ads around them and i said no bueno uh this is not fair and uh and they were doing it with tim ferris and some other folks and i just told tim and other folks like hey do you know this is happening and they were like this is bs and i said the person listen if you want to not put ads on it and you're using the original mp3 file i'm okay with it and you give credit and you link back to the original show in the interface and their interface was always like this week in startups next episode or this week in startup you know and no links back to us or they would make a tiny little link and this is where i think you know if you wanted to do an experiment okay do it but think about the person who put the effort into that content and how you could be fair with them link backs credit using the original mp3 file it's a very subtle point here but we get the data on that person listening to it if you rip the file and put it on your server i don't know how many views you're getting my advertisers don't know i don't i can't cookie the person if they're i have cookies turned on so i like your strategy here you can take a little risk you can do a little experiment i think turning off monetization on these things is also critical and framing them as an experiment and taking feedback in good faith um super important yeah all right listen this is incredible get a good lawyer if you're doing this because it's going to lead to a lot of discussions man i just sent you a link in the in the chat craigslist has gotten judgments against people i was reading this headline here craigslist garnered 60 million dollar judgment against rad pad and scraping dispute i think just be careful if you're doing this kind of stuff that this could be the end of your startup um pragmatically you raise a half million dollars or a million dollars you're coming out of tech stars or y combinator or launch accelerator and you get hit with one of these that's it game over for your firm you're you you have seen this happen startups get crushed with the legal bill freezes future funding so this could be existential be the end of the line for your startup yeah absolutely adam i was super fascinated by this one i could talk to you all day this is supposed to be like a short segment but i gotta bring it up andy warhol lost in the supreme court his uh tomato cans or whatever it was that were based on photos uh i can't remember which one it was that they actually yeah it was not transformed it was it was a yeah it was a photograph of prince right and and in the 80s, Andy Warhol had made a painting, a silkscreen or some form of reproduction based on this photograph, right?
24:06And what happened was his estate sold that, licensed that photo to a magazine as part of an article about Prince and the original photographer sued, right? And what was interesting here is the court really focused on the fourth factor, which is the impact on the market right they basically said this licensing of the warhol for use by journal like in a photo journal in a magazine is directly competitive with the intent and use of the original photograph right she took a photo for a magazine and so uh right direct competition right and it really took away the focus from the first factor which is whether this use is transformative or not And so if I read between the lines, I think this makes it a harder climb for companies who are scraping and training models on internet data to reuse fair use.
25:00because the real crux of that argument is this is transformative, right? There might be an image or data or there might be a conversation. What we're doing is changing all of that into this really complex computer system, this model, and then creating a predictable outcome later. That's like the original author had none of that in mind. And so historically, the Supreme Court and courts have really focused on transformative use as one of the prevailing and strongest factors whereas um you know this competitive you know is the market competitive or not uh this factor is um kind of they've risen its importance and i think that cuts against the the developers of generative ai models yeah you got to be really thoughtful about this if you're taking tarantino scripts and then you write a tarantino like movie and uh you know that's all clever and good except tarantino still making movies i think he's going to do one more so you literally it's not the intent now if you're inspired by somebody um and you know quentin tarantino was literally inspired by a lot of the exploitation films of the 70s uh and he is doing an homage to them in jackie brown in kill bill uh in inglorious bastards these are literal homages but it's so transformative that for you to know that this reference is from the killing and this reference is from this 70s film with pam greer that you would have never even heard of and it's such a minor inspiration of this one piece of dialogue or that piece of dialogue there's no like it's not interfering with the person who made that film in the 70s and in fact that person who made that film in the 70s when people do figure out that it had some inspiration to jackie brown would go seek it out and it would make more money so you're but the andy warhol one is like it became a cult thing it became a cult thing but the andy warhol one is like should i buy the photograph or should i buy this colorized version that andy warhol didn't transformed you're like it's an either or yeah it's so easy note that the supreme court was was really only focused on the licensing of the image the licensing of the work not on when he made the original work the derivative work from the photograph so what is that mean so they haven't ruled they haven't said that it wasn't fair use when warhol made the um work of art as a derivative work off of this photograph they just said using it licensing it versus licensing it is not fair use so he could make the painting he could put it on his wall perhaps even sell the painting one time sell it at auction one time no problem but licensing it too because if you did make the painting and sell it at auction one time did that really screw up the original photographers you might be able to make an argument if they were making prints maybe of prince prince of prince uh maybe but anyway rest in peace prince man and by the way just speaking of prince i know you're a music lover like me he does a guitar solo at the rock and roll hall of fame of while my guitar gently weeps oh yeah and if you haven't seen prince do this solo uh because you know we'll leave this in the episode he does just type in while my guitar gently weeps and you watch this with tom petty steve winwood and jeff lynn somebody had said like i don't know if it was rolling stone or something somebody had said prince was like overrated as a guitar player and he was like oh yeah we'll see you at the rock and roll hall of fame and he gives a mark knoffler you know level performance of a of a guitar solo that breaks the internet this video is so good i'm not gonna play it here because i don't have fair use i don't want to take away from the copyright of the rock and roll hall of fame but this video has got 123 million views it should have a billion all right listen if people want to get in touch with you they got this issue you got an email you got a way for them to get in touch with you adam yeah sure uh it's uh a cheveld at wsgr.com um a s h e v as in victor e l l at wsgr.com perfect thanks for doing this i appreciate you sharing all the wisdom and giving really pragmatic advice to the startup community and if you want more information on this this week in startups.com slash basics for more information on our basic series we do it every year uh thanks so much to my friends at Wilson Sonsini for supporting this and doing it with me on the most important topics for founders.
29:41And we'll see you next time on this week at startups. Bye bye.
From the publisher
Today’s show:
Wilson Sonsini partner Adam Shevell joins Jason to discuss gaining access to datasets (1:10), fair use in copyright cases (6:29), and more!
*
Time stamps: (00:00) Wilson Sonsini partner Adam Shevell joins Jason (1:10) Gaining access to datasets, avoiding copyright infringement, and why IP is protected (6:29) The concept of fair use and when it applies (11:26) What new founders should pay attention to (14:59) The Github case (17:42) Proper strategy for getting permission (22:56) How startups get crushed with legal bills (23:32) The Andy Warhol supreme court case * Check out Wilson Sonsini: https://www.wsgr.com *
Read LAUNCH Fund 4 Deal Memo: https://www.launch.co/four
Apply for Funding: https://www.launch.co/apply
Buy ANGEL: https://www.angelthebook.com
Great recent interviews: Steve Huffman, Brian Chesky, Aaron Levie, Sophia Amoruso, Reid Hoffman, Frank Slootman, Billy McFarland, PrayingForExits, Jenny Lefcourt
Check out Jason’s suite of newsletters: https://substack.com/@calacanis
*
Follow Jason:
Twitter: https://twitter.com/jason
Instagram: https://www.instagram.com/jason
LinkedIn: https://www.linkedin.com/in/jasoncalacanis
*
Follow TWiST:
Substack: https://twistartups.substack.com
Twitter: https://twitter.com/TWiStartups
YouTube: https://www.youtube.com/thisweekin
*
Subscribe to the Founder University Podcast: https://www.founder.university/podcast




