In short
Podcast Episode Notes: Practical AI - Legal Consequences of Generated Content
Episode Overview
- Title: Legal Consequences of Generated Content
- Description: Discussion with Damien Riehl, an expert in technology and law, about the implications of generative AI, copyright issues, and the evolving landscape of knowledge work.
Key Guests
- Damien Riehl: Technologist, lawyer, and creator of the All the Music project, which generated every possible musical melody.
- Chris Benson: Tech strategist at Lockheed Martin.
- Daniel Whitenack: Data scientist and founder of PredictionGuard.
Episode Highlights
- Introduction to the Guest
- Damien’s dual expertise in law and technology allows him to navigate the legal implications of AI.
- Brief background on Damien’s experience in litigation and software development.
- Current State of AI Regulation
- Regulatory bodies (like the EU and U.S.) struggle to keep pace with rapidly evolving AI technologies.
- Historical context of regulation has often resulted in ineffective measures due to lawmakers' lack of technical understanding.
- Damien expresses skepticism about the effectiveness of future regulations.
- The All the Music Project
- Damien and his colleague generated 471 billion melodies using brute force algorithms.
- All generated melodies were placed in the public domain to challenge existing copyright norms.
- This project raises questions about the nature of creativity and ownership in machine-generated content.
- Human vs. Machine Creativity
- Discussion on the ambiguity of creativity—whether generating melodies algorithmically is truly "creative."
- Comparisons of human-generated creative work versus machine-generated outputs.
- The implications of copyright laws on machine-generated content.
- Legal Implications of Generated Content
- Copyrightability of machine-generated content remains a contentious issue.
- The distinction between human creativity and machine output, and how this impacts legal frameworks.
- Future litigation may arise from the intersection of AI outputs and copyrighted works.
- Practical Considerations for Developers
- The ethics and legality of using copyrighted material in AI training datasets.
- Insights into transformative use in relation to large language models (LLMs) and their outputs.
- Encouragement for developers to remain engaged with legal developments and adjust practices accordingly.
- Future of Intellectual Property
- Speculations on the diminishing value of traditional IP in the face of AI advancements.
- Potential shift towards an abundance mindset as opposed to scarcity due to generative capabilities of AI tools.
- The Impact on Labor and Productivity
- Various scenarios are proposed regarding the future of work, AI integration, and potential job displacement.
- The importance of adapting to AI tools to enhance productivity and stay competitive in the workforce.
- Acknowledgment of the dual nature of AI as both an opportunity and a potential threat to job security.
- Conclusion
- The conversation underscores the importance of understanding and adapting to the legal landscape surrounding AI.
- Encouragement for individuals to harness AI tools to enhance their work and prepare for the changing dynamics in various professional fields.
Key Takeaways
- Embrace AI Tools: Lawyers and developers alike should leverage AI to enhance productivity and remain competitive.
- Stay Informed: Understanding the legal landscape surrounding generative AI is crucial for responsible use and innovation.
- Anticipate Changes: As generative AI evolves, so too will the implications for IP, creativity, and the future of work.
Closing Remarks
- Participants express gratitude for Damien's insights and encourage listeners to actively engage with the evolving discourse on AI, legality, and creativity.
Resources
- Damien Riehl's Talks:
- [Legal and Practical Consequences of Generative AI](https://youtu.be/vy_569-Tmt8)
- [Why All Melodies Should Be Free for Musicians to Use](https://youtu.be/rjpTBHjeZ_0)
Sponsors
- Fastly: Bandwidth partner providing fast and secure digital experiences.
- Fly.io: Platform for deploying apps and databases globally without operational overhead.
- Typesense: Fast, globally distributed search-as-a-service.
---
This summary captures the essence and key discussions from the episode, providing a structured overview for those interested in the intersection of AI, law, and creativity.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:06Welcome to Practical AI. If you work in artificial intelligence, aspire to, or are curious how AI-related technologies are changing the world, this is the show for you. Thank you to our partners at Fastly for shipping all of our pods super fast to wherever you listen. Check them out at Fastly.com. And to our friends at Fly, deploy your app servers and database close to your users. No ops required. Learn more at fly.io.
0:43Welcome to another episode of Practical AI. This is Daniel Whitenack. I'm a data scientist and founder of PredictionGuard, and I'm joined as always by my co-host, Chris Benson, who is a tech strategist at Lockheed Martin. How are you doing, Chris? Doing well, doing well. I'm really excited about today because there's so many questions you and I have brought up in the show without the ability to answer. And I know we might get some answers today. Yes. And actually this one, not only will it be super practical and interesting, but it's also a tip from one of our listeners who suggested this guest.
1:21So we're really excited to have with us Damian Reel, who is a lawyer and technologist with experience in litigation and digital forensics and software development. So welcome, Damian. Thank you so much for having me. I'm thrilled to be here. Yeah, I feel very selfish this episode because I just have like a million sort of like legal implications, copyright questions related to like generative AI, large language models, all sorts of things. But before we get into some of those specifics, I know over the course of this show, we have commented on various things that have come about, like GDPR and then California data privacy stuff.
2:04And now we have like the EU AI Act and and all of this sort of regulation stuff. And then you've got other things on the other side on the litigation front where companies are, you know, getting sued for code generation based on maybe questionable training of models and other things. So maybe before we get in, from someone who is an expert in this area and thinking about it all the time, how do you view where we're at in relation to AI technology and regulation and kind of the legal side of things? How are those things catching up to one another or outpacing one another? And where are we at now as opposed to like maybe a year ago?
2:48What's changed? Sure. And maybe before I answer that, a brief, I litigated for about 20 years. So I was a litigator for about 20 years. I did tech litigation. So I'm coming to this from a perspective as a lawyer, but I've also been a coder since 1985. So I have the law plus tech background. So for your listeners' benefit, I'm not just a stuffed shirt that doesn't know what he's talking about. I can walk the tech walk and talk the legal talk, if you will. So really, as far as regulation, having litigated since 2002, I've seen ways that the EU and the United States have tried to regulate technology.
3:22And of course, they've had various degrees of failure, I would say, largely because the three of us know exactly how lots of technology works. But sadly, the congresspeople and the regulators do not. And so it's really the law is by nature slow and trying to get up to speed on a fast moving area such as AI is very difficult. So I would say that, you know, if past is prologue, I don't anticipate much good things coming out of regulation of AI in the near future. Some of these things are related to like this generative wave of AI where people are generating a lot of content with AI. I know that also you have a background, you mentioned your sort of coding background, but you also have a background with generative technologies, you know, maybe not like some with large language models and other things, but I know you have a very interesting story of some generative things that you did with music.
4:20Could you describe some of that? Yeah, absolutely. So both my current state with my job that pays me money, that is with Vlex, where I'm doing lots with large language models right now. We have a billion legal documents that we're running large language models across and doing embeddings to be able to do outputs. For example, a legal research memorandum and eventually be able to provide motions, briefs, pleadings, that sort of thing. So that's my job that pays me money. For the job that doesn't pay me money at all, which you referenced, is my All the Music project. This is a project that I started with my friend Noah Rubin, who one thing your listeners might find interesting is that around 2018, I did cybersecurity.
4:55The biggest thing I did was that Facebook hired me and my company to investigate Cambridge Analytica. So I spent a year of my life on Facebook's campus with Facebook's data scientists and my former FBI, CIA, NSA people that worked with me to figure out how bad guys use Facebook data. Did that for about 50 some weeks in a row. The stuff we would do on Monday would make the New York Times and the Times of London by Friday. So at the end of a 14-hour day on the Facebook campus, I retreated to my hotel with Noah Rubin, my friend. And we were in the lounge and over a beer, I said, do you know, Noah, how we can brute force passwords by going A, A, B, A, C?
5:29I said, what if we could do that with music where we would go do-do-do-do, do-do-do-ray, do-do-do-me, do-do-do-fa until we mathematically exhausted every melody that's ever been and every melody that ever can be? And so he said F, yeah, but he didn't say F, yeah. So within a few hours, he had a prototype where he cranked out 3 ,000 melodies. To date, we've now cranked out 471 billion melodies with a B. mathematically exhausting every melody that's ever been and ever can be. We've written all those to disk. Once they're written to disk, they're copyrighted automatically. So we've copyrighted 471 billion melodies.
6:01And then we placed everything in the public domain to be able to protect to use to all my melody lawsuit defendants. And so the idea is that before my talk in 2019, every defendant in one of those lawsuits has lost. After my talk, which has been seen 2 million times, every defendant has used largely my arguments and has won. And my arguments are largely that maybe if a machine cranks this thing out at 300 ,000 melodies per second, maybe we shouldn't give one person a monopoly, which is a copyright, a monopoly of life of the author plus 70 years for what the machine cranked out in a millisecond.
6:33So that's largely my All the Music project, which goes to, it is in a sense generative AI. It's a brute force generative AI, but it's generative AI. But the real question is, if the output of that is copyrightable, essentially if carpet-bombed the entirety of every melody that's ever been, and if I were a megalomaniac, I would sue everybody, right? But I'm not. I put them all in the public domain. But if machine-generated works are copyrightable, these are the bad things that can happen. I love that story. I just want to say that. It's pretty fun. And because it's been seen so many times now, 2 million, I've been able to meet with some good friends now.
7:08There's, for example, the former chief economist of Spotify. I'm now friends with. And the guy who was responsible for the first commercial MP3 to be downloaded, Jim Griffin, I'm friends with him. So anyway, so it's opened a lot of doors for me. That's awesome. And one thing that comes to my mind as you're talking about that is, because I'm also a musician, and I think a lot of people would think of, oh, these sort of melodies or chord progressions or however you want to frame it. There's a very human element to that that involves creativity. And I think this would maybe be extended to if we think about knowledge work more generally, whether that's like you being a lawyer and writing briefs, or us being programmers and writing code, or marketers being marketers and writing copy, right?
7:55Now these generative models can generate a lot of those things in a very compelling, and even I think people would perceive it as a very creative way, however you think about that creativity and coherence. So in your own work, maybe it's on the lawyer side or the coding side. How is that project and maybe some of the work you do day to day with large language models, how is that shaping how you think about this sort of knowledge work and the output of humans versus the output of models? I would say two aspects to this. And I think that the word creative is ambiguous, as all words are, most words are.
8:34And it is, you know, when I generated 471 billion melodies, I was creating those melodies, but it was by new means creative, right? So in that way, you know, those are just mathematically exhausting everything that's ever been. So that is a very simple version of what large language models do, right? Just saying, what is the statistically next sentence, right? And so this really gets to the heart of what we think human creativity is, right? Is it creative as in brute forcing or is it truly lightning in a bottle creativity that we want to protect with intellectual property laws and other things like that?
9:13I think what we're learning from my project and from the large language models is that maybe human creativity ain't as special as we think it is. As an example, on a Tuesday, a jury found that Katy Perry had violated a copyright of a melody that sounded like this. That was on a Tuesday. On my talk on a Saturday, I said that that particular melody shows up in my data set 8 ,128 times. So Katy Perry got dinged for$2.8 million over something that I had thousands of times just through brute force. So was that melody creative? No, it was just a brute force. After my talk was made public, the judge actually went back and reversed the jury verdict, saying that melody was so unoriginal as to be uncopyrightable.
10:00Essentially what I was arguing my TED talk. I think going to the heart of what is human creativity, it's good that we have large language models and projects like mine to be able to go to the heart of there are some things that we should protect and there are some things that are just unprotectable because they're unoriginal. You're putting an obstacle against what was otherwise maybe the weaponization of IP. Is that a fair way of putting that? 100%. We're using it as a shield, not a sword. Right. Which I like because it's having seen a lot of, for a non-attorney as having seen a lot of IP concerns out there in business, sometimes you're just kind of like, there's not much there.
10:33Absolutely. And so I did that with the copyright side. My friend, Mike Bomarito, who was one of the guys who you may have heard beat the bar exam. They used GPT-4 to beat 90 % of humans on the bar exam. One of those was Mike Bomarito. Mike approached me and said, I love what you did with copyrights. Wouldn't be great to do that for patents. And so right now we're doing, I was doing all the music project. We're now doing all the patents project. What that project is, is going to be taking all the patents that have ever been filed, taking each of the claims for each of those patents, putting those claims together in vector space and clustering them that way, and then generating every possible combination of all of those claims in all those existing patents.
11:13So if anyone in the future tries to recombine any existing claim into a new thing, they can point to our thing as prior art to be able to say, no, no, no, Bobberito, Real Cats, and maybe Rubin. They already did that in 2023. You can't do that again because they did that as prior art. As both a coder and a lawyer who I'm assuming some of these things you're generating are generated, but I'm also assuming in your day-to-day work and in your day-to-day coding, There are portions of what you're doing still that are not completely generated, or at least that you're editing heavily. How has this type of work influenced how you think about your own job moving into the future, working alongside these models or at least in an environment where these models exist?
12:04So I will say, and anyone who works with me will agree, that I am a coder, but I'm a crappy coder. I'm probably one of the worst coders you're ever going to meet. So a lot of my work with large language models is in the textual area rather than the code area. But in the textual area, I'll give you an anecdote that answers your question. I was reached out by the editor of a large legal magazine to say, hey, Damien, I want you to write the cover story on GPT. I said, how long do you want it to be? He said about 17 double-spaced pages. And I was like, man, I don't have time because my rule of thumb is one hour per double space page.
12:35So that's 17 hours that I just don't have time to do. But then I realized, oh, wait, the topic is GPT. So what I did is I created an outline of headings and subheadings. That's about two pages worth. And I said to GPT, for each of the bullet points, give me four sentences, essentially a paragraph for each of the bullet points. And it out. That was my one for the day. Crapped out the 15 pages worth. And then I spent the next three hours editing, moving, adding, working with the text, not accepting the 15 pages outright, but working with the text and regenerating. And then I got it out the door three hours later.
13:06And the editor is like, oh, this is perfect. I don't need any edits. Let's get it out the door. So that took a 17 hour project down to three hours. That's thing number one, is that this is not just accepting the machine output as is, but it's really us using the output as an assistant, much like Copilot on GitHub is using it as an assistant, right? These are essentially pair coding, if you will, co-authoring with the machine. I was doing a talk with the U.S. Copyright Office Assistant General Counsel, and he was talking about the regulations that they're putting out, saying that if machine-generated, therefore uncopyrightable, if human-generated, therefore copyrightable, and if machine-generated, you have to be able to disclose what aspects of the thing is machine-generated.
13:47So if you think about if I were to file a copyright registration with my article that I just drafted. What extent was that machine generated? And what extent was that human generated? Because I spent three hours adding, editing. If the machine generated in a sentence three of the words that were unmolested by me, and the other 20 words in the sentence were actually mine, do I have to disclose what three words were machine generated versus the ones that I edited? And that's with text. How would I do that with music, right? If I said to the machine, hey, generate a melody and generate a chord structure and generate lyrics.
14:21and then I spent from 1 a.m. to 3 a.m. rearranging all those things and then getting out the door, if the Copyright Office said what aspects of that was human-generated and what's machine-generated, I would honestly say, I have no friggin' idea because there's no track changes with my DAW that I make my music on, right? There's no track changes. I didn't track changes on my lyrics that I messed around. So I think this idea of trying to bifurcate what is machine-created and what is human-created is a fool's errand, and we're really going to have to reckon with that. Well, Damian, I have some very selfish questions that I've been pondering over in my own life as I've encountered them.
14:59And hopefully this won't seem like popcorn questions because I think it's very related to what you're talking about. But I think practical developers are hitting these snags as they're developing apps with this technology that are entering a sort of gray zone. So I'd love to get your thoughts on a few of these. One example that I can think of is, you know, a lot of people are building chat interfaces. It's very popular now to say, oh, build a chat interface over a website or build a chat interface over documents or build a chat interface over data, something like that. But people are doing this very frequently.
15:37It's very useful. My question is that chat interface or those messages are generated content, right? And that's what the user is seeing. But I see this huge gray area where let's say that I want to chat with Harry Potter, right? I take the book of Harry Potter and I put it in my vector database and someone asks a question of Harry Potter and I go retrieve the content. I'm just injecting that into a prompt, right? And I'm sending the prompt with, I guess, the book content into a model. The model's outputting some generated answer, and I'm sending that to the user. Now, I'm assuming, you know, I'm no lawyer, but I'm assuming I can't sell a new copy of Harry Potter unless I have certain rights and agreements in place.
16:26But what if I put this, you know, interface up on the internet and I start selling access to it? So I guess with that kind of very real world scenario, what sorts of elements do I need to consider there? And what's known to have a good answer? What's gray area? What's kind of being maybe being litigated right now? I'm going to answer your question, and it's going to be a fun walk. So take a walk with me. Okay, perfect. So this walk, it's going to begin with the Google Books project. So you might remember the Google Books ingested every book that ever existed, perhaps also including the Harry Potter book.
17:02A bunch of publishers said, hey, no fair, because you can't ingest all these things because these things are copyrighted. Every one of these books is copyrighted. They sued. And then the district court and the appellate court, the second court of appeals, said, yes, all those things are copyrighted. But Google's use of that is actually fair use. And this particular type of fair use is called transformative use. That the use that Google was using was transformative to what the original purpose of the book was. Purpose of a book is to read it, enjoy it, et cetera. Google's purpose was to index it.
17:32to be able to create a word index, to be able to then search all of the books, and to be able to provide the end user with a snippet, say maybe a page or two of that. So because it was not, you couldn't, I as a user, couldn't use Google Books to be able to essentially replicate the book process, but instead I'm using it to search, that is a transformative use, therefore fair use, that is not an infringement of copyright. So that was back in the day. Now think about how large language models work. So a large language model, if you have the input, it is ingesting, say, the entirety of Harry Potter.
18:02But really what it's doing is placing those in vector space, right? It's saying that these words are similar to those words in vector space. And once that happens, largely it jettisons the thing, right? In copyright law, there is the idea expression dichotomy. Ideas are uncopyrightable. So if I have the idea of a man in a black hat fighting a man with a white hat over a woman who is tied to a railroad track, Those are ideas that are uncopyrightable. You've seen lots of movies like that. But the expression of the idea, any particular movie that has that in there, that is copyrightable. So ideas, uncopyrightable.
18:36Expressions of the ideas are copyrightable. So if you apply that to what's happening when the large language model ingests all the books, it's essentially putting all the words into vector space. So it's saying, you know, this is a Bob Dylan-ism, or this is an Ernest Hemingway-ism, or this is a Harry Potter-ism. Each of those are ideas, not expressions of ideas. And so really, it's taking the expressions and effectively jettisoning those in favor of the ideas. So that's on the input side. And then on the output side, one can imagine that it's taking those ideas, Bob Dylanism, Ernest Hemingwayism, and then it's outputting them in a new expression.
19:14And if you believe the Copyright Office, machine-generated output is similarly uncopyrightable. So we're kind of faced with an idea that inputs ideas uncopyrightable, outputs the expressions of ideas created by machines similarly uncopyrightable. To your particular point, this has not been tested in court. So, you know, a judge who, by the way, may not know what he's talking about or she is talking about might rule against what I'm about to say right now. But at least I would make the really good argument that the ingestion of the thing is extracting the ideas from the book. And that is a transformative use because if you think about Google Books, they were printed three pages or so verbatim of these books.
19:54That's way more bulk than just, you know, think about the vector space. It's not reproducing any expression, right? It's merely taking the ideas. So if Google Books is permissible, almost certainly the large language model should also be. And that's really what is being argued right now in the cases that are happening with the GitHub copilot case happening in the West Coast. and then the stable diffusion case in Delaware and there's others like it where if I were the lawyers in that, I would be arguing exactly what I just argued right now. So to your specific use case, you are kind of interrogating this copyrighted work, but I would make the argument if I were representing you that this would be a transformative use.
20:31And just like you as a human would have read the book and you can, as a human, could provide output based on that book. In the same way, a machine should be able to read that book and provide output that is just taking the ideas of the thing, not necessarily the expressions of the ideas. So with that new expression coming out that you're describing and it being assuming that the Copyright Office view stands, it's not copyrightable, that's a massively different way of producing content from before till now and in the future. What does that mean for business and the world at large, considering that's a major, major change in how everything works?
21:10How do you see the future rolling out if that were to stand? First, I think it should stand because otherwise my All The Music project would essentially carpet bomb all the music and we would just have machines creating new expressions that would essentially make human expression obsolete. So number one, I think it should stand because if not, we are in a world of hurt with a copyright. Thing number two is you're right that never before in human history have we had a machine that creates new things. We've had the printing press where we take my ideas and then I can replicate it a whole bunch of times.
21:43We have the digital revolution in this, you know, 70s, 80s, 90s, 2000s, where now we can replicate human stuff. But never before have we had a way that the machine itself is making new expressions of ideas. And so as a result of that, some of the smartest people I know that are thinking about this is said that the web, as it stands, is probably large language models are going to stop right around November of 2022. Because anything after that, you're going to have a whole bunch of machine-generated content that is going to be essentially large language model-created things. If you know about the tech, and I assume your audience does, because it's statistically likely, it is smooth.
22:19Humans are jagged in the way that they write text. And that's what detect GPT and others is see the jaggedness of humanness. Machine generated content is smooth, not jagged. So that's what detect GPT says. Could you define that real quick, what jagged versus smooth means in this context? Sure. Yeah. So jagged means random. Smooth means statistically almost deterministic. So the idea is that we as humans say random things and we put things in a way that maybe hasn't been said before. Whereas machine, at least an LLM machine, is going to be able to say, you know, what's the most statistically likely word?
22:49and therefore that is smoother than our jagged randomness. The idea is that as the machines are creating the smooth, deterministic, statistically likely next word, that is essentially as new large language model ingest that smooth text, it's going to further smooth the corpus and we're going to miss all of the human created jaggedness that is going in there. So some of the smartest people I know are saying that maybe the web as it stood in November of 22 is maybe that's the last time we're going to have a lot of human created stuff that is truly jagged because here and out, we're just going to have machine created stuff that is smooth.
23:24And one last thing I would add is that one of the last bastions of human created jaggedness that we have is the courts. Because it turns out that, you know, people have talked about us being in a post truth era and post fact era. There is a person that's literally called a fact finder and his name is a judge. They spend years trying to find facts and then they write things called judicial opinions that have found facts that have been battled over years in the courts. So one of the last places where we can find this jagged, almost certain to be human written thing that is actually based in fact in our post-fact world, maybe is judicial opinions.
23:58And my employer, Vlex, has about a billion of those across the world. So maybe as we think about what are new corpuses that the large language models can train on that are truly jagged, that are full of factual things and not bullshit that's on the Internet, that is unvalidated. This is truly validated, human-created content that is high quality that might be a source to be able to ingest. Maybe this gets a little bit back to your article example where you interacted with the GPT output to write an article. And at a certain point, it kind of morphs into its own thing. What portion of it is machine-generated, what portion isn't?
24:34I know this is also happening just from seeing things like people are generating, for example, like adult coloring books using AI models and posting those in an almost automated way to Amazon. And then like someone can literally order a book. I think of other examples maybe where, hey, this book was written a long time ago. And so the wording is really difficult. What if I used a large language model to rephrase it in modern English and then I just post that and start selling it? So how long will this be debated in terms of like the copyright around this? And what should be on people's minds as they're creating this kind of content that they actually want to commercialize?
Read the full transcript
25:18Maybe that's a more practical question. Sure. If I were to create this machine-created coloring book, for example, which under the U.S. Copyright Office today, that entirely machine-created thing is therefore uncopyrightable. This really goes to the heart of what is copyright in the first place. And all copyright is is a monopoly. It is a government-sanctioned monopoly giving you, the author, a monopoly of life of the author plus 70 years on the thing you created. But as an exchange for that monopoly, the government says this has to be original. that has to be your creative work that does this.
25:51And if it is truly original and it is truly creative, we will give you that monopoly of 70 years, a life of the author plus 70 years. So really the question is, is there anything copyrightable in the machine generated work? Well, probably not because there was no human creativity in that thing. So that's thing number one. But then let's look at another scenario. What if somebody else did a human created coloring book that was identical to what the machine had done? does that turn it from unoriginal, therefore uncopyrightable with a machine created one to all of a sudden if a human does it, it is copyrightable, even though they're identical?
26:26Yeah, essentially, because the machine wasn't copyrightable, and you're recreating it, even though there might be no creativity on the human's part, because they're literally looking at the output of the uncopyrightable machine output. But they can steal the idea without the creativity involved, and then copyright. Am I understanding you correctly? That's right. And we've dealt with this, the courts have dealt with this for a few hundred years. And you can imagine Shakespeare is in the public domain. It's not been in copyright for hundreds of years. You can build a top Shakespeare, say with West Side Story that was based on Romeo and Juliet, right?
26:56The writers of West Side Story don't get copyright in the underlying story of Romeo and Juliet, but they do get copyright in whatever they put atop Romeo and Juliet. So the fact that it's, you know, New York City and all of these things. So they just get what's called thin copyright on top of the public domain thing. So you could imagine going back to our coloring book example, right? If someone then copies that and then adds a little human touch on that, they don't get the underlying copyright thing because that's public domain. They only get what they've added atop the machine created thing.
27:26And really the question is how much can you really add atop that is really creative enough to make it worthwhile? And I would say, you know, if it's just another line here or there, that's not sufficiently original or creative to add copyrightability. I do want to get back to another couple of themes, but maybe one more selfish question, which is less related to, I guess, the inputs and outputs would be more related to the models themselves, what they're trained on and how they're released. Of course, we're seeing a lot of different approaches to how models are being released in the sense that, well, is a model code?
28:05Is it data? Do I use Creative Commons or do I use Apache 2? Also, the data that was used in the training, maybe that's a mix of copyrightable material, or maybe it's not even known. Maybe a model shows up on Hugging Face, and I don't know what the mix of the data set was that was used in training. As you've been working with these large language models and advising around this and thinking about these concepts, how do you see that side of, I guess, training data, fine tuning data, model release? What's on your mind as you look forward to this next season, which I assume will continue? We just had an episode with, I think it was titled the Cambrian explosion of models.
28:51There's so many being released. You know, this will continue. How do you see that side of things developing over the next season that we're entering? I think that what you're asking really about is the provenance of everything that comes downstream. That is, what is the provenance of the input? And what is the provenance of the essentially being able to manipulate that input to create a thing that's called a model? Yeah, of course, the output of the model could train new inputs to be able to go into new models, right? Yes, a cyclical sort of thing. That's right. It's like a snake eating its own tail.
29:24And so within the law, in criminal law, there's a thing called the fruit of the poisonous tree. The first act might be innocuous, but then it leads to a chain reaction, a bunch of dominoes that leads to the end. So this is the fruit of the poisonous tree is a legal concept that you can imagine is similar for the questions you asked. If there is input data that maybe has questionable licensing. So, for example, LLAMA, right? So it was released for open source, but only for academic purposes. So you could imagine if someone were to be able to create a model for commercial purposes that is based on that ostensibly licensed for academic purposes, that is maybe a fruit of the poisonous tree question to be able to say, is that model now tainted because it was ingested in opposition to the license?
30:07Yeah, I've wanted to use some of these LAMA-based models quite recently and just haven't because it makes me ask a lot of questions. So am I right with that hesitation or is it yet to be determined, but you would make certain assumptions or what do you think? Yeah, so I should clarify that I'm a lawyer, but I'm not your lawyer. So nothing I'm saying is going to be legal advice, okay? Yes, correct, correct. So I would say that, yes, anytime that you are ingesting items that are licensed and then you're using them in a way that is maybe against that license or not permitted by that license, I think anyone should be worried when that happens, speaking generally.
30:47I would also say that, yes, anyone who does that should be worried. Also, proving such things is tricky, right? Because there is the law and then there's what can be proved in a preponderance of the evidence in the court of law. Now, as the dominoes fall and as the snake keeps eating its tail, the provenance of what data did you actually use and where did it come from? It gets murky. It does get really murky. And so that's something I imagine litigation is going to happen a lot. So I'm absolutely fascinated by that and want to take it farther. So I'm thinking back on years of business and all the IP concerns.
31:25I work for a big corporation. Lots of other people work for various. is the snake continues to eat its tail, and you're seeing this happen over and over again. The value of current IP generally, I would argue, diminishes over time because its usefulness in business as things are progressing ever faster in the business plus technology world. You know, something that was a great piece of IP a few years ago, it might still be covered legally, but you're not necessarily going to use what was, you know, 20 years ago versus what you did yesterday. With that kind of utility of current IP diminishing, and with this sequence of snake eating its tail that you're describing, and you know, fruit of the poison tree, I believe you called it, that has to have massive, massive repercussions for how business uses IP in the large, in general, like your entire strategy about, because right now, you know, organizations, they will come up with an idea, they'll immediately go copyright, but whatever the appropriate mechanism is and get that in.
32:25They lock it in. It's part of their business strategy. That seems to me, from what you're saying, to fail in the future. It is no longer a good strategy. What does that mean in the large? I mean, that's a gigantic question, I think. What I think you've described and what I think we're seeing in a society is the intellectual property laws that we've created since the beginning of our founding history. The Constitution says that we will protect inventions, that's constitutional. So what I think we're seeing is our intellectual property regime that has existed since the 1700s creaking under its own weight with this new large language model generating in a way that's never been done in human history.
33:04I think you're right. The value of patents, what is the value of a patent if I can use a large language model to, much like I described with all the patents, you know, we with all the patents are saying every patent that's ever been done and every idea, every claim and every patent, Let's recombine those. But you can imagine, and I've heard, that there are companies out there that are doing not what's been done before, but new ideas. And then making a ton of new claims and filing those new claims with the U.S. Patent Office. And to be able to say, here's a new idea. And if you carpet bomb the U.S.
33:34Patent Office with all of these new things that are just machine generated, there's currently a case, the Thaler case, T-H-L-E-R, that he created, has said a machine created this patent. and the patent office said, ah, machine-created patents are not a thing. You can't do it. But they only knew that because Thaler told him that it was machine-created. How many of these things are being filed that nobody has told anybody that it's been patented? And then is that fraud on the patent office? Probably. But the question is, who's going to find out if it doesn't go to litigation? For a moment, I want to ask you to stop being an attorney and be a speculator here.
34:10I want you to kind of blue sky it. Like, where can this possibly go? Or what do the different paths, you know, what's your gut tell you in terms of how this plays out? Because you've described multiple ways in the last few minutes where the whole system can essentially collapse under its own weight. Not just one way. You've done it several different ways where that can happen, which isn't surprising because over the episodes of the show, we keep talking about the rapid change that all this is bringing in AI. It's the most fascinating moment in human history, in my view. and you're describing the weight of all the structure of the past in terms of the legal considerations, unable to keep up with what's going on now, and it's only accelerating.
34:50Where do we go from here? What does that mean? Sure. If past is prologue to what's going to happen, I would say that we've seen over the last decade business patents essentially going away. We've seen software patents almost pretty much go away with the Alice decision and others. So patents have already been diminishing in value, you know, over the last 10 years or so. And I think this is just going to accelerate that diminishment. Because, you know, if I'm going to compete in the marketplace, largely that, you know, anything I invent today is going to be obsolete in three years anyway. So what's the good in patenting a thing that is obsolete, you know, in three years?
35:26I think Elon Musk said, you know, I'm open sourcing all my patents, he said, because a patent is merely a license to sue. And that's true, right? is I spend a million dollars or two million dollars to get the patent, and then I have to spend millions on top of it to sue somebody over that patent. That license to sue often just doesn't make business sense. And I think it makes even less sense as the system is collapsing in its own way. So if you're looking for me to speculate, and you are, I would hope that the patent regime is going to fall away in importance and people are just going to innovate.
35:54As all of us on this call are knowledge workers and program or do legal work, that sort of thing. I think we're all benefiting in terms of productivity moving into the future. You know, outside of the IP stuff, the copyright things, all of those things, I'm coding much faster now, not because I'm necessarily a better coder, which maybe I'd like to think I am, but I'm probably not. It's because I'm using generative tools and suggestions in a much more robust way. And I was fascinated in one of your recent talks when you're talking about kind of the practical consequences of that. Like if I can work 50 % faster, do I still work the same amount or do I work less?
36:43And what are the implications of my employer's viewpoint on my work and that sort of thing? Could you talk us through a little bit about your thinking in that regard? Yeah. So I'm going to talk about four worlds. First world is 2022 world. before the large language bottles. And in that world, I would work 40 hours a week full-time, and I would give 40 hours a week of 20-22 productivity as a result of that. And as a result of that, an employer would hire a workforce like me to do that. So that's world number one. In world number two, I know of people anecdotally that are working three full-time jobs because they're getting at least 100 % or so productivity gains, maybe 10x productivity gains based on the code that you said.
37:25So they have three full-time jobs. So he's essentially working 30 % of the time for each, but still providing 100 % of the output for that. And their employer is saying, wow, that's great output. They don't care, right? So that's world number two, I think, what we're in today. World number three is probably the employer is going to say, hey, hey, hey, hey, don't give me 30 % of your time. Give me 100 % of your time and maybe give me 10x output of 2022 level output, right? I want that productivity gain from you. So that's world number three. But I think shortly thereafter is going to come world number four where the executives are going to say, wait, wait, wait.
38:02If we lay off two-thirds of the workforce and then still require them to work 40 hours a week with their 10x productivity, I can say to my shareholders, look at all the costs that we cut by laying off two-thirds of the workforce, and we're still getting 5x productivity on top of our 2022 productivity. We've cut costs. We've increased productivity. Aren't we great? I think that that's probably the world that we're headed for. And there's a world six beyond that, which that leads very obviously to that recognition of cut the workforce and stuff. I don't know if we want to go there or not to finish up, but we got some tough social issues to navigate there.
38:40Really, what we're describing here is there's a scarcity mindset and there's an abundance mindset. The scarcity mindset is that around 1979, accountants were really worried with this artificial intelligence that's called the spreadsheet. Because they said, wow, all we do all day is use ledgers and we add and subtract numbers and machines can do that in seconds. That's going to put us all out of work. But what happened was that when the clients realized, oh, it's not going to take me a week to get that ledger back, but it's going to take seconds. Let's do the scenario two and scenario three and scenario four and run more scenarios.
39:10And now we have more accountants than ever because the tools are actually a force multiplier that now there's more accounting work rather than less. So that is an abundance mindset that is not a scarcity mindset. So the real question in my mind, and maybe should be on all of our minds, is the scarcity mindset that I described with worlds one, two, three, and four, is that going to be our future? Or is there an abundance mindset where we just have 10x or 100x productivity and we keep growing and growing and growing? I think that's a great transition kind of as we get to the close here. Maybe one question that I'd like to ask you.
39:43We've talked about various interesting scenarios and maybe things that are honestly kind of uncomfortable for a lot of our kind of technical listeners around legal questions and lawsuits and copyright and that sort of thing. From your perspective, as you look to the future, kind of this next year, what are you encouraged by and or what, how would you encourage our listeners, maybe those practical developers or practitioners out there? Like, how would you encourage them to engage in this conversation and these topics moving to the future? And what are you excited about or encouraged by moving to the future?
40:23I think about AI as largely a tidal wave or a tsunami, and we are running faster than the tsunami. How do we run faster than tsunami? You learn how to use Copilot to be able to go faster. You learn how to be able to do things that the machine cannot yet do. That's running faster than tsunami. So really, I say to lawyers that are worried about AI that AI will not take a lawyer's job, but a lawyer that uses AI will take the job of a lawyer that does not use AI. And so really, I would say the same thing for coders who are listening, that learning to use the tool to run faster than the tsunami. There's another joke.
40:59You know, there was a bear at a campground and two guys and the one guy gets out of tennis shoes and the other guy says, you can't outrun a bear. And he said, I don't have to. I just have to outrun you. Right. So in that sense, I learned how to use a large language models to outrun your competition, because as the wave crashes over them, it's not going to crash over you. I think that we all have to reckon eventually the wave, I think, may crash over all of us. but for until then, I think we should be running as fast as we can. Awesome. Yeah, that's a great encouragement. And thank you so much for humoring us with all of our random questions, some of which were selfish on my part, but I've learned a lot and really appreciate your insights, Damien, and the work that you're doing.
41:39Look forward to seeing your future projects. And I'm sure that our listeners will find this super interesting. Thank you so much. Thank you. I don't often get to speak to an audience as sophisticated as yours. So I really enjoyed the really deep and probing questions. And I really am grateful for the opportunity.
42:04Thank you for listening to Practical AI. Your next step is to subscribe now, if you haven't already. And if you're a longtime listener of the show, help us reach more people by sharing practical AI with your friends and colleagues. Thanks once again to Fastly and Fly for partnering with us to bring you all Change Talk podcasts. Check out what they're up to at Fastly.com and Fly.io. And to our Beat Freaking residents, Breakmaster Cylinder, for continuously cranking out the best beats in the biz. That's all for now. We'll talk to you again next time.
42:48Game on!
From the publisher
As a technologist, coder, and lawyer, few people are better equipped to discuss the legal and practical consequences of generative AI than Damien Riehl. He demonstrated this a couple years ago by generating, writing to disk, and then releasing every possible musical melody. Damien joins us to answer our many questions about generated content, copyright, dataset licensing/usage, and the future of knowledge work.
Changelog++ members save 1 minute on this episode because they made the ads disappear. Join today!
Sponsors:
- Fastly – Our bandwidth partner. Fastly powers fast, secure, and scalable digital experiences. Move beyond your content delivery network to their powerful edge cloud platform. Learn more at fastly.com
- Fly.io – The home of Changelog.com — Deploy your apps and databases close to your users. In minutes you can run your Ruby, Go, Node, Deno, Python, or Elixir app (and databases!) all over the world. No ops required. Learn more at fly.io/changelog and check out the speedrun in their docs.
- Typesense – Lightning fast, globally distributed Search-as-a-Service that runs in memory. You literally can’t get any faster!
Featuring:
- Damien Riehl – LinkedIn, Mastodon, X
- Chris Benson – Website, GitHub, LinkedIn, X
- Daniel Whitenack – Website, GitHub, X
Show Notes:
- Talk - Legal and Practical Consequences of Generative AI (LLMs like GPT, Bart, PaLM, LLaMA, Alpaca, Codex)
- Talk - Why All Melodies Should Be Free for Musicians to Use | Damien Riehl | TED
Something missing or broken? PRs welcome!




