Art vs. AI: Unveiling the Tool That's Tainting Data

8 Apr 2024 · 10 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

AI Today Podcast Episode Notes: Art vs. AI: Unveiling the Tool That's Tainting Data

Episode Overview In this episode, the podcast discusses the emerging conflict between artists and generative AI technologies, particularly the challenges posed by new data poisoning tools like Nightshade. The episode delves into the implications for creativity and copyright, highlighting how these innovations could impact the AI landscape and its interaction with artistic integrity.

Key Topics Discussed

  1. Introduction to Nightshade
  2. Definition: Nightshade is a new tool that allows artists to alter the pixels of their artwork.
  3. Purpose: Designed to create corrupted datasets when AI companies scrape images for training, leading to nonsensical outputs from AI models.
  4. Target Audience: Primarily aimed at AI entities that utilize artists’ works without permission.
  1. Impact on AI Models
  2. Affected Tools: Notable AI models like DALL-E, MidJourney, and Stable Diffusion could face disruptions in their output due to data corruption.
  3. Example of Corruption: If AI systems are fed with altered images, they may generate bizarre results (e.g., morphing images of dogs into cats).
  4. Broader Implications: The tool targets the foundational aspect of AI—its reliance on vast amounts of internet-sourced data.
  1. Community and Legal Responses
  2. Current Legal Landscape: Major AI companies (OpenAI, Meta, Google, Stability AI) are facing lawsuits from artists regarding unauthorized data scraping.
  3. Artist Sentiments: While some artists support Nightshade, there is skepticism about its long-term effectiveness due to established data sets already in use by AI companies.
  1. Complementary Tools
  2. Glaze: Another tool developed alongside Nightshade, designed to protect artists' unique styles from being misappropriated by AI algorithms.
  3. Functionality: Glaze manipulates image pixels to confuse AI models while remaining visually unchanged to human observers.
  1. Academic Perspectives
  2. Support from Academics: Various professors, including Ben Zhao (University of Chicago) and Vitaly Shmetkov (Cornell University), praise Nightshade for highlighting vulnerabilities within AI systems.
  3. Potential for Abuse: Concerns about how data poisoning technology could be misused for sabotage or competitive advantage.
  1. Financial and Ethical Considerations
  2. Royalties Discussion: The possibility of Nightshade prompting AI companies to respect artists' rights and increase royalty payments.
  3. Opt-Out Provisions: Some companies propose mechanisms for artists to exclude their work from AI training datasets, although this is viewed as insufficient by many.

Takeaways

  • Vulnerability Highlighted: The introduction of tools like Nightshade emphasizes the need for robust defenses against potential misuse of AI technologies.
  • Ethical Debate: The ongoing discussion underscores the tension between technological advancement and the protection of intellectual property rights.
  • Future of AI: The implications of these tools could lead to significant changes in how AI is trained and how artists interact with AI companies.

Conclusion The episode sheds light on a critical intersection of art and technology, indicating a growing need for dialogue and resolution between artists and AI firms. The development of tools like Nightshade and Glaze signifies both a protective measure for artists and a challenge for the future of AI-generated content. The unfolding scenario will warrant close observation as it progresses, particularly in how major AI companies respond to these initiatives.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00What can 160 years of experience teach you about the future? When it comes to protecting what matters, Pacific Life provides life insurance, retirement income, and employee benefits for people and businesses building a more confident tomorrow. Strategies rooted in strength and backed by experience. Ask a financial professional how Pacific Life can help you today. Pacific Life Insurance Company, Omaha, Nebraska, and in New York. Pacific Life and Annuity, Phoenix, Arizona. the escalating tensions between artists and ai companies has taken i think a really interesting twist right now so essentially there's this new cool new tool called nightshade and it's come out and essentially it's enabling artists to um pretty much alter their art's pixels so if ai companies scrape these images for training data and the result is essentially just kind of like a corrupted ai model that produces like really wacky nonsensical outputs um so nightshade's primary objective is to essentially just kind of counter AI entities that are using artwork without the artist's approval, right?

1:02That's the whole idea. Not that they want to, you know, just permanently destroy all art generating AI models, but any that are using, you know, their content without their permission is essentially who they're targeting here. So AI models like Dolly, Mid Journey, Stable Diffusion, all of these, I think, could see their output severely disrupted by this tool. And so I think, yeah, it's kind of interesting. Essentially, what's happening is, you know, you could have like where it was once like you're getting a picture of dogs generated. Now, all of a sudden, it's going to be cats instead or cars are going to all of a sudden turn into cows.

1:35MIT's technology review, which recently covered some of this research, revealed that the paper has been slated for review at Eustynic's Computer Security Conference as a paper on how to do this. But there's a bunch of big AI giants, including OpenAI, Meta, Google, Stability AI, that are all kind of grappling with lawsuits from artists right now that essentially are alleging unauthorized scraping of copyrighted content and personal data. So Ben Zhao, who is the mind behind Nightshade and a University of Chicago professor, envisions a tool as a balance restore, essentially offering artists a way to counter infringements on their intellectual property rights.

2:12so responses from the you know the big companies that I just mentioned earlier I'm regarding this uh no one's really commenting on it right now uh for obvious reasons now here's the thing um I'll go into a little bit more of this um and I know that like a lot of artists are like I don't know super like rah-rah this is gonna like take down the big guys but like at the end of the day in my opinion this doesn't make that big of an impact because all of the big players mid-journey OpenAI, etc. have already scraped the data, they've already built their data sets. And so I'm pretty sure they have the data set.

2:48And if everyone just starts, you know, corrupting all the image files in the future, maybe it will stop them from being able to train off of like new stuff, but they already have like 100 years or 1000 years of paintings and art in their data set. So I don't know how much maybe like if new art techniques come out, and they're not able to access them, then there won't be as good at like that per se versus one that goes and gets the rights but at the end of the day I don't know how much of an impact this is going to have on a majority of of this content just being that the big guys already grabbed all of the data so it'll definitely stop any newcomers but I don't know if it's going to stop the ones that have already come out so that's just my caveat I'm going to throw on there so kind of complementing this new software though it's Nightshade it there's another one called Glaze which is also from Zhao's team and essentially this is allowing artists to protect their unique styles from ai's algorithms so glade functions essentially with nightshade but essentially is manipulating image pixels in ways that make it indiscernible to humans but it's wildly confusing to ai models this is very interesting so the images pixels you won't even notice them when you look at this picture but to an ai model it kind of wrecks what it's training on so plans to merge nightshade's capabilities into glaze glaze are underway additionally artists will have the freedom to essentially decide whether to use nightshade um making the tool open source as zhao's team intends i think could kind of amplify it's like how many people are using this thing as large language model data sets have billions of images um introducing more of these like tainted images could weaken the ai's ability to generate images, which means that, you know, we could start seeing a lot of these big AI models not pulling in new data if all of a sudden like this corrupted data is kind of like poisoning their data set pretty much.

4:40So Nightshade targets a generative AI model's Achilles heel, which is its dependence on, of course, vast internet source data. And essentially it's tampering with these images. So it is kind of undermining the AI's foundation. So for instance, artists can use glaze to mask their artwork under a different style. And then they also use nightshade at the same time. So when AI companies unknowingly incorporate these, you know, quote unquote, poisoned images into their training, it pretty much just wreaks havoc on the AI's performance. Experiments revealed some pretty, I would say startling results.

5:13So when stable diffusion pulled in, I think just 50 like messed up or you could call them tainted dog images, the resulting AI generated images were super bizarre they produced these like multi-limbed and character-like dogs with 300 of those images dogs began to completely morph into cats which was really really interesting so nightshades potency I don't think doesn't necessarily just end with direct word associations like dog it also goes into other terms like puppy husky wolf right so like a lot of the associated words so you could have like a bunch of messed up dog pictures or tainted dog pictures in a data set and it's now starting to um wreak havoc with like a bunch of similar uh or variations of that word so an ai um essentially that's been exposed to these images could be misled across a bunch of different prompts you know dragon or lord of the rings castle So like you put these words in and it's completely producing the wrong thing.

6:16So I think this is definitely a very interesting tool. Zao acknowledges the potential misuse of data poisoning for nefarious ends. of course you know like essentially sabotaging large language models is going to need like hundreds and thousands of corrupted images given their training on billions of samples like you're going to need a lot of these corrupted images in a data set to actually take it down so Vitaly Shmetkov who is a Cornell University professor and AI security researcher not associated with the project but just emphasize the urgency of developing defenses he seems to applaud this whole project.

6:57There's also someone named Gutman Kermath from the University of Waterloo. And he said that this is fantastic. It really looks like it's a bunch of academics and professors that are like really excited about this. And I think one thing it definitely does do is highlight the vulnerabilities in these AI models. So, you know, whether this project takes off or not, I think it's possible for someone to do this. And so I think it's important, And honestly, even between big AI companies, putting out something like this to sabotage their competitors, building a moat by going in. Like, here's a great way to build a moat for open AI.

7:33They could go buy 100 image websites, Unsplash, Shutterstock, whatever, right? Five, 10 of them that are just big data sets. They run the software on them. But when they use the data, they don't run it. But when they give it, but all their other competitors that are scraping it from the internet do get it and they could essentially cripple all their competitors. There's an interesting moat. So I think there's a lot of use cases, whether it's an AI company, whether it's a foreign adversary, like China doesn't want us to have as good technology as them. So they go and sabotage stuff. I think it just, you know, highlights a vulnerability.

8:07So it's important to be aware of. I think Zhu Feng Yang of Columbia University essentially says that Nightshade could compel AI companies to better respect artists' rights, and it would possibly be leading to increased royalty payments. So AI firms like Stability AI and OpenAI propose opt-out provisions for artists. Essentially, you can opt out of getting your content included in their data set. A lot of people find this insufficient. So for artists like Autumn Beverly, some random artist says that this is going to be awesome and likes it. So in any case, a lot of people are for it. I I haven't heard a lot of people against it necessarily.

8:47I haven't heard a lot of pushback on it. I don't think it's super common right now. Personally, I just think it'll be interesting to see what happens. I'm never for shutting down a specific technology, and I will just call this a new technology. It highlights a vulnerability in platforms. I would rather this be coming out of some universities in the United States that are trying to use it to get more royalties or whatever. not because I agree with it necessarily but because if it doesn't come out of if it doesn't come from here it's going to come from China or North Korea or you know different AI models are going to use it to sabotage each other so I'd rather at least everyone know about it and be a little bit more transparent I'm a big advocate if the technology exists it will get built and so the question isn't what happened like can we stop it from being built but like how do we what are our defenses against these kind of these vulnerabilities and whatnot So a very interesting topic that I think, of course, is going to be top of mind for some of the big companies that have billions of dollars to deal with this issue.

9:50So I'm not really concerned about this, but it's definitely going to be an interesting Sega to watch as this unfolds and see, you know, how OpenAI responds and how this kind of develops in the future.

From the publisher

In this episode, we explore the clash between artists and generative AI, examining the impact of data poisoning tools on creativity and copyright infringement.


See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

More from AI Today

All 897 episodes
Art vs. AI: Unveiling the Tool That's Tainting DataAI Today · 10 min
Listen in VO