Which AI Model is the Most Transparent?

19 Oct 2023 · 19 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: The AI Daily Brief - Episode on AI Model Transparency

Overview In this episode of *The AI Daily Brief*, the host discusses recent developments in artificial intelligence, focusing on a new methodology for assessing the transparency of AI foundation models. The episode also covers significant legal challenges facing AI companies and the implications of these issues for the industry.

Key Topics

  1. Lawsuit Against Anthropic by Universal Music Group (UMG)
  2. Overview: UMG has filed a $75 million lawsuit against Anthropic, claiming copyright infringement for using lyrics from numerous songs without permission.
  3. Accusations: UMG argues that Anthropic's AI model, Claude, provided lyrics from over 500 songs, impacting businesses that legally license such content.
  4. Broader Implications: This lawsuit reflects ongoing tensions about copyright in AI training and may contribute to future legal precedents regarding AI-generated content.
  1. YouTube’s Initiative with Music Labels
  2. YouTube is reportedly attempting to negotiate rights with music labels to develop an AI-powered tool that can replicate artists' voices for creators.
  3. This reflects an industry trend where companies seek legal frameworks for using copyrighted materials in AI applications.
  1. Additional Lawsuits in the AI Space
  2. Former Governor Mike Huckabee and other authors have sued major tech firms, including Meta and Microsoft, over unauthorized use of their works in AI model training.
  3. These lawsuits underline a growing concern about the ethical and legal standards of training AI on copyrighted materials.
  1. Amazon’s Robotics Advancements
  2. Amazon has introduced two new robotic systems aimed at improving fulfillment and delivery efficiency.
  3. Sequoia: Enhances inventory management speed.
  4. Digit: A bipedal robot designed for handling repetitive tasks alongside human workers.
  1. Innovations in Voice Cloning
  2. Companies like Eleven Labs and PlayHT are advancing voice dubbing technology for content creators, indicating a competitive landscape in AI-driven audio production.
  1. FBI and AI Security Concerns
  2. FBI Director Christopher Wray highlighted AI's potential misuse in amplifying terrorist propaganda and discussed the challenges of safeguarding AI applications against adversarial exploits.

Main Discussion

The Foundation Model Transparency Index

  • Introduction of the Index: Researchers from Stanford, MIT, and Princeton launched the Foundation Model Transparency Index to objectively measure the transparency of major AI models.
  • Motivation: The index aims to highlight the current lack of transparency in the AI industry, emphasizing the importance of clear information for developers, policymakers, and consumers.

Methodology

  • Indicators: The project uses 100 indicators to assess transparency across three dimensions:
  • Upstream: Inputs involved in model training (data, labor, compute).
  • Model: Characteristics of the AI model itself (architecture, capabilities, risks).
  • Downstream: The distribution and usage policies of the model.

Results

  • Top Scorers:
  • 1st: Meta's Llama 2 (54%)
  • 2nd: OpenAI's GPT-4 (47%)
  • 3rd: Stability AI's Stable Diffusion (47%)
  • Findings:
  • Significant disparities exist between open and closed models, with open models generally scoring higher in transparency regarding data and training practices.
  • Despite high scores, no model achieved adequate transparency, reflecting a broad industry shortfall.

Reactions and Critiques

  • Support for Transparency: Many experts advocate for improved transparency, linking it to enhanced safety and ethical AI development.
  • Concerns About Results: Some industry leaders found the scoring system inconsistent and questioned the absence of various notable models.

Conclusion The episode underscores a pivotal moment in AI development regarding transparency, legal challenges, and technological advancements. The discussions around the Foundation Model Transparency Index signify a crucial effort to foster accountability in AI, which could shape future regulations and practices within the industry.

Further Engagement Listeners are encouraged to participate in discussions about AI transparency and related topics through the AI Breakdown Discord community.

Links

  • [Subscribe to The AI Breakdown Newsletter](https://theaibreakdown.beehiiv.com/subscribe)
  • [Watch The AI Breakdown on YouTube](https://www.youtube.com/@TheAIBreakdown)
  • [Join the Community](bit.ly/aibreakdown)
  • [Learn More](http://breakdown.network/)

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:01Today on the AI Breakdown, we're looking at a new methodology for determining which AI model is the most transparent. Before that on the brief, Universal Music Group has sued Anthropic around the infringement of copyright. The AI Breakdown is a daily podcast and video about the most important news and discussions in AI. Go to breakdown.network for more information about our Discord channel, our newsletter, and our YouTube. Welcome back to the AI Breakdown Brief, all the AI headline news you need in around five minutes. Today, we kick off with yet another lawsuit around the way that AI models have been trained.

0:38This time, however, it's not a group of authors. It is the viciously lawyered up music industry, specifically Universal Music Group, filing a$75 million lawsuit against Anthropic AI. They argue that Anthropic has perpetrated mass copyright infringement in serving up their lyrics to people who ask. Alright, so let's get into some specifics. This was actually a trio of music publishers, UMG, Concord Music Group, and ABK Co., and the suit was filed in Tennessee Federal Court. The specific accusation was, quote, systemic and widespread infringement by copying and distributing lyrics from at least 500 songs, including Katy Perry, The Rolling Stones, and Beyonce.

1:14The complaint says Anthropic has neither sought nor secured publishers' permission to use their valuable copyrighted works in this way. Just as Anthropic does not want its code taken without its authorization, neither do music publishers or any other copyright owners want their works to be exploited without permission. Now in terms of what their evidence is, they basically asked Claude what the lyrics to Katy Perry's Roar were, which it provided not without a few errors, but mostly. They say that this undercuts music lyric aggregators and other websites who have explicitly licensed those works.

1:42And that by not licensing this content, Anthropic is, quote,

1:56Now, I will, for the sake of this piece, bite my tongue about what I view as the utter stupidity of having to license the ability to print song lyrics on the web. But it feels to me like this lawsuit isn't really about that. This is yet one more in the line of lawsuits that are frankly trying to get their way to the Supreme Court to figure out how to deal with copyright in AI training in general. Now, the Hollywood Reporter suggests that this specific case and the way that Universal Music Group has architected it is meant to cut directly out what they would expect the defense to be, which is, of course, fair use.

2:32The reason this might be harder to defend with fair use is that since other websites are licensing lyrics to be able to print them, this actually potentially hurts the business opportunities of the publisher. Now, so far, Anthropic hasn't responded to requests for comment, but it's clear at this point that a major part of the AI lab's job is going to be fighting these copyright battles until it gets actually sorted out, either in the courts or through some sort of regulatory policy, but frankly, much more likely in the courts. It was then interesting to see, however, that even as Anthropic was being sued by the labels, YouTube is apparently trying to work with the labels to specifically get access to rights for their songs to train a new AI-powered voice replication tool.

3:11Writes Bloomberg, who broke the news, YouTube is developing a tool powered by artificial intelligence that would let creators record audio using the voices of famous musicians. The video site has approached music companies about obtaining the rights to songs it could use to train the tool, although major label records have yet to sign off on the deal. Now, if you've listened to me with any regularity, you've heard me talk about exactly this before. I literally cannot imagine that the way that this gets resolved in part is that we end up with a handful of rights-approved venues or rights-approved models through which people who want to create music that sounds like Drake or Sia or Bob Dylan or whoever can actually do so legally.

3:48One of the things about the music industry is that they are extremely good at adapting to the threat of new technology and then co-opting it in a way that reinforces their supremacy at the center of the industry. That was the byproduct of Napster into the streaming era, and I would be shocked if anything different happened here. That doesn't mean, of course, that there won't be bootlegged versions and people training their own models on artists without their permission or without the rights approval. but the bet will be that the legitimate use case that's willing to pay for the privilege in some way or another will be a heck of a lot higher than the pirated use case, and that sort of will work out better for everyone.

4:20Meanwhile, yet another story of lawsuits against AI training. Former Governor Mike Huckabee, along with a group of other Christian authors, have filed a lawsuit in New York federal court arguing that their works were used unlawfully in the training of models from OpenAI and others. This particular lawsuit targets Meta, Microsoft, and even Bloomberg LP. Interestingly, this threads a little bit of a line. The plaintiffs here say, while using books as part of datasets is not inherently problematic, using pirated or stolen books does not fairly compensate authors and publishers for their work. So I guess if Meta and Microsoft had just bought a copy of each of these books, it would have been fine with them.

4:55I'm not really sure, but it is an interesting wrinkle in this larger conversation. Now let's move on to some very, very different topics. First, Amazon has announced two new types of robots that are being integrated into their delivery and fulfillment systems. This news comes directly from Amazon in a blog post written by Scott Dresser, the VP of Amazon Robotics. First robotic system is called Sequoia, and it's designed to improve how warehouses fulfill customer orders. So far, it's operating at one of their fulfillment centers in Houston, Texas. Dresser writes, Sequoia allows us to identify and store inventory we receive at our fulfillment centers up to 75 % faster than we can today.

5:30This means we can list items for sale on Amazon.com more quickly, and when orders are place, Sequoia also reduces the time it takes to process an order through a fulfillment center by up to 25%, which improves our shipping predictability and increases the number of goods we can offer for same-day or next-day shipping. Now, the other robot that they announced was a bipedal, quote, mobile manipulator solution, in other words, a type of robot that can move while also grasping and handling items. This one is called Digit and is a collaboration with Agility Robotics. They argue that robots like these are going to be best used in collaboration with humans, specifically on hyper-repetitive tasks.

6:03For example, they write, our initial use for this technology will be to help employees with tote recycling, a highly repetitive process of picking up and moving empty totes once inventory has been completely picked out of them. Moving on to another development in the product space, there is a ton of innovation right now in and around voice cloning. Eleven Labs, Wondercraft, and others have all come out with voice dubbing for YouTube and podcast content recently, and now PlayHT has announced PlayHT 2.0 Turbo. They call it the fastest conversational AI text-to-speech model. I think this is an area that is going to have a lot of impact in how content gets produced, and so I'm always interested to keep track of the updates in the space.

6:39Now, closing out more on the policy and macro side of things, FBI Director Christopher Wray discussed AI earlier this week in a conference for the US and its closest intelligence allies. He said that AI has already been successfully used to amplify terrorist propaganda, and that terrorist groups are very focused on trying to get around AI safeguards. Said his British counterpart Ken McCallum, quote,

7:11Relatedly, another New York Times piece on AI today tells the story of a group of researchers who found that AI safeguards are not so safe at all. The paper is called Fine-Tuning Aligned Language Models Compromises Safety Even When Users Do Not Intend To. And basically, they're pointing out that there are safety costs associated with the type of customized fine-tuning that many models are going through right now as they go into production for various real-world applications. They write, Our red-teaming studies find that the safety alignment of LLMs can be compromised by fine-tuning with only a few adversarially designed training examples.

7:42For instance, we jailbreak GPT-3.5 turbo safety guardrails by fine-tuning it on only 10 such examples at a cost of less than$0.20 via OpenAI's APIs, making the model responsive to nearly any harmful instruction. Disconcertingly, our research also reveals that even without malicious intent, simply fine-tuning with benign and commonly used datasets can also inadvertently degrade the safety alignment of LLMs, although to a lesser extent. The concern was summed up by ScaleAI's Riley Goodside, who said, This is a very real concern for the future. We do not know all the ways this can go wrong. But for now, I will leave you to ponder all the ways that it can go wrong.

8:18I, of course, appreciate you guys listening or watching as always. Next up, the main AI breakdown. One of the notable and frankly often griped about features of the artificial intelligence space is how opaque many of the leading models are. Even the labs who work on these models don't really know exactly how they produce the results they produce. And for people on the outside who don't have information about things like what data the models were trained on, the lack of transparency can be an even greater concern. Well, there is a new project coming out of Stanford, but with the participation of researchers at MIT and Princeton as well, that is designed to create a system for actually capturing and scoring the transparency of major AI foundation models.

9:04What we are going to do today is look at that new system, see how different models score, and then look at what some of the responses in the AI community have been. The name of the project is the Foundation Model Transparency Index. And the project was organized by the Stanford University Human-Centered Artificial Intelligence Institute, or HAI, and specifically within it, the Center for Research on Foundation Models. The motivation, says Rishi Bomasani from the CRFM, is that companies in the space are becoming less, not more, transparent. About how the name OpenAI looks a little bit ironic now, given how they sit relative to others in the space.

9:37In their announcement post, Stanford's HAI makes a bunch of different arguments about why it is important to have more transparency around these models. They write, Less transparency makes it harder for other businesses to know if they can safely build applications that rely on commercial foundation models, for academics to rely on commercial foundation models for research, for policymakers to design meaningful policies to reign in this powerful technology, and for consumers to understand model limitations or seek redress for harms caused. Now, of course, even if one has a lot of excitement about the goals of a project like this, the devil can really be in the details.

10:09So how was this group of academics going about trying to figure out which models were most transparent? They ended up deciding on 100 different indicators that would be used to give a transparency score overall. They linked to a PDF hosted on GitHub that has all 100 of those indicators that you can scroll through. But the major dimensions that they list include data, labor, compute, methods, model basics, model access, capabilities, risks, mitigations, distribution, usage policy, feedback, and impact. Now in terms of the methodology, for the sake of this starting point, they picked 10 major foundation models, which of course leaves a ton of projects out, but which they thought represented a meaningful slice of the industry.

10:47From there, they gathered all publicly available information, which of course is the only information that matters, given that this is a study about transparency. In fact, that's one of the things that I think makes the methodology the best, or at least the most reliable, is that the whole point is to see what they've said publicly. So it doesn't so much matter if there's information we don't have access to behind the scenes, that's kind of exactly the point. From there, for each of these 10 projects, two researchers scored the 100 indicators, and then compared scores with their fellow, and discussed to resolve disagreements.

11:15The last step was sharing the scores with leaders at the companies, which gave them a chance to contest scores they disagreed with, which could then be factored into the final scores. So let's jump to the high level first. First, the most transparent model in this index was in fact Meta's Llama 2. Now this will probably be gratifying to Meta, who of course have set out to be the real OpenAI project, given their open-source-ish approach to how they're doing things. Other notable entrants on this list include OpenAI's GPT-4 coming in at 3rd, Stability AI's Stable Diffusion coming in 4th at 47, Google's Palm 2 and Anthropics' Claude 2 coming in 5th and 6th, with scores of 40 and 36 respectively, and way down near the bottom, Inflection coming in at 9th with a score of just 21, and Amazon's Titan Text coming in at just 12.

11:58Now, even though Meta's Llama 2 was the highest scorer, it still scored only a 54%, suggesting that there is a lot more room for transparency within this field. Now getting into the specific indicators, they divided those 100 indicators into upstream, model, and downstream. Upstream refers to, quote, the ingredients and processes involved in building a foundation model, such as the computational resources, data, or labor used. The model indicators, quote, specify the properties and function of the foundation model, such as the model's architecture capabilities and risks. And the downstream indicators, quote, specify how the foundation model is distributed and used, such as the model's impact on users, any updates to the models, and the policies that govern its use.

12:37I think in some ways more telling than just the overall score is breaking it down in terms of these major dimensions of transparency. For example, one of the real standout low scores is that when it came to information about the data that models were trained on, these 10 projects scored an average of just 20%. The high was Blooms at 60%, and Llama 2 had 40%, but many had 0 % to 20%. The highest category on average was Model Basics, the basic information about what models can do. The average score there was 63%. Now, one of the things that one might expect is that there would be a fairly significant difference between open developers versus closed developers, and indeed, that's what we saw.

13:14Three of the top four projects were the open models, including Llama 2, Blooms, and Stable Diffusion, which scored respectively 54%, 53%, and 47%. OpenAI's GPT-4 came in at third again, just above Stable Diffusion 2. The researchers also noted that a lot of the disparity here had to do with that upstream category and a lack of transparency around the data used to train the model, what labor was used to train the model, and how much compute was used to build the model, which was a lot more clear with the open developers than with the closed. Indeed, it's even more stark when you look at the average transparency of the open versus closed developers on a dimension-by-dimension level.

13:49When it comes to what data models were trained on, open developers scored 47%, whereas closed developers scored just 9%. Labor and compute were similarly stark, with 43 % versus 6 % in both cases. Methods, model basics, and model access also had huge disparity. Open models scored a 92 % on methods versus just a 29%, for example, for closed developers. Now, there were a couple areas where closed developers did outscore open developers. Those include capabilities, risks, and mitigations, where in each case closed developers were slightly ahead of their open counterparts. Researchers speculate that this is because closed developers dedicate more resources to actually controlling and shaping the way the model is used after it's released.

14:31Usage policy, for example, is another area where closed way outscored open 49 % to 33%. One of the researchers involved in the project, Sayesh Kapoor, who also writes the AI Snake Oil newsletter, wrote, Developers of open foundation models scored higher in many axes of transparency despite many of our indicators being easier to satisfy for closed models. For example, many indicators assess policies for downstream use. Since closed model developers often provide access only through an API, they can share information related to downstream use more easily, whereas developers of open models need to collaborate with the downstream deployers to satisfactorily provide such information.

15:05In theory, this should mean a much higher score for closed models on these indicators, but we find no substantive difference. Now, in terms of how they sum up their findings, the researchers said, quote, no major foundation model developer is close to providing adequate transparency, revealing a fundamental lack of transparency in the AI industry. However, and I thought that this was a really interesting point, while the mean score was just 37%, 82 of the 100 indicators were satisfied by at least one developer, meaning that if these developers simply adopted best practices from their competitors, it would significantly improve transparency all on its own.

15:39Now, this study has been picked up by the news quite a bit. I think putting on my meta-narrative analysis hat for a minute, it fits a bit of the general skeptical bias that many news organizations have when it comes to technology in general right now, and AI more specifically. The New York Times piece about this was called Maybe We Will Finally Learn More About How AI Works. Author Kevin Roof definitely wants more, not less transparency. As he put it, we can't have an AI revolution in the dark. We need to see inside the black boxes of AI if we're going to let it transform our lives. Now some, like VC Vinod Khosla, felt that the approach was just ridiculous to start with.

16:13He tweeted,

16:23Interestingly,

16:26Roos tries to peel apart more specific reasons why AI labs say that they're not more transparent. The first category of answers he writes is lawsuits. Basically here, lawyers at AI companies are worried that the more that they say about how their models were trained, the more it opens them up to lawsuits around that training. Given that the brief today started with a story about Universal Music Group suing Anthropic, and then continued with an extension of another writer's lawsuit, this is not an unreasonable concern. The second response Roos hears is around competition. Quote, Most AI companies believe that their models work because they possess some kind of secret sauce, a high-quality data set that other companies don't have, a fine-tuning technique that produces better results, some optimization that gives them an edge.

17:04If you force AI companies to disclose these recipes, they argue, you make them give away hard-won wisdom to their rivals who can easily copy them. The last argument Roos says is around safety, and certainly this is one you've heard from people like OpenAI's Sam Altman. Roos says, basically the argument here is that if you give more information about models, the faster progress around creating new models will accelerate and the more likely it is that it gets in the hands of the wrong people or just creates an arms race from which we can't escape. As Roos puts it, it would give society less time to regulate and slow down AI, which could put us all in danger if AI becomes too capable too quickly.

17:36Roos says these researchers don't buy it and neither does he. Quote, If AI executives are worried about lawsuits, maybe they should fight for a fair use exemption that would protect their ability to use copyrighted information to train their models, rather than hiding the evidence. If they're worried about giving away trade secrets to rivals, they can disclose other types of information or protect their ideas through patents. And if they're worried about starting an AI arms race, well, aren't we already in one? Certainly when it comes to the lawsuits, this is happening whether they hide the information or not, so it might not be an issue for much longer.

18:04Now, the one other type of response that I saw that is worth noting are folks who were broadly supportive of this goal, but who weren't really sure about the specific results of this test. For example, Clem, the co-founder and CEO of Hugging Face, said, The scores and ranking look super weird and inconsistent to me, and a lot of interesting models are missing, but I love the message from Stanford. More transparency equals more safety for AI. Now, I think when push comes to shove, questions of transparency and more specifically expectations around transparency are going to be dictated by policymakers rather than shaped by some industry norms.

18:35But I do think that this sort of attempt to try to articulate the dimensions of transparency is going to be super useful in trying to actually make that policy more precise and more targeted at what it's actually trying to achieve without negative unintended consequences. consequences. I think in general, it's a net asset to have this sort of information, even if it inherently remains incomplete. And so good on these researchers for going about this work. Another great topic for discussion in the AI Breakdown Discord. Again, that link is bit.ly slash AI Breakdown. Come join us, chat about transparency and anything else.

19:06And until next time, peace.

19:19Thank you.

From the publisher

A group of researchers from Stanford, MIT and Princeton have come up with a new system for determining which AI foundation models are the most transparent. Watch to figure out who scores highest. Before that on the Brief: Anthropic is sued by Universal Music Group, the FBI says AI is helping with terrorist propaganda, and more.
ABOUT THE AI BREAKDOWN
The AI Breakdown helps you understand the most important news and discussions in AI. 

Subscribe to The AI Breakdown newsletter: https://theaibreakdown.beehiiv.com/subscribe

Subscribe to The AI Breakdown on YouTube: https://www.youtube.com/@TheAIBreakdown

Join the community: bit.ly/aibreakdown

Learn more: http://breakdown.network/

More from The AI Daily Brief: Artificial Intelligence News and Analysis

All 1,099 episodes
Which AI Model is the Most Transparent?The AI Daily Brief: Artificial Intelligence News and Analysis · 19 min
Listen in VO