Researcher Claims Google Bard is Training on ChatGPT Data 馃憖

23 Feb 2024 路 6 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT 路 Add to Claude

In short

AI Today Podcast Episode Notes

Episode Overview Title: Researcher Claims Google Bard is Training on ChatGPT Data 馃憖 Description: This episode explores a controversial claim that Google Bard may be utilizing data from ChatGPT for its training, raising pertinent questions about data privacy and the future of AI-generated content.

Key Takeaways

  • Allegations suggest Google Bard is trained on ChatGPT data.
  • The implications of such practices on data privacy and AI model integrity are discussed.
  • Insights from prominent figures in the industry, including Sam Altman (CEO of OpenAI), are included.

Main Topics Discussed

Background of the Claim

  • Jacob Devlin, an AI researcher and former Google employee, resigned earlier in the year, bringing attention to the potential data misuse.
  • He reportedly warned Google鈥檚 CEO, Sundar Pichai, about the training of Bard using data from OpenAI's ChatGPT.

Sources of the Claim

  • A report by The Information cites sources with direct knowledge regarding the data usage allegations.
  • Google and OpenAI initially did not respond to requests for comments; later, Google denied the claims, stating Bard is not trained on ChatGPT data.

ShareGPT and Data Usage

  • ShareGPT allows users to share their conversations with ChatGPT, which raises questions about the public availability of the data.
  • Allegations imply that Google might be utilizing data from ShareGPT to enhance Bard鈥檚 training process.

Implications of Data Sharing

  • Concerns arise regarding the quality and originality of AI outputs if models are trained on the same datasets, potentially leading to homogenized AI responses across platforms.
  • The discussion highlights the risks of biases being perpetuated in AI models if they are developed from similar sources.

Industry Reactions

  • Sam Altman acknowledged the claims on Twitter, expressing annoyance at how Google framed the situation rather than being upset about the data use itself.
  • He pointed out that OpenAI鈥檚 terms of service strictly prohibit using ChatGPT outputs to train competitor models.

Broader Context in AI Competition

  • The podcast discusses the competitive landscape between OpenAI and Google, emphasizing the ongoing evolution of AI technologies.
  • Google's Response: In response to competitive pressures, Google has combined its AI teams, DeepMind and Google Brain, to bolster its capabilities against OpenAI.

Key Figures

  • Jacob Devlin: The whistleblower and former Google employee, known for his significant contributions to AI research.
  • Sundar Pichai: CEO of Google; involved in the discussions surrounding the data usage allegations.
  • Sam Altman: CEO of OpenAI; vocal about the ethical implications and industry practices regarding AI training.

Ethical Considerations

  • The episode raises critical questions about:
  • The ethics of data usage in AI training.
  • The responsibilities of companies in maintaining data privacy.
  • The potential for innovation versus imitation in AI model development.

Conclusion The episode sheds light on the complex interplay of competition, ethics, and technology in the AI landscape. The ongoing narrative between Google and OpenAI sets the stage for future developments in AI capabilities and the ethical framework governing their usage.

Links and Resources

  • [Invest in AI Box](https://republic.com/ai-box)
  • [Get on the AI Box Waitlist](https://aibox.ai/)
  • [AI Facebook Community](https://www.facebook.com/groups/739308654562189)
  • [Learn more about AI in Music](https://musicalai.pro/)
  • [Learn more about AI Models](https://aimodelspro.com/)

*For privacy policy details, visit [this link](https://art19.com/privacy).*

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00A report recently came out that claims Google may be using data from ChatGPT to train its own competitor, Barred. this is a lot of implications a lot of people have weighed in on this including the CEO of chat GPT or open AI Sam Altman we're gonna dive into all of this on the podcast and talk about what exactly is happening with the story but essentially what you should know is that um AI researcher who used to work at Google his name is Jacob Devlin he resigned earlier this year I think it was back in January so this isn't like super new but the reason has finally come to light So essentially he said he warned the CEO of Alphabet or Google, Sundar Pichai, and he also warned a bunch of other top executives that the company's chat GPT competitor, Bard, was being trained on data that was coming out of OpenAI's chatbot.

0:51So it's interesting to know this whole story was broke by the information and apparently they cited sources with direct knowledge of the issue as well as a bunch of people who had been briefed on it. But what's interesting is spokespeople for Google and OpenAI didn't immediately respond for requests for comment on a lot of the news stories that were written about this. Eventually a spokesperson for Google appeared to deny the report when they were talking to The Verge. They said Bart is not trained on any data from share GPT or chat GPT So essentially what what this is all is and how you know I guess what this the claim and what gets to the bottom of this is essentially there's a website called share GPT and You can publish it essentially kind of screenshots the whole message or conversation you've had with chat GPT And you can publish it on a website and share the link with someone So like if you have a really cool conversation with chat GPT Maybe you discover some cool prompts or figure out some cool things it can do and you want to have a way to share that entire conversation.

1:53ShareGPT lets you do that. And what they're alleging is that Google was going to ShareGPT's database because it's a public thing and using that to help train Bard. Now the problem with this number one is that obviously We can we'll just end up part Google bar just end up with this with as something that sounds super similar to Chai Chibut and this isn't actually very far-fetched because Recently a bunch of researchers at Stanford, right? They used a new tool from Facebook that or meta that really allows you to train models a lot cheaper and with 35 ,000 questions and answers from chat GPT they essentially cloned chat GPT's entire model with 35 ,000 it cost about 500 bucks and it was 35 ,000 questions and answers they trained a model off that and it was almost as good as chat GPT on everything and it's such a tiny model so people are able to build these smaller so people are alleging that you know that's essentially what google did because all of a sudden they call google bard maybe if they're really far behind This would be a good way for them to catch up quick now what is interesting with this whole story is Google denies it and Sam Altman the CEO of open AI he actually said on Twitter recently he said that he wasn't mad that they were He's like I'm not mad at Google's training off of open AI chat gbt outputs He's like I'm just annoyed by the spin.

3:21They're putting on this whole thing So Google's you know having this whole spin they're throwing on it. So it's interesting he's he's pretty much calling them out and saying that they are doing this even though they've denied it and what's interesting is this is actually directly against openai's terms of service so openai's terms of service says that you are not allowed to use any outputs from chat gpt to train a competitor ai model it makes sense right they spent a lot of time to build out their model and they don't want people to just essentially clone it that being said a lot of people are were responding to Sam Altman saying well we're not mad that you trained your model off of our data that you got for free so you know it's kind of like Sam Altman used a bunch of publicly available data from people to train his model and now Google's using that anyways it does lead to a small issue where like all AI models if they really were just from the exact same source and let's say chat GPT or open AI had different biases or issues with them those are just going to get passed down to the next thing.

4:21So, you know, it probably would be best if people were using their own data and training it in their own way. And I mean, it's so interesting, right? This whole conversation because, you know, it's like, is Google stealing from OpenAI, ChatGPT? And like, at the end of the day, OpenAI is using the transformer model, which Google invented back in 2017 to build their whole model. So it's kind of like everyone's borrowing from everyone. There was a lot of open source stuff. Now things are getting a little bit more closed source but a lot of these tools that originally were out there are all being used by everyone so it's like the data is mostly I mean essentially from the whole internet the tools a lot of them were open source for anyone to use and so yeah it's just really interesting to see what's gonna happen with that it is interesting to know though the guy that was kind of I guess the whistleblower on this whole Google case he's been working he worked at Google for over five years and he was the lead author of a 2018 research paper on training machine learning models for search accuracy.

5:18And it was literally one of the big papers that helped to initiate the entire AI boom. So his research has since become a part of both Google and OpenAI's language models. And he is obviously a very well-known, very reputable guy. And so a lot of people are saying, you know, Sam Altman included, that this probably is pretty legit. So it's interesting, you know, OpenAI has actually hired dozens of former Google staff over the years. And since the company's chatbot made headlines in November for its availability to, or pretty much ability to do anything from our NSA to provide code, Google and OpenAI have been locked in a pretty tight arms race.

5:59So it's going to be interesting to see which one comes out on top. Google's ramping up the pressure. They've combined their DeepMind and Google Brain, their two kind of AI teams they have over Google. They combine them together to try to better compete with open AI and it's gonna be interesting to see who comes out dominant in the AI space in the future

From the publisher

In this episode, we delve into the controversial claim by a researcher suggesting that Google Bard is being trained on ChatGPT data, raising questions about data privacy and the future of AI-generated content.

See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

More from AI Today

All 897 episodes
Researcher Claims Google Bard is Training on ChatGPT Data 馃憖AI Today 路 6 min
Listen in VO