In short
Podcast Episode Notes: Eye On A.I. #148
Episode Overview Title: Exploring AI Generative Models, Model Autophagy Disorder & Open-Source Challenges Host: Craig S. Smith Guest: Ahmed Imtiaz, PhD student from Rice University and researcher at Google Sponsor: Oracle
This episode dives into the complexities surrounding generative AI, the risks of synthetic data, and the implications of open-source AI development.
---
Key Topics Discussed
- Ahmed Imtiaz's Academic Journey
- Background as a PhD student at Rice University.
- Focus on deep learning theory and generative modeling.
- Involvement with a nonprofit initiative called "Bengali AI" to enhance AI capabilities in the Bengali language.
- Challenges of Non-English AI
- Discussion on AI's limited capabilities in lesser-explored languages like Bengali.
- Importance of creating datasets to improve AI's performance in underrepresented languages.
- Model Autophagy Disorder (MAD)
- Definition: MAD refers to generative models consuming their own generated data, which can lead to a loss of diversity and increased artifacts (such as blurring and cross-hatching).
- Experiments indicate that reliance on generated data reduces the quality of output over time.
- The Role of Synthetic Data
- Synthetic vs. Real Data: Emphasis on the need for a balance of fresh real data and synthetic data to maintain model performance.
- Potential dangers of over-reliance on synthetic data, including implications for diversity and model integrity.
- Data Collection and Privacy
- Discussion on how AI models use data from the internet and the ethical considerations surrounding copyrighted content.
- Questions about data sourcing practices and the roles of data preparation companies.
- Open-Source vs. Proprietary AI Models
- Debate on the merits of open-sourcing powerful AI models versus the potential risks of misuse.
- The importance of transparency in research and the capacity of the community to improve understanding and performance of models.
- Future Considerations
- Predictions on the growth of AI-generated content and its impact on traditional data sources.
- Strategies to mitigate the effects of MAD, including watermarking synthetic data and ensuring a consistent influx of real data.
---
Key Takeaways
- Model Autophagy Disorder: A significant risk for generative AI, stemming from models training on their own outputs, leading to artifacts and reduced quality.
- Synthetic Data: While useful, it must be supplemented with fresh real data to maintain model diversity and performance.
- Open Source Dynamics: Open-source models allow for greater scrutiny and improvement but pose risks of misuse. The balance between innovation and safety is critical.
- Data Sourcing: Ethical data collection practices are necessary to avoid issues related to copyright and privacy.
---
Conclusion The conversation with Ahmed Imtiaz sheds light on pressing issues in the AI field, particularly regarding generative models and the implications of synthetic data and open-source development. As the landscape of AI continues to evolve, understanding these complexities will be essential for future advancements and safety measures.
For a deeper dive into these discussions, the full transcript is available on [Eye on AI's website](https://eyeon.ai).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Synthetic data could be generated algorithmically. Like you could have some rendering engine rendering the world, and then you're creating synthetic data. Like it could be like a 3D object that you design and you're synthesizing images. We can have some ratio of real and synthetic data, but the most important thing is we need fresh real data compared to just having some ratio of real data. So some ratio of fresh real data is what can actually help us. If you want to compare the amount of real data being generated and the amount of synthetic AI synthesized data being generated, there's more real data being generated.
0:33There's no question about it right now. Hi, I'm Craig Smith, and this is Eye on AI. In this episode, I talk to AI researcher Ahmed Imtiaz about the phenomenon of model autophagy disorder, also known as MAD, in generative AI models. Ahmed explains how consuming their own generated data can cause models to lose diversity and become trapped in artifacts. He discusses experiments on image and text models showing that this effect emerges quickly, though it's not yet clear how prevalent generated AI data is on the wider internet. We talk about the potential strategies to mitigate MAD, including using fresh data and watermarking so generated data can be recognized in training datasets?
1:26The conversation provides an insightful look at this emerging challenge. I hope you enjoy the conversation as much as I did. AI might be the most important new computer technology ever. It's storming every industry and literally billions of dollars are being invested. So buckle up. The problem is that AI needs a lot of speed and processing power. So how do you compete with costs spiraling out of control. It's time to upgrade to the next generation of the cloud, Oracle Cloud Infrastructure, or OCI. OCI is a single platform for your infrastructure, database, application development, and AI needs.
2:13OCI has four to eight times the bandwidth of other clouds, offers one consistent price instead of variable regional pricing, and of course, nobody does data better than Oracle. So now you can train your AI models at twice the speed and less than half the cost of other clouds. If you want to do more and spend less, like Uber and Coher, take a free test drive of OCI at oracle.com slash IonAI. That's E-Y-E-O-N-A-I, all run together. That's oracle.com slash ionai. Thanks so much for having me over here, Craig. So I'm a PhD student at Rice, and I'm also a student researcher at Google right now. So I study the approximation theory for neural networks.
3:12So what I do is I think of neural networks as piecewise affine functions. and then I use this theory to see if we can explain how generative models work if we can do neural network interpretability or even say something about why the neural network has different training phases why we see such training dynamics so I come from Bangladesh I'm originally from Bangladesh I did my undergrad there and after my undergrad I came here to do my PhD I also founded, like right after my undergrad, so this is also something that's really close to my heart. I have a passion project called Bengali AI, which is a non-profit base in Bangladesh.
3:59And what we do is we create data sets and we open source them for research to accelerate research in like Bengali language technology. So LLM and also like speech recognition. Those are very primitive in Bangladesh. so um yeah that's that's interesting because i've if i'm not mistaken i've read that bengali is one of the languages that is underserved in in the training data for the big and and consequently uh chat gbt and some of these don't do well in bengali yeah yeah absolutely um and it's not it's not only like chat gpt um like for bengali we don't even have good ocrs like it's in development like OCRs and DSRs.
4:43These are also like fundamental technology that's required, especially when you think of Bangladesh. It's like in South Asia, it's so densely populated. So if you have these like groundbreaking technology that can enable people to access technology like in your own language, like if you have a Bengali chat GPT, you get all the benefits of chat GPT, but in your own language, which makes accessibility like, it makes it easy to access, right? So that's sort of like the target that we have, of making these models more targeted towards the language-specific crowd so that we can increase accessibility for the speakers over there.
5:23Yeah. Yeah. And one thing I have to ask, your screen name is Imtiaz. What's the naming convention? Because I'm calling you Ahmed. Yeah, yeah. So I have a big name, if you've noticed. It's like Ahmed Imtiaz Mayun, right? and the funny thing is like Ahmed is a very common name in Bangladesh so if you look for an Ahmed in Bangladesh it's gonna be yeah who me so yeah that's why so Mtyaz is like what I go with Mtyaz is my preferred first name Mtyaz okay great then I'll start calling you Mtyaz yeah so you're now at Rice is that right and and how how did this paper come about and then we can start talking about it.
6:12I'm interested in, I don't know if I said it before, but the paper is Self-Consuming Generative Models Go MAD. And MAD is an acronym. Maybe you can first tell us what the acronym MAD stands for and then we can get into how you came about doing the research and writing the paper. So MAD is the acronym for model orophagy disorder. So orophagy is a term that refers to consuming someone, like itself, self-consuming. The self-consumption is something that is a keyword that has come up in a lot of different literature. I recognize that it was also mentioned in Greek literature and everything. So Adolfic is the term that we chose to sort of denote this behavior that we see when generative models, when they consume their own generated data, they start behaving like in a non-standard way or by non-standard, I mean like something that we do not want to happen.
7:21So the way this research, so actually we started thinking about this problem, like I think more than a year ago now. We were, so like we always travel to conferences, like this idea of increased accessibility to generative models would obviously increase the prevalence of generated data, AI synthesized data online, right? And we, that's what the company, or that's what we want in general. Like we want people to be able to access LLMs. We want people to be able to access with Journey or such technology so that they can generate images that they want. And that's consequently going to lead to more synthesized data online.
8:01So we've been thinking about this problem and then at a conference we've been discussing with people and we see that there's this general consensus in a lot of people working with large language models and working with big data that we might be running out of data in the future because we want to go exponential and to be able to maintain that trend we need more data but we don't see that much data being generated or like it's very easy to like uh use up all the data that we have right so maybe and there's also the implications of privacy so there's this whole research direction called membership inference attacks where basically you can like this there are these algorithms through which you can find out if some data was in the training set of another model or not so if you think about that, if there's like, if all the data in the world, even sensitive data is being used to train models, then through membership inference attacks, you could be like, you could possibly find out, okay, so this sample was in the training data.
8:59So this is like true. This is like, if I have like maybe some sensitive document ID, right? And then if I can infer that this ID was inside this model's training data, then we can say, okay, so this should be a valid ID because it was in training data that's ground truth. So there's this whole idea of these attacks or there is this whole idea of privacy when we're using these data sets. So there are people who want to train models in a private manner, so they might also be using synthetic data, like synthesized data instead of using the real data, so that when those models are being attacked, like when they send it out in the open and someone attacks that model using membership inference attacks, like it's going to turn out that this is like, these are synthetic data that's been trained on.
9:47So it's not going to matter much. So there's this, there are people who want to use just synthetic data to train. And there's also synthetic data going out in the open into the internet. That's like filling up the internet with like synthetic content as well. So, and people are running out of data. So there are like these couple of directions where we see this need for synthetic data. So that the natural question arises that what is the effects of how differently is synthetic data going to play a role compared to real data because it's not real. Synthetic data has been like, so there's, I want to make a distinction.
10:20So synthetic data could be generated algorithmically. Like you could have some rendering engine rendering the world and then you're creating synthetic data. Like it could be like a 3D object that you design and you're synthesizing images. But that's like one direction. We're speaking of AI synthesized data because these AI models are getting more prevalent. People are being able to use it and generate data. So we are thinking of AI synthesized data. What are the implications? So that's... Okay. Let me just stop to unpack some of that. First of all, on the synthesized and AI generated, do both have this autophagy effect?
11:00And I ask because I have a friend from the New York Times that then started a startup ai reverie and met about it and he was they were a synthetic data company uh visual data they were you know producing uh you know entire cities rendering entire cities uh so does does that kind of data have the same effect on a large model as ai generated data or would the effect be different? The effect would be, so in my opinion, the effect would be different. Because when you're synthesizing data using a renderer, then you wouldn't necessarily have this autofaggy loop because you don't generate those images and then use those images to retrain the renderer.
11:52So if there are elements in the renderer that do training, that do learn based on the images that are being rendered, so if there's a generative model over there somewhere, then that could be affected. by this feedback loop. But in general, like, like if we render data using these physics engines, it doesn't like, that is out of the scope of our discussion. Right. And, and the, before we, we go on with it, can you just describe for listeners, what is this, this loop and what happens? And my understanding is that you, you, you lose the long tail and the data over time and you end up with, with just the mean, but it is, yeah, just describe the, the, the effect.
12:40Yeah. So, so the, the interesting part is that, so this effect basically says that if I want to like shorten these, like the takeaways, the, the effects of like a self feedback loop in any format, like whether it's like just generated data that's being used, synthetic data that's being used to train or whether it's in some ratio with training data. So the effect that we, from our experience that we see is that we would have this loss of diversity. So like you said, it's going to be converging towards the mean. So we would see more generic features compared to more diverse features. So we would lose the tails.
13:21And another very interesting thing that we see is that we see this emergence of artifacts that are subject to the algorithms that are being used to generate the data. For example, I can talk about two different generative models that we studied. So we can talk about stable diffusion that has the diffusion model method that's different from just generative adversarial network that has a different way to train. And for both these methods, we've seen this autophaggy effect or these artifacts sort of appear. And what we see is that it's subject to the algorithm itself. For example, for StyleGAN or the generative adversarial network that we've studied, we see that there are these cross-hatching artifacts that are appearing that could be related to the way these models behave or the way these models learn compared to diffusion models, which sort of have these blurring artifacts that are coming up.
14:19So it's very subject to the model, but we do see this autofagy loop increasing the artifacts in the images. And it's not only just like something that happens like in the infinite limit. It happens like within maybe 10 or 11 steps, like in a short few steps of autofagy. Yeah. And this applies not only to image generation, but to text generation as well. Yeah, so there were some concurrent papers who were studying this effect for LLMs. So they've also seen that for large language models, if you retrain or fine tune a large language model with its own data, you would see a similar effect of converging toward the mean.
15:06And the generated data, slowly, it also loses quality and it loses semantic meaning as you continue this loop. So it's a general phenomenon. Yeah. And one thing that people have asked me, and I don't know, so I'll ask you, how much, I mean, the internet is, the volume of data on the internet is, you know, massive. Has anyone been able to measure or estimate how much of that data to date is generative or AI generated? And are there any predictions for how that percentage will grow? I can imagine today it's minuscule, but are there predictions about how that might grow as generative AI spreads through the global economy?
16:04Yeah. So two questions, right? The first one is what fraction of data is currently synthetic online? It's very hard to say what is synthetic and what is not. even like that we have these methods to detect deep fakes and like a lot of technology is being built to be able to detect synthetic data but it's not completely like infallible right so so there's i i'm not aware of any studies where they're claiming that such a fraction of the data on the internet is generated but we've seen that already uh in the data that's being used to train large models like stable diffusion for image generation we see that in lay on uh the the large like open source data set that's being used to train these large models, already contains some synthetic data.
16:51And it wasn't that hard to find. So if we say that... How did you find it? There's a website called Know Your Data. And there's also another called Have I Been Trained? So through these websites, you can explore these data sets that are being used to train to see if your image was there or not. So you can search on these data sets using queries. And it was very easy to just one or two queries away, we were able to find some synthetic data on those data sets, like AI synthesized data. What was the query that you would use? So we used like avocado on a chair. That's like the reason. Yeah. Yeah. Like an avocado chair, a chair shaped like an avocado, stuff like that, like something that people would use to generate these images.
17:41And then we were able to find it. Yeah. Yeah. And that actually, I don't mean to get sidetracked. I do want to get back to the paper. But another question that people ask me and that I can't answer, there are the model builders like OpenAI, Anthropic, Google and whoever, and they're training on data. And people ask me, well, where does the data come from? And I say, well, it's on the Internet. But there's a class of middlemen, data preparers or data prep companies that are packaging the data into data sets like Lion or others. I mean, Lion is an academic data set, but there are companies that are doing this.
18:33And how are they getting the data? And for example, I was talking to my wife this morning about this lawsuit against OpenAI by Sarah Silverman, the comedian. That's a copyrighted book. How did OpenAI get a digitized copy of that book if it's copyrighted? Did they buy a Kindle version? I mean, how does it actually happen? Do you understand? Yeah, in my opinion, I think it's just how humans work. Sarah Silverman's book is quite good. And I think that's why some humans have decided to put it on the internet. People do piracy all the time, right? There are a lot of websites where there are PDFs. And I think those are maybe sources from which it might have come into the training data set.
19:29So it's like when we are going out into the internet to just get all the data, it's very hard to say like the way it generally works is like we get an initial set to then scan through and select. So that's how like data set preparation, that's how we do it at Bengali AI as well. We crowdsource that initial set and then we annotate it or we clean it up. So my assumption is... And just let me stop that. That first step, when you outsource, are those just web crawlers scraping the internet? Or are people downloading all of the Gutenberg online library or all of the Google Books online? Or how the mechanics of getting that data into a data set, how does that happen?
20:25I would assume it's being crawled because it's actually, it's generally, given the scale at which we need to train these large language models, it's actually impossible for us to have anything manual in that. That's why even in the Leon data set, they collected all the images and then they had an automated method to score those images and say whether the captions are matching up with the images or not. So these were all automated. They were like, it's very hard to do anything manual there. So if it's like a smaller data set, then we can assume that there could be some manual intervention. But for these super large data sets, it's mostly like crawled from the Internet.
21:08Yeah. So OpenAI, for example, they don't have, I mean, Google, it's obvious. They have the data through the search engine. But OpenAI, would they have hired a data collection company to go out and scrape the web? And then I imagine there's some prep in putting it into a format that then the model can be trained on. Yeah, I'm not really sure. That's something. It's kind of specific to OpenAI. I'm not really sure whether I know what they did. But we know from the public records that for the RLHF data sets, they did outsource it to some other companies who did help collect the RLHF data sets or prepare it.
22:01So I don't know if it's exactly outsourcing or if it's part of OpenAI because these are big companies. But I assume that there needs to be, for RLHF, for example, there needs to be the human element because we need the data set to align our models. But when we just need a large chunk of good quality text, I think the manual intervention would be more towards how we can curate the text in an automated way. So even that's hard, because there can be so many different cases, right? So there can be cases that you need to handle, maybe put in failsafe so that you don't really scrape the insensitive part of the internet, where you have everybody saying whatever they want to, right?
22:46like 4chan maybe who knows so there there could be like those are the parts where there could be manual intervention coming in and maybe like uh like you would need a big team to be able to do that it could be that they're using like other teams as well yeah so there would be like a white list not a white list it would be too large a blacklist you know that you cannot scrape from these domains and then you set your web crawler free and it's gathering this data. And then is that data converted into vectors or is it dumped into a giant text file or what's the next step? Yeah, generally the way it should work is, so when you're collecting, when you're scraping data from the internet, at least at my company, what we do is if we create image data sets, we generally keep track of the URLs that are being scraped.
23:49And then I think even for Leon, when they release the data set, they basically release the URLs. So if someone from the internet, like someone from whosoever web page, one image has been scraped, if they remove that image, then that URL is not going to work anymore. So the data set sort of changes. So that way the data sets are dynamic depending on like the actual owners of the data. And the only thing that the, like Leon is releasing are the URLs. So that is like one method in which like data could be shared. And so, and I'm sorry, I'm saying Leon, you're saying Leon is, am I mispronouncing? It's L-Y-O-N, right?
24:32It's L-A-I-O-N, yeah. Okay, L-A-I-O-N, my mistake. Don't worry. Yeah. So Leon, when Leon is, it's a list, it's a set of URLs. How does a company or a lab convert that into training data? Well, they would then download the data, like download using the URLs that are valid and then have their own methods to pre-select which sample they want to use. It could be based on keywords or any other filter that they have. And is that data then downloaded in text format or is it immediately vectorized and put into a massive vector database? Yeah, these are like very case specific. So this would vary between like some labs, like for some people, it could be like more convenient to have like JPEG images, for example.
25:35And then so when we think about such large data sets, there are more efficient ways to store this data, as well as like transfer this between like memory or like between locations. Right. So I'm guessing like it would be very case specific. yeah okay so back to the paper uh the internet is has some fraction of of its content as uh ai generated uh presumably a very small fraction at this point but you're already seeing autophagy or this matte effect or is that only in the testing that you do with in a closed system? Great question. We don't see
26:33autophagy is very interesting as in the effects of autophagy or the symptoms of autophagy, they come up late even after autophagy has been happening for a while. So we don't necessarily see symptoms of autophagy or we haven't really, we haven't even very rigorously explored as well if stable diffusion has any autophagy effects. It's kind of non-trivial to see if there are autophagy effects over there. But we do see that in our controlled experiments where we are using the same algorithms used to train stable diffusion, the same algorithm behind generative models, when we do those control experiments we do see that autofag is happening in a few steps so there are three different settings which we explore, so one setting is where we're completely training a model using its own generated data, so we take some generative model, we create some data and then we take another we start from scratch so we retrain another similar generative model using that generated data and see what the effects are.
27:40So this is the most extreme case. And it would be relevant for, like I was saying, people who want to use just synthetic data for privacy purposes. So that is where we would see the strictest form of autofaggy effects coming up. So we would see very fast decay of the diversity, and then we would see these artifacts come up way faster. Another setting would be where suppose we have some training data set that we are always going So we always have this fraction of real data, but then we are going to create synthetic data to augment the original training data that I have. And the reason a lot of people actually have been thinking about this is because if you think of, suppose we have two different images, one is my image and another is your image.
28:27And then you want to, there are a couple of images between these two faces. Like if we think of interpolating between my face and your face, then we would see like there can be like a continuous change in the facial features that I have that maybe would slowly lead to your face. Right. So if we have a generalized generative model, then a generalized generative model by definition should be able to do this interpolation. It has learned the data manual so it knows like how to go from my face to your face. So a lot of people, what they've been thinking of doing is like using generative models to create these intermediary faces compared to what you have in training data.
29:09You have like when you generate synthetic data, you would get these intermediary faces compared to what's in the training data. So that might help in training a newer generation of models. But these people have been thinking of using synthetic data as well. So over there, what we see is that the autofaggy effect does exist, but it's a little it's not as sharp as just training on synthetic data. but it doesn't stop autophagy from happening. We do see that if we have a fixed training set, if we keep generating more synthetic data, it's going to slowly decay towards madness. One small thing that I wanted to mention is the term mad comes from, it's actually very much inspired from the mad cow disease.
29:51Because some person had the smart idea that, okay, let's feed cow brains to other cows or something, right? and then maybe it was a good idea at first but then it turned out to have the self-consuming loop led to some very bad effects so I mean it's also it's the same principle in
30:17incest or narrow gene pools or small gene pools where yeah the so first before we talk about how you might deal with this so you're seeing it already you were saying these hash marks in the style GAN generation yeah style GANs is is is there some metric about how much uh and i understand it's model specific and algo specific but uh you know that you once you reach uh two percent or five percent or ten percent or twenty percent of generated generative uh data uh you're going to start seeing this effect or is it can you is it impossible to say the percentage of generated data versus real data. Right, versus real data.
31:27Yeah, that is a very important thing that we want to explore in the future, and we've done some explorations as well. So the most important takeaway there is we can have some ratio of real and synthetic data, but the most important thing is we need fresh real data compared to just having some ratio of real data. So some ratio of fresh real data is what can actually help us sort of like nullify defects, like maybe asymptotically. We are doing more studies here, but asymptotically, we could have like a situation where like this percentage of fresh real data on every loop would basically keep delaying like the mad cow effect to like asymptotically, right?
32:12So we will never reach like autophagy. Yeah. Yeah. And in terms of the Internet, because, of course, the popular mind, the layman's mind goes to, wow, you know, all of this AI -generated data is going online. maybe in 100 years, 90 % of the data online will be AI generated. And then, you know, this autofagy will really be a problem. Is that something that you guys have contemplated? Is the growth of generated data so high with regard to the generation of real data that that's possible? If we want to compare the amount of real data being generated and the amount of synthetic AI synthesized data being generated, there's obviously more people.
Read the full transcript
33:20There's more real data being generated. There's no question about it right now. But the thing is that the direction that we want to go towards is people adopting these technologies to help generate, to help write better, maybe, or help with the image that they're trying to create. So we want people to be adopting this. So we want all the people in the world to slowly start using these generative models. And that would be beneficial in so many domains. I think that is why it's not like a scenario that's sort of sci-fi. It's almost here. We have a lot of different websites that are completely generated.
34:06I think there was an article, I don't know if it was New York Times or Guardian, where they were reporting that there are these fake news websites which are completely generating data using LLMs. So we already have these sources that are creating these websites, only creating synthetic data. And we have millions of people using ChatGPT to write better. So slowly, from this point on, from the ChatGPT going public for use and then onwards, we're going to see a shift in the language that's being uploaded online, in the way things are being written. and we want it to be adopted more. So we'll see this effect grow as time goes on.
34:51So that's why we need to think about this right now as in like, okay, so the new generation of GPT is maybe GPT-N. When we try to train it, maybe after training, we'll see that it's very formal. Like it sounds very formal because for the past 10 years, people have been using GPT-like technology to write like more formally or in a very nicer way, right? So those effects would start coming up. And for us to not meet that point where our performance is getting saturated, we want to start thinking right now how we can deal with this probable phenomenon that we're going to see. Yeah. So what are some of the strategies that you might pursue that could deal with this?
35:39Yeah. So like I was saying, that there's the idea of having fresh new data, which can be one strategy of mitigating the mad phenomena, right? Another way could be like another easy way that we should like try right now actually is just watermark the data that we are generating. Like we already see like when DALI was released, DALI had that like sort of the multicolored watermark block right at the bottom, right? and there's also recently Google announced that they'll be using SynthID on all of their generated images so SynthID is basically they put in some really perceptible watermark into the image that you can use to see that's something I think is also being used in YouTube and a lot of other where people are uploading their content to be able to monitor if it's been copied or something so SynthID is also going to be used in generated images.
36:38So that's like another thing. There's a lot of research going towards watermarks. So watermarking could be one way to detect these images. But then if we want to use any of these watermarks, so there's this dilemma. If we want to use these watermarked images to train generative models, we have forcefully added this noise into the image and then we're using that to train the model. So that's going to increase the autofaggy effects because autofaggy basically exaggerates these hidden artifacts. That's why we see this cross-hatching coming up in StyleGAN2. So there is this dilemma of, okay, watermarking, and then, okay, so now we can detect synthetic data, but autofaggy might still happen.
37:21So then we need autofaggy-aware watermarking, meaning that once we detect the watermark, we should be able to remove it as well before using it for training so that we have the actual information intact. so this can be like an interesting like this would be an interesting dynamic field to like sort of do research in. A lot like adversarial like attack research if you're aware like neural networks like for some imperceptible change in the input can behave completely differently and the way it became such a such a like a the way the field was explored or the way it flourished was there are a lot of people working on attack methods and a lot of people working on defense methods so then there could be uh like there can also be like in terms of these auto faggy effects there could be a lot of people working on watermarks and a lot of people working on like removing the watermarks to like mitigate the the effect of watermarking on like auto faggy right so um although the point of watermarking would be uh in order to exclude those things from the training Right.
38:26Like you might want to keep, like I was saying, there could be some benefit of having some synthetic data in your training as long as you have fresh, real data. So you might want to keep some synthetic data, maintain that ratio compared to your fresh, new data that you're getting, and then try to retrain. And that's how we'll keep getting better because we want to keep getting better. That's the target. That's right. And the fresh, but in any case, watermarking, and I just interviewed the, I can't remember who he's the CEO or CTO of Digimark, who's the biggest watermarking company in the world.
39:07And then Scott Aronson, who's computer science out of Texas, who's working on safety for open AI. He's working specifically on watermarking text, which is obviously much more difficult by creating statistical patterns in the generated text that can be recognized by another system. But watermarking would be a way of excluding data, synthetic data from training sets, or at least knowing what ratio or... what ratio you have. Without that, there's really no way to know, right? Yeah, without that, like there are, like without watermarks, there are going to be ways to maybe find out what synthetic, like there are ways that are there right now, but as we keep improving, it's going to keep getting harder, right?
40:03So that's like one aspect of it. So there could be like other methods to like control this as well, because this is like a new, like this is a new phenomenon that we will learn. and there are a lot of like research groups who are starting to like jump in and study this so yeah yeah because uh finding synthetic data to train a model on is fairly straightforward once the internet is polluted with synthetic data that's not watermarked it'll be very difficult to find fresh real data, right? And it's obviously not practical and would be too costly to generate real data just to train a model. Yeah, this is a very good question.
41:00I think there might be something in the... Maybe companies in the future, they decide, okay, we're going to generate our own data. And maybe when I think about this situation, the first thing that comes to my mind is Matrix. So in Matrix, we have humans generating something for the machines, right? So it could be that we have a situation where we want to generate, the companies might want to generate their own data, could be for privacy purposes, could be just to get the fresh real data that we need to keep on accelerating. Because we do see this exponential growth, but because of such phenomena, it could be that we saturate.
41:38And then we would need to do interventions to keep growing, keep getting better models. Yeah. And when you say saturate, this is another thing that somebody was asking me this morning, actually. Is there a point at which... So the LMs are right now training primarily on public data on the Internet. Is there a point at which all of that data would have been used to train a model or models? And does that create some sort of a convergence between models, even though they're owned by different companies, if they're all being trained at some point on the same data? Yeah, that's a very good question.
42:28Yeah, the convergence of models to one sort of, it's very hard to sort of say whether two different algorithms are going to converge to the same point or not. It is intriguing that using the same data, that could be a case that happened. But we've seen that for a lot of these, like the open source models that we have, that's why I think open sourcing is so important. So because we have these open source large language models and because people are doing studies on these models, we've seen that there are a lot of different tricks you might also have to apply to get the utility out of your network.
43:09So just training on the data might not be the be all to get the performance that you desire. So I think there would be the aspect of like a recipe as well, apart from having the data. The data and the recipe together gives you the perfect dish of an amazing large language model.
43:34I've started talking to people about other AI architectures or strategies beyond generative pre-trained transformer models. I mean, there is, that is, people, it's certainly in the popular mind today, that is AI, but there's, AI is a very large field. The, and just another question related to data collection, and probably you won't be able to answer it, but you certainly could take a stab at it in a more educated way than I could. But OpenAI, you know, GPT-4, for example, has that, what percentage of the internet has that consumed in its training? Do you have any idea? I don't think anybody has any idea, like, other than OpenAI.
44:31Because with the rising trend of, like, companies, they don't want to share the information of how they're building their flagship products. so like apart from meta who's like releasing everything like i'm completely pro open source so there's i'm gonna add something about this a little while later but to answer your question like it's it's very hard to say like how much they're using uh but i would assume they're trying to use all the good quality data that they can get um yeah so so we and let's talk about llama then because uh again you're pronouncing it different than me is it llama not llama it could be llama yeah that's that's yeah that's something yeah i'm not sure like what the perfect pronunciation for llama is yeah uh but uh is it possible at this point or or very soon it will have consumed all of uh the internet um the pub not the dark web obviously but but the public internet yeah um like there's gonna be this limitation like it also like requires some time to train these models So suppose you're in 2021 and you take all the data and you start training your model, and then it'll need some time to develop your model.
45:47So it'll take one more year. So you'll always be in a situation where you're maybe behind by a year. So that's what we see in chat GPT as well. We have data up until 2020. So I think that effect would be there. We would not always have everything consumed by a network. but yeah, I think any good amount of good text that exists, I think these models would be using it. The social media company, so social media data is very different from the data we have online in terms of the blog sites, like Wikipedia and stuff. Social media data is very unstructured and it can be very, it's very hard to quality control.
46:30From my personal experience, from our company when we worked with social media data for We generally, what we do is we crowdsource data and we open source them to crowdsource the solutions as well. The way we crowdsource solutions is through competition. So we have like a Kegel competition running right now for like$50 ,000 price money on Bengali speech recognition. So we basically crowdsource the data online through influencer campaigns and through targeted social media campaigns. And then we just crowdsource the solutions to this competition. So in these settings, when we are crowdsourcing data, if we ever try to crowdsource data through social media, we always see that it's very unstructured in the sense that in social media, it's very hard to control for.
47:14That's more natural text, maybe. But more natural text, it has a lot of bad words. It has transliterations. It has code switching and spelling mistakes. So many different things. right so it's harder to work with social media data so maybe like it could be that like the social media data that exists in out in the public to be used to be used for training maybe the companies will work towards like other maybe they already have methods to clean that data up like so it's going to be case specific like social media data would have to be handled differently compared to wikipedia data or something but i think people would be doing that to get more good quality data to train.
47:56Yeah. And then you were going to say something about LLMs. I have another question, but I don't want you to forget what you're going to say. Oh, I was talking about like the openness, like open sourcing, I feel like is so important. So there's an archive paper that David Donahoe, who's like a legend in the field, he posted a couple of days ago, where he says that, so he takes a stab at the existential crisis question. right uh and and what he says is like uh and which resonates with me so much is that there are like the three things that have accelerated this field and those three things are also gonna are also it's also gonna ensure that we don't have some like we don't have a situation when we are gonna like not exist anymore right and the three things is like uh codes code sharing uh sharing code and recipes between like people like for free coach coach sharing uh the the The second thing is data sharing, which we have in all the...
48:57Everything in AI has happened because of data sharing. We have image data, right? Like someone open source that has completely changed, like completely accelerated how computer we've been in progress. And the third is competitive competitions. So these three together, like sharing large language... Like if you think of large language models, sharing the code and data for large language models, as well as sharing the code and data for these, like for generating more, like stable diffusion as well. So these are going to, these I feel like are very important to be able to assess these networks so that we can actually tackle any existential crisis question that can come up in the future.
49:45right so um yeah that's interesting that's interesting the question i was going to act ask is about open source because i've been debating with some very smart people on this who take the view that open source will not dominate because there's too much value creation in proprietary models. And in order to create that value, you need tremendous financial resources. And the open source community will never have resources to match the proprietary models. And certainly Meta is the exception, but the question is whether or not Meta will continue to invest in open source models. And there's a big question mark as to why they would do that, given the cost.
50:48And then there's, you know, that's interesting what you say about the existential threat. But these models are extremely powerful. Once they're open source, people can remove the guardrails and do whatever they want with them. And, I mean, I think Jeff Hinton, or at least Joshua Benji, quoted Jeff Hinton as saying, you know, would you open source nuclear weapons? You know, it's too dangerous to hand to any potential malicious actor out there. And unfortunately, there are a lot of them. So how do you feel about that debate? First of all, do you think open source can marshal the resources to dominate or at least compete with proprietary models?
51:39And two, what about the dangers of open source? Yeah, like working in the open source domain like myself, I think I agree with this partially that it would be very hard for open source to match the performance of proprietary models because open source efforts is basically community efforts. so wherever you need capital for making something better the closed source models are going to be better so suppose we run out of real data on the internet only the big companies have enough capital to hire teams to create new data so the open source efforts so it could be that some big company does that and open sources the data so that would be completely on them but just for the open source models and the development that's happening with open source models right now, it might be like, obviously it could be hard to like compete with the proprietary models.
52:35But the benefit is not in performance when you open source. So the benefit is in understanding what these models are doing. There are so many research groups who are studying the open, like LAMA 2, for example, they're studying these large language models to see like when they're failing, why they're behaving this way, how do we interpret them better. And when something is open source, Like it's theoretically, you have like N number of heads who can work on this problem, right? Compared to when you're in a company, you have like, you have the number of heads is limited by the budget that you have to hire people, right?
53:10So it's always going to be like open sourcing models is always going to allow more people to solve the problems with AI that we see that could harm people or that like them like having bias. So there are people studying the bias in networks. That's why. And they're being able to do that with the open source models, right? So I think open source has the benefits in that direction compared to, like, maybe they won't be able to match the performance of proprietary data, proprietary models someday. But they'll be able to help the people working with proprietary models to make decisions on how to make them better as well.
53:47Like, it benefits everyone, benefits humanity. Like, better models help everyone. Yeah. Yeah. Yeah, it's interesting on the dangers of open source. I don't remember how long ago, a couple of years ago, I interviewed a guy named Connor Leahy. Do you know him? And he started Eleuther.ai, an incredible organization that is now incorporated. But at the time, it was a totally decentralized group of hackers on Discord. and I think it was when GPT-2, maybe it was when GPT-3 came out, they said, hey, let's build an open source. And they built GPT-J and now they have Neo, I can't remember the name of the models that they have now.
54:40But Connor became so freaked out by the dangers of these models being not only existing, but being open source, that he quit Eleuther AI, and he has a startup now that works on the alignment problem on AI safety. What do you think about the danger of open source, of turning these powerful models over to whoever wants to use them? Yeah. Without a doubt, there are so many risks involved. When we open source really strong models, like we see fake websites, full of text that's being generated, right? Fake news. These are going to be like super easy. Spamming people. Like you can like generate emails and just send millions of emails out.
55:27Like we had the Nigerian prince before. Maybe now we have like the AI overlord from Mars who is sending emails to everybody, right? So there could be like these situations can happen. But like, so this is my belief completely. So what I'm going to say now. So I believe that when we think of like, when we open source, we send it out to everybody, right? All of people. and when we have such big number of peoples the law of large numbers should come into play so we would have some modal tendency of doing something and then we would have tails and I believe that humans in general the mode should be towards good because humans we to survive in society we have to talk nicely with other people and we have to behave nicely we need to help others so we can get help so I think because of that reason, when we open source, the mode would be towards good, but there are always going to be people towards the tails who would want to do bad with it.
56:25And then comes the question, how much resources do these people have? So if we have bad actors with a lot of resources, that's when bad things can happen. If you're talking about open sourcing nuclear mishaps, so if you have someone who has the capital to build nuclear mishaps, like they have access to uranium and all the everything needed to build them that's when it becomes so so dangerous right but then if you don't have access to that then it's it's just yeah i have another piece of paper that says like how to make uh like a bomb a big bomb or something right so i feel like um it's gonna be subject to that person to us humans it's gonna be on us humans who have especially who have capital right who have the the capability to use this in a bad rate, we would then that's when we need to like, that's where we need to think about like, okay, so as long as the people who want to do good are like being able to like maybe triumph over the people who are trying to do bad, as long as FBI is like hunting all the people who's like using this to do maybe like the worst possible things on the dark web like the good will triumph that the mode will be towards the good That's my belief.
57:41That's what I think. Yeah, so this idea of self-consuming neural networks, we have these amazing collaborators on this paper, especially the, like I must mention, the first two authors, Sina Ali Mohamed and Josue Casco-Rodriguez, who literally spearheaded this research and brought these amazing insights up. So everybody, go check their papers as well. They're amazing researchers. AI might be the most important new computer technology ever. It's storming every industry and literally billions of dollars are being invested. So buckle up. The problem is that AI needs a lot of speed and processing power.
58:25So how do you compete with costs spiraling out of control? It's time to upgrade to the next generation of the cloud, Oracle Cloud Infrastructure, or OCI. OCI is a single platform for your infrastructure, database, application development, and AI needs. OCI has four to eight times the bandwidth of other clouds, offers one consistent price instead of variable regional pricing, and of course, nobody does data better than Oracle. So now you can train your AI models at twice the speed and less than half the cost of other clouds. If you want to do more and spend less, like Uber and Coher, take a free test drive of OCI at oracle.com slash IonAI.
59:20That's E-Y-E-O-N-A-I, all run together. That's oracle.com slash IonAI. That's it for this episode. I want to thank Ahmed for his time. If you want to read a transcript of the conversation today, you can find one, as always, on our website, ey-on.ai. And remember, the singularity may not be near, but AI is changing our worlds. So pay attention.
From the publisher
This epsiode is sponsored by Oracle. AI is revolutionizing industries, but needs power without breaking the bank. Enter Oracle Cloud Infrastructure (OCI): the one-stop platform for all your AI needs, with 4-8x the bandwidth of other clouds. Train AI models faster and at half the cost. Be ahead like Uber and Cohere.
If you want to do more and spend less like Uber and Cohere - take a free test drive of OCI at oracle.com/eyeonai
Welcome to episode 148 of the 'Eye on AI' podcast. In this episode, host Craig Smith sits down with Ahmed Imtiaz, a PhD student from Rice University working on deep learning theory and generative modeling. Ahmed is currently spearheading his research at Google, exploring the dynamics of text-to-image generative models.
In this episode, Ahmed sheds light on the concept of synthetic data, emphasizing the delicate equilibrium between real and algorithmically generated data. We navigate the complexities of model autophagy disorder (MAD) in generative AI,highlighting the potential pitfalls that models can fall into when overly reliant on their own generated data.
We also go through AI capabilities in lesser-explored languages, with Ahmed passionately sharing about his initiative "Bengali AI" aimed at advancing AI proficiency in the Bengali language. Ahmed introduces pioneering strategies to differentiate and manage synthetic data.
As we wrap up, Ahmed and I deliberate on the merits and challenges of open-sourcing formidable AI models. We grapple with the age-old debate of transparency versus performance, juxtaposed against the backdrop of potential risks.
Dive into the world of AI, synthetic data, and deep learning and join the discussion with Ahmed Imtiaz, as we tackle some of the most pressing issues the AI community is facing today.
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI
(00:00) Preview and Oracle ad
(02:56) Ahmed's Academic Journey
(04:14) The Challenge of Non-English AI
(06:34) Model Autophagy Disorder Explained
(14:40) Internet Content: AI's Growing Involvement
(21:08) The New Age of Data Collection
(26:28) AI's Role in Protecting Digital Assets
(38:51) Open-Source vs Proprietary Model Debate




