Dr. Olga Russakovsky: Shaping the Next Generation of AI Leaders

19 Dec 2024 · 45 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Notes: Generative Now | Episode with Dr. Olga Russakovsky

Episode Overview Title: Shaping the Next Generation of AI Leaders Host: Michael Mignano, Lightspeed Partner Guest: Dr. Olga Russakovsky, Computer Science Professor at Princeton University and Co-founder of AI4ALL Description: Dr. Russakovsky discusses the education of the next generation of AI talent, the field of computer vision, interdisciplinary research, and the importance of diversity in AI.

Key Themes and Discussions

  1. Introduction and Guest Overview
  2. Michael introduces Dr. Olga Russakovsky, emphasizing her significant role in computer vision and AI education.
  3. Discussion about her research background and contributions to AI, including her work on ImageNet.
  1. Olga's Career Journey
  2. Transition from theoretical machine learning to applied AI, focusing on computer vision.
  3. Passionate about understanding and explaining AI systems.
  1. Understanding Computer Vision
  2. Definition: Interpreting and understanding images and videos (e.g., object recognition, autonomous driving).
  3. Real-world applications:
  4. Medical diagnostics (e.g., skin cancer detection).
  5. Agricultural monitoring and conservation efforts.
  6. Space exploration.
  1. Generative AI and Computer Vision
  2. Generative models (e.g., GANs, diffusion models) have become integral to computer vision.
  3. Discussion on the intersection of generative AI and computer vision, highlighting the excitement in the field.
  1. Interdisciplinary AI Research
  2. Emphasis on the need for interdisciplinary approaches in AI research, beyond traditional computer science.
  3. Importance of integrating social sciences to understand the broader impact of AI technologies.
  1. AI4ALL: Diversity of Thought
  2. Overview of AI4ALL's mission to increase diversity in AI.
  3. Emphasizes that diversity in demographics is crucial for enhancing creativity and preventing echo chambers in AI development.
  1. Challenges and Bias in AI
  2. Discusses the various sources of bias in AI systems: data collection methods, representation issues, and value systems of those building AI.
  3. Importance of addressing and mitigating bias at all stages of AI development.
  1. Future of AI and Data
  2. Concerns about the over-reliance on data.
  3. The evolving role of computer scientists as they engage with ethical, societal, and philosophical questions surrounding AI.
  1. ImageNet and Its Impact
  2. Insights into the creation of ImageNet and its influence on the deep learning revolution.
  3. Discussion on the importance of large datasets for training AI models.

Key Takeaways

  • Interdisciplinary Collaboration: The future of AI research requires contributions from diverse academic backgrounds to address complex societal challenges.
  • Importance of Diversity: Increasing diversity in AI development teams can foster innovation and creativity, leading to more equitable AI systems.
  • Bias and Ethics: Continuous attention to bias in data and algorithms is essential for developing fair and responsible AI technologies.
  • Education and Training: Programs like AI4ALL provide pathways for underrepresented groups in AI, focusing on responsible AI practices.

Closing Thoughts

  • The podcast encourages reflection on the direction of AI research and the importance of questioning current methodologies and paradigms.
  • Dr. Russakovsky's work illustrates a commitment to shaping future AI leaders with a focus on responsibility, inclusivity, and ethical considerations.

Stay Connected

  • Follow Lightspeed on social media:
  • [Twitter](https://twitter.com/lightspeedvp)
  • [LinkedIn](https://www.linkedin.com/company/lightspeed-venture-partners/)
  • [Instagram](https://www.instagram.com/lightspeedventurepartners/)
  • Subscribe to the podcast for more insights on AI developments: [Generative Now](http://generativenow.co/)

---

Note: This markdown file summarizes the key points and discussions from the podcast episode featuring Dr. Olga Russakovsky, highlighting her insights on AI, computer vision, and the importance of diversity in the field.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:04Hey, everyone, and welcome to Generative Now. I am Michael Magnano. I am a partner at Lightspeed And this week on the podcast, we're talking to Dr. Olga Rysakovsky, a computer science professor at Princeton University. She is a leading researcher in the field of computer vision, having worked on many important breakthroughs in the field, including ImageNet. She's also the co-founder and board chair of the nonprofit AI for All. This was an honest and surprising conversation about who is shaping the future of AI, what today's computer science students are learning, and why the next innovation in AI may not come from where you expect it.

0:38So check out my conversation with Dr. Olga Rosakovsky. Hey, Olga, so good to see you. Hey, Mike, thanks so much for having me. Yes, thank you for doing this. I've been looking forward to this for a while. I have a bunch of questions for you. I want to learn from you in this hour that we have. Like so many people have probably in your career. I guess before we dive in, though, I think we should tell people a little bit about your career. And so, yeah, do you mind just sort of telling people about your work and your research and everything you've been working on these past couple of decades in AI?

1:17Yeah, absolutely. So I'm an associate professor in computer science at Princeton, also associate director of the Princeton AI lab. Let's see, I started my career at I would describe starting my career at math camp in high school, was a math major in college at Stanford. And then went to do a PhD in theoretical machine learning, sort of got very interested in machine learning and in AI and wanted to sort of do the more theoretical work. But over time, kind of drifted more and more into applied AI research or drifted into computer vision, which I think is just fascinating. I know there's a lot of hype and excitement about natural language processing these days and about text processing, but I still think computer vision is by far the most interesting one.

2:03Sorry, no offense to anybody, but that is very much my passion. And so I've been doing a lot of work both on building computer vision systems, but also kind of studying them, analyzing them, thinking about explainability of these systems. Recently, I've been thinking a lot about AI fairness and bias, particularly, you know, in computer vision systems and in AI systems more generally. So that's kind of, yeah, that's kind of my trajectory. Do you mind just explaining computer vision a bit? I mean, I have I have an understanding of it, but just for the audience sake, like dive into computer vision.

2:38When you say that, what does that what does that mean? Should we be thinking, you know, self-driving cars or more just sort of identification of objects within within images? Like tell us more about computer vision at a high level. Yeah, I think it's anything that's related to understanding pixels, understanding images or videos. I think you can think of autonomous cars for sure. You can think about when you upload a photo to any social media platform and sort of says, OK, here are the faces. Here is maybe who these people might be. It automatically tags your friends. You know, you upload, you create a photo album and it sort of sorts your photos by not just by date, But by sort of who is who is in the photos, what what is the scene like?

3:21So but also, I mean, you can think about a lot of applications like medical applications. You can think about sort of computer vision that's used to do to diagnose skin cancer. You take a photo of a of a lesion on your skin and it sort of is able to tell you, OK, should you see the doctor? Is this worrisome? So you can also think about applications to things like agriculture. So both in terms of building better systems that are able, you know, better robotic systems that are able to sort of operate on the ground and do some of the agriculture tasks. But you can also think about like aerial photography and trying to understand conservation efforts, trying to understand sort of how the Earth landscape is changing over time.

4:05And that's also computer vision. You're sort of analyzing these photos that are taken from space. So there's all kinds. You can also think about, you know, space exploration. You send a robot to Mars, right? It's trying to kind of navigate and move around in its environment. You're not going to have somebody steering it with a joystick, right? That's not going to operate just because of the latency, right? And so it needs to understand to be able to have a camera that looks around, understands what it's seeing, kind of help guide it to, you know, drive around safely. So, yeah, all kinds of things.

4:37Tons of applications of computer vision right now, probably more than ever before. in the field. And the other thing I think of when you start talking about computer vision is something that I don't think is the same, but I wonder if there's intersection, which is kind of the other side of images right now with an AI, which is more around the generative AI of sort of diffusion-based image models. Where does that sort of intersect with computer vision? Or are they sort of like two completely unrelated fields? Yeah. So it's very interesting because I I think initially we would think of that as a sort of computer graphics, as kind of generating, you know, kind of generating visual data.

5:17But it's very much become part of computer vision. So that was actually a very interesting transition. So I think when generative models, sort of GANs, you know, generative adversarial networks, diffusion models started coming out. So that's now very much part of computer vision. And that's been sort of interesting that it's kind of become both the images to some kind of understanding of these images, but also going back. And so, yeah, so we do a lot of work on diffusion models in my lab as well. And that's now we've kind of embraced that as part of the computer vision research. It's probably, I'm guessing, driven a lot of interest, like new interest in the field.

5:52And I wonder what that's meant for your work and your research and maybe the computer science department at Princeton. Yeah, so there's definitely, I mean, for the computer science department, we're the largest major on campus. And I think that that's that sort of says it all. A lot of undergraduate students were very interested and very interested in taking courses, very interested in doing research in this space. Then, you know, lots of applications from prospective PhD students, you know, graduate students. And I think we're trying to, you know, hire more and more faculty in this space to sort of keep up with the interest.

6:26I think the new Princeton AI lab is a good example of sort of trying to connect all of the folks doing AI research on campus and sort of grow that. I think it's when I think about AI research, I actually think of it as being broader than just computer science. I think there's a lot of I mean, like you say, it's sort of the advancement of this technology has driven a lot of interest, but it also raises a lot of more social questions. It raises a lot of questions about the impact of this. It raises a lot of questions about how people are actually interacting with these systems. It raises questions about impacts on different fields.

7:03And so I think when we think about computer science research or AI research, you know, in the past, AI research has been much more of kind of engineering or mathematical endeavor. And I think now it's becoming much more interdisciplinary. So if you need to really connect to folks who have studied things like societal shifts in other contexts. and now can help us kind of grapple with how should we think about some of these changes in AI? How should we be thinking about analyzing and understanding these models, which we no longer can analyze or understand using kind of standard methods? And so there's just a lot of opportunity to grow this field and grow the way we're thinking about it as well.

7:45Yeah, that makes a lot of sense. I mean, I feel like, you know, more and more you're seeing people within AI or tech more broadly think about, like you said, the societal impacts or, you know, get philosophical and compare and compare back to previous times or technological innovations in history. So, yeah, I've actually been thinking quite about that. Is the science, the computer science aspect of this, like over time, does that become like kind of less relevant and sort of the field of AI becomes, as you say, like a little more generalizable and sort of applicable to almost like more different like humanities related topics like it's super interesting like i maybe just to take that a bit further before i before i hand it to you but like i was a computer science major and um you know we didn't we didn't talk at all about philosophy or societal impact we just learned how to program in a bunch of different languages right and like learned about like this you know the the hardware and the with a lot of these generative models, it seems like we could get into a point where that stuff, the type of stuff that I learned with my computer science degree, actually becomes less and less relevant over time.

8:54And the stuff I think that you are talking about becomes more and more relevant. So, yeah, I mean, how do you think about the future impact on computer science as a result of AI? Yeah, so it's a very good question. I mean, so first of all, I think to your point about, you know, learning about the different, you know, coding language, you know, different programming languages and so forth. I mean, I think that is becoming somewhat less relevant by virtue of the large language model. I mean, we've seen a shift in how people code, right? We've seen a shift in sort of some of the models actually kind of assisting you in coding.

9:32And I have, you know, it raises a lot of interesting questions about how do you design assignments in computer science courses? And how do you think about, you know, the fact that historically a lot of the assignments are around, you know, write this code to solve this particular problem. But now, you know, you can ask Chad GPT and it'll write that code to solve that exact problem because it's been in computer science assignments all over the world. And now the model has learned from all of that data and it can easily, you know, solve all of these kind of classic problems that we teach students to solve.

10:02So I think it raises questions about what is the role of programming and programmers. And I think that's sort of one can of worms. I think to your point about that sort of computer science will become less relevant. So I actually don't think so, because I think building these systems and sort of the technical work around building, deploying, sort of developing levers of control and things like that, I think that's going to stay core technical, but I think a lot of it is going to be driven not just by kind of intellectual curiosity of can I do this, but driven by sort of concrete needs and applications and concrete questions that are being asked.

10:48And those applications are going to be very broad, but also the kinds of questions that we're going to ask about sort of what does it mean to understand the system or trust the systems or what kind of safeguards do we need around these systems or how are these systems actually influencing society and what do we need to sort of steer it more to influence it for the better rather than for the worse. I think a lot of those questions are going to be asked sort of by scholars who are much more interdisciplinary. And so I sort of I don't think the computer science role is necessarily shrinking. I think computer science role is kind of staying the same, but maybe the field is kind of growing to bring in more and more ideas and questions to kind of help steer the computer scientists.

11:37Yeah. How much is that actually happening right now? Like, you know, is curriculum changing even now to head more in this direction? Or is this something you see like further down the line? Yes, I think this is happening right now. I think the curriculum is changing in many ways. And I think the connections to sort of interdisciplinary work is very much, very much happening. So later today in a few hours, right, so I'm so my co-instructor and I are teaching the last lecture of the undergrad computer vision course. at Princeton for the semester. And what we're doing is we're hosting a fireside chat with Dr.

12:14Molly Crockett, who is a psychologist from the Department of Psychology. And they're going to come in and talk about one of their latest papers on sort of AI as used for scientific discovery and sort of what are some of the pitfalls of using AI in sort of scientific research or how it sort of creates an illusion of understanding, but really like, what does this mean for scientific research? And so that's sort of the conversation we're going to have in last computer vision class for the semester, because we think it's sort of important to ask some of these more fundamental questions in addition to, you know, a semester of teaching about convolutional neural networks and, you know, deep learning and different methods and different techniques for building the systems.

13:05Can you give us a little preview of what the pitfalls are of using AI to make scientific discoveries? That sounds super fascinating. I'm not going to give it justice, but there's a lot of sort of opportunities to speed up the research process using AI. My interpretation of this is that it can drive research towards sort of what is currently the most sort of promising direction. I think sort of the scientists will start using kind of similar tools to summarize the data in similar ways and will drive the field towards, you know, you're trying to sort of get a paper published. You're going to, you know, select sort of the current thing that sort of is best supported by current evidence or will kind of is the easiest to put together the data for.

13:58And you're going to analyze the data in particular ways. And it's ultimately it's kind of going to limit creativity. It's going to limit creativity and sort of some of the joys and potential of scientific discoveries that you are thinking about these things in different ways. ways you're going down the wrong path, you're failing and then you're succeeding. And some of that breadth of approaches and thinking about things, that's what's sort of leading to scientific discovery and sort of relying too much on AI models to steer that is going to sort of lead to, I think, short-term success, but not long-term success.

14:39Yeah, that makes sense. It's almost, you know it's gonna it's almost like gonna converge to the path of least resistance almost in a way whereas the more the the research is happening directly via humans like you said to use the word like creativity there's randomness random ideas that that get explored um that makes a lot of sense maybe good tie into ai for all uh this is a big part of of your work and your research. Well, I'll let you explain it. Tell us about AI for All. Yeah. So, so AI for All is, you know, what I do in my copious free time. So I'm a co-founder and now chair of the board of this nonprofit.

15:18We are working to increase diversity and inclusion in AI. So I think I personally see the big, you know, I know there's a lot of conversations about what is the biggest threat to AI? What is the sort of existential threat? There's a lot of this talk. And to me, the biggest threat of existential threat is the lack of diversity of thought in this field. You can, you know, it's harder to measure diversity of thought. It's easier to measure diversity of demographics as sort of a proxy for diversity of thought. And you can look at, you know, within AI, there's different statistics, but it's around 15 % women.

15:57It is very few black and brown folks. It's very few folks who are black, very few folks who identify as Latina or Latina. It's very few indigenous folks. And what that's doing is it's driving us towards echo chambers and decreasing the diversity of thought in the field, decreasing the creativity with which we're able to approach some of these problems. And ultimately, I think going to drive this field into the ground. I think we're not going to reach the full potential of AI if we keep going the way we're going. And, you know, just just to sort of, you know, I know the question, some of the questions about like demographic diversity get into very complicated sort of legal and ethical questions these days.

16:48And I totally get that. But one way one way that I like to think about it, which I think is a lot less controversial, is that the reality is folks who are working in AI these days come from particular schools of thought, particular types of training. particular sort of top schools that have produced these students. A lot of them read similar books as kids. A lot of them play with similar toys as kids. A lot of them run in sort of similar social circles. So they talk to folks who are like them. And then all of this kind of translates into how we think about building the technology. A lot of this translates into the kind of applications they care about.

17:30This translates into their value systems. This translates into kind of how they approach problem solving. And at the end of the day, this limits the creativity of the field as a whole. Yeah. You know, oftentimes I feel like when this topic is discussed, like, you know, what I gather and what I learn is that, but I would love for you to tell me if this is right or not, is that a lot of the bias in these models comes from the data that it's trained on. Sounds like what you're saying is it also comes from like the way that these systems like are actually designed. Can you can you like can you add a little more specificity on on like how they might be designed in such a way that it leads to to to a bias in one direction or the other?

18:17Absolutely. So I think the bias comes sort of at every stage of the pipeline. So you mentioned, you know, I mean, data is it's actually interesting because that was a controversial point a few years ago that people would sort of disagree that it comes from data. But but that's actually become sort of well accepted that. Yes. So so a lot of the bias will will there's bias in the data. There's bias in both sort of the way that the data is collected. So, for example, you know, I'll just give I mean, I know you want to talk more broadly, but let's just give one example on the data bias. So if you look at the common computer vision data sets, you can run analysis on geographic distribution of that data.

18:57So you can, you know, using a lot of this data comes with GPS tags, right? You just grab the GPS tags and you can plot sort of which countries does this come from. This comes primarily from the U.S., countries in Europe. And that's more or less it. So you can you can there's there's various papers, like including some of the work from our lab, but lots of other folks as well, where you have this map of the world and you have the U.S. highlighted, you know, bright colors and parts of Europe highlighted bright colors. And then sort of the rest of the world, in particular, sort of South America and Africa are just completely missing from from this data.

19:32And so then you get things like object recognition systems that are purported to sort of, oh, it can recognize all objects in the world. And then in practice, if you feed it a picture of a house in Africa or you feed it a picture of a plate from Africa or you feed it, you know, a picture of one of the classic examples is like bar soap. And it will refuse to recognize bar soap as soap, but it will recognize the US brands of liquid soap as soap. And so, you know, so that's sort of the data answer. But I think thinking more broadly, right, what are the big applications that people are working on now?

20:08Like thinking about sort of, I mean, we're working a lot on autonomous driving and I think there's a lot of power in autonomous driving. But if you think about, well, where is this coming from? We have a lot of folks in the Bay Area who are sitting in traffic for, you know, two hours a day, you know, each way. Right. And so, of course, we're going to work a lot on autonomous driving because this is sort of front and center on people's minds. And I want to be clear, there's lots of, you know, excess, like very important sort of accessibility needs and environmental impacts that and, you know, the 100 people that get killed, economic impact.

20:44And also, you know, I think in the US, 100 people a day die in car accidents, right, which can be remedied. But if you think about the relative lack of work on solving hunger crisis and solving sort of economic, and I think a lot of that sort of comes from who is the help, like who is driving some of this, some of this work. And if you look at, you know, if you look at startups, if you look at research projects and sort of different students undertake that there's you will kind of immediately see some of the connections where what people are passionate about, what they're choosing to work on is very much influenced by their values, by their culture, by their upbringing.

21:30Right. That's that's what they feel passionate about. That's what they're going to work on. And so if we want all of these applications to be actually sort of solved and tackled and get the attention that they deserve, then we need people who are going to be passionate about that kind of work, about sort of each of those. how does it happen though when we when as a society at least in the u.s we have the incentives that we do right where i mean built around you know the capitalistic structure that we all you know we all exist in and the motivation of you know of private private companies i mean yeah i definitely hear you on the autonomous driving thing like from a humanity standpoint there probably are like more important things but um you know uh a company or a corporation or you know an organization trying to to to maximize economic impact like you can see why it leads to things like um full self-driving or identification of the the you know the to use the example you used earlier of like to identification of the people in the photo for the social media site you know what i I mean, so like how does how yeah, how do you shift it away when the incentives like very clearly point in the other direction?

22:46So I don't think we know what will happen with more diversity of thought in the field. I very much hear you. Right. And I'm not, you know, I'm not sort of pollyish about this. Like I grew up in the Bay Area. I understand sort of economic. I get it. I don't think we know sort of the kinds of things that can be done that can be potentially sort of aligned with with the economic incentives as well. Right. I suspect there is, you know, like. I don't know, the food industry is is changing, I think, in many ways. And I think that's, you know, driven by various economic incentives. And I think, I don't know, kind of the random things that come to mind are sort of like development of of the like fake meats.

23:33Right. The like different kinds of like impossible. Right. And and that's something that, you know, would we have thought before this started happening, that this is something that would have, you know, enough economic incentive to to drive that. But yes, yet it seems to be happening. I mean, I think something I mean, there's a lot of money in the medical system. Right. So there's like economic incentives to drive some of the medical innovations. I mean, I think we're also kind of limited because we're seeing what's what's happening now. I don't think we have the like the full creativity that we need to reimagine, like what are the different kinds of applications that could be tackling and could be aligned with the with the economic incentives, but also aligned more with the full range of things that we could do with AI and that could have both sort of economic and social and sort of social good incentives.

24:21Yeah. So how do you how do you do that? How do you how do you get people building in this way? So what we're trying to do at AI for a little sort of train more students and provide pathways into the AI space and sort of concretely thinking about, you know, I mentioned in the beginning that sort of, you know, the students that are going into this space come from a small set of universities that have very good AI training programs. And, you know, I'm proud to be at one of those universities. Princeton students get a lot of training in AI, but there's lots of universities around the U.S. that don't, that don't have the AI faculty, that don't have the strong AI curriculum or that, you know, don't have the capacity to teach students, you know, how to pursue some of the, you know, their own passions in AI.

25:09Maybe they have sort of the basic courses, but not the opportunity to really do, you know, independent work or projects. And so at AFRL, we run this program called AFRL Ignite for college students, particularly targeting Black, Latinx and Indigenous women and non-binary students from around the U.S. that come into our program. It's a year long program. It combines both sort of education on AI, in particular on responsible AI. So sort of thinking about both teaching them some of the core machine learning skills, but also getting them to think about data, about evaluation of these models, about sort of all of the various decisions and values that get embedded into the systems.

25:50And then they get to work on a portfolio project guided by industry mentors. And there's, let's say there's lots of folks in industry who are very eager and willing to contribute their expertise. And maybe they don't themselves come from some of these backgrounds, but they're very passionate about volunteering with some of these students and sort of bringing in more younger, more diverse, passionate voices into this space. And so they guide the students through a, I just want to say a passion project of theirs. I mean, it's a portfolio project. They're working on something concrete, but it's driven by the student's interest of sort of what do they want to explore with this technology?

26:30And then we provide some, you know, career readiness and career prep workshops and sort of helping them put together their resume, helping them think about how they're going to apply for sort of their first paid AI internship and kind of get their foot in the door. And then from there, you know, hopefully the world is their oyster. Hopefully they've, you know, at least sort of gotten to experience some of the, you know, joys and hardships of building in this space and they can they can take it in in the direction that they want from there. So it sounds like it's about bringing in different and more diverse types of people into the field rather than, say, like deliberately changing the way the technology is built to maybe like correct for something.

27:17I'm just thinking of like this example that I'm sure you saw. I don't know. Maybe it was like six months ago with Google, Google Gemini, where, you know, Google was was was worried about like the bias of the model. And so they I don't know exactly how they did it, but they sort of overcorrected it in the other direction to the point where it was basically factually incorrect about a bunch of different things. You're not advocating for that. What you're advocating is for bringing in different perspectives, different types of people that typically are not involved in the field to solve for this.

27:50Is that right? Yes. Got it. Yes. And we've done, I mean, like a lot of the research in my lab has been about, you know, tackling bias and sort of how do you define bias? How do you think about bias in these systems? How do you measure it? How do you develop algorithmic methods to correct for it? And all of that is really hard. It's very hard. And then the root problems and the root cause is that we are not being creative enough or thoughtful enough or diverse enough in how we approach this, right? So some of the, you know, lack of geographic diversity in the data, I mean, why aren't we? I mean, we're sort of using the simplest data collection approach.

Read the full transcript

28:31We download things from the web and from social media platforms that we know. That's the cheapest source of data. But like, why can't we think about going to some of these countries and sort of collecting data? But all of that can be all of that can be done. If you think about some of the folks who have been leading voices in the AI fairness space and some of the people who have identified some of the early issues of bias in these systems that have been people who come from more diverse backgrounds and perspectives. And they've asked those questions that other researchers did not think to ask.

29:07And I think, I mean, I can give sort of one example of this, which is when I was starting as a PhD student. So we were building the lab that I was in. We were building this robot and the robot would understand, you know, follow commands based on based on language. So you would sort of give it a command and follow the command. Right. And this voice recognition system would understand everybody else in the lab except for me. And other folks were from, you know, many different countries with like different accents. in the room, but would have no problem understanding it, but it would not understand me.

29:38And like, what do you know? I was the only woman in the lab. And until you have that woman researcher who is in the room, who sort of tries to give the command to the robot and it just outputs complete garbage and has no idea, you don't realize this, like you don't ask these questions. You don't think to ask this. And so then you can't solve it if you don't think to ask this question. Yeah, super, super interesting. You know, back to your point about how in terms of how many of these companies models are collecting data and therefore like lacking data, maybe from certain parts of the world. you know the the counter the counterpoint i've heard to that is like well when they go you know when these models go out and they collect all all of the data and the entire entirety of the internet like that that actually represents the ground truth of the data because that's everything that exists and actually if we sort of try to go out and manufacture additional types of data to like unbiased the thing we're actually like skewing away from the ground truth like how do you respond to something like that.

30:44I think it represents the ground truth of what's been put on the internet. So then I would ask the question of who has access to the internet. Exactly. And who has chosen to upload photos to the internet and for what purpose, right? We upload a lot of photos for advertisement. We're trying to sell products. That's what we upload. Or social media. Social media, right? You're trying to get the most likes. You're going to upload the photos that have particular properties or you're going to upload videos to TikTok. Like, it's a carefully curated, like, is it the ground truth? Or is it a very carefully curated subset of data?

31:18Unintentionally carefully curated, right? Like, it's not intentional. Nobody's, like, intentionally trying to exclude data. But I think as a result of the use cases, you only end up with this, like, very selective group of data. No company wants to intentionally output a biased product out there, right? Of course. That is nobody's interest. Like, yeah, nobody wants to be that, obviously. So. We take the path of least resistance, we do. But I think sort of your point about, like, is this ground truth? I mean, I think it's important to remember this is not ground truth. Right. And that's more we're moving into like sort of SDS, like science, technology and society studies, which computer scientists historically don't know those spaces very well.

32:03Like we, when you get training, I mean, you said you're a computer science major, right? I was, you know, I was a math major and then, you know, graduate degrees are in computer science. And as a math major in college, I avoided like the plagal humanities courses. Like I tried to take the fewest number possible. Like I, you know, we had five that we had to take, but you could double up and take some that sort of counted for two requirements. And so I took three very strategically, just trying to take the fewest number. And in retrospect, I'm horrified by that. And I really wish I hadn't because I feel like I'm playing catch up and trying to learn all of these things that I just never got training in.

32:35And but this is why, you know, again, I think coming back to diversity of thought, like, well, we need to bring in AI researchers who are interested in doing some of the technical work, but also have this training in social sciences. Right. Who have some of that background in knowing how to ask some of these broader questions and how to think about the technical building of these from the from a broader perspective. Yeah, something that's coming up for me, which is super interesting, is, you know, there's a lot of talk right now. I'm sure you've heard it, especially from the companies that are building, you know, the large language models about how we're reaching or we may be reaching some sort of data wall or running out of data.

33:18We've already trained on the entirety of the Internet. And so like where we, you know, we need to go out and we get we need to find more data. Right. Or we need to generate synthetic data. We should get to the synthetic data thing. But maybe before that, it sounds like, you know, what you're saying is actually maybe the solution to the data wall problem is actually the same as making the same to the problem of making these models unbiased. I don't think there's a such a thing as an unbiased model. I think there's I think we can mitigate bias models and we can mitigate problematic behavior in models.

33:54But I don't think there's a ground truth of unbiasedness. But the other part of the of your question, I fully agree with. Right. I think the solution to a lot of the issues we're running into with AI is this lack of diversity of thought that, you know, we are taking like, for example, we're taking for granted. Right. that data is going to be what's driving it in the future. But if you think about AI 20 years ago, data was not at all valued. Like the shift towards valuing data and sort of viewing data as sort of data collection, data curation, and sort of engaging with the data, viewing that as a valuable part of the field, this shift has happened in the past 10 years, maybe 15 years.

34:43And this is with sort of ImageNet kind of being the first example and kind of, I would argue, sort of the transformative moment when folks started really focusing on data and thinking about data as being important. But even with ImageNet, even after that, sort of for many years, trying to get a paper on a new data set into a top computer vision or machine learning or sort of AI conference was a huge lift. Because people would say, why are you trying to publish a data set? Like, whatever, it's just data. You know, what is the algorithmic novelty and algorithmic innovation? This is something that many of us have kind of fought against for years.

35:22And if you look at one of the top ones, so NeurIPS, they just introduced the datasets and benchmarks track a few years ago to allow for sort of create a space in this community for publishing dataset papers, sort of papers focused specifically on data and on benchmarking and on sort of evaluation of these models. And the reason why this was needed is because papers were getting rejected from that conference that dealt just with data because we were saying, well, whatever the insights are in the technical and algorithmic innovations, it's not in the data. And so I think the long story short, what I was trying to get at is right now, suddenly everybody is saying, oh, you know, we need data.

36:11Data is the driving force. But first of all, that itself is a new idea. or sort of relatively new idea in the space. Second of all, I mean, that's what we know how to do now. We've sort of stumbled upon this solution that involves digesting tons of data and then outputting these models. And it's great. But like, is that the only path? We don't know. Like, are there other alternatives? I mean, I think, you know, right now sort of the alternatives, right, generated data and you use generative um generated data to composite synthetic right is that a real solution or i think to me that seems like that was a solution but wouldn't that actually just either wouldn't that further increase sort of bias and sort of repetitiveness of of output like we're just it's just the same stuff getting recycled over and over right so so so i thought that but um i i think you can be thoughtful about how like like yes and no i I think it's not going to help with issues of like geographic diversity.

37:17If there's no representation of data from particular countries in your original data set, no matter if sort of generated data will help with that. But I think it can help with things like sort of a lot of the data that's uploaded is, you know, again, coming back to your point about trying to get likes or it's sort of framed in a certain way or there's certain composition of, you know, it's always these three objects that sort of appear together or the person is always posed in a particular way in the photo. That kind of stuff, I think you could get generative models to help with so they can help break some of that.

37:55You can generate from a sort of different compositional distribution than the original distribution. So, for example, like the easy example is if you. So this is one of my students was saying this at a meeting recently. So it's top of my mind. So if you always have sort of apples on top of tables in your in your data set, very rarely do you have sort of apples under tables. But but if your model sort of understands the difference between like apples and tables, so now you can generate data that where the apples are also sort of under tables or near tables or there's. So you can kind of change the image composition and you can change sort of the distribution of appearances and connections.

38:39Sort of, yeah, no different word than composition. And so there's creative ways that you could go about manipulating this. It kind of hinges on having control over the generative model, sort of having control over what exactly it's generating. So it's not just generating from sort of uniformly at random from the distribution it's learned, but it's sort of generating more from particular kind of subsets of the distribution or generating kind of different compositions of objects or concepts and so on. Yeah. Do you think there are any companies or labs that are sort of best positioned right now to sort of in a scalable way, unlock new forms of data from maybe geographically or throughout the world?

39:29Like, what are the companies that are actually like going after this actively right now? I'm going to put in this and say that it's AI for all. Who's positioned to unlock all of these different types of things? I think it's going to be the companies that are founded by our students. I think it's going to be folks who are coming into the space who are right-eyed, who have not yet been sort of indoctrinated into the current ways of thinking and who are coming in with new ideas and who ask these kinds of questions of like, can we do something that's radically different from what's being done right now?

40:07Yeah. So you worked on ImageNet as a PhD student. Tell us what that was like. The original ImageNet was constructed by my colleague, John Deng, who's at Princeton, and then was advised by Faithe Lee, who is my PhD advisor. I sort of worked closely with them. I was around when this was happening, but I was a junior PhD student and I was sort of watching them in awe and watching this unfold. And so I was there sort of from the early days. But it's a very large scale data set that they collected of photos from the photos from the web sort of illustrating different visual concepts. There were a number of sort of key innovations, including the use of crowd workers, the use of sort of people in labeling all of these all of these images.

40:49The whole data set is more than 20 ,000 classes and it's 15 million images. Sorry, there we go. And so all of these images are sort of carefully labeled by people on the web. But the big point is that this is what jumpstarted the deep learning revolution. So it provided the data that allowed for some of these models to sort of harness all of the kind of learn all of the patterns from from this very large scale collection of photos and then really build these models that are so that are less engineered in terms of the actual architecture, but very, very reliant on really bottom up learning the visual patterns from from the data.

41:30And this is part of what kind of jumpstart the deep learning revolution. And then we were running sort of the ImageNet challenge for a number of years and kind of collecting more and more data. So that that works sort of the running the challenge and that paper about the ImageNet challenge. That's that's where I was sort of lead author on. And I kind of stepped in right at the boom of ImageNet and kind of helped lead the project when Ja was graduating and sort of moving on to become a professor himself. And I was kind of still a PhD student and very excitedly took on some of that role. So I played I played an important role, but I wasn't I wasn't sort of the early pioneer of this.

42:07And I want to I want to give sort of John Feifei and their collaborators lots and lots of credit for sort of going out on a limb and trying to really do something different. Like this was completely different than what folks were doing at the time. And this was the emphasis on data and on like sort of the bet that collecting more data will actually unlock sort of the next generation of AI models, which is in fact what happened. But I think watching that now, my question is, well, that bet had panned out really well. And now this is sort of taken as a given, as a sort of, you know, ground truth as the default in the field.

42:46But like, I would like to see somebody else make a different bet. Like this is the time for somebody else to make an equally ambitious bet that it's not, you know, data or algorithms, but it's sort of something else entirely. It's maybe it's models of, you know, maybe it's human brain inspired models. Maybe it's, you know, something like I don't know what it's going to be, but I would like to see those kinds of ambitious bets. Yeah. And I've definitely, you know, at least from sort of the private sector, like I've definitely heard of some impressive and noteworthy researchers going after exactly what you're saying and asking the question like there have to be more scalable ways to to to break through.

43:27Right. Without just needing like all of the data in the world. So it'd be fascinating to see what comes of that either, you know, from from some of these companies or from AI for all students. um yeah olga this has been fascinating i've learned so much i'm super confident the audience has as well so really really want to thank you for the time today and um yeah hope to do it again sometime yeah mike thanks so much for having me this was a blast thank you thank you so much for listening to generative now if you liked what you heard please rate and review the podcast that really does help and of course subscribe to the podcast so you get notified every time we publish a new episode.

44:08If you want to learn more, follow Lightspeed at LightspeedVP on YouTube, X, or LinkedIn. You can follow me at Magnano, M-I-G-N-A-N-O on all the same places. And Generative Now is produced by Lightspeed in partnership with Pod People. I am Michael Magnano, and we will be back next week. See you then.

From the publisher

Dr. Olga Russakovsky, Computer Science at Princeton University, joins Lightspeed Partner Michael Mignano to discuss what the next generation of AI talent is learning and where she expects to find the next big innovation in artificial intelligence. 

From her research in computer vision and human-computer interaction to her work in fairness, accountability, and transparency in AI, Dr. Russakovsky has earned many awards, including the MIT Technology Review’s 35-under-35 Innovator award and the Foreign Policy Magazine’s 100 Leading Global Thinkers award. Dr. Russakovsky is also the co-founder and Board Chair of AI4ALL, a nonprofit that aims to increase the diversity of thought in Artificial Intelligence. 


  • Episode Chapters

    00:00 Introduction and Guest Overview

  • 01:17 Olga's Career Journey

  • 02:30 Understanding Computer Vision

  • 04:43 Generative AI and Computer Vision

  • 06:36 Interdisciplinary AI Research

  • 15:00 AI4All: Diversity of Thought 

  • 17:44 Challenges and Bias in AI

  • 30:01 Future of AI and Data

  • 40:08 ImageNET

  • 43:38 Closing Thoughts


    Stay in touch:

  • X: https://twitter.com/lightspeedvp
  • LinkedIn: https://www.linkedin.com/company/lightspeed-venture-partners/
  • Instagram: https://www.instagram.com/lightspeedventurepartners/
  • Subscribe on your favorite podcast app: generativenow.co
  • Email: generativenow@lsvp.com
  • The content here does not constitute tax, legal, business or investment advice or an offer to provide such advice, should not be construed as advocating the purchase or sale of any security or investment or a recommendation of any company, and is not an offer, or solicitation of an offer, for the purchase or sale of any security or investment product. For more details please see lsvp.com/legal.

    More from Generative Now | AI Builders on Creating the Future

    All 90 episodes
    Dr. Olga Russakovsky: Shaping the Next Generation of AI LeadersGenerative Now | AI Builders on Creating the Future · 45 min
    Listen in VO