In short
NVIDIA AI Podcast - Episode Summary: Cleanlab's Curtis Northcutt and Berkeley Research Group's Steven Gawthorpe on AI for Fighting Crime
Podcast Overview Podcast Title: NVIDIA AI Podcast Episode Title: Cleanlab's Curtis Northcutt and Berkeley Research Group's Steven Gawthorpe on AI for Fighting Crime Description: This episode focuses on the innovative data curation techniques employed by Cleanlab and explores how AI can be utilized to combat economic crimes and corruption. Hosted by Noah Kravitz, the discussion features insights from Curtis Northcutt, CEO of Cleanlab, and Steven Gawthorpe, senior data scientist at Berkeley Research Group.
Key Themes and Takeaways
Introduction to Cleanlab
- Founder Background: Curtis Northcutt, the CEO, shares his journey from MIT's PhD program to creating Cleanlab, emphasizing the importance of quality data in AI applications.
- Cleanlab's Purpose: Cleanlab focuses on improving data reliability and trustworthiness through automated error identification and correction algorithms.
Innovative Data Curation
- Automated Processes: Cleanlab automates the process of labeling and verifying the reliability of data, effectively enhancing data quality for machine learning models.
- Use of GPUs: Cleanlab leverages GPU technology to efficiently analyze large datasets, making data curation scalable and faster.
Collaboration with Berkeley Research Group
- Steven Gawthorpe's Role: Gawthorpe discusses his transition from the United Nations Office of Drugs and Crime to Berkeley Research Group, emphasizing the need for effective data labeling and analysis in economic investigations.
- Data Sources and Challenges: The discussion highlights various data sources, including seized company data and open-source intelligence, and the challenges of ensuring data reliability.
Data Analysis and Machine Learning
- Error Identification: Cleanlab's system can identify errors in datasets and determine which data sources contribute most to inaccuracies, assisting in refining data collection and processing methods.
- Interdisciplinary Approach: Gawthorpe emphasizes the importance of collaboration across disciplines in tackling complex issues like fraud and corruption.
Generative AI and Future Trends
- Impact of LLMs: The conversation delves into how generative AI and large language models (LLMs) influence data science techniques and facilitate more efficient data analysis and crime investigation.
- Future Vision for Cleanlab: Northcutt envisions a future where non-experts can easily use Cleanlab's tools to convert low-quality data into actionable insights, thus democratizing access to advanced data analysis.
Recommendations for Aspiring Professionals
- Networking and Collaboration: Gawthorpe encourages young professionals to engage with interdisciplinary networks like the Interdisciplinary Corruption Research Network and to gain hands-on experience with organizations like AI Makerspace.
- Learning Resources: He also suggests reading recent publications on corruption and geospatial analysis for practical insights.
Conclusion The episode underscores the critical role of high-quality data in AI systems and its potential in combating crime and corruption. Cleanlab and Berkeley Research Group exemplify innovative approaches to data curation and analysis, showcasing the intersection of technology and social responsibility.
Further Information
- Cleanlab Website: [cleanlab.ai](https://cleanlab.ai)
- Documentation: [help.cleanlab.ai](https://help.cleanlab.ai)
- AI Podcast Website: [NVIDIA AI Podcast](https://ai-podcast.nvidia.com/)
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:10Hello, and welcome to the NVIDIA AI podcast. I'm your host, Noah Kravitz. We're coming to you live from GTC 2024 back in person at the San Jose Convention Center in San Jose, California. And I am joined by Curtis Northcutt, the CEO and co-founder of CleanLab, and Stephen Gawthorpe, senior data scientist at Berkeley Research Group. We're here to talk about data. We're here to talk about GPUs, obviously. I think we're going to talk about crime and corruption as well. This is going to be a good one. Let's get right into it. And if I might, to ask you guys to kind of set up the story here, Curtis, let's start with CleanLab.
0:49What is CleanLab? What is the role it plays in all of the AI stuff that's happening now? Data feeds AI. CleanLab is all about good data. And then we'll get into how the two of you met and what Berkeley Research Group is doing with all that data. Awesome. Good to be here. It's a long story, and I'll do the short version. Great. Short version is I was at MIT doing my PhD, and they tasked me with the first—I was the first research scientist at edX, and they were building the cheating detection system. And as we know, innocence until proven guilty. Right. So we assumed a bunch of zero labels for a bunch of education data, and I was doing this for hundreds of courses, for all of MIT and Harvard courses.
1:31And I wanted to train machine learning models to look at educational data and predict if someone cheated or not. and discovered something fascinating. At that time, there were no labels that we knew for sure. So we had a bunch of people who were labeled zero, but they actually were cheaters. And we had some people who were one, but they weren't cheaters. And we tried to train machine learning models on that real world data and they failed. So I went to the inventor of the quantum computer, Isaac Chuang, who's my PhD advisor. And I went to the head of AI at MIT, who's done, invented many of what now is common.
2:04Right. And I talked to them and they said, you know, AI is actually, we haven't gotten that far. And this was 2013, 2014. So I spent the next eight years inventing a new field called confident learning. And then I went to Amazon and I went to Google and I worked at all these places. I worked at Oculus Research, at Facebook, Meta. I went to Microsoft and I worked at every big tech and I saw how to implement this stuff for enterprise level. And that's where CleanLab came. Got it. And so how old is CleanLab the company now? The company incorporated about two and a half years. The research and the tech behind it, about a decade.
2:41Right, right. And so in a nutshell, what does CleanLab provide for customers? What CleanLab provides is an automated way to take every data point that's fed into data-driven systems, analytics, prediction, machine learning models, inputs and outputs, and be able to add metadata that tells you if you can trust the data, if it's reliable, what's potentially wrong with it, what's right with it, automated for every data point so you get the most value out of every data point. And so there's a lot in there. The labeling of the data, the trustworthiness of the data, determining if it's trustworthy or not, that's an automated process?
3:17It's automated now. That's 2024. Yeah. Okay. So how does that work? Yeah. So what we do is we look at distributions over the entire data set that that one data point is within. And we learn based on all the other data, what's typical, what's normal, What's the right label? What is the class? What is the type of class? We're learning advanced algorithms that we've invented in-house, and a lot of them came out of MIT when we did our PhDs. And we take those and we automate them on GPUs at scale. That's how we partner with NVIDIA. Sure. And it's incredible because what we have now is we have a workforce of data scientists who eat only electricity, which is one five hundredth of the energy cost of vegetables.
3:57Right. Who work all night, and they do exactly what our algorithms specify. And what they do is they check things like, is this data point an outlier? Is it incorrectly labeled? Is it correctly labeled and with what confidence? Does it have PII, personal identifying information? Is the image not safe for work? Is it blurry? Is it ambiguous? Would it confuse a machine learning model if it was trained on this data point? Would it make it learn faster and better? And we add all of this metadata for every data point. Are there humans in the loop verifying? Yeah. So what we do is we work with large enterprises that typically were doing this by hand.
4:31So you'd have 100 % of your data, which in some cases is hundreds of millions of data points, fed to a human-in-the-loop system where you have humans checking each one. And what we do is we take now 90 % of that and we automate it with high trust scores, high confidence scores. And then the 10 % of stuff is not able to be automated. We send that back into the traditional system. Gotcha. Okay. Is there a number? Is it for the accuracy of the CleanLab systems at this point? So it's domain-specific. So I'll give you an example. In terms of error checking, we have a minimum that we provide that we guarantee of 50%.
5:04And that's for finding errors. So that means if we throw out, say you have a data set that's a million data points, and you throw out, say, 10 ,000 of those data points, or you fix, you correct 10 ,000 data points, what we can guarantee is that you're not going to ever throw out more than half your data, but you'll get very high precision. And precision means, say you have 5 ,000 total errors, you might toss out 10K data, which in the grand scheme of your data set is not that much, but you've eliminated 100 % of the error. And what we're doing is constantly providing thresholds that allow you to say, look, if you look at higher confidence, then you're going to throw out less data, or you're going to correct less data, but that means that you're going to have more data still, errors in your data set.
5:43And so if you throw out everything CleanLab flags, you'll probably get most of the error, but you might throw out a little bit of extra data, and that's the automated aspect. Yeah, yeah. And so this is for all kinds of data across a variety of industries, or are you specializing in one or, you know, a couple of sectors? Awesome. So, Noah, that's a very key point. So I'm glad you brought that up. I want to ask you about popcorn and chicken, but people weren't in the room before we hit record, so they wouldn't know what that meant. So Noah's bringing up an allegory I use, which is that the microwave doesn't just cook popcorn.
6:16And I think it's easy to think of CleanLab that way. If you first learn about CleanLab and you hear, oh, this worked on image data, then you think, oh, this is a tool that helps automate and add value to image data. CleanLab is unique and differentiated in the fact that the way the algorithms work, the way the systems work, the whole software stack works, is we convert every type of data to the same representation. And that representation is then fed into our systems. So CleanLab is domain-specific, meaning it works for any particular type of use case, but it's also data-agnostic, meaning it doesn't care the type of data.
6:47It works for images, videos, texts, tabular, audio, so forth. Right. So, for instance, it could even work on data related to investigating crimes and corruption and all kinds of mistakes. That's my very slick transition. Yeah. Stephen, how did you guys meet? How did Berkeley Research Group start working with CleanLab? Yeah, so I got brought on to Berkeley Research Group coming in from the United Nations Office of Drugs and Crime. When I came in, I took the senior role and wanted to kind of get our tech stack up and running. And the work that I'd done prior to that, my professional career had always been about data labeling.
7:28Always trying to make sure that you take a good data-centric approach and make sure that, yeah, you know the ingredients that are going to go into the recipe. so to speak. So I had already done my homework, known what was out there, started to look into some other tools to help, you know, automate data labeling, model-driven data labeling, and went to a conference hosted by Snorkel where I saw Curtis's talk on CleanLab. And I think you were just getting started at that time, right? So, listened to that, went and downloaded, yeah, your open source package that was still going on. And concurrently, it had some work that I was doing on investigating price manipulation.
8:18And so, I already had some data that I was looking at, and it was kind of a bear to go through and manually label, as you can imagine. I used Curtis's open source version and got immediate impact on the results. I saw that they were putting together an enterprise solution and I called him right away and said, hey, I want to talk. Let's see what we can put together here. So this is when you were already at Berkeley? Yeah. And so, yeah, I was trying to get our tech stack up and running. And so to take a step back, just to set the stage for the listeners, What does Berkeley Research Group do? Yeah, so our main focus is working on economic damages, disputes, investigations.
9:05If you see, you know, a large-scale scandal in the news, high-profile cases in health or finance, chances are we're involved in it in some way, shape, or form. So when I get those notices about a data leak from my health provider, you're the guy I call. I might be connecting some dots that aren't really meant to be connected, but gotcha. And so that was your, you were in a similar kind of work when you were at the UN? The UN, I was investigating drug trafficking, human trafficking, and smuggling of migrants. And I put together what's called the drug monitoring platform. So I created an AI solution that could web scrape all of the news in the world on drug trafficking and create a supplementary data set where you could see how much drug seizures are for every single type of drug, illegal, illicit drug in the world in, I think, 30 languages.
10:03It's still going to this day. Yeah, cool. And so now you can see for free, publicly available, you can download all of the content about heroin trafficking, cocaine trafficking. A lot of that stuff is from the legacy that I put together. Gotcha. Very cool. So when you're investigating now, you work with Berkeley and you're investigating, you know, financial misdoings and that sort of thing without getting into too much detail. But where are you getting the data? And then, you know, kind of maybe before and after, you know, you met Curtis and CleanLab, like what was your process like and what's it like now?
10:41Yeah, yeah. That's a good question. So I get data two ways. Either a company is in trouble, just to put it plainly. Their entire data repository is seized either from servers or, you know, proprietary sources like Outlook files and all kinds of stuff. Right. I need to stop and emphasize here that the type of data that we see is all shapes and forms, which is, I think, where there's overlap with what Curtis is trying to do. So I have to be able to address every single type of use case and data file type and create something out of it, something useful for an investigation. Right, right. The other way that I get data is through open source intelligence, usually through web scraping.
11:26And yeah, I can talk about some examples on cases that I've worked for that. Yeah, I mean, juicy details are welcome, but I don't want to put you in a weird spot. No, no. Okay, again, and this is actually something I used with CleanLab. I was commissioned for work from a cryptocurrency insurance company to try to identify the scope and scale of fraud in cryptocurrency. try to identify the mechanisms, assess the economic damage, how much people were ripped off and so forth. So I scoured Twitter to try to scrape relevant content and then use natural language processing to identify the mechanisms.
12:09You know, such and such coin screwed me over for in this way. There was a fake website or there was a meta wallet attack or, you know, and then there was a value for the amount that they were saying that they were ripped off for. And so I put this together to map the risk landscape for the cryptocurrency. And so in doing that kind of work, how are you able, and again, maybe this is a before and after thing with CleanLab, how are you able to kind of determine if, you know, a particular piece of data you're looking at is reliable, is useful, is maybe, eh, this kind of sounds like, I mean, I don't know what, but, you know, maybe somebody planting misinformation on purpose or, you know, how do you go about determining the value of data?
12:58Yeah, the veracity. Yeah, the data. Yeah, I mean, there's all kinds of bots where they'll have multiple tweets on the same thing. Yeah, I think the way that I've always gone about it is just to try to identify who is the person that's disseminating that information. So if you can find reliable ways on the entity that's putting that out there, the more reliable the data. But, you know, at its face, you got to assume that someone claiming that they've been ripped off in a certain amount, you kind of have to go with that and assume certain levels of bias. There's bias in this type of – anytime you take on open source data and web scraping, there's issues with that.
13:49But yeah, so I would say that you got to take certain levels of acceptable bias. And then the way that you kind of refine is to look at who's disseminating the information and try to dig down further. That's how you target kind of the veracity of it. And so is CleanLab able to kind of do its thing on, you know, the types of data that Stephen's talking about in terms of, and I don't know if I'm approaching this the right way, thinking about it. But thinking about, you know, looking for outliers and the kinds of things you were describing before, Curtis, in a data set, is it the same basic principle if you're analyzing data for, you know, an insurance fraud investigation?
14:31Yeah, that's a good question. Let me try to do some quick specifics. Great. All right. Yeah. So say that you've got a data set, you've got a bunch of sources of information. Your goal is to figure out which source of information is screwing up your data set the most and screwing up the models that you train on that data the most. What you do is you just upload to CleanLab Studio that data. You already integrate it directly with whatever your VPC is, however you have it set up. You integrate with the data set. And then what we do is we'll say, look, here's all the errors. And now you have the column of the source.
15:03And you would just say, let's look at source one. How much error is in that? And then let's look at source two. And so now you have an automated way to figure out of all your sources, which ones are generating the most type of error? What type of error? Where is it coming from? And then you go back to your system that set up the data process in the first place and you correct that. And so that's a quick iterative way. It's a quick example of how you can get to where's the source of the error coming from so we can do better. And then a second thing that's worth emphasizing, and this is something Stephen's brought up to us.
15:32He didn't mention an example, but I think he can chime in at any time. I talk a lot about data because it's a unique differentiating thing that we do. But what we're doing with that data is helping you with two things. One is better data-driven analytics. So that's like Shopify's, e-commerce websites, anybody who just has data that the better it is, the more likely they are to sell product. And then there's another type of customer that just wants to train better models. So that's like your Mosaic MLs, your Databricks, your Big Tech. They're training LLMs, big AI models. For those companies, what we're doing is improving the data.
16:06And so I talk a lot about the data. But for smaller companies, SMBs, mid-size, mid-tier, we do something else that we don't talk about as much, but we do it really well. With one click of a button, we take that improved data set that we've created for you, and we let you just auto-generate a workable, reliable ML model that's deployed for you and that you can use. So that's another example. It's a concrete example that we know Berkeley Research has done, where instead of having to do a bunch of work and train a bunch of models and optimize them, they click one button on CleanLab Studio, and they have a model, and it can make predictions.
16:38Right. One of the things, Curtis, I would say that you don't take a lot of credit for here is that, yeah, you can help identify the anomalies, can help with the data labeling. But I think one of the other things is that you provide an interface to just simply look at the data, which is, you know, most people kind of shy away from. There's so many different instruments and mechanisms and diagnostic tools to help you look at, you know, your data distribution, everything. But at the end of the day, you should still be looking at your data. And there's quite a few cases where I didn't know what I didn't know.
17:13I didn't fully understand the scope and what the data labeling should be. And I think with the finding the errors, it kind of helped me better identify what my annotation policy is, what my labeling scheme should be. Just kind of helped me think about things by just looking at the data, simply put. you know i've got to i've got to add noah yeah thanks for saying that steven because when we started the company and this is a decision every founder has to make everybody else in this space made the decision to go api because it's data and data is a very technical thing so you just build a back-end layer and api you call everything with code it was a lot of work to build a whole front-end team sure my first executive hire was not a vp of engineering it was a vp of design right so i had to put a lot of work in and a lot of investment in order to actually make the front end so that because I believed in this I believe that you cannot have a data company and a data improvement company if you can't see the data and yet everyone pushed from all directions other founders investors all over the world saying look you should just do API that's what everyone else does yeah and I'm really glad to hear that you're like yes that is what everyone else does and that's part of the reason we buy your product so there's it's there's a lot of front engineers at clean lab that are going to smile when you hear that good you get them Excellent.
18:32I'm glad to hear that. Speaking with Stephen Gawthorpe and Curtis Northcutt, Stephen is a senior data scientist at Berkeley Research Group, and Curtis is the CEO and co-founder of CleanLab. We're talking data, we're talking investigations, and we're talking about the value of being able to see your data. I think for a lot of folks, I mean, that's interesting that it came up because for a lot of folks, it is scary, right? And this idea of these vast data sets and not even knowing where to start, let alone how to start making actionable sense out of these things. And so there's a lot of reliance on, you know, whatever it is, the business application that surfaces insights on its own and that sort of stuff.
19:16And so it's interesting to hear that about a data company putting that front-end interface first. You've both been at this for long enough, you know, that sort of before Gen AI started to become a big thing. And now, you know, we're sitting at this conference that it's Gen AI this, LLMs that all over the place. No disrespect to the robots, but, you know, we're talking about this stuff now. How has, I guess, Stephen, I'll start with you. How has generative AI and other, you mentioned NLP a little bit before. Or how have these newer technologies kind of merged with some of the more traditional data science techniques that you came up on?
19:59Or perhaps, you know, taken their place? Or sort of how have kind of the older technologies and the newer technologies come together in the work you do? This is one of my favorite questions. And I think I'm still trying to see how this landscape is taking shape. Yeah, we all are. You know, I came from this background when, you know, you didn't use the term data scientist, right? So that's kind of a newer term. And a lot of the work that I did in analytics is, you know, kind of, I guess, now thought of as more old school, right? And a lot of the tasks that I took pride in, you know, in professional development and, you know, named entity recognition and all this kind of stuff is now kind of taken over by large language models, keyword extraction, all variety of tasks can be done so much faster and more accurate in certain cases.
20:50I think that for myself, I've kind of come full circle and realized that there's some areas that maybe I don't have to do as much sentiment analysis. I think LLMs do that quite sufficiently. But I think data is still kind of an important backdrop for all of this. So I've been doing lots of retrieval augmented generation, RAG pipeline development. And Curtis and I were talking earlier this week on just some of the basic techniques of just, you know, parsing a PDF and just getting good quality content into, you know, a vector database. That's an old school technique, you know, just old school parsing and good quality data recognition.
21:31But it's paramount to get a good RAG pipeline up and running. So there are some areas where it's kind of give or take and, you know, some of the old school still applies. But some of the newer techniques, I'm saying, okay, well, I just don't have to do that kind of work anymore. Right, right. What do you think about that, Curtis? Yeah, I think that's on point. I think there's been an evolution that has been happening and is going to continue to happen. If we go to like, let's go to 2005. Everyone's doing logistic regression, random forests. These are really simple machine learning algorithms that just take data that's like, here's my heart rate.
22:10It's perfectly curated. it's not arbitrary text. And then there's a label and it's like healthy or not. And then you predict if the patient's healthy. Then we entered this wonderful era where Jan LeCun was able to get convolutional neural networks to run on GPUs, which allowed us to scale and distribute image classification. And that was one, it was incredible time because that was the first time that people who didn't really understand AI, didn't know what we'd been doing, could see, wait, that camera knows that's a dog. That's different. That gives me chills. that makes me see that this is possible.
22:42Cars are going to drive now. People started to think. That motivated a lot of money and research that was fed in. And then we saw another evolution, LSTMs. LSTMs were text models that could then take into account a history of tons of text and make predictions, whereas before it had to be very finely curated text. And then beyond LSTMs, there was something called GRU, which is very obscure, but it's actually pretty effective, gated recurrent units. And then Transformers came out, and the paper's called Attention is All You Need, and those blew up. That was my first existential crisis with the Transformers.
23:18Now on my second round of existential crisis with the LLMs. Sorry to cut you off. No, it's good. It's good. I think it's worth fleshing out this history because I had several friends at MIT who started companies that were based on LSTMs. And I had other friends who started companies that were based on convolutional networks. All of their companies failed. Yeah. And that's because they built companies that were model agnostic. It's a dangerous thing to do because we're in the fastest evolving market that's ever existed in humankind. AI and technology is increasing and changing now within a year.
23:51So where we are today is not something we can foresee, for example, even two, three years in the future. And what that means is if you're going to build technologies, if you're going to solve problems, what you want to do is you want to think, how do I do this in a model agnostic way? The only other thing I'd add is that if we think about how does classical ML interact with now transformers and the modern LLMs, I think a good model to look at that has happened before is how did quantum computing interact with classical computing? And that's an example where you have a new technology that's not well understood yet by the public.
24:26And yet we're able to interface with that using old technology that is well understood by the public. And so this idea of having old things mixed with new things has actually happened many times in technology. Sure. And we can learn from those examples. Yeah, well said. So forgive me for asking this question on the heels of you literally saying, you know, it's moving so fast we can't predict what's going to happen. But let's start with you, Stephen. The job of tracking down the bad guys, so to speak, right, investigating fraud and corruption and all of these things. How do you see your line of work continuing to change and be influenced and hopefully getting, you know, faster, easier, more accurate for you because of these things, because of LLMs, because of NLP, because of being able to build RAG pipelines and companies like CleanLab who are able to, you know, help you do more with your data faster?
25:20Yeah, I think that, let's see how I can say this. I think that, one, it's forced me and a lot of other data scientists that are part of my team to start thinking about more interdisciplinary ways that we approach things. For myself, I've been doing lots of LLM ops, thinking about stuff that I would have never considered, like UI and caching. I've said the word caching more now than I have probably my entire life, just thinking about trying to set up the entire ecosystem to make some of this work. And so it's become more interdisciplinary, I think, and that's a big challenge there. And then I think like with some use cases that are very, very interesting, that's been very difficult.
26:10But I see this like moving very quickly with entity linking. So you'll have lots of ambiguity in person's names and with large scale illicit funds being stolen and shifted abroad. It's hard to find and identify and link Mr. X with Steve or something like that. And I've been doing some stuff on my own with RAG pipelines in development, and it's extremely easy to now link persons together and show a sequential order of chain of events on transactions between multiple parties and just putting all of this information together. It's more at your fingertips, I think, than it's ever been before. And today, it has always been a very difficult challenge, but RAG is making it quite simple.
27:01Curtis, what's, this is maybe not the right way to say it, but what's next for CleanLab? What are you guys working on now? And, you know, kind of bigger than that, where do you see things? I keep thinking about, I think because you were talking about, Stephen, you were talking about scouring Twitter, looking for, you know, people talking about being scammed out of crypto and trying to, you know, find clues and talking about bots posting. And so with generative AI, the bot problem, if you will, is getting worse in a lot of ways. And so there's a lot of bad data being generated faster than ever along with useful data.
Read the full transcript
27:39Curtis, what does that mean for the future of data, for the future of a company like yours, and for customers like Berkeley Research Group who are looking for ways to separate the good data from the bad and do more with it? Yeah, it's a good question. So there are a lot of things that are changing, but there are some things that are remaining the same. So some of the things that are remaining the same we can work with and we can guarantee. So one of them is that there is no learning without data. You have to have something to learn. And so we know that as long as we keep doing more AI globally, then we're going to keep increasing our dependency on the quality of the data that's fed into those systems.
28:19Even though the quantity of the data that's being fed in those systems is ever increasing. And so that's one of those things where the quantity is changing all the time, but the dependency on quality is always going to be there. It's a constant, yeah. So that's good for us. Right. And it just creates a bigger and bigger market. Where I see us in the next three years is you have an idea in your head. You don't have to hire a bunch of PhDs. Stephen is unique. He's been an incredible pleasure to work with because he's very smart. He has a PhD. He can think about these things. Yeah. We want to get to a place where you have a team of people who want to get stuff done.
28:53who are creative. They don't have a PhD in machine learning and AI, but they do know, I've got some data here, and I know CleanLab integrates with everything. And all I want to do is I want to click a few buttons, and I want to go from low-quality data to high-quality, reliable, data-driven solutions. And when they think that problem, when that problem enters their headspace, they think one word, CleanLab. There you go. We usually end these conversations by asking for recommendations for listeners who want to learn more. And so we'll get to that in a second, but I'm going to put you on the spot, Stephen.
29:26For listeners who might want to learn more specifically about using data, data science, all of this stuff to fight crime, to, you know, all of that good work that you do, do you have advice, do you have recommendations for somebody who's, you know, a little on the younger end in university or just starting out their career and they're interested in the future of data-driven crime fighting? Yeah. Okay. So I would start with the Interdisciplinary Corruption Research Network. Okay. It's a group that I started when I started my PhD. It's a group of individuals that wanted, that didn't come from the legal background or economics background, which has kind of just been a mandatory requirement in the international space to participate in anti-corruption.
30:14We felt that, you know, corruption is everybody's problem. And so there's computer scientists, there's psychologists, there's engineers, there's people that have a stake in this. And so I would say, look them up, see the work that they're doing. And then I think if you want to get into, you know, some technical aspects of just learning the latest and greatest on LLM development, I did some work with AI Makerspace. So shout out to them. They're a fantastic organization, Very practical. They get you up and rolling. Nice. And then maybe the last thing I would say is to read one of my recent publications.
30:55There you go. Plug it. That's very helpful. Yeah. How to Identify Widespread Corruption, New Insights from Geospatial Analysis. Excellent. So, yeah. And Curtis, you get the easy question. For folks who want to learn more about what CleanLab's up to, where should they go? That easy. You go to cleanlab.ai. And if you want documentation, you go to help.cleanlab.ai. Perfect. All right. Well, we'll have to get together again and talk about the, was it the PhD rapper, the rapping PhD? Yeah, that's a whole other side. That's a whole other podcast for another time. But thank you guys both so much. This has been tremendous.
31:31Yeah, it's a, you know, we're sitting here at GTC and it's an audio podcast. People can't see, but the way the room's set up behind the two of you, I'm looking out at the show floor and everybody milling about and there's robots and there's models and there's all kinds of stuff going on. It really does feel like it's not the start of something new, but it's a moment. So thank you for coming on the pod and taking the listeners with us. Thanks. It's a pleasure. Thank you.
32:39¶¶
32:46Thank you.
From the publisher
Talk about scrubbing data. Curtis Northcutt, cofounder and CEO of Cleanlab, and Steven Gawthorpe, senior data scientist at Berkeley Research Group, speak about Cleanlab’s groundbreaking approach to data curation with Noah Kravitz, host of NVIDIA’s AI Podcast, in an episode recorded live at the NVIDIA GTC global AI conference. The startup’s tools enhance data reliability and trustworthiness through sophisticated error identification and correction algorithms. Northcutt and Gawthorpe provide insights into how AI-powered data analytics can help combat economic crimes and corruption and discuss the intersection of AI, data science and ethical governance in fostering a more just society.
Cleanlab is a member of the NVIDIA Inception program for cutting-edge startups.
https://blogs.nvidia.com/blog/cleanlab-podcast/




