In short
Podcast Summary: Zero-Shot Auto-Labeling: The End of Annotation for Computer Vision with Jason Corso - #735
Podcast Title: The TWIML AI Podcast Episode Title: Zero-Shot Auto-Labeling: The End of Annotation for Computer Vision with Jason Corso Episode Host: Sam Charrington Guest: Jason Corso, Co-founder of Voxel51 and Professor at the University of Michigan
Episode Overview In this episode, Jason Corso discusses the cutting-edge approach to automated labeling in computer vision, emphasizing the potential of zero-shot auto-labeling to reduce costs and improve efficiency. The conversation focuses on Voxel51’s work, particularly their open-source platform FiftyOne and the recent report on zero-shot auto-labeling which shows how it can rival human performance in labeling tasks.
Key Concepts
- Introduction to Voxel51 and FiftyOne
- Voxel51's Mission: To provide tools for visual AI that facilitate the co-development of datasets and models.
- FiftyOne: An open-source platform designed for visualizing datasets, analyzing models, and improving data quality.
- Zero-Shot Auto-Labeling
- Definition: A method that leverages foundation models to automatically label data without requiring traditional human annotation.
- Significant Benefits:
- Cost-Efficiency: Substantial savings in time and resources compared to manual labeling (e.g., $124,000 for human labeling vs. $1.18 for auto-labeling).
- Performance Comparison: Auto-labeling can achieve F1 scores comparable to human labels, especially when employing varying confidence thresholds in model predictions.
- Key Findings from Research Report
- Performance Variability: The effectiveness of auto-labels can depend significantly on the confidence thresholds applied.
- Human vs. Auto-Labels: While human labels can be flawed (e.g., under-labeling, mislabeling), auto-labels tend to avoid some egregious errors, highlighting an area of potential improvement.
- Verified Auto-Labeling Approach
- Quality Assurance Process: Utilizes a "stoplight" QA workflow (green, yellow, red) to categorize labels:
- Green: High confidence, likely accurate.
- Yellow: Requires human review due to uncertainty.
- Red: Low confidence, discard.
- Challenges and Future Directions
- Decision Boundary Uncertainty: Addressing the challenges associated with data that exists at the edges of model decision boundaries remains an ongoing research interest.
- Handling Out-Of-Domain Tasks: Acknowledgment that current methods may struggle when applied to data from domains not well-represented in foundation models.
- Integrating User Feedback: Future developments will focus on how to effectively incorporate human insights into the auto-labeling process to enhance model accuracy further.
Key Takeaways
- Shift in Annotation Paradigm: The episode argues for a shift from traditional manual labeling to zero-shot auto-labeling methods that utilize advanced foundation models.
- Importance of Human Review: While auto-labeling is powerful, human expertise remains crucial in refining the labeling process.
- Future of Annotation: The approach towards creating more effective and efficient annotation systems is evolving, with the potential of agents that interact with humans in decision-making processes.
Conclusion Jason Corso presents a compelling case for the future of computer vision annotation, emphasizing the transformative capabilities of zero-shot auto-labeling. The ongoing challenges and research directions suggest a promising future for both automated systems and human-machine collaboration in model training and dataset creation.
For detailed insights and additional resources, listeners can refer to the complete show notes at [TWIML AI Podcast Episode 735](https://twimlai.com/go/735).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Wait a second, what we've been doing as annotation, just blindly sending everything out for labeling is not the future. we're seeing foundation models come along that could indeed actually replace a lot of those typical case human labels. So if what was annotation 1.0, blind send me data and get me labels, annotation 2.0 might only be human answers questions asked by the agent, if you will, agentic labeling or something like that. To me, we're on that timeline and we're probably somewhere in the middle of it.
0:41All right, everyone, welcome to another episode of the TwiML AI podcast. I am, of course, your host, Sam Charrington. Today, I'm joined by Jason Corso. Jason is co-founder of Voxel51 and a professor at the University of Michigan. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Jason, welcome to the podcast. Hey, Sam. Thanks for having me. It's great to be here. Big fan. Absolutely. Thank you so much. I am looking forward to digging into our conversation. We're going to be talking about some of the work you're doing on automated labeling for computer vision.
1:19I'd love to have you start us off by sharing a little bit about your background. Right on. Sounds good. Yeah. So as you said, I'm a co-founder at Voxel 51, where we make a software dev tool, kind of like VS Code for computer vision or VS Code for visual AI. My background, I go back maybe 10, 15 years in research in computer vision. Most of my work focuses on high-level computer vision, video understanding, vision language, the relationship between the physical world and vision, and so on. And through that work at the University of Michigan, I realized that there was a huge need for more analysis tooling to help support the co-development of models and data sets together.
2:00And ultimately, that's given me rise to the last five or so years of my life at Voxel 51 and the university together. Awesome. So with Voxel, who are the intended users of that tool? Yeah, Voxel is pretty much the IC as a technically deep creator of either data sets for visual problems like image, video, 3D, point clouds, meshes, and so on, or model developers. At our bigger customers, these tend to be separate teams where you have a data curation team and you have a model development team. Whereas at our smaller customers, typically a small team does both of that. And what we've learned is over the years, generally, you don't just get data dropped into your lap like a student does in my computer vision class.
2:49And then you have to train the model. The real world, life in the real world is making the data set alongside making the model and really understanding how these things coincide together. You know, and so that's ultimately that's though that is why we we released 51, which is the name of our primary software. It's both commercial and open source. So open source for single single users, local data, local compute. That's really how we got started in 2019. when we began this direction toward like DevTool. We released it open source, both in the computer vision community and the machine learning community.
3:25And it's become a pretty widely adopted tool for everyday usage. Our initial charge was, for the first three things someone does when they're starting a new computer vision project, PIP install 51 should be one of those things. One of those first three things. Awesome. And are there particular types of projects that you're seeing, you know, the most traction around or where 51 is one of those first tools that someone is using? I mean, it's a good question, Sam. I think the, you know, 51 is a pretty flexible piece of software. It's not intended to tell you how to do your work or what to do. It's, think of it kind of like a set of building blocks.
4:06And if, you know, if you want some annotation or you want like model performance analysis or you just want to visualize your embeddings or whatever, it can kind of do all these different pieces for you. So we do have a pretty broad user base. I mean, that said, I think the object detection is definitely, at least amongst our commercial users, object detection really shines as like a core task in computer vision that our users are doing. And most typically people are starting with an existing model that either they've trained or they've inherited, say, an off-the-shelf foundation model or something like that.
4:41And their goal is to render that useful in their domain for their problem, whether or not that's quality assurance on the manufacturing line or pedestrian avoidance for mobility or healthcare problems. From a vertical standpoint, getting to that, there's no one domain actually of users. We have quite a few different domains in play. Maybe the most exciting one for me, I have young kids, but we have some users that are monitoring marine life through models built in 51. And they use that across various oceans in the world. Pretty cool project. I often hear that one of the biggest lift investments that teams can do is to build custom annotation tools.
5:32because every application or use case is kind of a little bit different versus using something off the shelf. Do you get that kind of or how do you look at that kind of comment? And is 51 meant to be used like off the shelf or can it be integrated into more application specific workflows? Yeah, I have I have personally fallen prey to that to that to that direction as well. when I was a postdoc at UCLA, I built my own 3D, I was working in 3D medical imaging, so there was no 3D medical annotation tool out there. So I built my own one at the time. In fact, I think that notion was so sort of prevalent in our strategic thinking early on that within 51, from the very beginning, we explicitly did not support annotation directly.
6:29So in order to annotate in 51, we provided integrations with common annotation tools like CVAT and Label Studio, Label Box, other annotation tools, mostly because there were so many of them out there. And we felt like it was going to be hard to distinguish any new tool, you know, from any new tool like ours from another tool that was already out there. We really wanted to emphasize the notion. And at the very core of what we did is like, you know, you train, you get some data and train a model and then it doesn't work as well as you think. So you're like, what do I do next? Let me get more data and let me try.
7:08You kind of iterate this loop and we inject 51 into the middle of that as like the analysis work. Right. So kind of like turning the lights on in a dark hallway. Right. So like you get some data, train some models, then you do analysis within 51 and then you kind of repeat this loop. um so um that's why we made it flexible though right because like that loop is very is very team specific in fact you know we i used to equip or like you know for every 100 machine learning engineers out there there are 1000 ways of working right because like everyone has their own way of doing things really so um so you know but but to your point um we indeed we work with um we still maintain the integrations with other annotation tools now even though you can annotate within 51 now, at least through this auto-labeling stuff.
7:54And actually coming out in the summer, you will be able to do some annotation, let's say classical annotation within 51. But we still integrate with all the other tools out there. And I think it's primarily because 51 has become a little bit of a Rosetta Stone for visual data formats. You know, it's very easy to get your data into 51 format and it's kind of like this universal format and then get it out in whatever format you want. But many of our commercial clients do indeed have their own homegrown annotation toolbox. And so it's typically like the first month of the engagement with a new commercial customer will be, you know, format merging, if you will.
8:37Or normalization. Normalization, yeah. Good word, yeah. So, but we don't, we do, and the other, maybe the other asterisk to add is that one thing you can do in 51, which we're very proud of, is you can add your own panels. So 51 is both a Python SDK and a web-based app, right? And so in the app, you can extend the capabilities of the app through what we call panels, which are essentially plugins that have a visual component to them. So we have actually seen some of our users and customers porting their existing annotation capabilities directly into a 51 panel so that it's an integrated experience.
9:21I think that's the way forward I mean I think the indeed many instances of problems are nuanced so that you do have slightly different changes and you want to be able to have your annotators next to your QA people, next to your consumers downstream as model developers, so having them all under one hood is important And so you talk about inserting analysis into that loop, talk a little bit about the nature of that analysis and how it aids the user? Sure. So like one key use case for that is, right? So you have a data set, you have a model. Obviously it doesn't perform perfectly. So what can you do?
10:05So one question you might want to ask is what are the corner cases, right? Like where is it performing poorly? Kind of analysis. Exactly. Yes. Like, and are there commonalities or is there structure to that lack of performance? So we have capabilities in the system to, for example, compute what we call hardness or mistakenness that take your model logits on the data and manipulate them in certain ways to then provide to you clusters of outputs that you can then visually analyze. Because in our view, the human expert, like the actual domain expert or the MLE in the loop, is the only one who really can make the decision.
10:48Oh, you know, whatever, we're training a pedestrian avoidance algorithm. And it turns out there's no, there's very few examples of suburban parking lots for some reason. You know, only the human really is going to identify that that was the failure mode of the problem. So that's one typical analysis. Another analysis we use is visualization of embeddings. So you can kind of, you can compute your embeddings either with a like 51 approved embedding model or you kind of bring your own embedding model for that. and then you can visually interact in in the app to find situations for mislabels on your actual data set right so like you visualize your embeddings you can project the class labels on those embeddings and ultimately we'll see like well this this big big mass of red labels here with some blue labels in them let me see what those blue labels are so let me turn the red mask off the red label off and then i can get a view only into those blue labels.
11:44So it's really kind of this notion of interactive analysis that is imperative to cleaning up your data and finding corner cases where you might then go need to add more labels for you for the task. Talk a little bit about the evolutionary path towards automating labeling or auto labeling. Was this something that you started out intending to do or is it something that arose as a result of powerful foundation models, for example, or something else? Yeah, I think it's, I mean, I'd be lying if I told you in 2018 when we were first starting the company, or 2007 when I started as faculty, I was going to predict like, oh wait, we're not going to need labeling anymore.
12:30That couldn't be true. But it's not necessarily strictly a function of just recent development. So about maybe like two years ago, we began to see a shift in the types of conversations we would have with users of 51, where they were going from like, I'm strictly sending all of my data to annotation, to humans to annotate, and the spending, just say round number, I'm spending a million dollars on this. and then I'm getting it back. And instead of being able to use it all, I'm finding out that actually I can only use 10 % of it or 20 % of it, which I mean, dollars for dollars means they're throwing away a significant amount of money.
13:13And was that because of poor labeling, mislabels or? No, I think it's, I mean, obviously there were some mislabels, but I think there's just a lot of similar cases being sent for labels. And it wasn't clear that the right subset of the data, right? Like data, it's easy to find like typical plus, plus minus a couple of sigma cases. But what really matters is finding data along the decision boundaries. And that's hard to do if you don't know what the decision boundaries are, right? So the conversation shifted over into sort of that direction. And we began to try to sort of think about ways of how are we going to help users, even ourselves?
13:53It's a great problem technically, right? Like how do you find data? How do you find the boundary without knowing the boundary? How do you get there quickly, right? So I wrote a blog, maybe January of 2024, with the title Annotation is Dead. And it was intended to start this conversation generally around like, should we really be doing what we've been doing? It wasn't so much that like, okay, no one's going to need annotated data. Obviously, it's a well-oiled machine. We know how supervised machine learning works. We can measure somewhat predictable performance, things like that, right? But it was really intended to catalyze this notion that wait a second what we've been doing as annotation just blindly sending everything out for labeling is not the future the future cannot be that way and it just so happened that like okay we also we're seeing at that time like foundation models come along that could indeed actually replace a lot of those typical case human labels um and so you don't even need to do that anymore and so so if if what was annotation maybe we call that like annotation 1.0 and annotation 2.0 is like you there's never blind send me data and get me labels annotation 2.0 might only be a human answers questions asked by the by the agent if you will agentic labeling or something like that to me that that's we're on that timeline and we're probably somewhere in the middle of it not just we like voxel 51 but i think the community is generally we're moving down the line of, okay, we can expect to get decent labels for common classes.
15:28What can we expect for less common classes? Or, you know, I have a automotive scenario. Can we, can you predict my model performance without actually labeling the data beforehand? Things like that. And I think we're right squarely in the middle, moving in the direction of fewer, you were sort of just here's the media give me that give me the metadata um toward more like i think this is a teddy bear is this really a teddy bear you know and these are not technically speaking like those are not new directions right like people in the community of machine learning and computer vision have been looking at this for at least a decade or two right semi-supervised learning is not new active learning is not new but i i think what what we're seeing is that with the creation of these semantically enriched foundation models and embedding spaces, there's more of a structure, underlying structure, to the way data is represented in those embedding spaces that we need to crunch on and understand more so that we can then go and do a better job of asking the right questions.
16:33It strikes me that one of the linchpins in kind of fully getting there, if you will, is being able to better characterize or quantify uncertainty in these models. And that's been a challenge in the community for a really long time. Can you talk a little bit about how embeddings and kind of the semantic, you know, modeling that you just mentioned helps us maybe overcome or sidestep that challenge? It's a good point. I do think uncertainty is at the core of a lot of this, a lot of this discussion. Indeed, I agree. You know, in some sense, if we really want a true measure of uncertainty, right, like a classical, you know, like P of X type uncertainty, I don't think we have, I still don't think we have made much progress, right?
17:24We don't, we still don't really know how to do that, but there is enough structure in these embedding spaces that we can start to even leverage some classical ideas. Like, for example, we have a tech report that we've been working on. I can send you the link after the podcast here that is able to take a classification problem and embed it using these contemporary foundation model embeddings. Say, you know, like a ResNet 50 or ResNet 18 for a perceptual space. And then something like, why am I forgetting the name of the model now? Not Coco, but the very classic vision language embedding model that everyone uses.
18:20Clip, sorry. So take a classic perceptual embedding like ResNet 18, concatenate that with a clip embedding space and model within that space. And the problem we set out to ask is, can we measure the difficulty, the expected classification difficulty in that space without training the downstream classifier, right? Assuming, say we only had labels on a small subset of the data, like 100 samples per class, for example. And we use very simple, shallow auto encoders like two two three layers auto encoders they're trainable in seconds per class and if you take ratios of reconstruction errors which is a measure of uncertainty to some degree you can there are certain certain classes of those ratios that correlate very strongly with downstream classifier performance how do you train a full classifier on the whole data set so that's what leads me to say that this is evidence if you will that there seems to be some structure in these more rich semantically enriched embedding spaces that even though we still don't have the right machinery to go and fully compute uncertainty we at least can get say this is using the term loosely kind of like marginal uncertainty in that space and does that presume to some degree
19:51I'm thinking about like, you know, the big challenge in the object detection scenario is going to be like, or in the classification scenario, rather it's going to be like, you know, outliers or, you know, rare things that you're trying to classify. You know, the kid running out in the street relative to the street sign. And are there assumptions made on having examples of that rare data in your data set? Yeah. So there are no assumptions that we know that they occur in the data set. But we also can't deduce them. There's sort of no way to figure those out, right? Those are essentially the Bayes' risk or the Bayes' uncertainty, right?
20:39Like, I guess, and essentially what we've observed is that, you know, again, there is, the decision boundary is always going to be the challenge here, right? But like, and since we're not, in that work, for example, we're not updating the embedding at all. We don't modulate the embedding to try to pull them apart. if indeed a data set is dominated by those rare cases that are hard to disentangle from the truth, then it would be a kryptonite for this modeling that we did. But in typical situations, the mean is the mean, right? Typical samples are typical samples. so um i guess what what i'd like to see though we haven't done this is you know i've been hearing this notion of like internet scale data sets or real scale data sets uh coming up in the vision community of a lot in the last couple years i still don't have a great sense for how to quantify what is like a an internet scale data set versus a canned off-the-shelf data set But my intuition says that the closer we get to large-scale, internet-scale data sets, the more the mean behavior is going to dominate.
22:01And I think that's critical. Yeah, and I guess to maybe answer my own question or comment on my own question, or comment rather, part of the core value prop going back to that initial observation you made is, hey, we spent a million dollars on labeling and it wasn't all that useful. it would be useful just to auto-label the stuff that's the same. You don't necessarily have to auto-label the outliers to create value. That is the observation I think many of us are making collectively, right? Like, why am I spending this money on things that I'm pretty sure even my own models may do well with already?
22:50and I think the challenge is I think there are two key challenges really one is how can we build a model of using the term model loosely here, like a model of predictability, right? How much of that auto labeling can I expect to be useful and I think that's a key challenge and then another one is given that I've just auto labeled things how do I minimize the human work needed to do QA on the output of that? Because if I have to have humans go and validate 100 % of the auto labels, then I've not saved any time at all, right? So how do we get that number down? I think ultimately those are two things that I've been thinking about carrying out lately, so yeah.
23:36So you guys recently released a report on Archive called Auto Labeling Data for Object Detection, as well as a blog post. zero-shot auto-labeling rivals human performance that kind of captures all of these ideas. Can you, you know, let's maybe start by talking about like how you are approaching auto-labeling. Is there a, you know, generic setup that you're using? Is it very use case specific? Yeah, sounds good. So I think that this initial report and our initial work, I guess, in it is probably the simplest setup that you can imagine, simply because we want, I mean, I've been around for some time, and I think simple works.
24:27We need to get in the simplest case before we can talk about more complex settings. So basically the assumptions are you have a foundation model that is relevant to the domain you're operating in, general natural images, medical images, whatever. And then you have a... A VLM in particular. A, I guess a VLM in particular, because we do prompt it with labels. I mean, the data sets we work with are VOC, COCO, BDD, and Elvis. so all of our all of our prompts were pretty simple so you know that's where that's that's why i'm sort of hesitating at vlm in the sense that you know the foundation models we used are like yolo world yolo e grounding dino concretely so so they're not really vlms but but i mean i think more generally speaking you could say vlm um uh and you have a large corpus of unlabeled data basically right and the idea is okay what would you have done in annotation 1.0 right you would have then sent all of those images to humans to label.
25:31You would have taken the output from them and trained your model, your object detector in this case, so we're only looking at object detection, and then, you know, computed some validation performance on a holdout set that was part of those annotations as well. In our auto-labeling scenario, take the same un-labeled data, take a foundation model, generate those labels, auto-labels, you know, we call them now, similarly train a apples to apples sort of object detector we use like rtdter yolo yolo 11 i think we use and then compute the same validation performance we wanted to make it as apples to apples there's nothing special about the domain or the use case or what have you aside from the fact that we're only using only doing object detection other assumptions we made are only one foundation model for input there's no voting you know like there's there's some recent work like there was a paper um the florence 2 model at cvpr 24 uh they had this interesting notion of like a data engine where they're like waiting over multiple foundation model outputs um we don't do any of that again simple simplest case here uh and we we wanted to assess given these simple settings um you know like um what is the cost comparison in terms of money and time how well do the auto labels match human labels and how well to do the downstream object detector detectors match object detectors that were trained on those human labels those really were the three things and i should say concretely there's no new new architecture here there's no there's no technical contribution that way the only parameter we vary is the threshold on the model confidence per object label or object auto label uh and so we wanted to keep it as simple as possible and just because i think there's a lot that we could do downstream right like in the future but we don't have a good baseline uh you know and i think concretely we approach it this way not just because it's important to do things simple first but we weren't aware of a baseline that actually did this experiment.
27:40That compares cost to ultimate gain in downstream performance or what have you. So, yeah, so that's the way we set it up. Is that clear? Any questions about that setup? No, that's clear and simple. Okay, great. So, again, I can enumerate the datasets we used, VOC, COCO, BDD, ELVIS. the foundation models were YOLO-E, YOLO-World and Grounding Dino. Downstream models that we trained were different sizes of YOLO 11. So YOLO 11N, S, M, L and X varies from like 2.6 million to 57 million parameters. And then RT-Dieter, which has 33 million parameters. So in some sense, like, you know, we wanted to take one model architecture, vary its capacity and then completely different model architecture at roughly the mean capacity of the first one is the way I would think about those.
28:39And so what did we find? Actually, I guess before that, so to do all this work, we ultimately had to train on the order of 445 different models. And so our compute node was running for some some month or something like that to generate all these training experiments and uh the models that you know per data set uh we we we only use the standard prompts so we didn't we didn't do any prompt expansion or generalization in any way uh so it was basically you know like you generally have to prompt these models with every class label the period after it ultimately that's all we did uh so what do we find from a cost standpoint so um using off the shelf established numbers for how much things cost to to annotate so seven cents per box for example um we estimate that collectively these four data sets would cost about 124 000 to have humans annotate them just one pass of human sanitation so no qa nothing like that um and that the comparable cost in auto labeling like gpu rental on aws uh was a a dollar 18 uh for for an nvidia l40s a dollar 18 total for the 400 something models dollar no dollar 18 total to do yeah a dollar 18 total to produce all the labels so for all the models well actually no to produce the labels which is only done once not not not so So inference, not the creation of the models, not the training of the models.
30:18Okay, got it. Yeah, in fact, I don't have the number for how long it took to train all the models. We have a box. We did this on a box of six L40Ss, and it took about a month, so whatever that cost is. But that's the full research experiment, not the cost analysis. So it's six orders of magnitude, more expensive to have humans label than have models label. So that's, but that's expected. So it's not surprising that the cost is greatly different, right? 124 ,000 to a dollar, right? That's expected. But I think quantifying it is still something that we weren't aware someone had done, right? So I think even putting a pin in the ground is important to do it in science.
31:00Ultimately, it only matters if the labels are good. Exactly. Ultimately, it only matters if the labels are good. And I will get there. But one One other aspect is the time, though, it took as well, right? So we estimate that it would take about 6 ,000 hours of humans to label these four data sets. And this is based on a paper from Kristen Grauman's group in, I think, ICCB 2017 or 2019, where they were kind of like measuring the value of labeling or the cost of labeling. So 6 ,000 hours to one hour, to one and a quarter hours. So about four orders of venture difference there as well. So you can do it fast and you can do it cheaply, but does it work?
31:45Like, does it matter? So you said 6 ,000 hours to one hour. That one hour is spent labeling or is one hours? 6 ,000 hours of human labeling to 1.27 hours of GPU labeling, of GPU inference labeling. Okay. So we're just talking about wall clock time as a comparison, as opposed to 6 ,000 hours as the measure of some cost versus an hour as the measure of some other cost. Exactly wall clock time. Yeah. Very straightforward. Okay. So how do the models perform? We looked at it along two different axes, right? Like one axis is how well can we reproduce the human labels, which is important, right? Just as a measure of they are in some sense a gold standard, although they're not a perfect gold standard.
32:48For example, one of the data sets has an image of donuts. And there are these two trays of donuts in the image. I actually forget which data set it comes from, maybe Elvis. And in the human labels, it's a mess, actually. So if there are something like six dozen donuts per tray across both of them, roughly, that's probably about what it is. Maybe a little less than that, four dozen. Humans sometimes, this is real data, right? They will only label three of the donuts, a handful of the donuts, right? That sounds about right. This stuff is tedious. It's exactly tedious. It's so tedious. Similarly, they might group together six donuts as one donut.
33:34And actually, one example we have is the whole entire tray, essentially the whole image, is labeled as donuts. Which is just like this is a complete ontological failure of the domain. Right. So whereas the auto labels, OK, they miss some of the donuts, but it doesn't make the egregious error of like grouping them together and so on. That was that was a great finding, I guess. But so but if you do assume that human labels are perfect or are the gold standard, we have found that performance does vary drastically based on the configuration that we call it configuration. But in this case, it's essentially what confidence threshold you use to say it's a it's a it's a we keep this all label or not, basically.
34:21So confidence threshold is coming from the model. Coming straight from the model. So YOLO or whatever, what have you. Right. Without without doing any calibration at all. So we just, again, took it took it as is black box. We did we do notice that grounding dyno has the greatest variation in performance of the of the three foundation models. Whereas YOLO-W and YOLO-E, maybe at the boundaries of the curve, it does fall off pretty drastically. But from confidence 0.3 to 0.7, it's a pretty flat, it's a pretty plateaued performance in terms of F1 to the human labeling. So that's a good finding, right?
35:00It's like in some sense, for each of the three foundation models we used, if you're able to find the right confidence threshold, you can essentially get almost perfect F1 score. Maybe that's a little bit reaching. You can get very good. Let's say comparable F1 score to make it usable. The challenge is, though, in practice, you don't necessarily have the you can't calibrate that well. you'll have to get some data labeled to do your calibration of your competence and then go and hope that that extrapolates to the rest of the data set. So we really think the right measure is the other axis that we measured, which was, OK, take the auto labels and predict their downstream or not predict, measure their downstream model performance.
Read the full transcript
35:51So this is where we train the 445 models and do the analysis there. And what we found is that in all cases, for each of the different three variants of foundation model across all four data sets, we can get within relative numbers in the 80s and 90s performance of whatever like that, you know, which in say YOLO 11S human performance is, these numbers are not from the paper, but like say 60 % accuracy in auto-labeled trained YOLO 11S. can get like 55 % performance or something like that. Very close performance. And interestingly, there are some interesting findings here. One is that if you're willing to save the money you would have spent on annotation and instead train a bigger model for compute and spend that money over time, you can, apples to apples, you can actually outperform the smaller model, right?
36:50So YOLO 11S has 9 million parameters. Yellow 11n has 2.6 million parameters. If you're willing to train the yellow 11s, you can outperform the yellow 11n in many of the results we have as well, which is a cool result, I think. But even more interestingly, that the best downstream model performance comes for confidence thresholds filtering in and out the auto labels that are not aligned with the top F1 performance. So in the sense that when you're generating auto labels, if you try to maximize your match to human performance, you won't necessarily get the best downstream object detector performance.
37:34Instead, what we found is that lowering your confidence threshold to somewhat egregiously low numbers like 0.1, 0.2, where you will have noisy outputs in the auto labels, ultimately maps to better downstream performance. and I thought this was really exciting. It's not necessarily a - It's super counterintuitive, but also consistent with other results we're seeing recently. There was a paper within the past couple of weeks about RL fine tuning. Like it doesn't really matter what the answers are. Like you can just, and the general idea is consistent with, you know machine learning just seems to perform better when there's the right amount of noise like just like drop out and all these other things that we've done to like create noise when we're training models you know so from that perspective maybe not surprising but counterintuitive absolutely i think it's just another chip or you know another tile in the mosaic of of messaging around like machine learning models need data and the data does not have to be perfect but more is better than hike better more is better than better in some sense i don't know the right like mantra yet but but but it's indeed it's pointing in that direction um and i think you know in some sense what i what i hope for this particular report is that it gets recognized amongst in the community as rather concrete and up-to-date evidence for not only that counterintuitive nature, but like a way of measuring how well these contemporary foundation models perform in the context of auto-labeling.
39:22Yeah, thinking about more is better than better. It's more is better than better if better is a metric of like human label performance, But if better is a measure of like having the right data, meaning outliers and these other things identified, then, you know, that is, you know, better. and so I'm kind of using that as a segue to to prompt you to kind of circle back on this idea of like so now that you know we understand that zero shot auto labeling you know works well like if I'm building you know these kind of object detection models or I've got an application or systems that's using it like how do I take that knowledge and kind of build a system around it that It helps me zero in on, you know, what's really important and identify, you know, these outliers that we've talked about and overall, like, you know, build a better system, like in light of all the tradeoffs that you've kind of mentioned.
40:24I think it's a great, great, great point or great, great segue. Noting that I think it's the experiment you hinted at, right? If we could get better data at the boundaries, then we could really measure how much more is important. How to do that is really hard. But I would love to spend the next year thinking about that problem. But anyway, so what has this taught us in some sense? right like okay these experiments are typically like these data sets that we've that we've used in this experiment are are not they're mostly like let's say kind of like in in domain data sets if you will right for the foundation models we used right so voc coco even bdd like these most or all of the classes in those three data sets are are at least covered in some way or another in the training data for the foundation models that were used now the exact images were not used, but the concepts were there.
41:27LVIZ, I think, is a little bit different where it has something like 1 ,200 classes. And we incorporated LVIZ to measure the limitations of this case, where we find clearly, if there are no labels predicted, then downstream object detector performance is not going to be there. So that's all measured in there. So a limitation of this notion of auto labeling is like, how do you handle the decision boundary classes or how do you handle complex classes that are out of domain? And I think that like basically the key, the key nugget here is being able to identify either automatically or semi-automatically with the human in the loop, like when these auto labels are, should be accepted and when they should be rejected.
42:18and so that's actually conveniently gets back to that method i was mentioning to you earlier the way i think at least the way voxel 51 is going to be doing it within 51 is this notion of what we're calling verified auto labeling so like you generate your auto labels and it's still the onus is still on the human whether or not it's a qa person that you're that's a contract that you're paying from outside the company or a team member an actual trained mle like someone really should be saying yes no yes no yes no but again like you know it's a problem if they have to do that for every image right so so our our approach is to what we think will work and we're still working on this is admittedly but what we think will work is being able to identify uh what we what we're thinking of is like green yellow and red light think of it like a stoplight right like clearly these samples are in cluster in class right like they they don't vary at all from the mean of the cluster.
43:12And whether or not that's measured with the shallow autoencoder model I talked about earlier or some other, there probably are dozens of ways of doing that, as long as you can do it fast and effectively. And then similarly, okay, the red light, right? These cases are clearly at the decision boundary. We're not confident about these. We don't know what they are. So let's just throw away those. And then the challenge is the yellow, the cases in the middle, right? Like, are there cases where it's unclear if it is a decision boundary problem, like a child near the stop sign example you gave earlier, or if it's a mislabel from the auto label or something like that.
43:53So let's have humans look at that. So ultimately, that's the way we're approaching the problem. So humans are still, I think humans are needed to verify. If you can minimize that work and quantify and predictably quantify how much work there would be, I think that's like when we get to annotation, you know, next version of annotation, whether it's 1.5 or 2.0 or whatever. Do you have any intuition around how this approach would perform for out-of-distribution domains? So like, I don't know, fault detection and electron microscopy images or, you know, whatever some industrial use case that's not represented in COCO and some of these other data sets?
44:36I don't have a great answer. I mean, I think we just haven't measured it really, right? I think that's something we want to do. I think one way that we would measure that is the first step we probably would do, and I'd be glad if someone else who's listening does this work before us. It's fine. This is not a competition. It's a cooperation, right? I would take a data set domain like medical imaging for which there are in-domain foundation models and then use out-of-domain foundation models like the ones we've used here as well as those in-domain ones. Interesting, yeah. and kind of like predict the performance compare the performance see if there is a their distribute predictable distributions over them almost like a sort of like a foundation model t-test if you will that that might be able to do that um for for truly out of domain or like like even if it's in because out of the main can mean a lot of things right what if it's in domain but challenging you know like the the old like visual verbs work from ali for haughty like a decade ago or something like that right like it's better if you have you know a person you have person on horse, you have a person on bicycle, yeah, person climbing or something like that, right?
45:42Like, it's actually, what they found was it was better to train separate detectors for, like, the compositions, because these models were training, these spaces were training on, and these models, these features were leveraging are not compositional. I still don't think the current semantically enriched embedding spaces we have today are compositional, sufficiently compositional still so but that's that's one direction of complexity whereas like truly out of domain like fault faults or like manufacturing defects like one of our customers is looking at that notion right like how do we find defects in production of like glass for example or like microchip microchip type things and um you know the the way the way i joke about that problem is in some sense, like it's very, the way things can go well is, you know, in terms of performance and like building things on the manufacturing lines are very few, like it's correct or it's not correct.
46:40It's right. Whereas like the way things can go wrong is like this really massive and hard to predict space. So I don't have a great answer for how we would assess this auto labeling of that domain. I mean, in general, though, I think the way I envision, you know, the future of annotation, if you will, or dataset building really, is models will be more, like we will have agents, like almost like embedding space agents who are trained, whether or not it's with RL or I don't know, but like we'll train to be asking the domain experts when they're not sure directly, right? Like generalizing the notion of uncertainty as use in active learning for the situation where you're not necessarily training one model downstream.
47:30You're just kind of enriching the embedding space to ensure that decision boundaries are well separated. I think that's where the space is going to be going technically. And is there something unique about the agentic-ness of that? Or are you envisioning both a technical approach and a user interface? Like, is what you're describing any different than, you know, you have a bunch of low-certainty images and you just send that back to the user for review? Is there some specific thing you're envisioning, you know, this core agentic behavior? I don't know. We're overusing collectively the term agent or the notion of agent right now.
48:21But I am thinking of it in the context of a human in the loop like UX for sure. But I do think that the underlying optimization function or cost function that's being optimized may need to be evolving based on what the feedback is. And so there may actually be this notion of a meta model, if you will, that's exploring the space with other models. And we may need to be refining how the space is being characterized. But that's a very high level. I don't have much details about that. No, certainly interesting to think about how agents would play into this, since we're all thinking about agents all the time anyway.
49:08Right. Right. And so we've talked about a handful of kind of next steps for this research. Any additional ideas that you're excited about? I mean, I am really excited about how to bring this into practice. You know, like I keep hearing about, I mean, I see a lot of papers that try to do some form of this or like, you know, or the whole world of semi-supervised ML. Like these are not really new ideas. Yet I still don't, I don't hear people like singing from the mountaintops that like, oh, look, I don't need humans anymore for labeling this. My need for human labeling is 10 % as it was last year.
49:53So either they've discovered it and they're not talking about it, which is a possibility, or it's still just in this experimental stage. And I'm really excited about the impact that these types of results or these types of systems will have in practice. I think we all want to see better AI models or better vision models being trained and being deployed. I'm just hoping that this is one way in that direction for practical attack. I think from a near-term technical or more researchy direction, I think that notion of how you minimize human work but still provide some guarantees on it. You know, like this is, you know, there's a whole area in like controls and other spaces in engineering that have models for mechanisms, say not models, but like mechanisms for measuring those types of uncertainties or performance guarantees and so on.
50:58We just haven't done that in our space. And I think we, as the field matures, we will need to mature the types of guarantees that we can offer in the space. So I'm pretty excited about that direction as well. On the topic of, you know, seeing people out talking about the application of these types of approaches, it does strike me that there's a lot of that happening in the text domain and language with the use of LLMs like, you know, a very popular topic is evals. and a lot of that eval work is being done nowadays with the generation of synthetic datasets for evaluation. And I wonder if there's a parallel here.
51:46I'm not an expert in the language domain specifically. I think there probably are similarities, but I think there are also interesting differences, right? Like, I guess, first of all, I'm a little worried about using VLMs or LLMs to generate synthetic data sets because they assume knowledge of the underlying, like, embedding space or, like, the manifolds in that space. And if we already know that we don't have full understanding, like, those manifolds are not yet sufficiently well-defined necessarily, right? And I guess evidence to that point, even in the tech domain, is that although we see LLMs being used to generate data sets and so on, many annotation companies have left the vision space and moved into the tech space because of the greater demand for more human expertise.
52:42And I think, why is that happening? Because there's this rich semantic structure in language that is sometimes hard to tease out. And I think there's still some learning to be done there. whether or not it's strictly like when you get when we get to trillions of tokens that will be enough i don't know i'll just i'll just shrug my shoulders with that i don't really know maybe maybe we need a different model i'm not sure um whereas in the visual domain the types of ambiguities you can find like those types of semantic ambiguities frankly i think are less i don't want to be on on record as saying computer vision is easier than nlp most of my colleagues which should be for saying that.
53:22But I think it's easier in some ways and it's harder in other ways, I think is a simple way to put it, right? Like there are some like lingual, visual, like some semantic vision ambiguities like donut hole. I gave this in a talk, I got to see the PR Decade or something like that, right? Like what's a donut hole? Is a donut hole actually the gap, the part of the donut that's missing? or is a donut hole the like the pop-up or whatever, right? Like the actual thing that came out of there, right? And until you've seen some examples, you don't necessarily know what that concept means. So there are certain ambiguities in the visual world like that and it's just, the visual world is just much higher dimension and very complex to process.
54:07You could even argue that that's not actually a visual ambiguity, but it's a linguistic ambiguity. You probably could, If you compare the hole and the donut thing, it's clear that those are different. It's just the language is weird. Yeah, I think you're probably right there. But I guess that's a characteristic of this notion that why do we perceive? We perceive to have some semantic grounding of what we perceive. And I think it's that bridge that makes it hard. Whereas I think the visual world is hard in other ways, right? Like the way shadows are created or the way like reflections happen or even like, you know, like my eyeglasses, for example, are highly reflective.
54:50In fact, I have a new pair waiting at the optometrist because of that problem. Like aspects like this, right? So anyway, yeah. So good to be in research in these problem spaces these days because it's such a rich environment. Awesome. Awesome. Well, Jason, thanks so much for jumping on and sharing a bit about what you've been working on. Thanks a lot, Sam. It was a great conversation. Thanks for having me.
55:42Thank you.
56:13Thank you.
From the publisher
Today, we're joined by Jason Corso, co-founder of Voxel51 and professor at the University of Michigan, to explore automated labeling in computer vision. Jason introduces FiftyOne, an open-source platform for visualizing datasets, analyzing models, and improving data quality. We focus on Voxel51’s recent research report, “Zero-shot auto-labeling rivals human performance,” which demonstrates how zero-shot auto-labeling with foundation models can yield to significant cost and time savings compared to traditional human annotation. Jason explains how auto-labels, despite being "noisier" at lower confidence thresholds, can lead to better downstream model performance. We also cover Voxel51's "verified auto-labeling" approach, which utilizes a "stoplight" QA workflow (green, yellow, red light) to minimize human review. Finally, we discuss the challenges of handling decision boundary uncertainty and out-of-domain classes, the differences between synthetic data generation in vision and language domains, and the potential of agentic labeling.
The complete show notes for this episode can be found at https://twimlai.com/go/735.




