Self-Supervised Contrastive Learning is Approximately Supervised Contrastive Learning

28 Jan 2026 · 15 min · 6 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Self-supervised contrastive learning (no labels) can approximate supervised contrastive learning, because “false negatives” vanish as the number of classes grows; the learned representations exhibit neural collapse geometry (augmentation collapse, within-class collapse, simplex ETF) that improves few-shot learning.

Guest backgrounds

No guest names or bios are provided in the transcript; only two speakers are present.

Key claims

Standard contrastive learning treats random samples as negatives, but NSCL (“Negatives only supervised contrastive loss”) shows the gap to the supervised objective goes to zero roughly like 1/dollars (number of classes). Training on unlabeled data implicitly minimizes supervised loss.

Notable examples

Dog-image augmentations forming positive pairs; random “strangers” as negatives; experiments on CIFAR-10/100 tracking unsupervised vs supervised loss in lockstep; augmentation similarity ~0.9–1.0 and different-image similarity ~0.1; few-shot learning explained via “pancake” directional variance and linear probes.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Contrastive Learning

0:50 to 4:00

Exploration of contrastive learning and its implications in AI modeling.

“Yet when we look inside their brains, their vector spaces, they've organized the world almost exactly like models that were supervised by humans the whole time.”

The Flaw in Contrastive Learning

4:00 to 5:00

Discussion of the false negative issue in contrastive learning models.

“It introduces a concept to explain the why called NSCL.”

NSCL: A New Perspective

5:00 to 9:50

Introduction to Negatives Only Supervised Contrastive Learning (NSCL) and its advantages.

“And the magic ingredient that makes it all work is the letter of dollars, the number of classes in your universe, the number of distinct concepts, dogs, cats, trucks, planes.”

The Impact of Class Size

9:50 to 10:40

Explaining how increasing class size reduces false negatives and improves model accuracy.

“But, and there's always a, but why do we care about this geometry?”

Neural Collapse and Its Properties

10:40 to 14:00

Exploration of the phenomenon of neural collapse and its geometrical properties.

“And if it's a big fuzzy cloud, we assume it's hard to separate from other clouds.”

Exploring Learning Distinctions

14:00 to 14:34

The conversation delves into the boundaries between supervised and unsupervised learning.

“So does the distinction between supervised and unsupervised learning even exist in the limit?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Imagine for a moment that you've been locked in a room, a really, really big room. and all over the floor are, say, a million photographs. Okay. You have absolutely no context for them, no captions, no labels, nothing. Your only instruction is organize these. That sounds like a nightmare, or at least a task that would take a human a lifetime to do and probably do very poorly. Right. And yet, if you were forced to, you would eventually start noticing things. You know, put all the fezzy, four-legged creatures in one pile. You'd put the metallic four-wheeled objects in another. Sure. You wouldn't know the words dog or car because no one gave you those labels, but you'd get the basic concept.

0:38You would have reinvented categories from scratch. Yeah. And that essentially is the holy grail of what we call self-supervised learning. But it comes with a paradox. A huge one. Yeah. And that's what we're tackling today. We have these massive AI models that are never shown a single label. They are blind. Yet when we look inside their brains, their vector spaces, they've organized the world almost exactly like models that were supervised by humans the whole time. It really shouldn't work as well as it does. That's the mystery. If you don't tell the model what a cat is, how does it build a perfect mathematical representation of catness?

1:15Well, today we are deep diving into a theoretical breakthrough that claims to have solved this. It argues these models aren't really unsupervised at all. They're secretly approximating supervision. It's a case of fake it till you make it, but, you know, backed by some really rigorous math. We're going to talk about something called contrastive learning, a geometric phenomenon called neural collapse, and why the size of your data set matters more than we ever thought. So let's unpack this. We have to start with the mechanism itself. If the model doesn't have labels, how does it learn anything? The main technique right now is called contrastive learning.

1:51And it's deceptively simple. The whole idea is just similar things should be close together in my mind and different things should be far apart. Okay, but similar is the key word there. Without a label saying this is a golden retriever, how does the computer know what similar even means? It's just pixels. It's sort of. It cheats. It creates its own similarity. The model takes an image, let's say a photo of a dog, and it creates two distorted versions of it. How so? Maybe it crops one, turns the other black and white, rotates it slightly. These two messes are called the positive pair. So the AI tells itself, okay, I don't know what this thing is, but I know this cropped version and this rotated version came from the same source.

2:34So they're family. Pull them closer. Exactly. That's the attraction part. But you can't just have attraction or everything would just collapse into a single point. You also need repulsion. You need to know what you aren't. Right. So the model takes that dog image and compares it to a random bunch of other images from the data set, a truck, a salad, a bird, whatever. And it says, these are strangers. Push them away. That's it. We call this contrast of loss. Pull the positives in, push the negatives out, do that a billion times, and you should get a smart model. But wait, I see a massive flaw in that logic.

3:06This has always bothered me about this method. I think I know where you're going with this. It's the strangers. The model assumes everything else is a negative. But what if I grab a random image and by pure chance, it's another golden retriever? You've hit it. The model has no idea it's also a dog. It just knows it's a different file. So it treats that second dog as a negative. So it's actively trying to push two dogs apart in its mental map. It's saying these two things are opposites, even though they are the same concept. Which feels catastrophic, right? You're punishing the model for being right.

3:40We call that a false negative. And logically, that should break the system. It's creating contradictory signals. It should completely confuse the model. It should, but we know that it doesn't. All the big models, MCLR, MoCo, they use this exact method and they work brilliantly. So why doesn't this false negative problem destroy everything? And that's where the new explanation comes in. Exactly. It introduces a concept to explain the why called NSCL. Which stands for something complicated, I'm sure. Negatives only supervised contrast of loss. Wow, that's a mouthful. Let's break that down. The name is complex, but the idea is actually pretty simple.

4:17Imagine a hypothetical world where we do have labels, but we only use them for one thing, to prevent those accidents we just talked about. Okay, so in this NSCL fantasy land, the model grabs that second golden retriever. And a little referee steps in and says, Whoa, wait a minute, check the label. That's also a dog. Don't push it away. Just ignore it. So the referee just removes the bad data. In NSCL, you never treat a classmate as an enemy. Precisely. It's the perfect version of contrastive learning. It's what we wish we could do if we had all the labels. So NSCL is the ideal. Standard contrastive learning is the messy, blind reality.

4:52And the breakthrough is proving that the messy reality is destined to become the ideal. Yes. The math proves that the gap between the two, the error from those false negatives, it vanishes. And the magic ingredient that makes it all work is the letter of dollars, the number of classes in your universe, the number of distinct concepts, dogs, cats, trucks, planes. So the argument is that as the number of classes goes up, the probability of picking a false negative by accident just drops to zero. Exactly that. I think the stadium analogy is perfect for this. Go for it. Okay. Imagine you're in a tiny room with only four people.

5:29One is your brother. You're blindfolded. you have to pick one person at random to be your opponent. What are the odds you pick your brother? In a room of four. Pretty high. It's a one in three chance you pick him as the opponent. You're very likely to get it wrong. That's your false negative. Right. Small data set, few classes, high error. But now let's move to a massive football stadium. 50 ,000 strangers. Your brother is still in there somewhere. Just one of them. You pick a random person. The odds are astronomically low. It's practically zero. Exactly. The stadium size is the number of classes, or dollars.

6:03As the complexity of the world grows, the chance of making a mistake by acting blindly just disappears. That is the vanishing gap. The math shows the error decays at a rate of roughly one over dollars. So if you train on ImageNet, which has a thousand classes, the error is tiny. If you train on the whole internet with millions of concepts... The noise from those false negatives is just completely drowned out. Right. It's such a cool inversion because usually we think more complexity makes a problem harder. Here, more classes make the unsupervised method better. It's like the diversity of the world itself becomes the supervisor.

6:39The model implicitly minimizes the supervised loss without ever seeing a label, purely through the law of large numbers. Okay, so that explains how it learns without breaking. The math saves it. But what does it actually learn? What does the AI's brain look like when it's done? This is where we go from probability to geometry. And the geometry is, honestly, it's beautiful. We're talking about a phenomenon called neural collapse. Neural collapse. That sounds bad. Like a medical emergency for a robot. It does sound dire, but in machine learning, collapse is usually a good thing. It means removing variation.

7:13It means tidying up. Things are falling into their proper place. So the messy data collapses into a clean structure. Exactly. The analysis shows that as this loss gets minimized, the data points, the representations, arrange themselves into this very specific, almost crystalline structure. It's broken down into three properties, right? Let's walk through them. First up, augmentation collapse. This one's the most intuitive. Remember those distorted images we started with? The crop, the black and white version. A positive pair. Right. Augmentation collapse just means that the model learns to map all of those distorted versions to the exact same point in space.

7:50So to the model, a photo of a dog and that same dog, but upside down in purple, are mathematically identical. Identical. The model becomes invariant to the noise. It learns that style doesn't matter, only content. Okay, that makes sense. Which brings us to the second property, within class collapse. This is the one that really surprised me. It is surprising for an unsupervised model. In a supervised model, you tell the AI, all these thousand different photos are dogs, so of course it tries to group them. But here, we never gave that instruction. Yet the math shows it does it anyway. It takes all the different golden retrievers, different lighting, different poses, and sucks them all into a single tight point.

8:30Or very, very close to one. All the variation within the category just disappears. It leaves only the pure essence of the class. So you end up with these tight, dense dots of meaning. One dot is dog, one is cat, one is airplane. So how are those dots arranged relative to each other? That's the third property, and it has this fantastic name, the simplex equiangular tight frame, or just simplex ETF. Okay, that sounds like we're back on the bridge of the Starship Enterprise. I know, it's a mouthful. Yeah. But the concept is all about efficiency. Imagine you have a sphere, and you want to place, say, 100 points on its surface, so that they are all as far apart from each other as possible.

9:09Like maximizing social distancing? Exactly. You want maximum separation, but you also want it to be fair. You don't want dog and cat to be far apart, but dog and wolf to be close. You want every concept to be maximally distinct from every other concept. And that perfect spiky shape where everything is equidistant, that's the simplex EPF. Correct. It's the optimal geometric structure for telling things apart. And what this research shows is that the loss function naturally pushes the representations into this exact shape. I love that. You start with this messy, label-free soup, apply a simple rule, and it collapses into this perfect crystalline jewel.

9:47The geometry just emerges from the probability. But, and there's always a, but why do we care about this geometry? It's elegant, sure, but does it make the AI better at doing useful things? It absolutely does. The killer application here is something called few-shot learning. Okay, so this is a scenario where I have a huge pre-trained model, and I want to teach it a new trick. Like, recognize a rare bird it's never seen, and I only have five photos of it. Exactly. Five-shot learning. Usually, five examples isn't enough to learn anything. But with these collapsed models, it works incredibly well.

10:20And the explanation for why comes down to pancakes. Ah, that's one way to put it. The formal term they use is directional class-distance normalized variance. Let's stick with pancakes. It's much catchier. Okay, pancakes it is. Imagine two clusters of data, dogs and cats. Traditionally, we'd look at the total variance like how big and fuzzy is the whole cloud of points. And if it's a big fuzzy cloud, we assume it's hard to separate from other clouds. Right. But this research says total messiness doesn't really matter. What matters is directional messiness. So imagine the dog cluster is shaped like a giant flat pancake, and the cat cluster is another pancake.

11:01Okay, I'm picturing two Frisbees floating in space. Perfect. Now, even if those Frisbees are miles wide, if they're facing each other flat side to flat side, you can easily slip a piece of paper between them. Because they're really thin in the direction that actually matters. Exactly. The directional variance along the line connecting them is super low. They are compact where it counts. The simplex ETF structure naturally creates this. So these self-supervised models are basically prepackaging the world into these easily sliceable segments. Yes. That's why we can use something called a linear probe, which is just a fancy term for drawing a straight line or a plane to classify things so well.

11:41The hard work of shaping the data is already done. That's the real aha moment for me. Yeah. It explains why these models can adapt so fast. Learning a new task isn't about reshaping the universe. It's just about finding where to slide the knife. That's a great way to put it. So we've covered a lot of theory. Did they actually prove this happens in the real world, or is this just a nice set of equations? Oh, they ran the experiments. They trained models on standard data sets, CIFAR-10, CIFAR-100, and they tracked the invisible supervised loss the whole time. So while the model was training blind, they were secretly calculating how well would this model be doing if it did have the answers.

12:16Exactly. And they plotted both curves, the unsupervised loss and the supervised loss. And they moved in perfect lockstep. One went down, the other went down, almost a mirror image. So optimizing one really is accidentally optimizing the other. And they validated the stadium analogy, the one over dollar rule. They showed that when they went from a 10 class problem to 100 class problem, the performance gap between the blind model and the supervised model shrank. The stadium got bigger, the errors got smaller. And they found that augmentation collapse is very real. The similarity between two different distortions of the same image hit about 0.9.

12:54And 1.0 is a perfect match. Right, so they're almost identical to the model. Meanwhile, different images stayed down near 0.1. The signal is incredibly clear. Okay, this has been a heavy lift, but I feel like we connected some huge dots. We started with the paradox. How does a blind model learn to see? And we found the answer in probability. If the world is complex enough, the errors just wash away. We saw that this process sculpts the data into that perfect simplex ETF structure, packing information as efficiently as possible. And we learn that this structure is shaped like, well, pancakes to make future learning incredibly efficient.

13:31You know, it's actually comforting. It's nice to know that unsupervised learning isn't just wild magic. It's implicit supervision. It's supervision by the structure of reality itself. I think that's the key takeaway. We tend to think we provide the meaning when we add labels to data, but this suggests the meaning is already there in the data structure. If the data set is big enough, the truth is just unavoidable. Exactly. And it leaves us with a really provocative question to end on. As we build models that ingest the entire internet, billions of concepts, that gap we talked about, it effectively vanishes.

14:05So does the distinction between supervised and unsupervised learning even exist in the limit? Or is the universe of data just supervised by its own complexity? If you see enough of the world, maybe the world just teaches you what it is. No teacher required. That is a thought I'm going to be chewing on for a while. the idea that data is its own teacher that changes everything. It really does. It suggests intelligence might just be an emergent property of scale and geometry. Well, on that note, we're going to wrap up this deep dive. We hope you learned something new about the hidden geometry inside these digital brains.

14:42It's been a pleasure unpacking it all with you. Thanks for listening. We'll catch you on the next deep dive.

From the publisher

This research explores the theoretical alignment between self-supervised contrastive learning (CL) and supervised learning, specifically investigating why label-agnostic training produces organized semantic clusters. The authors prove that standard CL objectives implicitly approximate a negatives-only supervised contrastive loss (NSCL), with the gap between the two vanishing as the number of dataset classes increases. Their analysis identifies that global minimizers of this loss exhibit augmentation collapse, within-class collapse, and a simplex equiangular tight frame structure, mirroring the "neural collapse" found in supervised models. The paper introduces a new few-shot error bound based on directional feature variability, which explains how these models support high-accuracy label recovery with minimal supervision. Empirical tests across diverse vision datasets confirm that minimizing the unsupervised CL loss effectively drives down the supervised NSCL loss. Ultimately, the study provides a robust mathematical framework to justify the success of contrastive pre-training in downstream classification tasks.

More from Best AI papers explained

All 475 episodes
Self-Supervised Contrastive Learning is Approximately Supervised Contrastive LearningBest AI papers explained · 15 min
Listen in VO