In short
Explains why self-supervised contrastive learning (CL) without labels can still produce semantic categories, and shows a theoretical/empirical link to a supervised objective called NSCL (Negatives-Only Supervised Contrastive Learning).
Guest backgrounds
No guests mentioned; it’s a solo “Deep Dive” host discussing a paper.
Key claims
(1) CL’s “false negatives” (e.g., poodle treated as negative for retriever) are corrected by NSCL by removing same-class negatives rather than pulling positives together. (2) As number of classes grows, unsupervised CL’s loss approximates NSCL. (3) Weight space diverges strongly, but representation geometry aligns closely.
Notable examples
Shared-randomness experiment: supervised vs unsupervised models had ~85.7° weight-vector angle but ~27.8° representation-vector angle (CKA/RSA). Temperature stabilizer: higher temperature (1.0 vs 0.1) improves alignment. Neural collapse: NSCL/CL preserve structured class clouds instead of collapsing to single points. “Modern merging”: averaging encoders trained on different data/variants improved performance, showing compatible geometry.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Contrastive Learning
0:45 to 2:13
Explains the basics of contrastive learning and its role in AI.
“Push this specific picture's duplicates together, push everything else apart.”
The Mystery of Unsupervised Learning
2:13 to 3:14
Discusses the puzzle of how models learn categories without labels.
“But before we get deep into the math, I want to tease what I thought was the craziest part of this study.”
NSCL: A New Approach
3:14 to 6:06
Introduces Negatives Only Supervised Contrastive Learning and its significance.
“So in standard CL, you have an image, say a golden retriever.”
The Tale of Two Spaces
6:06 to 8:13
Explores the divergence of weights and similarities in representations between models.
“So the unlabeled objective starts to look just like the labeled one mathematically.”
Stabilizers in Neural Learning
8:13 to 10:50
Examines factors that stabilize the alignment in learning models.
“Centered kernel alignment and representation similarity analysis.”
Neural Collapse and Its Implications
10:50 to 12:57
Discusses neural collapse and how it relates to classifier performance.
“Okay, I want to pivot to something the paper brought up that I found fascinating, mostly because of the name.”
Merging Models: A Surprising Discovery
12:57 to 14:01
Details an experiment on merging different trained models and its outcomes.
“So we already established the weights are completely different, right?”
Understanding Contrastive Learning
14:01 to 14:22
Learn how contrastive learning relates to supervised learning and its implications.
“It's implicitly optimizing a geometry that aligns with a very specific and very gentle type of supervised learning.”
Key Takeaways for Engineers
14:22 to 15:00
Discover the paper's insights on stability, representation geometry, and training methods.
“I think the biggest takeaway is about stability.”
The Mysteries of the Mind
15:00 to 15:43
Explore the implications of different neural wiring and shared realities in minds.
“That looseness, that cloud structure is a feature, not a bug.”
Show all 11 chapters
Final Thoughts on Consciousness
15:43 to 16:08
Reflect on the nature of consciousness and representation in learning.
“Are there infinite ways to wire a brain to see the same reality?”
Transcript
Automatic transcript. May contain errors.0:00All right, welcome back to the Deep Dive. Today, I've got a paper in front of me that I think really gets to the heart of what feels like a magic trick in modern AI. It's a puzzle that, you know, has been bugging researchers for a while now. It really is the central mystery, isn't it? We're talking about how these huge models can learn so much without, well, without a teacher. Exactly. So just to set the stage for everyone, we're talking about contrastive learning. This is the engine behind so much of what we see. You take a neural net, you throw millions of images at it, but, and here's the catch, you give it zero labels.
0:35No dog, no cat, nothing. Right. You just give it a very simple game, take two versions of the same image, pull them close together, and everything else, push it far away. That's the whole game. It's a purely instance-level task. Push this specific picture's duplicates together, push everything else apart. There's no idea of a category given to the machine at all. But then, and this is the magic trick, after it plays this simple push and pull game for, you know, billions of cycles, you look inside its brain and it's organized the entire world by category. All the dogs are together. All the cats are in their own little cluster.
1:10The toasters are hanging out with the microwaves. It creates this deep semantic meaning. It figures out what a dog is without ever being told the word dog. How? Why does just pushing random images apart create a structured understanding of reality? That's been the question. For so long, we've just kind of accepted it. We had intuition, sure. We'd say, well, it looks like a duck. It's grouped with ducks, so it must be learning duck. Yeah. But we haven't had a really rigorous mathematical reason why. Until now, maybe. We're diving into a paper titled On the Alignment Between Supervised and Self-Supervised Contrast of Learning.
1:45And I have to say, this isn't just another benchmark paper. This thing reads like a geometric detective story. It really does. The authors wanted to find that hidden bridge, you know, between learning on your own self-supervised and learning with a teacher. And they found this this kind of mathematical translator to do it. They call it NSCL. NSCL. Negatives only supervised contrast of learning. OK, that's a mouthful. It is. But the idea is actually pretty elegant. And it's the key that unlocks the whole mystery. We are definitely going to unpack that. But before we get deep into the math, I want to tease what I thought was the craziest part of this study.
2:20To prove all this, the researchers went like full on scientific control freak. They use this method called shared randomness. A fascinating setup. They trained two models, one with labels, one without. But they use the exact same starting point, the exact same batches of data, even the same random augmentations. Every single variable was controlled. They wanted to see if the models would, you know, evolve the same way. And the result, well, the spoiler is their brains, the actual wiring, the weights, they end up looking completely different. Exponentially different. But their thoughts, how they saw the world, were almost identical.
2:55It's what they call the tail of two spaces. You have weight space and representation space. And that distinction, I think, fundamentally changes how we should think about what a neural network even is. Okay, let's back up. Let's build this bridge properly. Let's start with that core puzzle again. We mentioned contrastive learning, CL methods like SimCLR or MoCo. Right. So in standard CL, you have an image, say a golden retriever. You make two tweaked versions of it. Maybe you crop one, change the colors on the other. The model just learns to say, hey, these two came from the same source, so their vectors should be identical.
3:30And everything else in that batch, the trucks, the birds, even other dogs, those are negatives. The model's job is to push them away. Exactly. So I get why it learns to recognize that specific golden retriever. But here's the part that's always felt like a leap of faith. Why does it then group that retriever with a poodle? The model was never told they're both dogs. In fact, it was told the opposite. Under the rules of the game, that poodle is a negative. It was explicitly told to push the poodle away. So why does an instance-level task create semantic boundaries? That's the mystery. And the paper argues that as you scale up, the math of that unsupervised loss function starts to, well, it starts to mimic a very specific kind of supervised learning.
4:17And that gets us back to NSCL. Right. Negatives only supervise contrastive learning. So how is this different from normal supervised learning? Usually if I have labels, I'm telling the model, hey, all these dogs, they belong together. Pull them all close. That's it. Exactly. In standard supervised contrastive learning, SCL, you are explicitly forcing that cluster. You say, maximize the similarity between this dog and all other dogs. It's aggressive. It's like a sheepdog herding all the sheep into a tiny pen. That's a great analogy. NSCL is way more passive. In NSCL, you only use the label to clean up the negatives.
4:52Clean up the negatives. What does that mean? Well, think about that unsupervised game again. You have your golden retriever. You're pushing away everything else. But statistically, everything else might include a poodle. So you're accidentally telling the model that the retriever and the poodle are totally different concepts. Ah. I see. Which is technically true. They are different images, but it's semantically false. They're both dogs. So you're introducing noise. You're creating what's called a false negative. NSCL just fixes that one error. It looks at the batch and says, okay, we're pushing things away.
5:22But hang on a second. That poodle is actually the same class as this retriever. So just don't push it away. Take it out of the negative pile. So it doesn't force them together? It doesn't say you two should be friends? No. It just stops telling them they have to be enemies. That's a huge distinction, but it's so subtle. It really just sounds like contrastive learning with the mistakes erased. That's exactly what it is. And the big theoretical finding here is that as the number of classes grows, the math of the unsupervised model of CL just naturally starts to look like this NSCL objective. Because if you have a thousand classes, the odds of a poodle randomly showing up in the same batch as your retriever are pretty low.
6:03Exactly. The accidental pushing away of other dogs just becomes rare noise. The signal dominates. So the unlabeled objective starts to look just like the labeled one mathematically. So the unsupervised model is accidentally doing supervised learning. But this specific negatives only type. That's right. The ghost in the machine is this NSCL process. It's unknowingly chasing the same goal. Okay, that makes sense as a theory. But proof is better. This is where we get to that tale of two spaces. This part honestly just blew my mind. This is the empirical heart of the paper. They set up that shared randomness experiment we talked about.
6:41Just to be clear how strict this was. Two training runs. One is unsupervised CL. The other is supervised NSCL. They start with the exact same random weights. They see the exact same images in the same order. You'd think intuitively that their parameters, the numbers on the neurons, would stay pretty similar. I mean, they're learning from the same world. Right. If we're both solving the same math problem with almost the same method, our scratch paper should look kind of similar. But that is not what happened. Not at all. In weight space, they just diverged completely, and they measured the angle between the weight vectors of the two models.
7:14And for anyone trying to picture this, zero degrees means identical. 90 degrees means totally unrelated. The angle between the weights was 85.7 degrees. 85.7. That's, I mean, that's basically orthogonal. It means the internal wiring of the two models had almost nothing in common. It shows us that parameter space coupling is just inherently unstable. Even that tiny difference in the loss function, just removing the poodle from the negatives, causes the weights to drift apart exponentially. It's like the butterfly effect inside the network. So under the hood, there are two totally different machines.
7:47Strangers, but, and this is a huge but. But when you look at their output, the representations, the thoughts... They measure the angle between the representation vectors. And that angle was only 27.8 degrees. Wow. That's incredibly close. It is. So despite having completely different brains, they arrived at basically the same understanding of the data's geometry. And they used some fancy metrics for this too, right? CKA and RSA? Right. Centered kernel alignment and representation similarity analysis. Without getting bogged down in the math, there are just rigorous ways to compare the geometry of two high-dimensional spaces, and they both confirmed it.
8:25The geometry is preserved. The supervised ghost guides the unsupervised machine to the same conclusion, even if the path it took, the weights, is totally unique. That's wild. It kind of implies there are infinite ways to wire a brain to see the world correctly. The wiring doesn't matter. The reality it represents is what's fixed. It really does suggest that. The truth of the data is what's dictating the representation. But this alignment isn't just magic. The paper found there are conditions, right? They call them stabilizers. Basically, what knobs can you turn to make this happen reliably? Exactly.
9:00And they found the biggest knob is temperature. Temperature, tau. In AI, this usually controls how confident a model is in its predictions. Yeah, that's a good way to put it. A low temperature makes the model very opinionated, very sharp. It picks a winner. A high temperature softens everything out. It spreads the probability around. And what do they find? High temperature is a stabilizer. Models trained with a temperature of 1.0 had a way higher alignment than those trained at a sharp 0.1. Why? Wouldn't you want it to be more decisive? You'd think, but not for alignment. The math shows that higher temperature just smooths out the optimization.
9:36Remember those false negatives? The poodle we accidentally pushed away? With a low temperature, the model basically says, I hate that poodle. Get it away from me. It makes a really harsh, incorrect judgment. But with a high temperature, it's more like, push it away a little, but let's not go crazy. It softens the penalty for those semantic mistakes. So chilling out helps the model find the real structure. It makes it less confident in its own errors. You got it. What about batch size? We always hear bigger is better for contrastive learning. And that holds up. Larger batches just give you a better sample of the true data distribution.
10:11It's the law of large numbers. So it stabilizes the alignment because you're less likely to get thrown off by a weird batch of images. And they also said the number of classes matters, which goes back to the whole poodle thing. Yep. It validates the theory. The approximation to NSCL gets better as the number of classes grows. If you only have two classes, this alignment is kind of a mess. But if you have thousand, ten thousand, the unsupervised method lines up almost perfectly with the supervised one. Which explains why these huge foundation models trained on the whole internet are so good. The scale itself is what creates the alignment.
10:49Precisely. Scale isn't just more data, it's better geometric alignment. Okay, I want to pivot to something the paper brought up that I found fascinating, mostly because of the name. Neural collapse. Neural collapse. It does sound like a sci-fi disaster movie. A neural collapse coming this summer. It sounds bad, but it's often what you want in a classifier. Ideally, if you're training a model to recognize cats, you want every single cat picture to map to the exact same point in space. You want the representation to collapse to a single dot. Right. A cat is a cat is a cat. The model should just say cat.
11:20Exactly. And methods like standard supervised contrastive learning, SCL, and cross-entropy are very aggressive about forcing this collapse. They crush the data down. But this paper says NSCL, and by extension, unsupervised CL. They don't do this. No, and that is the key. NSCL creates a looser structure. It doesn't force all the cats to one point. It keeps the instance level structure. A tabby cat is still different from a Persian cat in the representation space. So it's more a structured cloud than a single point. A structured cloud. And this is why it tracks the unsupervised model so well. The unsupervised model can't collapse the classes because it doesn't know they exist.
11:58NSCL just happens to mimic that exact behavior. It respects the data's geometry instead of crushing it. That feels like a philosophical point almost. Yeah. By not forcing the collapse, you're preserving more of the world's nuance. You're preserving the richness. And they showed this holds up across architectures too. ResNet 50, Vision Transformers, same story. The Vision Transformer proof was really cool with the attention maps. Right. They looked at what parts of the image the model was looking at. And the heat maps for the unsupervised model and the NSCL model were structurally almost identical.
12:27So if one was looking at the dog's ear, the other was also looking at the ear. But the standard supervised model might be looking somewhere else entirely. Exactly. The standard model is just looking for the most efficient shortcut to get the right label. The CL and NSEL models are actually looking at the object itself. Okay. To really hammer this all home, they did an experiment that I frankly didn't think would ever work. The modern merging one. That was the mic drop moment for me. Walk us through it. So we already established the weights are completely different, right? 85 degrees apart. Different universes, right.
13:04They took an encoder trained with unsupervised learning on all the data and another one trained with NSCL on just 30 % of the data, and they just averaged them. They just mixed the representations, like plugged one half of a brain into a totally different brain. Pretty much, yeah. They just linearly interpolated the signals. And the result wasn't just garbage. I would assume that would create chaos. You would. But no, the merged model was actually better than the individual models it was made from. That is absolutely mind-blowing. It proves the geometries are compatible. They're like plug-and-play pieces.
13:36It improves that the truth they found is stable. Even though they had different weights, they built a map of reality so compatible you can just overlay them and get a sharper image. It's like we both drew a map of a city from memory. My handwriting is different. Maybe I used a blue pen, you used red. but the streets align so perfectly that if we stack them, we just get a better map. That's a perfect analogy, and it validates the entire premise. Contrastive learning isn't magic. It's implicitly optimizing a geometry that aligns with a very specific and very gentle type of supervised learning. Oh, the mystery of the supervised ghost is, well, maybe not solved, but we have a really strong suspect.
14:15We have a very clear picture of it, yeah. The ghost is just NSCL in disguise. So let's wrap this up. Why should someone listening to this, maybe an engineer, training these things, why should they care about this paper? I think the biggest takeaway is about stability. We worry about AI being a black box, you know. We see weights diverging and we think it's all chaos. But this paper says maybe stop worrying so much about the weights. They're just implementation details. Exactly. The truth is in the representation geometry and that geometry is surprisingly stable and predictable. It also gives us a recipe for better training, doesn't it?
14:51Absolutely. You want better alignment with semantic reality. The paper says, high temperature, large batch sizes. And don't worry that it doesn't collapse the classes. That looseness, that cloud structure is a feature, not a bug. It keeps the nuance. It keeps the poodle distinct from the retriever, even while knowing they're both dogs. It preserves the richness of the data. And that might be why these models are so much more flexible than the old supervised ones. I want to leave our listeners with one final thought this sparked for me. We talked about how the weights can be completely different, 85 degrees apart.
15:25But the understanding is the same. If you apply that to us, to biology, if you and I can have totally different neural wiring, different weights from our life experiences, yet we can agree on a shared reality. What does that imply about where our mind is? That's a heavy question. Are there infinite ways to wire a brain to see the same reality? And if there are, is consciousness in the physical wiring or is it in the abstract geometry of the representation that floats above it? That is a question for a much, much longer deep dive, but it's a great one. Makes you look at your own thoughts a little differently.
16:01It does. We'll leave you to think on that one. Thanks for joining us on this dive into the geometry of learning. My pleasure. See you next time.
From the publisher
This research explores the mathematical and empirical relationship between Contrastive Learning (CL) and Non-Contrastive Supervised Contrastive Learning (NSCL). The authors demonstrate that CL and NSCL converge toward highly similar structural representations, a phenomenon they validate using metrics like Centered Kernel Alignment (CKA) and Representational Similarity Analysis (RSA). Their theoretical framework identifies key variables—such as temperature, batch size, and learning rate—that determine the proximity of these two methods in similarity space. Experimental results on datasets like CIFAR and ImageNet confirm that these training dynamics lead to nearly identical attention maps and feature distributions. Ultimately, the paper provides a formal proof that unsupervised contrastive models inherently approximate their supervised counterparts under specific optimization constraints.




