In short
Netflix researchers argue that cosine similarity between learned embeddings may not reflect true semantic similarity because hidden scaling introduced by regularization can make cosine scores arbitrary, even when dot-product rankings look correct.
Guests
Harold Steck, Chaitanya Kanadem, and Nathan Callas (Netflix research team; paper authors). The hosts discuss deployment in recommendation engines and semantic search.
Key claims
In linear matrix factorization, L2 regularizing the product (Scheme 1) is invariant to rescaling user/item latent dimensions by an arbitrary diagonal matrix D and D inverse; dot products stay unchanged, but cosine similarity breaks because normalization happens after training and scaling/normalization don’t commute. Extreme cases: item-item cosine becomes identity (only self-similar) while recommendations remain unchanged; user-user cosine collapses to noisy raw data.
Notable examples
Simulated 20,000 users/1,000 items with five ground-truth clusters shows three different cosine similarity heat maps from the same model under different valid D rescalings; Scheme 2 (separate weight decay) removes the ambiguity but may not guarantee semantic correctness. Remedies: train for cosine directly (e.g., normalize before final dot), compute similarity in reconstructed feature space, and reduce popularity bias via inverse propensity scaling; Word2Vec’s negative sampling (frequency^0.75) is cited as evidence.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOIntroducing the Problem with Cosine Similarity
0:45 to 2:12
Exploring the implications of cosine similarity and its role in AI.
“That's Harold Steck, Chaitanya Kanadem, and Nathan Callas.”
Understanding the Research Approach
2:12 to 3:19
Discussion on the Netflix research team's methodology in addressing cosine similarity.
“And we confidently declare those items semantically similar.”
The Distortion Origin in Training
3:19 to 4:43
Unpacking how regularization impacts cosine similarity during training.
“And their methodology is particularly elegant here.”
Exploring Regularization Schemes
4:43 to 6:10
Examining two distinct regularization schemes and their effects.
“Since we rely on L2 regularization to prevent overfitting, to stop the model from simply memorizing the training data, the specific implementation we choose is critical.”
Vulnerabilities in Scheme 1
6:10 to 8:13
Identifying vulnerabilities in regularization scheme 1 and its implications.
“We will refer to this arbitrary scaling factor as matrix D.”
Case Studies: High Stakes Examples
8:13 to 9:38
Analyzing extreme cases of cosine similarity and their disastrous implications.
“It means the similarities your model is spitting out are not unique.”
Simulating the Ideal Scenario
9:38 to 11:49
How simulations were used to test cosine similarity against a baseline.
“The unnormalized dot product was serving up highly relevant content.”
Introducing Scheme 2
11:49 to 13:04
Presenting an alternative regularization scheme and its benefits.
“Those three valid scalings produce three drastically different item similarity heat maps from the exact same model.”
Caution Despite Stability
13:04 to 13:50
Highlighting the limitations of stability in semantic similarity.
“So, on paper, that sounds like a definitive fix.”
Warning for Deep Learning Practices
13:50 to 14:02
Discussing the implications of the research on modern AI practices.
“Which bridges us to the most urgent warning in this paper.”
Show all 17 chapters
The Shift from Linear to Deep Learning Models
14:02 to 15:17
Learn about the transition to deep learning and the challenges it introduces.
“But the reality of the modern AI industry is that almost no one relies solely on linear models anymore.”
Risks of Cosine Similarity in Deep Learning
15:17 to 16:02
Understand the complexities and risks associated with using cosine similarity in deep learning.
“We are measuring distances and angles in a high dimensional space that has been stretched, skewed, and squished by implicit scaling factors we can't even calculate.”
Practical Remedies for Measuring Similarity
16:02 to 16:57
Explore actionable strategies for effectively training models with cosine similarity.
“Let's break those down for the listener.”
Bypassing the Embedding Space
16:57 to 17:54
Learn about the innovative approach of avoiding the embedding space for better results.
“And their second proposed remedy is a fascinating structural pivot.”
Addressing Popularity Bias in Data
17:54 to 19:22
Discover methods to normalize data and mitigate popularity bias when using cosine similarity.
“And their third remedy addresses a pervasive issue in all statistical modeling.”
Word2Vec's Impact on Semantic Similarity
19:22 to 19:44
Learn how Word2Vec effectively handled popularity bias to achieve semantic similarity.
“They aggressively mathematically squashed the popularity bias during the training loop.”
Rethinking Semantic Similarity in AI
19:44 to 21:02
Reflect on the implications of the research on how AI interprets semantic relationships.
“It is a massive wake-up call for the entire industry.”
Transcript
Automatic transcript. May contain errors.0:00Welcome. It is it's genuinely great to have you with us today. Yeah, thanks for having me. I'm excited to get into this one. So imagine for a second that you are relying on a highly precise laser measure to build the foundation of a skyscraper. Okay. A high-stakes construction project. Exactly. You have used this exact tool a thousand times. It is the industry standard. But today, someone taps you on the shoulder and proves mathematically that the units on your laser measure can stretch or shrink entirely based on the ambient temperature. Wow. And it does this completely invisibly while still telling you the foundation is perfectly square.
0:37You would immediately halt construction. You'd have to. Your ground truth is gone, right? Well, today we are taking a deep dive into a brand new paper from a team of researchers at Netflix. That's Harold Steck, Chaitanya Kanadem, and Nathan Callas. And the paper is titled, Is Cosine Similarity of Embeddings Really About Similarity? It is a genuinely striking paper. Yeah. And quite provocative. It really is. Yeah. And what they found might just be the equivalent of that broken laser measure for the entire artificial intelligence industry. We're going to explore why cosine similarity breaks down, the hidden math causing the chaos, and a major warning for deep learning.
1:14As someone who spends most of my time analyzing the theoretical architecture and the mathematical bounds of these models, I find it vital to regularly look under the hood of the tools we just take for granted. Definitely. And the tool in question here is cosine similarity. The implications of this research really challenge some of the most fundamental assumptions we make in modern machine learning. And as someone who actually builds and deploys these systems, I'm talking about recommendation engines, semantic search pipelines, the systems you interact with every day. I can tell you firsthand that cosine similarity is our default semantic yardstick.
1:49It's everywhere. It's everywhere. When we take discrete entities like words in a large language model or movies in a catalog and we map them into dense vector spaces, we need a way to compare them. So we measure the angle between those vectors using cosine similarity. It is the absolute industry gold standard. Right, because if the angle is small, the vectors point in the same direction. Exactly. And we confidently declare those items semantically similar. It's a clean, intuitive metric, and virtually everyone in the field assumes it maps perfectly to human intuition. But this raises an important question.
2:24If that metric maps perfectly to human intuition, why does it empirically and quite frequently, I might add, perform significantly worse in AD tests than the simple, unnormalized dot product between those exact same vectors? Yeah, that is the open secret we rarely discuss at industry conferences. Right. In production environments, I have seen countless models where we just multiply the raw vectors together, the raw dot product. We entirely ignore the normalization for the angle, and we actually generate better real-world recommendations, higher click-through rates. Which feels completely counterintuitive.
2:58It really does. Yeah. Because the raw dot product incorporates the magnitude of the vectors, which usually just biases the model toward overall item popularity rather than true semantic meaning. Yet it often wins. So the Netflix authors recognized that anomaly, and they set out to rigorously investigate the mechanics behind it. And their methodology is particularly elegant here. Because they didn't just look at a massive neural network. Exactly. They avoided the trap of trying to dissect a massive, multi-billion parameter neural network. In deep learning, the layers are black boxes. You simply cannot trace the exact mathematical transformations with absolute certainty.
3:38There's just too much noise. Right. Right. So instead, they anchored their proof in linear matrix factorization models. Which is brilliant because matrix factorization, or MF, is incredibly well understood. It's foundational. Yeah. You start with a massive sparse data matrix, say millions of users and their interactions with thousands of items. And the objective is to estimate a low-rank matrix that approximates that massive data set by taking the product of two smaller dense matrices. The user embeddings and the item embeddings. Exactly. And because it is a linear model, there are no nonlinear activation functions muddying the waters.
4:13We have closed form mathematical solutions for these linear models. There is nowhere for the variables to hide. This allows us to see the exact downstream effects of our architectural choices. So what did they actually find when they looked at those closed form solutions? Well, by analyzing them, the authors discovered that the root of the problem isn't the cosine similarity formula itself. Oh, really? Yeah, the distortion is actually born much earlier in the white line, specifically during the training phase. And it stems entirely from how the model is regularized. Okay, let's unpack this. Regularization.
4:47Right. Since we rely on L2 regularization to prevent overfitting, to stop the model from simply memorizing the training data, the specific implementation we choose is critical. And the paper focuses heavily on two very distinct regularization schemes. Yes, let's look at what they call Scheme 1. This applies the L2 norm penalty to the product of the user and item matrices. Just the product. Right. Scheme 1 is pivotal to this entire discussion. Mathematically, the objective function is evaluating and penalizing the reconstructed matrix as a whole. And in my pipelines, we almost always default to Scheme 1, or at least its deep learning equivalents.
5:27Regularizing the product is mathematically akin to learning with denoising or applying dropout. Exactly. You're masking things. Right. We randomly zero out certain data points during training, forcing the model to work harder to find robust underlying latent patterns. And it is incredibly hard to convince an engineering team to drop a regularization scheme like Scheme 1 when your empirical test data consistently shows it improves predictive accuracy. Because it does improve prediction on held out data sets. It just works better for predicting the specific items a user will click on. It definitely moves raw predictive accuracy, but, and here's the catch, the researchers analyzed the closed form solution for this specific objective function and found a glaring vulnerability.
6:09What's fascinating here is that scheme one is entirely invariant to being rescaled by an arbitrary diagonal matrix. We will refer to this arbitrary scaling factor as matrix D. Let me bridge the math to the practical application here to make sure I'm following. If I take the learned user matrix and scale its latent dimensions by this random matrix D, and then I simultaneously take the item matrix and scale its dimensions by the exact inverse of D, so D inverse, the fundamental predictions of the model remain completely untouched. Untouched, yes. The raw dot product of those two scaled matrices is identical to the dot product of the original matrices.
6:49So the model is just completely blind to it. Completely blind. To the unnormalized model, that arbitrary matrix D and its inverse completely cancel each other out during the multiplication process. Because 5 times 1 fifth is just 1. Exactly. The loss function only cares about the final reconstructed product, not how the individual latent dimensions were scaled to arrive at that product. You can introduce literally infinite variations of this arbitrary scaling, and the unnormalized predictions will not change by a single decimal point. The raw predictions don't change. But cosine similarity isn't an unnormalized metric.
7:25We normalized the row vectors to calculate the angle after the model has finished learning and generated those embeddings. And that chronological order is the fatal flaw. The after-the-fact part. Yes. Because you apply cosine similarity after the fact, you introduce row-wise normalization matrices. matrices. And mathematically, diagonal matrices, unless their diagonal values are perfectly uniform, do not commute with normalization matrices. You cannot simply swap their order in the equation. Precisely. Because they don't commute, the neat little cancellation of D and D inverse completely fails to happen in the embedding space.
7:59So when you normalize those vectors to find the cosine similarity, you permanently bake that arbitrary matrix D into your results. The cancellation is destroyed. Your final similarity score is now entirely dependent on whatever arbitrary scaling happened to be floating around in the latent dimensions during training. That is wild. It means the similarities your model is spitting out are not unique. Because there are infinite, valid choices for matrix D that satisfy the training objective, there are infinite, completely valid, but wildly different cosine similarities for the exact same data set.
8:32Here's where it gets really interesting. If the math allows for infinite arbitrary matrices, what happens if we plug in the absolute worst case scenario matrix? Oh, the edge cases are fascinating. The paper walks through some bizarre extreme cases in a full rank model to prove just how disjointed this semantic space can become. Case A is a perfect example of this. Right, KSA. By choosing one mathematically valid scaling for matrix D, the authors demonstrated that the item-item cosine similarity becomes the identity matrix. And the implications of that are staggering for anyone building a search engine.
9:06An identity matrix means that the off-diagonal elements are all zero. Yes. So in the context of semantic embeddings, an item is only similar to itself. It possesses exactly zero semantic similarity to any other item in the entire database. Yet, ironically, the user item recommendation rankings remain exactly identical to the unnormalized dot product. I have actually run into the ghosts of this problem in production without realizing what the math was doing. Really? Yeah. I once deployed a model where the recommendation rankings were phenomenal. The unnormalized dot product was serving up highly relevant content.
9:42But when we queried the raw embedding space to build a more like this carousel based on cosine similarity, the results were chaotic. Completely disconnected. We were suggesting niche horror films to users watching baking tutorials. The semantic space was a disjointed mess. But because the dot product was still optimizing for the click, the primary recommendation metrics looked flawless. The model was working, but the embeddings meant nothing. That is a perfect real-world example of what they proved. And case B in the paper is just as destructive. Right. Case B approaches the distortion from the opposite direction.
10:16If you choose the exact inverse for D, suddenly the user-user cosine similarity completely collapses and reverts to the raw, noisy data matrix. It effectively erases all the smoothing benefits of matrix factorization. Exactly. The entire point of pushing data through a low-rank bottleneck is to smooth out the noise and find latent patterns between users. Case B mathematically strips all of that away. Putting you right back at square one, as if the model had learned absolutely nothing about user relationships. To solidify these theoretical extremes, the researchers built a highly controlled simulated experiment.
10:54Because in real life, we don't really have ground truth. Right, right. We rarely have a perfect metric for ground truth semantic similarity in real world data sets. We don't objectively know how mathematically similar two movies are, so they simulated 20 ,000 users and 1 ,000 items and strictly grouped those items into five ground truth clusters. They engineered a block diagonal structure, so they possessed the absolute mathematical truth of which items belonged together. Building their own universe allowed them to test the metric against an undeniable baseline. Figure 1 in the paper visualizes this beautifully.
11:28Yeah, they show the ground truth clusters as a pristine heat map with five distinct bright blocks along the diagonal. Then they train the linear model using Scheme 1, the product regularization that mimics denoising. Because the model allows for that invisible matrix D, they simply extract three different mathematically valid rescalings. And the visual evidence is undeniable. Those three valid scalings produce three drastically different item similarity heat maps from the exact same model. From the same model. One heat map is a blurry, washed-out mess where the clusters are barely perceptible.
12:05Another shows vague hints of the original structure but introduces massive amounts of noise. And the third wildly overemphasizes certain connections while suppressing others. You are looking at three entirely different structural realities, all equally valid under the loss function, generated simply by tweaking that invisible scaling factor. But Figure 1 also visualizes the outcome of their second regularization approach, Scheme 2. Right, Scheme 2. Scheme 2 represents standard weight decay, where you apply the L2 norm penalty to the user matrix and the item matrix separately, rather than to their product.
12:38Because Scheme 2 penalizes the matrices independently, the underlying mathematics simply do not allow for the arbitrary matrix D to exist. The separate penalties lock the latent dimensions down. Exactly. It prevents them from scaling inversely against each other without incurring a massive loss penalty. Consequently, the visualization shows that Scheme 2 produces one unique heat map. The arbitrary distortion is eliminated. So, on paper, that sounds like a definitive fix. Practitioners should just migrate to Scheme 2, use separate regularization, and secure unique, stable cosine similarities across the board.
13:16Well, the researchers are incredibly cautious about drawing that conclusion. Why is that? Scheme 2 yields a unique mathematical solution, yes. But it remains an open question whether this unique solution actually captures the best possible semantic similarity. Oh, I see. Stability does not automatically equate to ground truth accuracy. Precisely. It simply means the model is consistently producing the same representation, but that representation could still be a poor reflection of true semantic meaning. It is uniquely stable, but potentially uniquely unproven in its alignment with human intuition.
13:50Yes. Which bridges us to the most urgent warning in this paper. We have spent this entire time dissecting linear matrix factorization models because they offer mathematical transparency. But the reality of the modern AI industry is that almost no one relies solely on linear models anymore. No, everything is deep learning now. Right. We are building massive, multilayer deep neural networks. Large language models, deep sequential recommenders, complex retrieval augmented generation pipelines. If we connect this to the bigger picture, the Netflix researchers offer a primary cautionary warning. Deep learning is a minefield of implicit regularizations.
14:27Building deep models is inherently messy. We utilize dropout layers to prevent co-adaptation. We apply weight decay. We insert batch normalization or layer normalization at different stages to stabilize gradients. You're constantly tweaking layer-specific hyperparameters just to force the model to converge. And the paper mathematically proves that all of those messy, layer-specific interventions implicitly apply unknown, opaque scaling across the network. Just like our arbitrary matrix D. Yes. By utilizing dropout or varied weight decay across different layers, practitioners are essentially injecting invisible, arbitrary D matrices throughout the entire architecture.
15:07So if linear models, which are mathematically pristine and fully observable, can have their semantic spaces completely shredded by a single regularization choice. Applying cosine similarity blindly to deep learning embeddings is highly risky and opaque. We are measuring distances and angles in a high dimensional space that has been stretched, skewed, and squished by implicit scaling factors we can't even calculate. You are trusting a measuring tool that has been warped by invisible architectural forces. The resulting similarities are not just opaque. They are quite possibly entirely arbitrary byproducts of how the network decided to route its gradients.
15:43So what does this all mean for the engineers and data scientists deploying these models today? We cannot simply abandon vector databases or stop relying on semantic search. No, you can't. And the authors don't just leave us staring at a broken system. They outline several highly actionable practical remedies. Let's break those down for the listener. Their first recommendation is essentially, train directly for the metric you want to use. It's the most direct solution available. If you want cosine similarity, train the model with it from the start. Right. Rather than hoping the network magically organizes the space correctly, you mathematically mandate it.
16:20Exactly. For example, using layer normalization just before the final dot product in the architecture, this forces the model to learn representations that are inherently normalized. Meaning the arbitrary scaling factor d cannot take root because the vectors are constrained to a unit hypersphere during the optimization process itself. There are computational trade-offs, of course. Calculating exact cosine similarities across massive batches during training can be computationally expensive. It often requires batch-wise approximations. Right. However, the theoretical guarantee of a stable semantic space often outweighs the computational overhead.
16:57And their second proposed remedy is a fascinating structural pivot. Avoid the embedding space entirely. When I first read that, it sounded completely counterintuitive. Same. The entire premise of representation learning is to utilize the embedding space. But the logic is sound. The embedding space is exactly where the arbitrary matrix D wreaks havoc. So to bypass the mathematical trap of the latent dimensions, you project the data back into the original feature space. You take your learned user and item matrices, multiply them together, and reconstruct the full, smooth data matrix. Yes. And then you apply your cosine similarity metric to those reconstructed feature vectors rather than the raw embeddings.
17:38You avoid the flawed embedding space entirely. You still reap all the benefits of the model's complex pattern recognition and noise reduction, but you execute the similarity measurement in an observable space that hasn't been warped by internal regularization scaling. It is a brilliant bypass. It really is. And their third remedy addresses a pervasive issue in all statistical modeling. Popularity bias. They strongly advocate for pre-normalizing your data. Right. Addressing popularity biases before or during learning. When you apply cosine similarity at the end of a pipeline, you are normalizing vectors after they have already been skewed by the frequency of the training data.
18:17Because in real-world data sets, some items are wildly popular. Blockbuster movies? Common stop words. And if left unchecked, their vector magnitudes explode. They aggressively pull the entire semantic space toward themselves simply due to sheer frequency. The default statistical approach is to standardize the raw inputs to zero mean and unit variance. But the paper highlights more advanced techniques, like inverse propensity scaling or IPS. IPS is highly effective here. By weighting the training samples by the inverse of their probability of being observed, you mathematically flatten the statistical artifacts caused by hyperpopularity.
18:53You force the model to focus on the actual contextual relationships rather than just memorizing which items appear most frequently. The researchers actually cite a legendary example of this mechanism in action, Word2Vec. Oh, Word2Vec. Anyone working in natural language processing reveres Word2Vec because it achieved uncanny semantic word similarities early in the deep learning boom. And it turns out much of that success wasn't just the architecture, it was how they handled negative sampling. They sampled negative examples with a probability proportional to their frequency in the training data, raised to the power of 0.75.
19:29They aggressively mathematically squashed the popularity bias during the training loop. Word2Vec proves the central thesis of these remedies. You cannot treat semantic similarity as a passive byproduct of an optimization function. You have to actively architect for it. It is a massive wake-up call for the entire industry. To synthesize what we have uncovered today, cosine similarity is arguably the most ubiquitous metric in artificial intelligence for measuring semantic relationships. Yes. But, as these Netflix researchers have mathematically proven, the underlying regularization of your model, especially if it mimics due-noising or dropout, can secretly introduce infinite degrees of freedom.
20:05These hidden scaling factors can render your resulting cosine similarities completely arbitrary. You cannot simply extract a vector, apply the metric, and blindly trust the results. Because the dimensional space itself might be fundamentally warped, this research demands a paradigm shift from assumption to rigorous mathematical verification. And that leaves us with a final, rather provocative thought for you to mull over. We rely on these embeddings to form the cognitive core of modern artificial intelligence. They dictate how large language models parse nuance, how visual models categorize our digital world, and how algorithms curate the information we consume daily.
20:45It's the foundation. But if the spatial relationships in our most advanced AI embedding spaces, the very blueprints of how AI understands how words or concepts relate to one another, can be so easily distorted by invisible mathematical scaling. What does that mean for the worldview of the AI we are building? That's the real question. Is the AI's map of reality genuinely grounded in true semantic meaning, or is it just an accidental, arbitrary byproduct of a regularization parameter we chose to lower the loss function? It is a profound question and one that every researcher, engineer, and practitioner needs to be actively interrogating.
21:22Absolutely. Thank you so much for joining us on this deep dive. Keep questioning your assumptions, keep looking under the hood of your architectures, and we will catch you next time.
From the publisher
This paper investigates whether cosine similarity accurately reflects the semantic similarity of learned embeddings, particularly in linear matrix factorization models. The authors demonstrate that the metric can produce arbitrary or non-unique results because certain training objectives allow for the random rescaling of latent dimensions. While some regularization methods yield a unique solution, others leave the final similarity scores dependent on opaque modeling choices rather than the underlying data. These findings suggest that the common practice of applying cosine similarity to high-dimensional vectors may lead to misleading conclusions. Consequently, the researchers advise against the untested use of this metric and suggest alternative normalization or projection techniques to ensure more reliable measurements.




