In short
Why CLIP-style image-text retrieval fails on specific attributes (e.g., “zero people,” “brick vs stone,” “morning vs afternoon”), and how promptable (question-conditioned) image embeddings fix it by making the embedding dynamic and attribute-focused.
Guest backgrounds
No guest names or bios are provided in the transcript; two speakers discuss the paper and experiments.
Key claims
Standard embeddings optimize “global alignment,” so dominant scene cues (forest, horse race, beach) drown out subtle attributes (car, zero people, texture, time of day). Text captions omit mundane details, training the model to ignore them. Promptable embeddings reweight attention based on a question; LLM-generated question templates guide retrieval.
Notable examples
“beach with zero people” returns crowded beaches; “car” returns horse race; counting fails (2 vs 5 people); morning vs afternoon confusion; bird-on-cup causes text hallucination (“two birds flying”); chicken ambiguity (animal vs dish). Performance: recall@5 improves from ~59% to ~75% (counting/materials biggest gains). Acceleration: offline preprocessing plus “linear approximation” (lightweight projection) for low-latency deployment.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Model Limitations
0:47 to 1:20
Exploring why AI models like CLIP struggle with specific image attributes.
“And that is exactly what we're digging into today.”
Global Alignment Challenges
1:20 to 3:04
Discussion on how models prioritize dominant signals over specific details.
“I think most people assume the model sees an image like a human does, but that's not really right, is it?”
The CocoFacet Benchmark
3:04 to 3:52
Introduction to the CocoFacet benchmark and its implications for AI performance.
“And to prove this, the researchers built a new benchmark, CocoFacet.”
The Problem of Reporting Bias
3:52 to 6:24
Examining why AI models ignore specific details due to human input biases.
“They both just map to group or crowd or take materials.”
Promptable Image Embeddings Explained
6:24 to 7:37
Overview of the concept of promptable embeddings and how they function.
“This paper proposes promptable image embeddings.”
Performance Improvement with Promptable Embeddings
7:37 to 8:36
Analyzing performance improvements achieved by utilizing promptable embeddings.
“If I'm running a search engine, I can't have a human sit there and write, what is the texture for a billion images?”
Text-Only vs Embedding Approaches
8:36 to 10:19
Comparing the effectiveness of text-only methods to embedding techniques.
“In retrieval, you're usually fighting for 1 % or 2%.”
Acceleration Strategies for Querying
10:19 to 11:52
Discussion on strategies to enhance model efficiency during querying.
“Running a question and answer loop for every single image sounds like it would just burn a hole in your GPU budget.”
Visual Reasoning and Its Implications
11:52 to 13:14
Exploring the shift towards visual reasoning in AI and its potential.
“It's a very elegant engineering compromise.”
Transcript
Automatic transcript. May contain errors.0:00You know, there's this very specific, very relatable frustration with image search. Not the easy stuff. Oh, right. Like you type in sunny beach and it's perfect. You get a million results. The model has seen that a billion times. Right. It's the hello world of computer vision. Exactly. But the moment you add a negative constraint, the whole thing just falls apart. You search, find me a photo of a beach with zero people. And you get a beach party. A huge beach party. Crowded boardwalks, kids everywhere. It sees beach and just completely ignores the zero people part. It's the classic failure mode.
0:36We are incredibly good at what you might call vibes, the general idea, but surprisingly bad at specific attributes, the details. And that is exactly what we're digging into today. Why these big systems like CLIP, which are basically the engine behind every image tool we use. Yeah, they're everywhere. Why are they so blind to these details? And we're going to look at a new approach called promptable representations that, well, it seems to fix it by changing how we talk to the model. It's a really fascinating pivot. It moves away from this idea that an image has one single static meaning. To what?
1:13To the idea that an image's meaning is dynamic. It depends entirely on the question you ask it. Okay, so let's start with the status quo. I think most people assume the model sees an image like a human does, but that's not really right, is it? Not really, no. Standard text-to-image retrievers, and we're mostly talking about models like CLIP here, They work by mapping images and text into a shared space, a latent space. Think of it like a huge high dimensional map. Exactly. And the whole goal is to push the vector for an image of a dog and the vector for the word dog into the same little neighborhood on that map.
1:49So they have a high cosine similarity. They're close. Right. But the problem is how they get there. These models are optimized for what we call global alignment. Meaning the loudest signal in the noise, the biggest thing it sees. The loudest signal, yeah. Yeah. If an image is, say, 90 % forest and there's a tiny red car in the corner, the global vector is going to scream forest. And the car signal just gets washed out in the math. Completely. In the data we're looking at, they talk about this exact problem. If you search for car, the model might give you back a picture of a horse race. A horse race?
2:21Why? Because there's a blurry vehicle way in the distant background. The model isn't seeing the car as an object. It's just seeing a scene where a car might exist, but the horse signal is just so, so dominant. It feels like it's prioritizing confidence over accuracy. It grabs the biggest concept it can find. That's a great way to put it. I like to think of CLIP as an over-enthusiastic tour guide. You walk into a museum and he just shouts, Art Renaissance. He gets the genre, he gets the vibe. But if you ask him, hey, which of these paintings used oil versus acrylic? He'd have no idea. Useless. He didn't look at the brush strokes.
2:59He just glanced at the room. Because he's optimizing for the general theme, not the attribute data. Right. And to prove this, the researchers built a new benchmark, CocoFacet. They realized the old tests were just too easy. You know, is this a cat? A low bar. A very low bar. So CocoFacet is a stress test. It's over 9 ,000 queries across eight really specific categories, counting people materials time of day gestures stuff that requires actual inspection and the results I'm guessing they weren't great cracks in the foundation is a mild way of putting it I mean they were failing on things a toddler could do day counting right these massive multi-billion parameter models cannot reliably tell the difference between two people and five people that seems insane given the compute power we're talking about is it that they can't count or that they just don't care it's that in the embedding space, two people and five people are basically the same general concept.
3:55They both just map to group or crowd or take materials. Distinguishing a stone wall from a brick wall. If the shape is similar, the model just encodes wall. It completely glosses over the texture. There is that one failure case in the notes that I found really telling the morning versus afternoon light thing. Yes. It confuses time of day all the time. And this brings me to a different analogy. Think of the AI as a movie reviewer. Okay. This reviewer writes a summary. It says, this is a high octane action movie starring Tom Cruise. It gets the genre. It gets the star. But if you ask, hey, was it raining during that car chase?
4:31The reviewer has no clue. Draws a blank. Because that detail wasn't necessary to establish the main vibe of the movie. It ignored the non-dominant attributes. Okay. But this is where I get stuck on the why. We train these models on like the entire internet. Surely someone somewhere has labeled a photo brick wall or morning light, why is the reporting bias so bad? This is the crucial insight. The model isn't just mimicking images. It's mimicking how humans describe images. I know we're lazy. We're efficient. If you upload a photo of your friend at a cafe, you caption it, brunch with Sarah. Right.
5:05You do not caption it. A wooden table with four metal legs, natural lighting, roughly 10 a mere featuring one human female. No, because people would think I was As a robot, the table and the lighting are just their assumed context. Exactly. Yeah. We mentioned the extraordinary, a person skydiving, but we omit the mundane, a person breathing. And because the text data constantly leaves out these details, the model learns that those features are statistically irrelevant. It creates a blind spot. So it's not that the model is blind. It's that we've effectively trained it to ignore the background because we ignore the background in our text.
5:41Precisely. And this leads to that dominant versus wrong behavior. The model will actually prefer an image that has the wrong attribute, but a loud, clear subject, over an image with the correct attribute where the subject is subtle. Wow. It chooses the strong signal every single time. So we've got the problem. The tour guide is just skimming. The obvious engineering fix, and I know they tried this, is just make the model bigger, throw more parameters at it. Which is the standard Silicon Valley answer to everything. Right. But it didn't work. I mean, they tried scaling up the resolution. They tried using the really heavy hitter of multimodal large language models.
6:18And it helped a little, but the blind spots were still there. You can't scale your way out of a fundamental lack of attention. So what's the actual fix? This paper proposes promptable image embeddings. How is that mechanically different from what we do now? So right now, an embedding is static. You feed in an image, you get out a vector. That vector is supposed to represent everything in the image forever. This new approach says, no, let's make the embedding dynamic. Let's condition the visual encoder on a specific question. So instead of just saying process this image, you say look at this image and answer.
6:52How many people are there? Exactly. It's like giving that tour guide a specific mission. Instead of him just wandering around shouting, art, you hand him a clipboard and you say, count the people. And his attention shifts. Suddenly, the model's entire attention mechanism re-weights the features. The vibes fade into the background, and the count signal becomes the dominant part of the vector. I love the spotlight analogy for this. A standard embedding lights the whole stage equally. Everything's visible, but nothing pops. Promptable embedding is like manually telling the operator, shine the spotlight on the drummer.
7:26And suddenly, the drummer is the only data point that matters. The rest of the band is still there visually, but for the vector representation, they're just noise and the drummer is the signal. But practically, this implies we need a question for every image. If I'm running a search engine, I can't have a human sit there and write, what is the texture for a billion images? No, of course not. And that's where the LLM comes in. They use GTT4O to automate the curiosity. They just had the language model generate templates for every category. So for materials, GBT-4 would automatically spit out queries like what materials is that object made of or is the surface smooth or rough?
8:05So you use an LLM to generate the questions, which then guide the vision model to pull out the right features. What did that do for performance? The jump is massive. They used a metric called recall at five, which is basically does the correct image show up in your top five results? Which is a pretty reasonable bar for a search engine. Yeah, totally. On these really hard tasks, the baseline models were sitting at around 59%. With the promptable embeddings, that shot up to 75%. That's a 16-point swing. In retrieval, you're usually fighting for 1 % or 2%. It's a step change. And the biggest gains were in the hardest categories, the counting and the materials.
8:45Once you've forced the model to look at the texture, it suddenly becomes an expert on brick versus stone. I want to play devil's advocate for a second. If we have these super capable multimodal LLMs that can look at an image and answer questions, why do we even need the embedding part? Why not just ask the LLM to write a really detailed caption and then search the text? Ah, the text only or dense captioning approach is a very logical question. But it fails for two very specific reasons. The first one is hallucination. The model just inventing things that aren't there. Yes. There was a great example.
9:20They showed the text model an image of a cup with a painting of a bird on it. Okay. They asked, are there birds in this image? And the model said, yes, two birds flying. Oh, wow. So it completely failed the ontology check. It couldn't tell the difference between a bird and a picture of a bird. Exactly. It saw the semantic concept bird, and then it hallucinated the action flying because, you know, birds fly. It decoupled the object from its physical reality. But the second problem is ambiguity. Take a word like chicken. Right. If I just search the text chicken, do I want a live animal walking around in a coop, or do I want a roasted dinner on a plate?
9:56Two very different images. They're very different. And text compresses both of those visual realities into the same six-letter word. Whereas the promptable embedding keeps the visual data, the feathers versus the crispy skin, it keeps all that in the vector space. Precisely. It preserves the visual nuance that language just trips away. Okay, so we have a much better method, but there's always a catch. This sounds computationally heavy. Running a question and answer loop for every single image sounds like it would just burn a hole in your GPU budget. It absolutely would. You can't run a massive MLLM for every single query at inference time.
10:31It's way too slow. The latency would be terrible. So how do they get around that? They came up with two acceleration strategies. The first one is pretty straightforward. Pre-processing. Okay. If you're an e-commerce site, you know your users are going to search for material, color, and fit. So you just run those questions offline ahead of time, store the embeddings, and they're ready to go. That's the cheat sheet approach. It works great for structured data where you know the questions, but what if I'm searching for something totally random? That's where the second strategy is really clever. They call it linear approximation.
11:03Right. This is the part that felt a little bit like magic. Can you break down the math for us, sort of simply i'll try so think of the heavy question answering model as an oil painter it takes hours but it creates a perfect portrait okay got it the researchers found they could train a tiny lightweight transformation matrix a filter basically that just mimics what the oil painter does that's a projection a shortcut yes it's like a sketch artist who watches the oil painter and learns a 30-second shortcut to get a similar result. It's not quite as detailed as the oil painting, but it captures all the essential structural changes.
11:41And the performance actually holds up. Surprisingly well. This sketch method still got an 8 % improvement over the baseline, and it runs at basically the same speed as a standard search. That's the holy grail, isn't it? High accuracy, low latency. It's a very elegant engineering compromise. It means you can actually deploy this stuff at scale. Stepping back, this really feels like a shift in how we think about machine vision. We're moving from just pattern matching to something that looks a lot more like reasoning. That's the term they use, visual reasoning. We aren't just matching the color blue to the word sky anymore.
12:16We are asking the AI to actually understand the state of the world, you know, the physics, the quantities, the materials. It's the difference between glancing and actually observing. Yes. And the really provocative thing is the intelligence was sort of already there. We didn't need a new brain. We didn't need a new data set. We just needed to change how we talk to the brain we already have. That's the part that sticks with me. The brick wall information was sitting in the model's latent space the whole time. It knew it was brick. It just didn't think that was important enough to tell us. Until we asked.
12:50It implies there's all this latent intelligence just sleeping in these models, waiting for the right prompt to wake it up. It makes you wonder what else they're missing. Causal relationships? Maybe even emotions? I bet if you prompted, why is this glass broken? Instead of just show me broken glass, you'd get completely different results. One looks for shards, the other looks for a ball or a hammer nearby. We're just scratching the surface here. So the limitation isn't just the silicon anymore. It's our own imagination and how we query it. Couldn't have said it better myself. Well, next time you're yelling at your search engine because it gave you a beach party instead of solitude, just remember it's not ignoring you.
13:27It's just an overenthusiastic tour guide who needs you to point the spotlight for him. Keep asking the right questions. That's it for this deep dive. Thanks for listening.
From the publisher
This paper propose using promptable image embeddings guided by questions generated by an LLM, which help Multimodal models focus on specific visual attributes. They also implement a linear approximation strategy to reduce the high computational costs associated with using multimodal large language models (MLLMs) for large-scale searches. Experimental results demonstrate that these techniques significantly improve retrieval precision on complex queries compared to traditional baseline methods. Ultimately, this research aims to bridge the gap between global semantic understanding and the recognition of non-dominant visual details in digital images.




