In short
Argues “RAG is dead” as a misleading label; retrieval still matters, but teams should replace demo-driven “alchemy” with “context engineering” to reliably assemble the right information for LLMs.
Key claims
LLM performance degrades as context length increases (“context rot”); distractors worsen degradation; logically structured “haystacks” can perform worse than shuffled text; Claude tends to refuse (“conservative collapse”), GPT tends to hallucinate confidently, Gemini is variable, and even with context extension some models degrade.
Notable examples
needle-in-haystack tests with semantic ambiguity, distractors, and haystack structure; model-specific failure modes.
Guests
Jeff Huber, CEO of Chroma (background: Google planet-scale ads/maps systems). Episode also references Chroma’s “context rot report” and “generative benchmarking” approach.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Claim: RAG is Dead
0:45 to 2:12
Exploring Jeff Huber's statement about RAG and its implications.
“Yeah, and show you how it's actually shaping, you know, the next wave of AI applications.”
Understanding Context Engineering
2:12 to 3:56
Defining context engineering as a systematic approach in AI.
“Okay, so if RAG, or at least the term ag, is leading us down this alchemy path, what's the alternative?”
The Context Rot Report
3:56 to 6:01
Discussing the findings of the context rot report and its significance.
“And for Chroma, pushing this idea makes strategic sense, I guess, moves them beyond just being seen as a database component.”
Practical Tips for Context Engineering
6:01 to 7:49
Jeff Huber's five actionable tips for implementing context engineering.
“But here's the really counterintuitive one.”
Evaluating Context Engineering Success
7:49 to 12:01
Understanding generative benchmarking and its role in evaluation.
“What he means is be explicit about the building blocks.”
Chroma's Technological Journey
12:01 to 14:01
Exploring Chroma's development process and architectural choices.
“Okay, so with all this focus on context engineering and evaluation, how does Chroma itself, the company pushing these ideas, actually build its underlying technology?”
Exploring Chroma's Unique Positioning
14:01 to 17:02
Learn how Chroma differentiates itself in the vector database landscape.
“Object storage generally has a higher baseline latency than local SSDs.”
Chroma's Strategic Insights and Market Position
17:02 to 18:42
Understand Chroma's strategy and its implications for AI development.
“21 ,000 GitHub stars, over 5 million monthly downloads.”
Rethinking RAG and Context Engineering
18:42 to 19:55
Discover how context engineering transforms AI systems beyond traditional RAG.
“We've got to move beyond thinking about RAG in that simplistic, naive way.”
Transcript
Automatic transcript. May contain errors.0:00Welcome back to the Deep Dive. Today, we're digging into something that definitely caught my eye. The claim, a rag is dead.
0:09Jeff Huber:Yeah, quite the statement. It is. It comes from Jeff Huber, the CEO of Chroma. And, you know, it makes you stop and think, what does he really mean? And what does that mean for, well, for how we're all building AI right now? Exactly. It's not about just throwing retrieval away, right? It's more fundamental. Seems like it. More about rethinking how we actually get information into these models effectively. That's right. So for this deep dive, we've looked at an interview Jeff Huber did on the Latent Space podcast and also a pretty detailed investigative report digging into Chroma's claims and where they're headed strategically.
0:43Jeff Huber:So our mission today is really to unpack that RAG is deadline, then introduce this idea he talks about called context engineering, which sounds crucial. Context engineering, right? Yeah, and show you how it's actually shaping, you know, the next wave of AI applications. We'll get into Chroma's role, some of the tech underneath, and definitely some practical steps you can actually use today. Perfect. So, okay, RAG is dead. First reaction. Whoa. Retrieval augmented generation. I mean, that's been the bedrock for so much stuff. It really has. Is he saying just scrap it? Or is it more nuanced, like a critique of how maybe we've been, I don't know, misusing it?
1:19Jeff Huber:It's definitely the critique, not a literal death sentence for retrieval itself. Huber's argument is that the term R, the label, has become a bit of a liability. How so? Well, he argues it sort of mashes together three really distinct complex things. Retrieval, then augmentation, then generation. And it often gets boiled down to just one narrow idea, like a single dense vector search. Ah, okay. Oversimplified. Exactly. And that oversimplification, he says, leads to these apps that look great in a demo. Right. Easy to demo. But are a nightmare to actually put into production reliably. It feels less like solid engineering and more like, I think he called it alchemy.
1:59Alchemy. Yeah, that resonates. Things just sort of work sometimes, but you don't always know why.
2:03Jeff Huber:Precisely. And that confusion, that sort of conceptual fuzziness, it really gets in the way when you're trying to build robust, dependable AI systems. It hinders clear thinking. Okay, so if RAG, or at least the term ag, is leading us down this alchemy path, what's the alternative? This is where context engineering fits in, right? He's positioning it as a big shift. That's the idea. But what is it exactly? Is it just like a fancy name for prompt engineering, something we already talk about a lot? Good question. It's definitely distinct from prompt engineering, though related. Context engineering, as they define it, is more like the systematic discipline of designing AI systems.
2:42Okay.
2:42Jeff Huber:Systems that dynamically figure out, then retrieve, and then assemble the optimal set of information, tools, data, whatever is needed into the LLM's context window for a specific task. So it's much broader than just writing the prompt itself. Much broader. Prompt engineering is like crafting the perfect node or instruction within the context. Context engineering is about designing the entire room the AI is working in. Okay. I like that analogy. The whole information environment. Exactly. Making sure it has the right library, the right tools, maybe even the user's past conversations, all perfectly organized before the LLM even starts working on the immediate task.
3:18And you mentioned loops, inner and outer loops.
3:21Jeff Huber:Right. There are two key cycles. The inner loop is the real-time bit, pulling together the context for just one single generation step. That could be user history, docs you just retrieved, instructions, tool outputs. Got it. The immediate stuff. Yeah. Then the outer loop is the bigger picture, the meta process. It's about constantly improving that whole context-filling system over time. How? Like looking at performance? Exactly. Analyzing performance, getting feedback, tweaking your retrieval strategies, maybe updating your ranking models. It's about refining the system so it consistently delivers only the best, most relevant stuff to the LLM.
3:56And for Chroma, pushing this idea makes strategic sense, I guess, moves them beyond just being seen as a database component. it.
4:05Jeff Huber:Absolutely. It elevates the conversation. It directly tackles the real messy problems developers hit when they try to build complex, multi-step AI agents that actually work reliably in the real world. Okay. It sounds logical, very systematic, but you know, is there solid evidence backing this up? Why do we need such meticulous engineering? Which brings us to that context rot report from Chroma. Yes. The context rot report. This is where it gets really concrete. This research seems to directly challenge that whole idea that LLMs can just handle like infinitely long context windows perfectly. It really does.
4:39Jeff Huber:And that's the fascinating part. The core finding is pretty stark. LLM performance consistently goes down as the input context gets longer, even on simple tasks. Wow. OK. So much for those million token window marketing claims, meaning perfect recall. Exactly. It directly undermines those claims of perfect utilization. It's not just a minor glitch. It's just a fundamental issue. It forces us to confront this maybe surprising truth. Just stuffing more context in isn't always better. In fact, it can make things worse. It can actively degrade performance, which kind of flips the common wisdom about huge context windows on its head.
5:18So how did they test this? What was the methodology?
5:21Jeff Huber:They did something quite clever. They took the standard needle in a haystack benchmark, you know, hiding a key piece of info in a load of text. Right. See, the model can find it. Yeah. But they extended it. They varied the semantic ambiguity, deliberately added distracting information, and even tested different structures for the haystack, sometimes logically ordered, sometimes just randomly shuffled text. Okay. And what came out of that? Any big surprises? Oh, yeah. Several. First, distractors really amplify the rot. If you put in information that's semantically similar to the right answer but is actually wrong, performance degrades much faster.
5:57That makes intuitive sense for real-world data, which is rarely clean.
6:01Jeff Huber:Totally. But here's the really counterintuitive one. Structure can be a trap. Models actually performed worse when the haystack text was logically structured. Worse! Not better. That is surprising. Isn't it? It suggests that the LLM's attention mechanisms might get sort of sidetracked by the surface coherence of the text rather than zeroing in on the actual semantic meaning it needs. Huh. Okay. Anything else? Yeah. They also found distinct failure patterns depending on the model. Claude models tended towards what they called conservative collapse. They just started refusing to answer more often.
6:36Okay. Played it safe.
6:37Jeff Huber:Right. Whereas GPT models showed confident confusion. they were more likely to give confident hallucinations or just make stuff up. Oh, that's dangerous. Confidently wrong is the worst. It really is. Gemini's performance was more variable, needed careful testing, and Quinn, even with context extension techniques, still showed degradation. So knowing these patterns is actually useful for developers, then? Incredibly useful. If you know, say, that GPT might confidently hallucinate more with long contexts, you can implement stronger fallback logic or more rigorous post-generation checks. It helps you tailor your safety nets.
7:13Right. Build more robust applications based on the model's specific weaknesses. That structure as a trap finding is still blowing my mind a bit. Really hammers home the need for careful curation, not just dumping data in.
7:26Jeff Huber:Absolutely. Quality over sheer quantity. So if context rot is the diagnosis, what's the prescription? You mentioned Jeff Huber gave some actionable tips, the kind of blueprint. He did. He offered five really practical tips that basically lay out how to implement this context engineering idea. They turn retrieval from this vague thing into a proper workflow. Okay, let's hear them. This sounds like the how-to part. Exactly. Tip number one, don't ship rag. Ship retrieval. What he means is be explicit about the building blocks. Name the primitives. dense search, lexical or keyword search using filters, adding a re-ranking step, context assembly, and having an evaluation loop.
8:06So break it down, make it modular, transparent, less alchemy.
8:10Jeff Huber:Precisely. More engineering. Tip two, win the first stage with hybrid recall. Don't rely on just one method. Combine vector search with traditional full-text search and metadata filters to cast a really wide net initially. How wide? Aim for maybe 200 to 300 candidate documents. Huber's point is that LLMs can sift through way more information than a human researcher ever could. So get all the potential ingredients first. Okay, gather everything potentially relevant. Then what? Then tip three. Always re-rank before you assemble the context. Take those 200, 300 candidates and distill them down. You need the top, say, 20 to 40 most relevant items for the final context.
8:51And how do you do that distillation?
8:52Jeff Huber:Often with another LLM or a specialized model called a cross-encoder, These are really good at comparing a query to a document and scoring relevance. It's an extra step computationally, but it massively improves the quality of what goes into the final prompt. And interestingly, using LLMs themselves for this re-ranking is becoming more common and cost effective. Right. Using the AI to help refine the input for the AI. Makes sense. What's next? Tip four flows directly from the context rot research. Respect context rot. Tight structured contexts beat maximal windows. Don't just try to max out the context window.
9:27Jeff Huber:Focus on giving the model concise, well-structured, highly relevant information, quality over quantity, again. Minimize the chance for confusion. Got it. Be selective. Don't just dump. Exactly. And finally, tip five. Invest one evening and create a small gold set. This is all about evaluation. A gold set. Like perfect examples. Yeah. A small, high-quality, domain-specific data set. Questions relevant to your data and the ideal answers. Then crucially, wire this into your CICD pipeline for automated testing. Ah, so you can automatically check if changes break your retrieval performance. You got it.
10:03Jeff Huber:It turns evaluation from this thing you maybe do once in a while into a continuous engineering practice. You can quantitatively measure improvements and, just as importantly, prevent regressions. Those are genuinely practical steps, especially that last one about continuous evaluation. Yeah. But okay, how do teams actually measure if this context engineering approach is working day to day? That brings us to generative benchmarking, right? Yes, generative benchmarking. This is Chroma's answer to the evaluation problem. What is the problem they're trying to solve there? Well, the issue with a lot of standard public benchmarks is that they often use artificially clean data.
10:42Right.
10:42Jeff Huber:And they lack domain specificity. They don't necessarily reflect the kinds of questions your users will ask about your specific documents. So performing well on a public benchmark doesn't always mean your system works well in production. Okay, so how does generative benchmarking fix that? The idea is clever. Use a powerful LLM like a GPT-4 or a CLOD-3 to actually create a high-quality domain-specific evaluation set from your own documents. How does that work? The LLM reads your docs and makes up questions. Sort of. It involves having an LLM act as a judge to filter relevant documents and then generate realistic user queries about those documents, along with the ideal answers derived from the text.
11:22Jeff Huber:So you end up with a gold set of question-answer pairs that are perfectly tailored to your data. That sounds powerful, but maybe complicated to set up. They emphasize the practicality. Even a small gold set, maybe just a few hundred examples, something you could potentially create in an afternoon or like a pizza night labeling session with the team can establish a really valuable baseline. OK, a small set is achievable. Definitely. And once you have it, you integrate it into your CICD pipeline. Now, evaluation isn't some sporadic chore. It's an automated part of your development process. You get continuous regression testing for your retrieval system.
11:57Turning evaluation into an ongoing engineering discipline. I like that. Okay, so with all this focus on context engineering and evaluation, how does Chroma itself, the company pushing these ideas, actually build its underlying technology? What's their database engine look like?
12:12Jeff Huber:Yeah, it's had an interesting journey. Chroma actually started out as a Python tool using things like DuckDB or ClickHouse under the hood for analytics. Okay, common tools. Right. Great for getting started. Good for demos. But they ran into production issues. A big one was the Python GIL, the global interpreter lock. Ah, the GIL. Limits true parallelism in Python. Exactly. It bottlenecks performance, especially with lots of concurrent requests, and can cause instability. So that led to a really pivotal decision around their version 0.4, a complete rewrite in Rust. Rust. Known for performance, memory safety, concurrency.
12:50Jeff Huber:All of the above. It got rid of the GIL bottleneck entirely. So now for their single node open source version, it uses Squelite for metadata, which is super common and reliable. And HNSW, a popular algorithm for vector indexing stored in simple binary files. Very streamlined. OK, that's the single node version. What about Chroma Distributed, the cloud offering? That's built on more modern distributed systems thinking. Two core principles really guide it. First, separation of storage and compute lets you scale them independently. Second, separation of read and write paths. So heavy data ingestion doesn't slow down your queries.
13:26Smart. And how do they handle the actual data storage?
13:29Jeff Huber:This is key. They made a strategic bet on using cloud object storage, think S3, Google Cloud Storage, as the primary persistence layer. Object storage, not traditional databases or SSDs. Right. The benefits are huge. Massive scalability, incredible durability, and it's significantly cheaper, especially for data at rest. They claim up to 10x cheaper than SSDs. It also allows their compute nodes to be stateless, meaning they can scale up and down rapidly, even scale down to zero when not in use. Scale to zero, that's a big cost saver. Are there downsides? Latency? There is a trade-off, yes. Object storage generally has a higher baseline latency than local SSDs.
14:09Jeff Huber:And there can be a cold start latency penalty when things scale up from zero. But they argue they mitigate this with aggressive caching. It's a deliberate choice favoring simplicity, scalability, and overall cost-effectiveness at potentially massive scale, which fits the AI workload pattern well. Interesting architectural choice. So, okay, you've got Chroma with its focus on Debex and this object storage architecture. How does it actually stack up against the competition? Other vector databases like QDRANT, Milvis, Weeviate, it sounds like it's not just a raw speed race. Definitely not just about speed, though performance matters.
14:42Jeff Huber:Quantitatively, they each have sweet spots. Q-Drant often wins on raw query speed and requests per second, especially with filtered searches. It's built for low latency, high throughput. Q-Drant for speed. Milvus tends to have really fast ingestion times and offers a huge number of different index types, making it suitable for very large, complex enterprise setups. Weeviate offers solid, well-rounded performance and has a rich feature set. And Chroma. Chroma shines in other areas. Its index build times are typically very fast, and it uses significantly less memory than some competitors. Query latency is competitive, and throughput is decent, maybe around 500 queries per second in benchmarks.
15:22Jeff Huber:It's really good for getting started quickly, prototyping, and situations where resource efficiency is important. But the qualitative side, the developer experience, that seems to be Chroma's main angle. Absolutely. That's where it's arguably the leader. The feedback is consistently about its simplicity, the intuitive API, that whole pip install chromab, and it just works feeling. Yeah, that's huge for developers starting new projects. It really is. It's often recommended as the first vector database to try. QDRANT and Milvus are sometimes seen as the next step up when you need more raw power or complexity, but they definitely have a steeper learning curve.
16:00And what about something like PG Vector, using PostGras school?
16:03Jeff Huber:PG Vector is pragmatic, especially if you're already heavily invested in Postgres. But you generally sacrifice performance and specialized vector features compared to dedicated databases. Chroma seems to be carving out this distinct developer-first simplicity path. A third way, almost. Between integrated convenience like PG Vector and specialized performance like QGVM Milvus. Exactly. And a key part of their strategy is owning that entire developer journey, making it seamless to go from experimenting on your laptop with the open source version to scaling up in the cloud, all using the same API.
Read the full transcript
16:37Jeff Huber:That lack of friction is a major draw. Okay, so pulling it all together, what's Chroma's grand strategy here? Where do they fit in this absolutely exploding vector database market? You mentioned it's projected to hit, what, over$10 billion soon? Yeah, it's a huge market growing incredibly fast. Chroma seems to have successfully positioned itself as the in facto starting point for a massive number of AI developers. The numbers speak for themselves. 21 ,000 GitHub stars, over 5 million monthly downloads. It's clearly resonating. So the ease of use is paying off. Hugely. And then Chroma Cloud provides that essential frictionless path to scale when those projects get serious.
17:15Jeff Huber:Their target audience is crystal clear. It's the application developer. And the value prop is built around simplicity, that great developer experience, and easy scaling. How are they differentiating beyond just ease of use? Well, there's the thought leadership aspect, actively defining the conversation with things like the context rot research and generative benchmarking. That builds credibility and shapes how people think about the problem space they solve. True, they're setting the agenda in some ways. They also have some neat developer-centric features like native reject search support, which is apparently very useful for code search applications, and index forking, letting you experiment with different versions of your data, almost like Git branches.
17:55Index forking, interesting. Like A-B testing your retrieval data.
17:59Jeff Huber:Kind of, yeah. Makes experimentation much faster. And architecturally, that serverless object storage native design we talked about gives them a strong argument on total cost of ownership, especially at scale. And the founder's vision. Jeff Huber's background seems relevant, too. Definitely. His experience at Google working on massive planet-scale systems like ads and maps brings a culture of building robust, reliable infrastructure. There's an emphasis on intentional design and engineering excellence. He talks a lot about wanting to do work one loves with people one loves and for customers one loves, and sees his role partly as a curator of taste for how AI systems should be built.
18:39Okay, quite a comprehensive picture.
18:41Jeff Huber:So to really wrap this up, I think the main takeaway from our deep dive today is pretty clear. We've got to move beyond thinking about RAG in that simplistic, naive way. Right. It's not dead, but the old way of thinking about it might be. Exactly. Context engineering, supported by research like the context-wrought findings, isn't just a nice-to-have. It's becoming absolutely crucial for building AI systems that are reliable enough for production. It's less about chasing infinitely large context windows. And more about intelligent curation, getting the right stuff to the model. Precisely. Smart curation, precise delivery, making sure the LM has the best possible, most relevant context to work with.
19:19This whole deep dive really feels like it maps out a shift, doesn't it? From just feeding data to LLMs to actively managing their attention span, their focus.
19:30Jeff Huber:That's a great way to put it, managing AI attention. So as these AI systems inevitably get more complex, more autonomous, here's maybe a final thought for you, the listener, to chew on. How much of this detailed context engineering work will eventually be handled by the AI models themselves? And if that happens, what kind of new human roles might emerge around guiding, tuning, and overseeing that AI attention and memory? Something to ponder as you navigate what's next in AI.
From the publisher
Today, instead of discussing a research paper, we review the interview by Jeff Huber, CEO of Chroma, discussing the evolution of AI search and retrieval systems. He champions "context engineering" over the widely used "RAG" (Retrieval-Augmented Generation) concept, arguing that the latter is vague and often misunderstood. Huber highlights the importance of efficiently curating information for Large Language Models (LLMs) to combat "context rot," where model performance degrades with increasing input length. The conversation also touches upon Chroma's distributed database, strategies for code indexing and retrieval, and the significance of generative benchmarking for evaluating AI systems.




