The timelessness of vector databases | Pinecone’s Ram Sriharsha

14 Oct 2025 · 48 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Dev Interrupted - The Timelessness of Vector Databases

Episode Overview In this episode of Dev Interrupted, hosts Andrew Ziegler and Ben Lloyd Pearson sit down with Ram Sriharsha, the CTO of Pinecone. The discussion revolves around the significance of vector databases in the evolving landscape of AI technology, particularly as it pertains to search functionalities.

Key Themes

  • The Importance of Vector Databases
  • Vector databases are critical in AI because they facilitate search, which is fundamental to AI operations.
  • Ram emphasizes that as AI applications grow, the need for effective search mechanisms becomes even more pressing.
  • Retrieval-Augmented Generation (RAG)
  • Engineering leaders are encouraged to start with simple RAG applications.
  • Establishing evaluation frameworks is crucial to manage issues like hallucinations in language models.
  • Skills for the AI-Powered Future
  • Curiosity and the ability to operate as a generalist are essential skills for engineers in an AI-dominated environment.
  • Engineers should be encouraged to ask the right questions and not take AI outputs at face value.
  • Building the AI Stack
  • Ram suggests starting with simple implementations and gradually building more complex systems, incorporating evaluations and guardrails along the way.
  • Using AI as a tool for testing can enhance productivity significantly.

Detailed Insights

The Role of Vector Databases

  • Core Functionality: At the heart of all AI operations lies the search capability. Vector databases play a pivotal role in enabling and enhancing this feature.
  • Externalizing Knowledge: Vector databases allow organizations to externalize knowledge, which can be audited and secured, addressing critical needs in AI deployment.

Practical Recommendations for Engineering Leaders

  • Starting Simple: Leaders are advised to begin with straightforward implementations of AI technology, focusing on functionalities like RAG.
  • Evaluation Frameworks: Implementing strong evaluation frameworks helps mitigate risks associated with AI outputs, particularly regarding hallucinations and misinformation.

Essential Skills for Engineers

  • Curiosity and Generalism: Engineers should embrace a wide-ranging knowledge base and remain curious about the tools and technologies at their disposal.
  • Collaborative Problem-Solving: The conversation highlights the need for engineers to work collaboratively, leveraging AI tools to enhance their problem-solving capabilities without losing sight of clarity of intent.

AI in Practice

  • Testing and Verification: AI can serve as a "good junior engineer" for testing purposes, enabling teams to generate tests and verify code efficiently.
  • Simplicity in Engineering: Ram advocates for simplicity in design and engineering practices to support effective collaboration and problem-solving.

Industry News Highlights

  • AI Adoption at Deloitte: A notable expansion of AI deployment within Deloitte, introducing AI tools to 470,000 technologists, emphasizing the importance of integrating AI in large organizations.
  • Coinbase's AI Code Generation: An insight into how Coinbase utilizes AI to generate code, primarily for testing, demonstrating practical applications of AI in the tech industry.
  • Internet Archive Milestone: The Internet Archive's celebration of one trillion web pages preserved, underscoring the importance of digital preservation.

Conclusion The episode emphasizes the critical role of vector databases in the AI landscape, advocating for a foundational understanding of search and retrieval mechanisms. Engineering leaders are encouraged to start small, build robust evaluation frameworks, and foster a culture of curiosity among their teams to thrive in the rapidly evolving technology landscape.

For further insights and resources, listeners are directed to Pinecone's website and Ram Sriharsha's LinkedIn profile.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:05Welcome to Dev Interrupted. I'm your host, Andrew Ziegler. And I'm your host, Ben Lloyd Pearson.

0:39from Sean Godecki talking about work politics for engineers. Another thing that caught our attention in the news is celebrating one trillion webpages archived on archive.org, an amazing achievement, but also a month of celebration for all things archiving and recollecting digital history. And a Coinbase engineer shared more details about their AI use case for code generation, which taps into a story we covered a few weeks ago. And last but certainly not least, the largest Anthropic enterprise AI rollout to date just happened. A lot of really interesting things happening in the news. What do you want to dive into first, Ben?

1:18Well, first of all, I'm super excited for our guest because I just started using Pinecone this week to build some agents for here or workflows that Devinterrupted. Super fun tool to use. But let's talk about the other tool, the other AI tool that I really love as well, and that's Anthropic. I want to hear what's going on here. Yeah, so the largest anthropoclod rollout to date just happened for Deloitte. And Deloitte's making clod available to 470 ,000 technologists across its global network and organization. Talk about a huge adoption wave of folks picking up this AI tool for what's really the first time inside of this large traditional enterprise.

1:59eyes. This marks a really powerful transition of AI being an everyday tool for knowledge workers. And this was covered in an article we read from Hithy Cyborg, which has actually quite a number of really interesting articles from this week. So we're going to make sure we embed his really great roundup in this post so you can go check it out. Also featured in there is a Google bug hunting bot called Codemender who hunts bugs in your code base overnight and OpenAI being valued more than SpaceX. So be sure to go check out those articles. They're really great. But the Anthropic and Deloitte announcement is an expanded alliance that brings to AI closer to knowledge workers in a way never done before for its type of industry.

2:39And as part of this, Deloitte understands the challenges that it's going to be heading face first into. So they've actually established a Claude Center of Excellence. And this is a center of trained specialists that are going to create implementation frameworks for how the Deloitte technologists are going to adopt and roll out AI. So this represents a really fascinating learning opportunity for the rest of the industry to pay attention to what's going to matter to Deloitte at their Claude Center of Excellence, because this collaboration is representing, you know, the work of both sides, both Anthropic and Deloitte.

3:15So you're talking about folks in the seat using AI to have massive impact for a huge company, global company at scale, and a technology firm making that model possible. So there's so many fascinating things happening in the story. Ben, what do you think? Yeah, you're right. I mean, there's a lot of reasons that I love this story. And part of it is, you know, our producer, Adam, he was showing us the Claude's new ad campaign that they're out with right now. It has slogans like keep thinking, like the AI for problem solvers. And I feel like those two slogans like pretty much describe like the role of consultants, you know, like Claude really is for consultants, in my opinion.

3:54But there's like two kind of things that sort of stand out about this to me. The first is that, you know, Claude is the most verbose GPT platform that I use, like by far. You're absolutely right. And in the hands of a consultant, I feel like that could be very dangerous. Like, I'm sure there's a lot of people at Deloitte that are going to be like introduced to this tool for the first time, probably going to be making all the same mistakes that we all did when we first started with GPTs. One piece of advice, if there's anyone from Deloitte listening, don't ask Claude to write your emails. It just always writes a novel, and no matter how many times you ask it to make it shorter, it will never be short enough.

4:29But more importantly, Claude is this incredibly powerful artifact generator. And when it's connected to useful data sources, as I'm sure consultants at Deloitte have, it can produce really useful components in seconds. You can build UI and UX components. You can build data visualizations. You can summarize things, any type of thing in any way that you would want to. Create automation scripts based on a database schema. Like, you know, there's so many different things that you can do with Claude because of how it's designed to generate artifacts, particularly if you're like connecting it to an MCP.

5:04So, you know, I think smart people with this power in their hands can do incredible things in an extremely short timeline these days. So like, first of all, kudos to Deloitte. Like, you know, I think you made the right decision picking Claude for your organization. Like it is probably the best tool out there for an organization like Deloitte. Yeah, totally agree. Yeah. But I like that you brought up the Claude Center of Excellence. Like in the article, it says it's staffed with experts tasked with developing internal implementation frameworks, sharing best practices and providing ongoing technical support.

5:40So like, in other words, they have AI consultants for their AI consultants now, which is kind of crazy. But Andrew, I'm wondering, like, you know, how long is this like we have our own internal AI consultant expertise? Like, how long is that going to be a marketable trait? Like, is that on a limited lifespan? Do you think? No, I don't think that's limited at all. I think we're actually just, that's why I'm paying such close attention to the Claude Center of Excellence because I think this is a model that you're going to see replicated across a lot of huge organizations and standardized organizations, also highly regulated ones, as they begin these same size rollouts.

6:17Also, I think you're going to see organizations and groups like this inside of governments relatively soon as they start operationalizing around how they and their knowledge workers are using AI. And I think that a lot of this is going to come down to a continued practice of learning and development. And also probably to, you know, to everyone's misfortune, probably a very robust certification ecosystem from all model providers and the folks that are using them at scale. I'm looking at like some of the largest tech companies in our industry use this model. You know, if you have a platform and tooling that hundreds of thousands or millions of engineers are using at scale every day and on top of all of your documentation and resources, you also probably have certifications and training and levels of excellence that you can accredit to the people that are using your tools.

7:09These are the AWSs and sales forces of the world where, you know, you go on LinkedIn and you see folks in these ecosystems with, you know, 50x cert trained like in their LinkedIn title. I think we're actually going to see a lot of this around AI skilling and specialization, especially around applying and using it at scale. So I think that this is just the beginning of a trend, in my own opinion. Yeah, that's actually a really great point. You know, I'm just waiting for like that man behind the curtain moment, though. Like eventually we're going to find out one of these organizations that are setting up their own like AI center of excellence.

7:43It's just going to be a bunch of AI agents behind it and secretly, you know, and it will eventually be. Hey, AI is great at telling you how to use AI. I mean, once again, it's like if you're not using AI to write your AI prompts, then you're already missing out on a flywheel. So it makes sense to me. Yeah, I actually think I just discovered for the first time this week documentation that was written for AI, not for me. So let's move on to our next story then. And this one's kind of a continuation of a discussion we started last week. They kind of got a lot of attention on LinkedIn about politics and software engineering.

8:15And I love this because it's almost the exact opposite of the discussion we had last week. And it's practical advice from someone on how they influence politics at their tech company. So what do we have here, Andrew? Like you just aptly summed up, well, last week we covered Matthias Lima's article about don't avoid workplace politics, giving advice to engineers about why they need to participate in the influence, you know, network at their jobs and both in a soft and hard way by being contributors, by also showing up for conversations where they need to weigh in and give their best advice. And this stirred up a lot of conversation on the Hacker News, got a lot of attention and in fact influenced in a response article from Sean Godecki.

8:55He's a writer whose articles we've covered here on DI a few times, And he was prompted to write this after reading that article. And his opinion, you know, is that software engineers, they're often fatalistic about company politics. They think they can't influence or change anything, or it's pointless. And they're just kind of along for the ride. And that kind of nihilism and cynicism can drive a lot of negative work around work politics. So he really breaks down how stakeholders, organization leaders view and work with software engineers to achieve their goals, to help better educate an engineering reader about how the mind of their manager counterpart might even work.

9:38And he also gives some practical advice on how to actually navigate that power network with your own passion projects. And some of the advice he gives is, I would say, starts to actually veer into a little more of a cynicism itself around how engineers work and typically view their work inside of an org. Something think he called out, for example, is that if you want to make a high profile project successful, then one way you could do this is to with like a pet project, making it available for what he labeled as a political campaign within your company, even calling it as much as the flavor of the week for what engineers want to get done.

10:17So if you want to migrate the billing code or rip out a build pipeline, or rewrite like a Python service into Golang, or, you know, he gives these examples of what he labels pet projects for engineers, then he gives advice on how to dress that up for the flavor of the month that your executives are worried about. That way that you can go after what they're trying to achieve while still achieving your pet project. And, you know, he's not wrong that org politics, they shape those technical outcomes, but his framing does make it sound like the only rational move is to become like a savvy opportunist rather than like a principal planner with your innovation.

10:55And ultimately, you know, those politics are inescapable. But his framing does assume that engineers are kind of tools within the political game and not agents in their own rights. I would really challenge engineers in response to this article, you know, to understand how their work aligns with the impact of their organization. But, you know, Ben, what did you think after reading this article? Yeah, I'll push back a little bit on the notion that he's saying you have to be a savvy opportunist as much as I myself like to view myself as an opportunist. Because he does point out, there's one section in this article that the easiest way to actively work to improve your political position is just to help high profile projects be successful.

11:36And yeah, sure, sometimes it's hard to like, be on be chosen to be on a high profile project. And, you know, it can be difficult to accomplish great things. So it's, you know, the work to do good work at a high profile project is difficult. But from a political perspective, simply just putting your best foot forward and always doing what you can to push projects you are involved with in a positive direction, that is the simplest way to make positive political pressure on an organization. But beyond that, yeah, he does focus a lot on opportunism. I think it is really worth pointing out, particularly, I think this is a challenge a lot of engineers struggle with, is that you might have the best idea, but it may not be the right time within your organization for that idea.

12:22And really what you should be focused on is like saving that type of project or any project that you care about for when there's a top-down mandate that aligns with it that you can kind of just include it with. So yeah, you brought up some of the great examples that were in that. And, you know, one thing to keep in mind is the stuff always kind of comes in waves. Sometimes it's cyclical. You know, the longer you're with an organization or the longer you've been around organizations of different types, you can sort of learn how to sense these things. And then, you know, at the end of the day, when it comes to opportunism, luck is kind of a controversial topic as well.

12:54You know, sometimes it just strikes completely at random. Sometimes you have to work to actively make it. But the way I've always kind of approached that luck, you know, to generate an opportunity is, you know, really luck is only like 10 % of the equation. Like you could be the luckiest person in the world, but if you're not ready to take advantage of that luck when it happens, then none of it matters anyway. So, you know, put in the 80 % and when you can, and just be ready for that moment where you were luck may strike or opportunity may strike that you can, you know, apply additional political pressure beyond just doing the right thing.

13:28I really do like that he's breaking it down for engineers just to help them understand like this doesn't, you know, because of the discussions I got on LinkedIn around this, there was a lot of mixed feelings. Like some people felt very strongly with that previous article on, you know, getting involved with politics and other people just felt like there's no way they could do it. And I don't think that's really true. I think everyone can do something to positively influence their organization politically. Yeah. I think that goes back to the fatalism viewpoint that Sean rightly called out in the beginning that that is, you know, a lot of the response.

13:59So I'm interested to see what the third article in the saga will be, you know, maybe next week that we cover because it does seem to be a heated and divided topic. I also like that you called out luck. I think that's really self-aware, but also ties back to an article we covered recently from Charity Majors where she acknowledged her own luck as an engineering manager and how it got her to the places and the opportunities she had. And she called out the same thing that you did. You know, luck is part of the equation. It's also about showing up and building every day. Yeah. Yeah. So let's move on to this next story, which I've been salivating to cover.

14:30related to Coinbase. And how much of the code is written by AI? So what do we have here, Andrew? Yeah, so there was an article we covered, or rather a tweet we covered from the Coinbase CEO, Brian Armstrong, earlier this year, bragging about how 40 % of the code at Coinbase was AI generated. And so we recently actually scooped this podcast interview from Syntax Podcast. And this is with Kyle Sessmet of Coinbase talking about the types of code at Coinbase that are generated by AI. And what he talked about is how this code is mostly tests, particularly in TypeScript around various parts of the infra there at Coinbase.

15:09And so, you know, Ben, I think that's right nail on the head of your prediction here. So I'm going to hand this over to you because I know you probably have a lot to say about this one. Yeah, for our dedicated listeners, if you remember back, I made some guesses about what Brian Armstrong actually meant when he said 40 % of their code was written by AI. And I love this interview because this is exactly what in that episode, what I said we need more of. Less CEOs with big numbers that they just tweet about and more engineers on the ground telling us exactly what they're doing with AI. So it turns out I'm right about a few key things about how Coinbase is using AI.

15:46First is that most of the code that AI is writing, as you mentioned, is for testing. It's not for product innovation. It's handling toilsome test work that they actually weren't doing before. To that point, Coinbase also has extra requirements that AI written code must be fully tested. So this creates this feedback loop where if you're using AI to write code, you're required to write more code because you're required to do additional testing that your non-AI peers don't have to do. And then the third thing that I was right about is they don't set goals around the number of lines of code that are changed by AI.

16:22They only suggest that as a best practice, just around the concept of, you know, keeping your code simple and your pull requests small. I really, I listened to a bunch of sections in this, in this podcast is really nice to just get like on the ground reporting firsthand of how Coinbase is actually using AI within their organization. So yeah. You absolutely called this one too. When we, when we covered this originally, this was exactly what you said that you thought it would be. And then we are learning that that is what it is. And And not even necessarily in a bad light, you know, we're just calling it what it is.

16:55It's code that you rightly have labeled as it didn't exist, didn't need to exist beforehand. It's great that they're doing that now. They probably should have always been doing that. But the fact of the matter is, is that it creates all of this extra framework around actually having to do it. So that just as a clear reminder to anybody that's adopting AI coding tools within their code base is to understand how it does impact all of that downstream code. and to understand what it really means when you have more code. Is that good or bad? Was it code that needed to exist yesterday? It's important to be aware of that stuff.

17:28Yeah, and it's one of those social media effects. If you see someone on Instagram that seems to be having a better life than you, you're just seeing one version of them. You're not necessarily seeing all the realities that are behind those pictures and videos. The same goes with executives tweeting about how much code is written by AI at their company. Don't judge yourself based on that. But if we got it all wrong, and maybe you know more than us at Coinbase, we'd love to talk with you about it. So please, reach out to us. Absolutely. All right, I think we have one more story, and it's about the Internet Archive.

18:01What's going on here, Andrew? Yeah, so this month, the Internet Archive, you know, that very diligently archives all copies of all pages on the Internet that it can get its hands on, is celebrating one trillion web pages preserved and available for access this month on the Wayback Machine. And, you know, the Internet Archive has been around since 1996. It's a crucial part of the Internet and how we can access documentation, things from before. It's basically the archaeological record of the digital age. And so celebrating One Tree and WebPages Archive is a huge celebration. And they have a month-long lineup of events, both virtual and in person, as well as lots of opportunities to get involved in the preservation efforts.

18:44We just wanted to call that out here. So if that's of any interest to you, please definitely go to the Internet Archive and check out their event lineup, which we'll also include below. You can continue to pay it forward so the next generation of Internet users can continue to benefit from the same kind of preservation effort. Yeah. Are there any other websites out there from the 90s that we still use today? I have to wonder. From the 90s? Yeah. I mean, are you still not logging in to feed your Neopets? I think that says a lot about you, Ben, than anything. Back when Flash was everywhere. Yeah, I mean, it's extraordinary to see this organization last so long.

19:23And it's a cause that I deeply believe in. I think information preservation is an often overlooked thing. It's often why I have difficulties with the current state of Western intellectual property law, because a lot of information just disappears from the world long before there's the monetary incentive to preserve it. So yeah, support the Internet Archive. They do a wonderful service for the internet for all of us, and we should help them back. So, Andrew, who's our guest this week? This week, we're sitting down with Rams Raharsha, the CTO of Pinecone. So stick around for this discussion. Are you tired of slow, inconsistent code reviews?

20:02Meet LinearBee AI, your new AI-powered workflow assistant built to supercharge your team's review process. With LinearBee AI, you'll get automatic PR descriptions, AI-generated review suggestions, and instant insights that help your developers fix issues before a human even looks at the code. No more waiting, no more guesswork. Just faster, smarter, and higher quality reviews powered by AI. Visit LinearBee.io to bring AI into your code review process today. Hello and welcome back to Dev Interrupted. I'm your host, Andrew Ziegler, and today we're sitting down with Ram Srirharsha, the CTO at Pinecone.

20:43Prior to joining Pinecone, he was the vice president of engineering at Splunk, where he began as a senior principal scientist. And before that, he was a product and engineering lead at Databricks for their unified analytics platform for genomics. Ram, welcome to The Dev Interrupted. We're so excited to have you here. Thanks, Andrew. It's great to be here. We're here at ELC, the Engineering Leadership Conference here in San Francisco, and you gave a talk today about vector databases. Why are vector databases still important? That's a great question. In fact, my talk goes into exactly this topic.

21:15The reason I wanted to give the talk in the first place was, first of all, there's a conference where people are talking about AI. Agents are a hot thing. It's an emerging technology that people are starting to get familiar with. Through all this, people may not realize that at the core of AI is search. And at the core of how AI does search is vectors. So I wanted to give a talk to frame the problem from that perspective and to kind of maybe educate the audience a little bit that if you take this to the logical conclusion and think about how do you now scale search, you have to have vector search and vector databases at the core of those.

21:49And it's a new way of motivating vector databases for a new audience. I think people who are familiar with recommendation systems, people who have done information retrieval, they kind of get it. But people who are working on agents, people who are trying to build generative AI applications and so on, may not necessarily realize at first glance that, oh, the thing that underpins AI, emerging AI, is actually search. So I just wanted to give a new framing to that problem. And so would you say that it's like framing a building block that folks should be using and getting them educated about how they're used to build higher level software?

22:22Yes. You know, vector databases are really important in the AI space. And I really want to understand more from your perspective of why their prevailing importance in search will probably never go away. and understanding the role that the vector database plays within the engineering world for people that are building agents or working with AI tools. And most importantly, what are the ways that engineering leaders can think about those kinds of tools, those primitives, those catalysts for change and use them within your own organization? What kind of perspectives do you hear from leaders? Yeah, that's a great question.

23:01So first of all, the reason why vector databases don't go away is that If you're building agents, you need knowledge. And you want to be able to externalize this knowledge. You want to be able to audit it. You want to be able to redact it. You want to make sure it's secure. And the moment you start thinking about that, you need vector databases, and it becomes a foundational component. And now it becomes the question of how do you take this foundational vector database, foundational language models, and put them together to build a stack for your agent. What I see people often do is they start simple.

23:36There are good recipes out there today called RAG, Retrieval Augmented Generation. There's a lot of knowledge by ourselves. We have lots of blogs. We have lots of templates of how to build RAG systems. We even have what we call an assistant, which allows you to just drop unstructured data and ask questions. So we tend to try to lower the barrier to entry to this kind of building agents as much as possible. So that's where I would start. If I'm an engineering leader who's trying to empower a team to build agentic workflows, I would start with something like an assistant, start consuming it, start building knowledgeable agents, and then slowly go deeper into the stack as you kind of gain more expertise and also start putting more guardrains.

24:25So when you work within that iterative space, it's about just starting and then measuring what works and what doesn't. What else matters as an engineering leader? How can they equip their engineers to really be thinking about unlocking more value? It's a great question. I think to me, starting simple matters. Putting evaluation frameworks matter because LLMs hallucinate. No retrieval is perfect. You kind of need to know how to tweak or tune your retrieval steps. You kind of need to know what to do about hallucination with language models. You need to know when your agent is actually returning the right relevant results.

Read the full transcript

25:00and so on. So there needs to be like an evaluation framework that's probably something you think about first. And then kind of go deeper and deeper and start building more sophisticated agentic workflows from there. So I generally would suggest start simple, but start putting guardrails, start putting an evaluation framework. And what would you say is like a prevailing skill that leaders should be inspiring or otherwise teaching their engineers to pick up in today's engineering age? I think curiosity is super important. Being a generalist is very important, right? Because now you have language models and AI tools that you can use for kind of specific tasks and they're going to just keep getting better.

25:44But you need to be a really good generalist now, more so than you had to be before, I think. So I think breadth of knowledge and curiosity is what I would really kind of look for in engineers and kind of guide them towards that. and to be able to ask the right questions out of your language models and prompt them correctly and to know when they're hallucinating, know when they're giving you the wrong results and not just accept what they're doing. Having that clarity of intent, knowing what you're asking for when you go into the situation, which is way harder, by the way, than it looks on the surface.

26:18Everyone thinks that they're coming in with the best possible way of explaining something, but until they actually try to convince another engineer to build it or an LLM to help them do it in an afternoon, You know, that's where the rubber really meets the road. And that clarity is really important, I think. When teams are ultimately coming together to solve problems right now, AI is kind of like a force multiplier. And like for all the good and all the bad, right? If you don't have process, it's going to make things really messy. If you have a lot of process, things can get really stratiated and rigid, right?

26:48I'm curious within Pinecone, how do y 'all approach that curiosity? off. And how do you create the space for your engineers to reinvent on top of what you're building? It's a great question. I mean, first of all, we are very heavy users of cursor. And, you know, I myself use cloud, for example. So we are very big users of AI. We also do kind of take advantage of Gemini and tools like that for code reviews. So there's a lot of first level kind of reviewing and the first level kind of cut at problems that AI is very good at doing today. You can think of it as a good junior engineer. Like AI tools are getting pretty good at and we leverage it for that a lot.

27:35What we typically tend to do though is to take it from there and refine it. And that's a process that, you know, has purely left to software engineers today. AI cannot really help you as much there. You need to know exactly what to take out of what it gives you, how to iterate on it, and when do you think it's good enough. So I think that's where we do find really good use of it, though, is in testing. So, you know, particularly Pinecone is building vector databases, is building kind of AI infrastructure and so on. And if you think about traditional databases, they are very complex systems that need to be accurate across a spectrum of things, right?

28:17Yeah. So correctness is a very challenging problem in database literature. And typically, when you have to build correctness testing and things like this, it would take months and teams of engineers to build it. Today with AI, I think it acts as a huge force multiplier in the testing and the first testing process of databases. So we use it extensively for that. And it's also a good problem because that's a problem that's in the wheelhouse of AI today, which is well-specified, it's code, it has structure to it. You are asking it to write tests and verification stuff, and it's really good at that.

28:53And the risk to having a test that is maybe not correct is less than the risk to having a production code that's not correct. So we tend to leverage it a bit more, I would say, with a bit more flexibility in testing. Obviously, with production code, we have far more guardrails and we really use it in a very... Yeah, we're just talking about the space we were playing and creating and trying things out because we're all trying to figure these things out right now. There's a lot to learn. Yeah. What I would do is, you know, it's a great way to start by having AI write your tests or in particular what's called property testing, which is it's like writing a test that captures a scenario of testing.

29:36It was complex to write before, but it's one of these things that you can easily teach an LLM how to do and it can do it for you. So these things can unlock, for example, things that used to take days, take minutes now. And that's like a really interesting unlock too, because you get this world where you're able to like nail down the determinism while you're working within a probabilistic space. Like you're engineering in this like latent area where nothing is yet defined until you or the LLM talk about it or write it down. And so by working with a really rigid structure of like having that conversation first, everyone's talking like the spec based coding, right?

30:13Getting that really good markdown document, creating the artifacts of the of that coding process that we did have a few years ago. So now whenever I start a new project, I'm getting, it's often something I spun up in an afternoon, maybe cursors working on it while I'm eating lunch, sometimes generally. And so you come back and then you push it and you want others to try it out and play with it and see what it is. And I find myself constantly checking in those artifacts, those documents. So that spec.md, that agent.md that we wrote together, and that becomes just as important as the code that came with it.

30:45And so it's like a new way of working with engineering knowledge work and sharing it with others. And I want to double click on something you said about the rise of the generalist. I think that's a really powerful message. I think we've talked about a bit before on Dev Interrupted. We had Lee Robinson here from Vercel, and he talked about not so much the end of specialization, but rather the rise of these hybridized engineers that just don't accept roadblocks. when they have a problem, it's a team in the past where, oh, you know, in order to figure this out, I need to go talk to the UX team. I need to set up a meeting.

31:20I need to go do whatever. And you get these fragmented bursts of coding and working. But now with AI, you can encapsulate all that in an afternoon. The things that used to help make an engineer get stuck where they'd have to swivel and ask that backend engineer, they can now ask the LLM. So, yeah. So how do you see that force multiplier affecting folks even within your own organization? Yeah, yeah. I think the reason I think generalists a very, you know, I think being a generalist is particularly a force multiplier right now is because you can think of AI as kind of having like a small team right next to you, right?

31:55So you could be a team of size one, but with language models and with agents and with AI, you suddenly have a team of fairly substantial size. And now if you are a generalist, you can fully utilize them. As a specialist, you're limited in some sense from utilizing that. Now, that doesn't mean specialists go away. I think specialists are extremely important. But being a generalist has like a special power now, which in some sense wasn't there before. I think there's a big unlock from AI from that perspective. So we do encourage our, first of all, encourage our engineers to tinker. All of our engineers tinker.

32:29We have them work on very different aspects of systems. So we generally, in fact, in the database team, for example, we don't really have special, I mean, we have some specialists, but we have more generalists than specialists. And generalists tend to work on very different parts of the stack. And in some sense, AI is helping them really unlock and kind of bring them all together in a way that's just harder to scale before. So I find that as a pretty big multiplier for the team itself. There's also this new experience happening for engineers. This is something I lived very vividly recently. I sat down for a vibe coding hackathon with block open source block goose platform.

33:11And I had a partner. So he and I were working together in this hackathon and we're using exclusively goose to basically command and write our project for together. And it was exceptionally hard to bring our worlds together in a way that I had never encountered before in a hackathon because I've done lots of them. And when you're all working with people, you're just shouting things out. It's like line chefs in a kitchen, right? You just keep everyone updated. But in this world, me and my partner were both commanding our own siloed teams of AI engineers that were all doing all sorts of stuff. And then we had to find a way to marry those worlds together.

33:47And it was exceptionally hard. I'm wondering, how do you think about that problem for engineers that work on separate problems and then bringing their AI collaboration together? I think that's where being a really good generalist helps, right? because you kind of need to know how to combine the two kind of AI teams, so to speak, which is more natural if you're kind of used to putting these parts together and being a really good generalist and so on. The other thing we do is another thing that I think is helpful is if you have kind of dabbled in those areas to begin with, right? It's much harder to put things together if you have actually not had some amount of expertise in those areas.

34:23So it's like building an agent end-to-end, but I haven't really built a UI before. So I'm far more trusting of that language model and I'm far less capable of, you know, maybe course correcting it. Yeah, you lean on it. Exactly. I lean on it far more than I would if I was a generalist. Right. So I myself, I'm not a true generalist in that sense because user interfaces are some area I've not worked in. Yeah, exactly. You're using it to fill in the gaps, like we talked about a moment ago, of just not getting stuck. Exactly. Not accepting that you're going to get stuck. Exactly. Keep going. Right.

34:56and I'm wondering about your own background from other engineering roles that you've had and you have a science background, right? You approach this from a very methodological way. How does that influence your role in your leadership as the CET? It's a good question. I find that, I think in general, kind of a physics or a scientific background, I actually find it really helpful because it makes things a little bit simpler for me in how I think about problems, which is instead of, you know, trying to remember a lot of facts and sort of trying to kind of keep a lot of things in my head, I try to build a model of the system.

35:35I try to think about the model of this thing that I'm building. And then I use that model to kind of infer about, oh, if this is the mental model of the system, then this is how it should be. And if it's not behaving like that, then I can go back to the basics to figure out what's happening. And that's something you're trained as a physicist, right? Like in physics, for example, you cannot really understand the universe. So you really boil things down to very simple components that you can piece together and build models out of. And that stayed with me forever. And I find that's extremely powerful and useful, even in computer science, even in AI, and every other field that I work on.

36:09So I try to instill that in engineers when I'm mentoring them, is always form a mental model of the system. And systems have a good mental model if they are simple. right the moment your system becomes complex your metal model is no longer metal models so simplicity is in some sense forced on you if you try to want to keep a simple mental model so this is what i try to inculcate in engineers yeah and it's it's worked i think uh when it works it's very satisfying because you can keep a lot of things in your head while keeping just a simple mental model yeah simplicity in engineering is is always best if it's easy for people to understand And especially because we're experiencing right now this democratization of engineering and people getting access to building that before never really had ability to do so.

37:01And in that world, you know, you see a lot of stuff and you read a lot of stories about people repeating or learning really fundamental engineering mistakes that are really that we have used to ship and create really great software in the last 10 years, right? And then it's like all of a sudden AI is on the scene. Everyone threw all that away. And then they're making all those same mistakes again. And then we're relearning how those primitives build and go together with AI. And in that world, do you see a place where, you know, non-engineers are agentically coding, going something like Supabase or V0 and spinning something up?

37:39And when they do that, is there a companion path for things like vector databases to be so simplistic? Like what you just said, that even they could pick, even grandma could vibe to that app that has a RAG database in it. Yeah, that's a great question. So I think that's one of the things we're working on is to make vector databases so simple, so easily integrated with the kind of general tooling and framework that people are using for agents and whatever they're building. that they don't have to think twice about it, which is you can start small, scale up, and never have to think about the database at all.

38:17That is one way to do it. The other approach of, you know, I just pick the tools that I have and I kind of just build something and go with it. I think that's fine as well, which is obviously we would like people to kind of use Pinecone and kind of scale up easily. I think the other approach also what happens is that people kind of do that. you get to some traction and you realize that, oh, now I have to scale and I'm meeting scale challenges. And then at that point, we want to be seamless to be able to switch, right? So there is the other aspect of it as well that we care a lot about, which is, yes, we would like to catch everybody very early, but we also want to be seamless kind of to help transition people to scale.

38:58And when you take a system like that to scale, a really important part of that that we dug into recently on this podcast is evals. And you mentioned that earlier about making sure you're evaluating things as you iterate, because when you're working in this like latent space of AI, right, it can be really easy to miss those small details that are going to compound and compound and compound. And that's what evals are really great at catching. We sat down with Andrew McNamara. He's their head of applied ML at Shopify. And he talked about their eval framework that they created for their agents and about how it was so fundamentally important in order to take any kind of agentic system to production.

39:36You know, what are your own perspectives on emails and how engineers should be using? That's a good question. So in the case of Pinecone, we are like one level below in the stack, right? So when we think about evals, we think about evals for the Vectors database itself. That's usually called recall. It's things like that. And we care about that. We measure that. We have benchmarking systems and tools that are open source that allow you to measure it against any database, including us. But then one level higher, we build these things called Assistant, which is kind of a layer on top of the database that acts as a knowledge system, which is you throw all of your unstructured corpus and you can search over it.

40:14Now, at this layer, your search quality has a different meaning. The question is, am I returning the relevant results? Is my search returning the relevant results? Secondly, if I'm taking this information and feeding it into a language model, does the language model find it useful? And we have evals that allow you to measure that as well. So for us, that's the level at which we operate because we are one level below where a Shopify or any customer is going to be using their evals. Now, of course, if you are somebody who's building Shopify, say, on top of using PlainCore and using SS10 and so on.

40:51And they have their own evals of how those agents are performing off of that data. But just as important is the eval one layer down of the actual data. Because that tells you whether the search system you're using or the vector database you're using is doing its job. Right. Because you have to evaluate where the problems are. Right. And that's like the great thing about using evals is that you're catching blind spots and things that you can't really write tests for that you're only going to be able to see by iterating and then seeing what it throws back at you. It's like a game of catch, right?

41:20100%. And what machine learning and generative AI has done over the last several years has made us from having to be machine learning experts to becoming evil experts. See what I'm saying? I think that's where we need to be spending more time. Another thing I want to ask about is understanding intent of your users and what they're trying to do, and then connecting the dots with what's available for the LLM. I think we're in this scenario where it's really easy to overwhelm your LLM or your tools or your systems with just way too much stuff. Like we've heard about like in the MCP world, like if you hook up 100 MCP servers to your LLM, it's going to get confused.

41:59It's going to use the wrong tools. There's 10 tools called update. How's it going to know what to pick? And it fills the context window, all of these problems, right? And so it's about being precise and about understanding the intent of the user and meeting them in their place every single time. And this is something we talk about a lot on Dev Interrupted. I'm curious how you think about that intent problem. And is it different for people that aren't of an engineering background versus are? How do you think about it? Yeah, I mean, I think when I think about intent, for example, I go back to kind of why we built a system in the first place, which is this is about a year back, slightly more than a year back, when people started building Gen AI applications.

42:42and they started using vector databases and so on. And then there were all these questions. Should I use just a vector database? Should I use some other document database together? Should I use a graph database? How do I chunk? What do I do? Right, what do I chunk? How do I summarize it? Like there's a whole bunch of little steps and not just turn on a server. And when we looked at that, in fact, we looked at a stack that had hundreds of components, right? But then we were starting to ask the question, what do people really care about? The problem that they were trying to solve at that point is hallucination, which is I have a corpus and I'm asking a question, I'm getting an answer.

43:20I just need to be able to attribute this correctly and to know that I'm not giving wrong answers. So we focused on that problem. We're like, okay, how does vector databases help us solve this problem? If you need to argument it with some other things, what does that argument stuff look like? And I talked about some of that in the talk today, which is vector databases are evolving. from what they used to be a few years back to the search system that's needed to be able to answer this question. And that's the way we started looking at this problem. We did just the bare minimum required at that point to be able to solve the problem that really mattered without having to deal with all of the complexity and start kind of going down the rabbit hole of all the things you need to tweak and tune.

44:03So that's an example of kind of how we approached this. If you were to ask or rather talk to an engineering leader right now about what you think is the most important thing for them to focus on in the next few months as we enter into our new year. The news cycle and the development cycle around AI is so fast-paced that any kind of prediction you'd have about a month for now would be wrong by the end of the week, just at the speed at which things come out. But what are some big rocks that you just keep coming back to as you evolve in the AI space and you see your organization to it that You think everyone should be sailing?

44:41Yeah, I would say that I still find a lot of organizations hesitating to build Gen AI applications, for example, because of the kind of probabilistic nature of this stuff, because there's still wrinkles, there's still unknowns and so on. What I generally advise people is that start building, start kind of putting these things together. Even in its current form, it is massively useful. And once you start, you will get familiar and you get comfortable with this kind of thing that gives you probabilistic answers. Things that can potentially hallucinate sometimes, but it's still very useful in applications that can kind of tolerate it.

45:24And then start putting guardrails together and just start kind of working away from there. I think it's time to start using it and time to start kind of getting familiar with it. rather than waiting in the sidelines and hoping that this thing gets perfect, right? Yeah, exactly. And so there's a lot of opportunities in the immediate future, even just right now, for folks to be digging into and taking advantage of. And as you talk with engineering leaders, is there a recurring question that they constantly, you find them constantly asking you about what they should be doing or thinking? I think it's mainly about how do you deal with the fact that these things don't give you answers that are, you know.

46:07How do you deal with the fact that they're not deterministic? Exactly. Which is their, by the way, their superpower, which is why, by the way, we turn to them. It is. Because it unlocks that ability to build things that before were uncommercially available to do. Exactly. So to me, that's what it often keeps coming back to. And some applications, it just doesn't work. In some applications, it's fine. As long as you kind of prompt things correctly and you kind of learn how to deal with it and so on, it does work. So this is what keeps coming back. And I don't know if I have a great answer here except to point out that that's a very common.

46:39It's a common problem that people encounter. And it really kind of like highlights again the importance of like context engineering and understanding what goes into the system because it's garbage in, garbage out, you know, and you want to understand and iterate on the inputs to your AI systems versus the outputs because the outputs, you're on the wrong end of the assembly line. You need to be at the beginning and understanding the source and the root and using search and your database is a core component of that. Yes, absolutely. And, you know, this has been a really enlightening conversation about how you think about vector databases, but also the opportunities you see for engineering leaders.

47:13And you gave a talk here today about why vector databases are this important primitive building layer for AI and why folks should be using a building with it. Where can folks go to learn more about what's coming next for you and Pinecone and Get Plugged In? Yeah, pinecone.io. So our website has a lot of information. We have very good tutorials on pretty much all parts of the stack, whether it is assistants and higher level parts of building agentic workflows or deep dives into a vector databases. And feel free to reach out to me if any of the topics I talked about today is interesting. Yeah, no, we'll include all of that in the show notes so that folks can go check it out and we'll put it in our newsletter as well.

47:52I'm a daily practitioner of all these tools, so I'm going to go in deeper on some of the stuff that you resource for us and see how I can use it. So thank you so much for sitting down with us today. It's been really great to have you here on Devon's Run. Thank you so much. I had a lot of fun. Thank you. Yeah, thank you.

From the publisher

With massive context windows and new agent frameworks, do vector databases still matter? Ram Sriharsha, CTO at Pinecone, joins the conversation to make the definitive case that they're more critical than ever. He explains that at the core of all AI is search, and externalizing this function is non-negotiable for security, auditability, and control.

Ram offers a clear starting path for engineering leaders: begin with simple Retrieval-Augmented Generation (RAG) applications, but immediately implement a robust evaluation framework to manage hallucinations and ensure quality. He shares his perspective on the skills that matter most now, arguing that curiosity and the rise of the generalist engineer are critical in an AI-powered world. This episode is a guide to building the AI stack from the ground up, from using AI as a "good junior engineer" for testing to cultivating the engineering mindset of tomorrow.

Bring AI into your code review process with LinearB

Follow the hosts:

Follow today's guest(s):

Referenced in today's show:

Support the show:

Offers:

More from Dev Interrupted

All 208 episodes
The timelessness of vector databasesDev Interrupted · 48 min
Listen in VO