In short
Podcast Summary: Production-Grade AI Systems with Fred Roma
Podcast Overview Title: Software Engineering Daily Description: Technical interviews about software topics. Episode Title: Production-Grade AI Systems with Fred Roma Episode Description: This episode discusses the complexities of deploying AI applications in production, touching on the evolving AI development ecosystem, challenges with data layers, and MongoDB’s advancements in AI-ready database capabilities.
---
Episode Highlights
Introduction
- Guest: Fred Roma, SVP of Product and Engineering at MongoDB.
- Host: Kevin Ball, Vice President of Engineering at Mento.
- Focus on building effective AI applications, integrating AI features, and the state of AI application development.
Key Concepts Discussed
- The Complexity of AI Production
- Despite advancements in AI prototyping, moving AI applications into production remains complex.
- Modern AI stacks require:
- LLMs (Large Language Models)
- Embeddings
- Vector search
- Observability tools
- New caching layers
- MongoDB's Approach to AI
- MongoDB is transitioning from a document database to a full AI-ready database platform.
- Recent acquisition of Voyage AI to improve embedding models and ranking capabilities.
- Data Layer as a Foundation
- The data layer acts as both a foundation and bottleneck for AI applications.
- Importance of simplifying the data stack, ensuring accuracy, and providing the ability to evolve quickly.
Key Challenges and Solutions
A. Evolvability
- The need for quick adaptation of schemas and components due to the fast-paced AI landscape.
- The significance of LLM-derived schemas, allowing flexibility in data handling.
B. Data Management for AI
- The data stack must be simple, accurate, and allow rapid evolution to accommodate changing AI models and requirements.
- The importance of combining keyword and semantic search for optimal results.
Importance of Search and Vector Search
- AI applications require more than just LLMs; effective search mechanisms are critical.
- Vector Search: Necessary for semantic understanding, improving relevant information retrieval.
- The integration of search and ranking models enhances the accuracy and relevancy of AI responses.
Best Practices for AI Application Development
- Quality of data should be prioritized: clean, relevant, and well-structured data is essential for performance.
- Understanding context in embeddings is vital to avoid outdated or irrelevant information.
- Cost-effectiveness and accuracy are key metrics when designing AI models and applications.
---
Key Takeaways
- Flexibility is Crucial: As the AI landscape shifts rapidly, leveraging a data platform that allows for easy changes and integration is vital.
- Quality and Cost Management: The quality of information retrieval can greatly impact the user experience, and cost management is essential for sustainable AI operations.
- Development Collaboration: Close collaboration between product and engineering teams enhances the development process and aligns objectives.
Conclusion The episode underscores the dynamic nature of AI application development, emphasizing the need for robust data management, a flexible infrastructure, and an understanding of evolving technologies to create successful AI systems.
---
Further Information For more insights, you may listen to the full episode on [Software Engineering Daily](https://softwareengineeringdaily.com/2026/01/27/production-grade-ai-systems-with-fred-roma/).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOIntroduction to Fred Roma
0:45 to 2:04
Fred Roma shares his journey in software engineering and current role at MongoDB.
“with integrated capabilities for operational data, search, real-time analytics, and AI-powered data retrieval.”
The State of AI Application Development
2:04 to 3:08
Discuss the complexities involved in building production-grade AI applications.
“Yes, I'm really excited to dig in with you on this.”
Core Components of AI Data Stack
3:08 to 4:32
Explore the essential components needed for an effective AI data stack.
“We're talking about data management and AI.”
Evolvability of AI Applications
4:32 to 6:46
Discuss the importance of adaptability and evolution in AI schemas.
“First, what we think about the data stack, I think there are three things we may want to, you'll tell me which one you want to dive on first.”
MongoDB and AI Integration
6:46 to 8:14
Analyze how MongoDB can scale with AI applications and what advantages it offers.
“You will have LLMs that are really good at some specific industry or use cases and others are really good at other things.”
Implementing Effective Search in AI
8:14 to 10:00
Understand the critical role of search and vector search in AI applications.
“I've been a customer of MongoDB far before even considering joining the company.”
Components of Vector Search
10:00 to 12:28
Break down the elements that contribute to successful vector search in AI.
“Depending on what you are building, you may need some stream processing.”
Multimodal Models and Their Importance
14:01 to 15:10
Learn about the significance of multimodal models in AI and context awareness.
“You have images and text and videos combined, and you will have PDF, and you probably don't want as a developer to break that things down and call different models.”
Simplifying Embedding Processes
15:10 to 18:19
Explore how a multimodal model can streamline the process of handling various data types.
“People sometimes don't realize, but these embeddings, they can be even bigger than the data that they represent.”
Combining Search Techniques for Better Results
18:19 to 21:06
Understand how combining keyword and semantic search enhances application accuracy.
“When you are running an embedding model on a text, you should cut.”
Show all 25 chapters
MongoDB's Aggregation Pipeline Functionality
21:06 to 23:49
Learn about MongoDB's aggregation pipeline and its advantages for developers.
“Sorry, that's a very important point that you are touching.”
Security Considerations in AI Applications
23:49 to 27:32
Discuss the importance of security when integrating AI and LLMs into applications.
“And just so that we're clear, those things can be defined on the fly.”
Managing Data Security in AI Use Cases
27:32 to 28:05
Explore how organizations are navigating data security in AI implementations.
“What is important is that there is nowhere the blending, if you want, of how the LLM is trained and your private information.”
AI Integration and Data Management Choices
28:05 to 29:15
Explore how businesses are deciding on AI and data hosting options between cloud and on-premises solutions.
“But I would say that with this AI-specific application, we see more and more, I would say at least even more discussions about this LLM integration, the one you spoke about before.”
Challenges of Speed and Team Organization in AI Development
29:16 to 31:09
Discuss the intricacies of team organization and the speed of development in AI environments.
“So I want MongoDB Enterprise Advance, and that's how I manage it.”
Navigating Prototyping and Production in AI Systems
31:10 to 33:11
Learn about the balance between rapid prototyping and the challenges of moving to production in AI systems.
“But from a product, I'll go back to if you want to go fast, but you have to move your data around and to...”
The Role of Expertise in AI and Engineering
33:12 to 35:38
Understand the importance of expertise and the limitations of AI tools in engineering practices.
“And if you put some restrictions in place, some things are then, if I was hearing you correctly, you can then actually take out to production.”
Organizational Changes Due to AI Integration
35:39 to 37:59
Examine how AI tools are reshaping team structures and improving collaboration in organizations.
“We use AI more and more, but we don't VibeCode as a core MongoDB database for sure.”
Future Challenges in AI Application Development
38:00 to 42:00
Identify the future challenges and considerations for building effective AI applications.
“No, I think there is definitely something and have been in a lot of conversations kind of talking about this convergence between engineering and product.”
Addressing Hallucinations in AI Models
42:00 to 42:52
Learn how minimizing hallucinations in AI enhances user experience and reduces costs.
“that means that you are totally reducing hallucinations.”
Best Practices for AI Data Preparation
42:52 to 44:08
Understand the importance of data quality and best practices for AI model preparation.
“In terms of accuracy then, I mean, I think you've talked some about, we've talked about what it takes in terms of searching, in terms of re-ranking and what you surface.”
Balancing Speed and Accuracy in AI
44:08 to 45:14
Explore how to balance speed and accuracy in AI applications based on user needs.
“Or maybe if you're an e-commerce company, you will want to go super fast and make sure that you have 10 or 12 good results and it doesn't have to be the only one.”
Embedding Models and Their Trade-offs
45:14 to 46:26
Learn about the trade-offs in choosing embedding models for AI applications.
“So I think a lot of our listeners are now familiar with in some of the AI coding tools, like you can dial up your budget for, oh, I want you to think longer.”
Re-ranking and Intermediate Results in AI
46:26 to 49:18
Discover the role of re-ranking and how to manage intermediate results in AI workflows.
“That's pretty fast and you optimize for that.”
The Future of Data Platforms in AI
49:18 to 51:18
Discuss the rapid changes in data ecosystems and the importance of flexibility in AI platforms.
“Or you do that because for a specific application or workload, accuracy is so important.”
Transcript
Automatic transcript. May contain errors.0:00Engineering teams around the world are building AI-focused applications or integrating AI features into existing products. The AI development ecosystem is maturing, which is accelerating how quickly these applications can be prototyped. However, taking AI applications to production remains a notoriously complex process. Modern AI stacks demand LLMs, embeddings, vector search, observability, new caching layers, and a constant adaptation as the landscape shifts week to week. Increasingly, the data layer has become both the foundation and the bottleneck to AI app productionization. MongoDB has been expanding beyond its core document database into a full AI-ready database platform with integrated capabilities for operational data, search, real-time analytics, and AI-powered data retrieval.
0:54The company also recently acquired Voyage AI to provide accurate and cost-effective embedding models and re-rankers to its users. Fred Roma is a veteran engineer and is currently the SVP of product and engineering at MongoDB. He joins the show with Kevin Ball to talk about the state of AI application development, the role of vector search and re-ranking, schema evolution and the LLM era, the voyage AI acquisition, how data platforms must evolve to keep up with AI's breakneck pace and more. Kevin Ball, or KBall, is the Vice President of Engineering at Mento and an independent coach for engineers and engineering leaders.
1:37He co-founded and served as CTO for two companies, founded the San Diego JavaScript Meetup, and organizes the AI in Action discussion group through Latent Space. Check out the show notes to follow KBall on Twitter or LinkedIn, or visit his website, kball.llc.
2:04Fred, welcome to the show. Hey, Kevin. Great to be here. Thanks for having me. Yes, I'm really excited to dig in with you on this. So let's maybe start with a quick background of you and how you got to where you are today at Mongo, and then maybe we can use that as a way in. Yeah, absolutely. So I started as a software developer. It was a long time ago. In France, in Paris, I evolved. I became a manager and worked also on the product side. And I worked in different continents, small startups, big companies. The biggest one was AWS, a small startup. It's probably a startup you haven't heard of, a French startup.
2:43But yeah, no, no. And I would say most of my career, I've been in the cloud, even before we call that cloud, like more application service provider, but plugging some servers and giving access to these servers to customers. and mostly around security things like payment, identity encryption, and more lately at MongoDB on data management and AI. So I'm having some fun here. Yeah. So let's dive straight in there. We're talking about data management and AI. I think everyone's trying to figure out what is it to build an effective AI application right now with these newest tools. What's your take on the core pieces you need?
3:20Yeah. I mean, it's never been easier to vibe code an application that's for sure and that's that's exciting i mean i don't know about you like i just love it like it's going so fast and that's so thrilling but so what we see though is that when you want to build something that is really for production be it large you know customer consumer application or large enterprise application it's still very hard and i think we could go to a couple of reasons the first one when you build i think there are other blockers later on when you want to launch it in production. But the first one when you built is just how complex the stack is now.
3:56Like you need an LLM, you need a vector search, you need a different kind of AI models, you need an AI framework, you need new caching mechanism. And that can be a little bit, when you leave the vibe coding piece and you really want to build that in a professional manner, that can be a bit scary for a developer. Yeah, absolutely. Well, let's talk about the angle that I think I understand you all are addressing, which is the data side of it. Because to your point, right, I have my database, I might have to do some embedding, I need a vector search situation, I need all those different pieces. So how do you think about the data stack for AI?
4:32What needs to be in there? First, what we think about the data stack, I think there are three things we may want to, you'll tell me which one you want to dive on first. But like the first one, we think it should be simple and simplified as much as possible. The second one is you need to make sure it's really accurate and cost effective because information retrieval can be pretty expensive if you don't take care of it. And I think the third part is you want to make sure it can evolve quickly. Things are going so fast. You never know which is the best LLM model. You just unplug for two days and that's a new one.
5:04You never know which tool you need to use. You never know how fast your data will grow. So I think it would be the three things. I mean, we want the stack to be simple. We want the accuracy of the information retrieval to be really, really good. And we want to make sure that if you change your mind or if the ecosystem is changing or things like that, you can touch your application and make it evolve easily. That would be the three key points. Let's actually talk about that last piece, which is the evolvability piece, because this is one of the things I see a ton as we're doing AI agents internally, the things that I'm working on.
5:35Like schemas are not nearly as durable as they once were. Yeah. No, absolutely. I mean, they're not durable either because you change your mind or you pivot your application as a developer, and that's totally fine. They're not durable as well because the ecosystem is changing so fast. When you want to connect to a new partner or new integration points or even using a new LLM framework, maybe you will have a new field to account for. Or maybe you want to adopt this new observability for LLM to do real good evaluation. And so, yeah, that is changing super fast. Absolutely. I think there's even an aspect that I'm starting to see, which is LLM-derived schemas, right?
6:17Instead of having one mega schema that the developers come to, let the LLM choose it. No, absolutely. Yeah, it's definitely a trend right now. Plus, when you say let the LLM choose it, the reality is that you may have to play with several LLMs. So you may have to be able to handle several schemas either, again, because the LLM that was best this month is different than the one that was best last month. Or because for some tasks, you still need some specialization. You still have LLMs. And I think that would be more and more of a trend. You will have LLMs that are really good at some specific industry or use cases and others are really good at other things.
6:52But no, absolutely. And we touch, I mean, it's been MongoDB value proposition forever. Like go with a document model. You don't have to stress about any change you will want to do. So, I mean, we love it, just to be clear. We love that the full AI world is speaking JSON, and we love that the full AI world is coming with all of these changes, because then we can say, yeah, I mean, that was already something important before to be able to evolve quickly, but that's even more the case in the AI world. So let's talk a little bit then about the second step that you talked about in terms of, okay, we want to take this to production, we want to be able to scale, we want to be able to deal with all of these different things.
7:30because I think Mongo has long been, at least in the front-end world where I used to live a lot, like database of choice for rapid prototyping. And then at some point, sometimes people would say, oh, well, now we've got to switch over. We've arrived where we're going. But I think at this point, you all scale all the way up, yeah? Oh, yeah. Now we have like, I think it's 75 % of the Fortune 500. I've been at MongoDB for one year and a half, and it started far before me. But I can tell you, we always speak about big for security and durability, availability, performance, being able to speak to large enterprises.
8:03And we see that more and more. We see more and more of this very large workload. You can totally manage asset transaction. You can totally manage this very strong transactional financial thing. So I think it's a, I mean, MongoDB is the open source. I've been a customer of MongoDB far before even considering joining the company. And so I also had initially, oh, yes, I remember we are going so fast, But that's a very serious database for scale, for sure. So let's then talk about how to effectively use it in an AI application. Because I think this is a space where, okay, the model layer is pretty well understood if changing very, very rapidly, right?
8:41You have these LLMs, you can throw text at them, they love JSON, all these different pieces. And then you have this kind of, okay, there's these cool applications being developed. But all those middle pieces, and as you highlighted, the complex stack that's going into that is very much in flux. Absolutely. Yeah, absolutely. So even past what we discussed before, this document model and JSON that is, again, very well adapted, optimized for AI, you still need a search and vector search. Because, I mean, you don't have any serious AI applications that will really give you a lot of value if you just plug an LLM.
9:15The value of these AI applications is, okay, how do I connect an LLM, what the LLM knows, with what my company knows? And if you want to do that, you need a search usually and a vector search, by the way. Maybe we can come back to that as well. But if you are looking for, I don't know, you are building an optimized e-commerce website and you are looking for red shoes, you probably want to see these burgundy sneakers as well. So you need search, you need vector search. You need AI models, embedding models and re-ranking models, because this is how you can really have very good information retrieval and really make sure that if you are building your right application, if you are your semantic search, your agentic system, to make sure that the right information is provided, grounded results are provided to your user.
10:00Depending on what you are building, you may need some stream processing. If you have some events, you may need many of these things. And so as a customer, you can choose to stitch together different solutions, a database, VictorSearch, a search, a re-ranker, an embedding model, et cetera. I don't think you should do that. I don't think you should tell your best friends to do that. But you could stitch these things together and connect multiple times your identity providers and create this pipeline for the data to transfer. What we are really betting on and what we see is bringing value to customers.
10:30It's just making it super simple. You have a database. We don't transform the world with all the stack you need for AI, but on the data layer, you have a database that is becoming a data platform. And so you can do, yes, storing of your information, querying of your information, but also information retrieval and data in motion and all of these AI model optimizations to make sure that you will get good results. Let's go a little bit deeper in what that takes on search, because I love that you brought up search. I feel like to me, one of the things that I'm seeing is everybody's thinking of these as chat products.
11:02They're not chat products. They're search products at their core. They're about surfacing the correct information. And LLMs help you interpret that and put it into context for someone. So it's not as simple as throw it at an open source embedding model or OpenAI's embedding model, run naive queries and just go. There's a lot of pieces that go into effective search. Absolutely. Absolutely. We can maybe take an example, a concrete example, a simple one. Let's say you are a bank and you are building an application for your customer support. And your bank's customers, they will maybe be able to ask questions about their account and their credit card, etc.
11:40And so obviously, you need an LLM because the full interaction, the conversation is an LLM interaction. So you need that for sure. But if you only have that and you don't have access to the internal documents of your bank, probably a bit weak as an experience. So you need search. You need search. I would say you need both. You need probably vector search that are really the semantic search. Because if I'm asking you, okay, how much money do I have on my account? Maybe you will look at it from not maybe how much money, but like what is the sold of my account? What's very different words in very different languages.
12:13languages. So you will need search, vector search. You will probably need as well, maybe some kind of reaction to that event. Because if you see that something has been paid in the last two minutes, and we are discussing for five, you want to take that into account. So you need all of these pieces to come together, the search, the vector search, the stream processing. So let's talk about the pieces that go into vector search. Because I think, you know, once again, for folks who are coming into this to the beginning, they think, okay, what do I need to know? Right? Like, I mean, I started the first interactions I had with search were back in the days of solar and okay, you've seen right in these things.
12:50Absolutely. No vector search. It was all keyword based, but you had some synonyms and things like that. But to me as an end user, I was like, okay, throw it all at solar, do a couple configurations and I'm good. It's golden. It serves things up. I think there's, there's more nuance than that. If you just throw all of your documents at OpenAI as embedding model and assume it's going to work. It's not going to just work. So what are the different pieces that go into that? Yes, there are a couple of different pieces. I will start with a basic. Different models will have different accuracy. So the quality of the results, and it's always a trade-off between how fast you want the results and how good you want these results to be.
13:30And the best embedding models, and that's exactly why we acquired Voyager a bit than six months ago now. I mean, they were doing and they're still doing the best embedding models out there. They are more accurate than this OPDI that you mentioned. So you actually want the best result with a reasonable latency because you have a user probably maybe behind a chatbot or maybe behind an agentic system application waiting for results. So accuracy is a big one. Multimodal is a big one as well. Most embedding models will be able to do a good job comparing text or maybe pictures. But the real world is messy.
14:04You have images and text and videos combined, and you will have PDF, and you probably don't want as a developer to break that things down and call different models. So that's also one, like what are the formats that you can support? I would mention two more things. So yeah, accuracy, multimodal. The ability to understand context is also very important. Like for instance, if I'm asking you, okay, let's say you have this support application, chatbot. but maybe now you are like a networking company. And I say, okay, how do I configure this router? If you find somewhere in your corpus of data an exact sentence or blurb or text that explains how to configure that router, most embedding models will be super happy.
14:45Oh, I found the information. The best embedding model will be able to say, well, wait a minute, I found a line, a sentence that looks exactly what the user is asking for, but is it part of the recent documentation? Or is it part of a ticket that is five years old and the setting is totally outdated. So we'd say context is very important. And last, and that's last in my list, but probably sometimes first in the customer's mind is like the cost of it. People sometimes don't realize, but these embeddings, they can be even bigger than the data that they represent. If you have a great model that is able to do all of that and you don't have many of them, good accuracy and multimodal and good context, but the embeddings are so big to be able to achieve these results, It would just be an awful ROI for your application.
15:33So you also want a model that is able to do that with really short embeddings, as that would be cheaper to store and cheaper to query. Yeah, so that's interesting. I kind of want to explore a few different of those aspects, and maybe we can explore them from the context of Voyage, because that is the recent acquisition. And that was kind of, as I understand it, the secret sauce there. So first, starting with this multimodal piece, right? Because if I think about an application I'm building, I am probably doing a fair amount of pre-processing right I'm like oh this is an image I've got to translate this image to text and now I've got to send this text over to my embedding model and now I've got to take that and do my vector search and it's like this whole pipeline of things absolutely but it sounds like what you're saying is they've got a multimodal model that you can just throw whatever at it and it's going to translate it yeah no you nailed it that that's exactly what customers were doing and when we speak to customers and they are describing that they say oh I have my, let's go back to the PDF example, and say, oh, I have all of this pipeline, and I will extract my pictures and my text from my PDF, and then I will run them through different kind of embedding models, and I will try to reconcile the results afterwards.
16:42Yeah, with a Voyager multimodal model, you just throw your PDF in the embedding models, and you will have an embedding. By the way, the result will be even better than when you are doing all the pipeline, because when you are breaking your document down, you will lose some context, you will lose some interaction. Where was this picture exactly? Was it above this text or below this text kind of thing? So the result would be better, but the big benefit is also as a developer, you can go super fast. Just use your document and your embedding model and you're done. Yeah, no, the simplicity definitely appeals to me.
17:11So let's explore the context piece a little bit more because what I kind of hear you describing is your embedding right now is taking into account not just the text or not just in the case of the PDF, like text and images, but it sounded like things like metadata, updated timestamps, like all these different things. Like what does this API call look like? What do I pass to it? Yeah. Or it's even more than that. So what you described, by the way, is the fact that you also want to take into account the metadata and some other information in addition to purely semantic search. It's super important, top of mind for customers.
17:45And that's why they're using, by the way, vector search and search combined. And that's why it's so important for search and vector search to be where your operational data is. Like if you're using a separate vector search, you will have to, oh, what is all this data or metadata? You have to also synchronize with this. No, when your database is there, you can do all this stuff that you're mentioning. The context is a bit different. It's part of the embedding model. It's just the way the model is trained. And then at inference time, instead of just isolating a chunk of text from your document, that's how you, again, maybe I should step back.
18:19When you are running an embedding model on a text, you should cut. You don't give it like two pages of document. You are chunking the document in small sentences or blurbs or things like that. We call that chunks. Then they will have an embedding for each of these chunks, but they don't really know what is a chunk before and what's the chunk after. And what the voyage model is doing, and the voyage context model, that's a specific one that we released, is that it will parse the full document and that it will preserve some context in addition to the specific chunk. So yes, you will know, for instance, yeah, this sentence really explains how you can configure this router, this security configuration maybe, but that's how you know that you are part of an old ticket because you also see maybe three or four chunks above that it looks like a support ticket and that it is six years old.
19:04And that probably you shouldn't give it too much importance. Got it. So conceptually, if I'm going to just try to like map this out, if I were building this with a much more naive model, it would look something like, okay, I have a summary of the whole document with maybe some additional things. And then I have each chunk and then those two things are getting put together kind of linked in each set. Interesting. Yeah, that's a great way to look at it, yeah. Fascinating. Well, and you alluded to another piece of this, which is combining search, and that gets us into this topic of re-ranking and all of that.
19:35So can you maybe like lay out what that looks like just broadly for context for folks who haven't built these applications before and then what the Voyage take is on it? Yes, it really depends about what you are trying to achieve in your application. But what we see most of the time, when you want to have the best results, and I'll give you one example, but when you want to have the best results, like the most accurate results, like your users asking for something through a chatbot, through an agent, etc., and you want to give the best results, combining a keyword search, like looking for a document with the exact keywords that were part of the query, plus also the semantic search, meaning that documents that may have very different keywords, but are speaking about the same topic, pick, combining the two is how you get the best results.
20:21Like the example would be, let's say you say, oh, I want to, I'm interested in a, I'll go back to my red shoes. I'm looking for Nike red shoes. Maybe red shoes is totally okay to go with burgundy sneakers because that's almost the same. Now, if as a user, you made the effort to mention Nike, it may be very important to you and you want really to make sure that you are looking at the keyword Nike. So let's look at the keyword Nike, but let's only look at the semantic meaning of All of these red shoes and maybe burgundy sneakers are perfectly fine. So this is just a very simple example. And it doesn't fit to all use cases, but most of them, the best accuracy would be to combine both of them.
21:00And that's why having search and vector search and the database at the same place, it's a big deal. You remove a lot of round trips to do that. Absolutely. You're right. Sorry, that's a very important point that you are touching. And I'm not saying that you couldn't do it with multiple pieces. You could, but then you have to run your search on a keyword search. you have to run your semantic search, and you have to build your own algorithm to see how you are ranking those results. Yeah, so that's exactly a lot of hurdles that you are removing. So implementation-wise, if I then was using your API, am I able to specify, run this against these two searches, re-rank in this way?
21:38What are the knobs that I have available as a developer? Yeah, so we didn't reinvent the wheel, by the way, with a nicer search or vector search, you use a MongoDB aggregation pipeline, the one you are using with your database to just query data. But we created new operators, score fusion, rank fusion. So I may not go into these details because these are just slightly different ways to merge the results. But you have full control about... So first, you have one operator where you can combine keyword search and vector search, but you have full control about how you want to do that. I mean, many customers are just happy with a basic kind of way to combine, but some say, okay, I want to have different weight.
22:17I want to over-index a little bit, maybe in this keyword and a bit less on this one. So it's up to you, but you just use the MongoDB aggregation pipeline. Maybe it's worth actually stepping back and talking about that aggregation pipeline a little bit, because that is a capability that doesn't exist in all databases. Yeah, absolutely. Yeah, yeah. No, no, that's really the... I mean, if you have used MongoDB, that's something that customers usually really like. It really gives you the ability to make several operations on your database, one after the other, and the result of an operation can be used as an input for the next operation.
22:52So it can be really, really powerful. And that's the case for this, again, search. You can use combined operators, as I was describing, but you can also decide to do some search and then you will do some re-ranking somewhere else and you will do some other stuff. So it's really a pipeline that you can implement to play with your data. Yeah. So just to sort of echo back, right? Like if you're using another database, you might have an external pipeline tool where you're defining a series of stages with dependencies and data transfer and kind of moving that around. And with the aggregations pipeline, you can do that all inside of the database.
23:24You can do that all inside the database because this is natively integrated. Like search, vector search are natively integrated with the database. You don't have to move the data around. You don't have to use different aggregation pipeline indeed, but you don't have to use different CLI. You can really have all of that as a single experience. It's not like an extension where you have to be careful about how that will be supported or how you can plug things together. That's the same tool is giving you access to all of this. And just so that we're clear, those things can be defined on the fly. If someone, for example, was giving the LLM tools to the kingdom, it could write its own aggregation pipeline and run it?
24:00The developer. Oh, so yeah. So we did release an MCP server. That's really trendy these days. And it actually is really effective as well. So customers are, I mean, the adoption is pretty nice. If you use this MCP server as an example, I'm mentioning MCP server because you mentioned LLM. And usually that's how developers more and more are going with interaction with our database. you can absolutely for sure create clusters and do operation, but also configure your aggregation pipelines. Let's talk a little bit about security. You mentioned security as an area that you had dealt with. And I think this question of what are LLMs allowed to see, what are they not allowed to see, all of this is definitely top of mind for those of us building applications here.
24:42What are the primitives that are, I'm guessing, baked into the database to allow you to build secure AI applications on top of Mongo? So I would say I want to step back on the overarching principle because you touched on LLM and what an LLM can see. That's exactly because you don't want probably to train an LLM on your private data. That's the pattern, the architecture that is winning out there is no. I will not. I mean, there are exceptions, but I will not fine-tune or post-train my LLM on my data. I will use an LLM, really good at, again, everything they are great at. But for my use case, when a user wants something, then I will connect what this LLM knows with what my company or my application knows.
Read the full transcript
25:27That's not about giving anything to the LLM. It's about your application being able, with very good information retrieval, to say, okay, my user is asking me again for what is my credit limits on my credit card. A lot of things can be handled by the LLM in terms of how to answer to a user asking question. But if I really want this information, I also have to go and find my private information of my bank about what is a real limit. And I will provide this information to my user combined, but the LLM will never see this private information. So maybe just to clarify the overall pattern before going into the security details, that's super important.
26:03I think that is very important in terms of what sets of data are, I think the term I sometimes use is moderated through the LLM. So even if the LLM, the LLM may be the UX delivery, but do I load this data, pass it to my model and have it presented? Or do I sidestep around the LLM because this has to be right and I can't count on it not hallucinating something about it? Yeah, yeah, absolutely. And there are different patterns there. Most of the time, what customers will do is, again, if a user, an agent, want to do something that does require information retrieval, you will first look for the information that will be relevant.
26:42I'll stick with the same example. What is the internal document that explains what are the limits on credit cards? And then we'll insert that in the prompt of the LLM. So, yes, that will go through the LLM from the prompt and the answer of the LLM, but it's not stored anywhere on the LLM side. It's not been part of the training of the LLM. Sure, yeah. And if you decide, and then that's up to you as a customer, where do you want this LLM to be hosted and where do you want these queries and this token to be served? And so you can decide depending on your security sensitivities. Many customers say, okay, I'm totally fine with having my LLM on AWS or OpenAI or Azure, et cetera.
27:20Or some customer will say, no, I want to do that, but I want some specific security agreement with these providers to make sure that my data is never shared. And some customers say, you know what, I want to host my own LLM. You can do that if you want. What is important is that there is nowhere the blending, if you want, of how the LLM is trained and your private information. Yeah. So coming back, though, to building applications with these pieces, I think I'm curious to understand how you're seeing people defining these lines or barriers. Is it changing at all in terms of how you're managing security at the data layer?
27:58I mean, security has always been a very, I mean, customers trust us with our data as a data platform. So it's always been top of mind anyhow. But I would say that with this AI-specific application, we see more and more, I would say at least even more discussions about this LLM integration, the one you spoke about before. And one of the big value of MongoDB is you can run it anywhere. And when you say run anywhere, sometimes people tell us, oh, you are cloud agnostic. You can run it on AWS and GCP and Azure and etc. Yes, that's true. We can also run it on premise. You can also, we have an enterprise and you can do that in your own data center.
28:37And we see customers that are doing that. What is very interesting, I think, in this AI world, we see customers that are saying, you know what, for this use case, I'm totally fine if it's in the cloud. But for this use case, I really want to make sure that my data is never in any cloud provider and never touched by any LLM provider either. And so I will run it on-prem. And you start to have this kind of... And again, I think it's so early and things may evolve, et cetera. But the fact that they have a choice and they can decide what to rely on as a cloud provider, as their own data center, I think is pretty interesting.
29:11And I do believe it's more and more a discussion topic. That's fascinating. and when you're seeing that are you seeing them doing this within the context of the same application so you're having to kind of have security boundaries and federation and however that's working or like yeah how does that work i'm not sure i can extract one pattern one one answer to that you have customers that will say okay i will really have all of my data will be on-prem and then some application will be in the cloud and some won't i see some customers saying you know what, I want my data to be in my data center. So I want MongoDB Enterprise Advance, and that's how I manage it.
29:47But I'm still okay to do a call to an LLM outside of my boundaries, because they believe that they can control, and they are right in many cases, they can control the prompt. So it's okay to send some information as soon as it's something that your application control, but at least nobody's seeing the raw data. So I really see different kind of patterns there. And I wouldn't be able to tell which one wins. I think what is very important though, is that overall, I think there is really this intent to remain as flexible as possible. I think, say, you know, I'm back to the previous point. Maybe this LLM is good for me right now, but in six months, that will be another one.
30:22Maybe this cloud hosting is good for me right now, but then I will want to go on-prem for any regulation or specific concern later on. I think there is really a willingness to remain as flexible as possible and to have options on the table, I would say. That gets to kind of a somewhat different topic, but you know i think one of the things that these tools are doing is they're changing the speed at which people are operating they're changing absolutely kind of how fast we're moving they're changing how adaptable we need to be how are you both internally and with your customers like rethinking the way that we organize the teams doing this work oh yeah yes that's that's uh okay i mean there is a product angle i thought we were initially going to the product angle and and uh we just go there super uh super quickly but i want to go to the team organization i think that's super important as well.
31:10But from a product, I'll go back to if you want to go fast, but you have to move your data around and to... I mean, any of us who have been software developers or architects at one point knows that when you have to optimize for latency, performance, iterate quickly with network layers and data to transfer, it's just a nightmare. So we say that that's one of the key arguments, not just for MongoDB. I think that's a big value of MongoDB, But overall, I think the database vector search and search are the same place. And you don't have to do this ETL as this kind of difficult network configuration.
31:44I think that's a big one for your question as well about the speed of development and iterating. Now, if your question is going more to the team organization, I think that's a very good question as well. Even so, I mean, two things. I think what's happening right now is that it can be also very thrilling when you are a product manager or business person say, oh, I can build it myself. I can beat fast. And personally, there's something I love about that. What I love about that is that instead of trying to debate about maybe some text and some PowerPoint, you can really show. I love that. I think the risk is to believe that it's easy.
32:19Yes, it's easy to show something. That's the old engineering and product manager. Sometimes they have to align on that. I think it's great if you use that well. If you use that as a way to say, oh, and just put it to production. Well, for some stuff, yes. But how will you scale? And because that's all of these tools, right? They're making the code, writing code super fast. Well, reviewing code is not faster. And making your security assessment of this code is not faster. And defining the right architecture is not faster yet. So I think it can be awesome if you know what it is. It can be a bit dangerous if you believe that, oh, then that's it.
32:54And I can just push it. So I don't know if I maybe pivoted a little bit to your question, but that just made me think of this point. No, I think it is key. and it kind of gets to a couple of different questions, some of them related to the product piece and some not. So like one of these things is going from zero to a prototype I can show is now very fast. Absolutely. And if you put some restrictions in place, some things are then, if I was hearing you correctly, you can then actually take out to production. Some things actually do, are able to ship relatively easily, but others are not. So I guess the place I would start to ask you is like, how do you draw those lines?
33:34And are there ways in which having, for example, a database that can handle all these different pieces together makes it easier to bridge that gap? Yeah. And I would obviously be biased because of working at MongoDB right now database. And we are not vibe coding a database. I mean, there's too much at stake. You're not? No, we are not. We are not vibe coding. No, that doesn't mean that we are not leveraging AI for many things. for prototypes, 100%, for this product and engineering alignment. By the way, I didn't answer your question about team organization, but maybe later can put a pin on that.
34:07But also how you handle tickets. Are you on board? I think people, even very senior engineers, when you ask them, say, oh, I have to discover this new part of my code base that I haven't touched in a while. I can go so much faster on that to understand. So I think there are many, many benefits as well for real product. But the code that is written for production, I think that we are still doing it very manually for the core database, for sure. And then I would say even for more internal things or stuff that are maybe a bit less sensitive where you can, for sure, go faster in your... I don't know if Vibe coding is the right thing, but like AI-assisted coding, I still believe like, yeah, the security audit, these observabilities, there's still a lot.
34:48It's not ready yet for this IoT, in my opinion, but they can help you, yeah, prototype and align on tickets and align on requirements. I think that's pretty impressive. So do you have a line in your product of like, within this must be handwritten outside of this AI assist okay? Yeah, all of the core database and core product of MongoDB, we don't vibrate it, we don't AI it. The code is written manually each time by a developer and we are doing the code review as we did and we are doing the security review as we did. That is for sure, yeah. We can use again some AI tools for some security finding kind of things.
35:23I mean, we can get some help, as many companies do. But I would say most of what we experience with AI are a lot of more management tools around that. There's still a lot for, you know, but like all the management tools around that, all the internal tools, all the before coding and after coding piece to go faster. We use AI more and more, but we don't VibeCode as a core MongoDB database for sure. So that's obviously probably pretty different than a very new, not so security-focused startup. So like thinking about that, has it changed anything about how you're internally organizing your teams? Or does it still take just as many people because the core has to be so kind of solid and locked down?
36:06Yeah, yeah, no, that's a great one. I would say so. There's something first as maybe as a cultural kind of stuff that I came to really enjoy about MongoDB. And I saw that in my previous company at AWS as well. Even people that are not engineers are pretty deep, technically. So we have product managers trying to build their MCP experimentation. I mean, I think that's culturally speaking, that's the case. And you can see that because you have people moving from engineering management to product management role easily. And I think I do. So I just want to see that in this context, because I think that if you are in a different context, which can have also pros and cons, but I think maybe that's different.
36:45But I think in that case, it can really fast track the alignment on what you want to build. Because the example you took before, like instead of just describing the long document, oh, that's the UI I want, you can just do it and then you can align and then you can discuss about how to do that. So yes, it does fast track the alignment between product and engineering, no doubt. We went pretty fast in terms of organization. We made the decision a few months ago now to we don't even have an engineering and a product organization anymore. We have a product and technology organization. That's part of it.
37:16It's not just because of AI, but that was part of it. I think two big objectives. One, how to make sure we can fast track the decision making, the alignment, the sharing of information between the product decision and the engineering decision because they are the same eventually. So that's number one. And number two, how to make sure that everyone is customer obsessed. And even if you're an engineer, you should be customer obsessed. And I do believe AI helps with that, by the way. You can really see what your product manager has in mind. We are using that to really show some reports from customer discussion, advisory board, these kind of things.
37:48So we really made this leap because you are touching on the organization. And I do believe that I guess we probably would have done that as well. But with these AI tools and ways to collaborate, I think it's helping product engineering to be in the same kind of smaller teams than before it was a product and engineering organizations. Yeah. No, I think there is definitely something and have been in a lot of conversations kind of talking about this convergence between engineering and product. Whether that looks at more technically minded product people, whether that looks at more product-minded technology people, as code becomes, maybe not in a database core, but in many contexts is more commodity, it means that product mindset is more and more important.
38:27I think that's spot on. I would have a bit of a nuanced take on this one, which is like, I love that product people can be a bit more engineering or engineering people a bit more product or the wording you used. I love that because again, I think that's a great way to be customer obsessed and to be focused on the outcome and what you are trying to achieve more than just trying to articulate what you even think kind of thing. However, the expertise doesn't go away. Like if you're a product person and you are meeting many, many customers a week, you will have an expertise about reading between the lines and understanding the reading the room in a meeting.
39:03If you're an engineering person, it's not just about this prototype that your product manager colleague can build. It's really about how will that scale? How will that evolve? What do we expect my growth to be and my new changes to be and my next security audit to require. So I think, yes, it's great to bring people a bit closer to the other side somehow, under brackets. But I do believe we should respect expertise. There's still a lot of expertise in what it is to build a system at scale for real production usage and what it is to really understand customer intimacy and what is the need of a market.
39:35And we should respect that. We shouldn't believe that because we have tools that are helping a little bit, this expertise doesn't matter. Yes. I think that that is, you know, if I were to summarize one of my big lessons about LLMs is they're incredible tools, but you cannot turn your brain off. Your brain as an expert is super necessary still. Oh, I love that. And to me, I don't remember who said that. I would love to have been smart enough to say it. But someone says something like, oh, I don't understand. My LLM is super smart on all the topics I don't know. But when it's a topic I really know, it's not that smart.
40:11I love that. Because your expertise is still important. When you know deeply a topic, you do realize that the LLM is wrong sometimes. And you do realize that grounded that with, back to the MongoDB case, but grounded that with real data, real knowledge is important. But even overall, I mean, expertise matters. Even within using it, I find I'm better able to guide these tools in areas I know well than in areas I don't. Yeah, I love that. So coming back a little bit to this piece around the data layer under LLM applications, what do you see as the big unsolved problems? What are the things that your team is working on looking forward for the next, I don't know how long we're allowed to project in the AI era?
40:58Two weeks? Six months? Something like that, right? Yeah, it may be. No, I think, you know, everyone out there is developing an AI application, right? I mean, it's pretty rare to see a company that... But you don't have so many... I think we are at this turning point right now where they are really going to production. I mean, you have this famous MIT paper like three or four months ago saying 95 % of these AI applications, they don't make it to production. Or when they make it to production, they disappoint. They don't bring the ROI they were expecting to bring. But I do believe it's changing.
41:32I do believe more and more of these AI applications are getting production ready. and to me the key to your question about what we see, I believe that people, customers really understand more and more how the quality of the grounded response is, how the quality of the information retrieval to make sure that your LLM will not just be LLM smart, it will be company smart, it will know what the company knows. I think the accuracy for use case just being 2 % or 3 % more accurate, that means that you are totally reducing hallucinations. So I think people are really realizing that even a little bit of impact on the hallucination is like a big deal for the user experience.
42:10And I do believe as well that people realize that, well, actually these stuff are expensive. This AI model can be pretty expensive. And even right now, where they are subsidized by a lot of VC money and all of that, that's still expensive. So if you are able to only call the LLM when you need to call the LLM, and when you're able to optimize the length of your prompt because you did a good job at finding the relevant information in your corpus of data before. If you are using embedding and ranking models that are pretty short and cost efficient, that can totally change the ROI of an AI application.
42:44So accuracy and cost of this AI application, I think that would be a big topic for the years to come, in my opinion. In terms of accuracy then, I mean, I think you've talked some about, we've talked about what it takes in terms of searching, in terms of re-ranking and what you surface. Are there any other kind of best practices you've seen or you recommend to folks in terms of what you're actually putting in that? Are you doing pre-processing? How are you navigating those different levers? Yeah, absolutely. I mean, that could be a very long discussion, but I would say, one, the quality of your data.
43:24if you don't clean your data regularly, handle metadata, know what's important. The quality of your data is key. How you are preparing your data for AI? We spoke about shanking. How do you cut your long text into smaller text that makes sense? I think that there is a science, there is a science to an art to that. How do you prepare your data for? So it has to be clean, it has to be prepared for AI. And then that's really the, what is the right information retrieval strategy for you? Do you want results that will be super fast? Do you want results that will be super accurate because you are a legal company?
43:58And when you are providing advice or financial companies that can impact the tax return or legal document, and then you will want to use the best embedding models with all the context and all the lengths, et cetera, and that's okay if they are more expensive. Or maybe if you're an e-commerce company, you will want to go super fast and make sure that you have 10 or 12 good results and it doesn't have to be the only one. So I think it's about your strategy, about what is the right trade-off for you between this usual quality and cost and speed and all of that. Yeah, the interactivity is an interesting one.
44:29Even within an application, you might have different workflows. Like I was talking to someone who was doing an agent and they were saying, yeah, when I know that the user is right there, I bias towards interactivity and speed and getting it up in front of them. And if it's an async workflow, now I care more about accuracy. Now I care more about, you know, I can take my time. Oh, absolutely. And we see that as an example, not a public reference, but we have an e-commerce customer of us and how you make your trade-off for your e-commerce piece, like where the user will go to find their red shoes or Pogbundi sneakers and how you will manage your stock behind the scene.
45:06These are different kinds of requirements and latencies and costs and things like that. So absolutely, depending on the workload, you will have a different sensitivity. So I think a lot of our listeners are now familiar with in some of the AI coding tools, like you can dial up your budget for, oh, I want you to think longer. I want you to reason more. I want the, you know, mini versus high versus whatever. What are the equivalent knobs that you have in the embedding models and the Reranker and all these different parts that you have with Voyage? Oh, that's a great, great one. So I didn't think about this parallel before, by the way.
45:38So I love it. The thinking and the one with the LLM. And you can definitely go fast and cheaper with like, I would say basic, but even the basic, you can have really good ones versus not so good one, but like a basic text embedding model. And we, I mean, one of the value of Voyage model is they come with different sizes. So you can decide how long your embeddings will be. And even the type, like do you want to go with a float or like to a binary kind of, so you can decide how much. So there is definitely a first decision there about a cost versus accuracy trade-off. Then you can discuss whether, you know, we discuss multimodal and context.
46:16These are heavier models. They can take a bit more time, a bit more compute, but they will give you better results. So is your use case worth this investment? And last, we didn't speak much about re-ranking, but re-ranking is an additional layer that more and more customers are using, which is like the embedding model will basically give you, oh, these are like the 10, 20, 100 best documents in your corpus of data for this query. That's pretty fast and you optimize for that. If you want to know the best model, sorry, the best document for this query, then it has to be compute intensive. You have to go beyond the embedding.
46:53You have to go back to the document itself. And that's what re-rankers are doing. So that's also another layer of your thinking versus the LLM analogy, which I like. You can also decide whether you need a re-ranker to have a very optimized ranking of your results. Now, you mentioned for a lot of these in MongoDB, you would put them as an aggregation pipeline. I'm thinking about use cases that I've had in building these things. Oftentimes, I'll kind of do things in layers where I'll show them something quick, fast, but then I might redo behind the scenes. Okay, I'm going to re-rank. I'm going to resurface this.
47:27I'm going to bump this. I'll do things like that. If I were to do that in your system, Can I get those kind of intermediate results streamed out to me in some way? Or like, how does that end up working? Yeah, you can have, well, what you described is like in a single query, you could first go with a quick search and a bit of a longer one. I don't have many use cases in mind doing that. What I have, though, is like a developer that will start maybe the first iteration, they will go with an embedding model and that's about it. And then when you really want to go to production and you will have real users and real data, et cetera, then they will upgrade their model to a more powerful one to improve the accuracy of the result, or they will add a re-ranker.
48:09But thinking out loud, what you described is totally possible. You could totally do a first search and then do a re-ranking in parallel. Like for instance, as an example, I haven't seen it, but it's such a fresh space that I kind of like the idea still. You could still show, oh, these are the 10 pair of shoes that are looking like your red shoes, Nike kind of thing. And you show them immediately. But after a few seconds, you can re-rank them to make sure that the one that is very accurate comes to the top. You could totally build something. The technology doesn't prevent you to build something that dynamically.
48:40I don't know if the financial aspect or balance is there to do that. Yeah, it'll depend for sure on the application involved. But yeah, I think these types of latency trade-offs are all over the place in these types of applications. And so then there is this question of, okay, how much value is there in showing something to the user versus getting the right answer in front of them? And maybe there's value in each. Yeah, absolutely. Absolutely. But at least what's important, I think, is to have options. Whether you use that in the same flow as you were describing, or you use that because you have different phases of your project.
49:14And at one point, you just want to improve accuracy and et cetera. Or you do that because for a specific application or workload, accuracy is so important. And for another kind of workload, maybe you have a free tier and you're okay that your customers have good results. But when they're paying and you have a premium tier, you want your customers to have excellent results. So I think having the freedom of doing that easily because you can change your schema, you can change your AI model, you can optimize these things is what customers are looking for, right? Yeah. And having that flexibility within the same API, same interaction, I don't have to, now I have to go and get a different thing.
49:50No, that's definitely. And same document model. I'm coming back to that. We didn't invent anything new to store the embeddings, for instance. They are part of your document model. They are there. Like that works. And we didn't reinvent the wheel about, because what people love about MongoDB usually is a horizontal scaling and the fact that you can have these shards and replication. Oh, we did the same for search and vector search. You can have your own search nodes and you can decide that they will be a bit more memory intensive and they will not impact your database. So just using the same principles of the document model, of the distributed architecture, just apply to these embeddings, as you mentioned.
50:24Awesome. Well, we're coming close to the end of our time. Is there anything we haven't talked about that you think would be important to discuss before we wrap? No, I think we touched on it, but I just want maybe to double down on it. Things are changing so fast. The quantity of data is changing so fast. It's not just even humans now generating data and consuming data. It's agents generating data. So the quantity of data is changing so fast. The ecosystem is changing so fast. There are new players and some of them are amazing. And some of them are, looks amazing, but won't be here in six months.
50:57It will go super fast. The LLM race is like, I love it, by the way. I love when like, oh, Google is coming with this great one and then OpenAI, but like it's changing so fast. Your sensitivity to, oh, that should be in this cloud provider and that should be on on-prem is changing so fast, I think it's super important to go with a data platform that can handle this flexibility. And that, you know, if something changes, you can change. You don't have to, oh, I have to rebuild this data pipeline and I have to change my schema and I have to integrate this new identity system to these new players. I think it's really important to at least over-index on flexibility.
51:32I love it. Let's wrap there. Awesome. Thank you, Kevin.
51:43Thank you.
From the publisher
Engineering teams around the world are building AI-focused applications or integrating AI features into existing products. The AI development ecosystem is maturing, which is accelerating how quickly these applications can be prototyped. However, taking AI applications to production remains a notoriously complex process. Modern AI stacks demand LLMs, embeddings, vector search, observability, new caching layers,
The post Production-Grade AI Systems with Fred Roma appeared first on Software Engineering Daily.
