In short
Episode Summary: #264 Amr Awadallah: Vectara's Mission To Make AI Hallucination-Free & Enterprise Ready
Podcast Overview
- Title: Eye On A.I.
- Host: Craig S. Smith
- Episode Release: #264
- Guests: Amr Awadallah, CEO and Co-founder of Vectara
- Focus: The ongoing challenges of AI hallucinations in large language models and how Vectara aims to address these issues to make AI suitable for enterprise use.
Key Themes
- Challenges of AI in Enterprises
- The episode discusses why large language models struggle to be truly enterprise-ready.
- Hallucinations (inaccurate or fabricated information produced by AI) pose significant risks for businesses.
- Vectara's Solutions
- Vectara's core mission is to develop AI agents that are accurate and do not hallucinate.
- Introduction of the Hughes Hallucination Evaluation Model which assesses the accuracy of AI responses in real-time.
- Importance of providing accurate and explainable AI technologies to maintain trust and security, particularly in regulated industries.
- Retrieval-Augmented Generation (RAG)
- RAG is emphasized as a vital approach for enhancing the accuracy of AI responses by retrieving relevant factual data before generating responses.
- Discussion on RAG sprawl, where multiple teams within a large enterprise create their own inconsistent DIY RAG systems, leading to management inefficiencies.
- AI Security and Transparency
- Vectara addresses security threats such as prompt attacks, ensuring AI systems remain resilient against manipulative queries.
- Emphasis on the need for transparency and explainability in AI outputs, especially for industries with strict regulatory requirements.
- The Future of AI Integration
- Vectara envisions a future where AI agents can assist humans in mastering specific tasks (referred to as "I know Kung Fu" era).
- Examples of customer success stories where Vectara's platform empowers users (e.g. radiologists and manufacturing workers) by providing real-time assistance based on extensive databases of knowledge.
Key Takeaways
- Real-time Hallucination Detection: The Hughes Hallucination Evaluation Model helps in identifying inaccuracies and ensures that responses are grounded in facts.
- Importance of RAG: RAG serves as a foundational method for improving AI accuracy and should be standardized across enterprise applications to reduce operational inefficiencies.
- Security Measures: Protecting AI systems from prompt attacks and ensuring ethical guidelines in AI use are paramount.
- Enterprise-Ready AI: Vectara is focused on creating a reliable AI framework that can be utilized across various sectors, including finance and healthcare.
Conclusion Amr Awadallah's insights provide a comprehensive understanding of the current landscape of AI, particularly regarding the challenges posed by hallucinations and the potential solutions offered by Vectara. The conversation encourages developers and enterprise leaders to consider the importance of accuracy, security, and transparency in AI systems, setting the stage for safer and more effective use of AI technologies in the future.
Links
- Vectara: [vectara.com](https://vectara.com)
- Eye On A.I. on X: [Eye On A.I.](https://x.com/EyeOn_AI)
- Craig Smith on X: [Craig Smith](https://x.com/craigss)
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00So Victara, what we're about is we allow our customers, which tend to be very large enterprises, to build AI agents and AI assistants that are accurate, so they don't hallucinate. When you're a large enterprise and you have 20 or 40 different teams all building their own DIY rag in a different way, you have a number of problems. In real time, we have a model that's called the Hughes Hallucination Evaluation Model that in real time, in a matter of milliseconds, reads the response that came back and correlates that response back with the needles in the haystack. This movement is all about that. How can I do that over and over and over again for any task that humans are doing?
0:42Build the future of multi-agent software with agency, A-G-N-T-C-Y. The agency is an open source collective building the internet of agents. It's a collaboration layer where AI agents can discover, connect, and work across frameworks. For developers, this means standardized agent discovery tools, seamless protocols for interagent communication, and modular components to compose and scale multi-agent workflows. Join Crew AI, Langchain, Llama Index, Browser Base, Cisco, and dozens more. The agency is dropping code, specs, and services. No strings attached. Build with other engineers who care about high-quality multi-agent software.
1:36Visit agency.org and add your support. That's A-G-N-T-C-Y dot O-R-G. Join other engineers who care about building the Internet of Agents. So go ahead, Amr, introduce yourself and tell us about Vectara. Sure. Hey, so I'm Amar Awadallah. That's my name. I'm from Egypt originally. I came to the US in 1995 and got my PhD from Stanford University in computer science and engineering. And then after that, I started a company actually during that, to be more honest, I started a company in 1999 that was acquired by Yahoo in the year 2000. That was a comparison shopping database engine. And we became a part of Yahoo shopping.
2:27And then I spent eight years at Yahoo, in which my career trajectory shifted more towards big data, business intelligence, and data science. And I left Yahoo in 2008 and started a company called Cloudera. Cloudera was one of the very first big data companies. We went public on the New York Stock Exchange in 2017, if I recall correctly. And then about four years ago, Cloudera got acquired by KKR, a private equity firm, for about$5.3 billion. And then after Cloudera, I joined Google Cloud for two years. I was vice president of developer relations at Google Cloud. And I was fortunate at Google and Google Cloud to experience CHLGBT two years before anybody else.
3:15Because inside Google, they had it already. They had something called MENA, M-E-N-A, short for meaning, essentially. And that said, they were very cautious about launching it because it hallucinated a lot. It made up facts all the time. So they didn't have the courage to launch it because they were afraid of the reputational risk. That said, it was super impressive. Like, just remember the first time when you tried JGBT, how impressed you were. And me trying it out, I right away saw, wow, this thing is going to change the world. However, it's not going to make any impact on the business world until we fix these problems with accuracy.
3:53So there needs to be a company that can make this technology enterprise ready. And that's exactly what Victara is about. So Victara, what we're about is we allow our customers, which tend to be very large enterprises, to build AI agents and AI assistants that are accurate. So they don't hallucinate. And when they do, you know when they are, so you can suppress or delete that result. And even correct it. We have some correction algorithms. So accuracy is our core, core. It's really our core differentiator, if you were to ask me. Second thing is security, making sure that these systems are resilient to what's called prompt attacks.
4:31Prompt attacks, you know, they are like when you try to psychologically engineer the LLM to reveal something to you that you should not be seeing. Sure. Like the same way somebody can trick your grandma to give them her pin code for her bank account, right? So you can do that with LLMs using front attacks, and we suppress that. So we prevent it from doing that. Security is very, very key. And then last but not least, for any regulated industry, finance, health, government, manufacturing, accounting, law, you need to have transparency. You need to have explainability for how did the AI come up with this result?
5:09How do I make sure this result is not biased? How do I make sure this result is not unethical? How do I make sure this result is not going to piss off my customers? Too negative in terms of the negativity content in it. How do I make sure this result is not a copyright infringement of somebody else's content? So these are all additional things we provide as part of this core platform. Now, there is another very key, I'll end with this last part. There is another very key value that we provide as well. And that is the notion of a standardized platform. So there is, and this is actually the talk I have at this conference on Wednesday.
5:46It's going to be partially about that. It's an issue called RAG Sprawl. So what is RAG first? Are you familiar with the concept of RAG? Yes. Or maybe for your listeners, just to remind them. We can go over it. It's an external database that's... Yeah. It's retrieval augmented generation. Right. And the idea very simply is when you generate a response, first retrieve the needles in the haystack necessary to give an accurate response. This is really what this is about. And we call that RAC. And there is lots of companies now doing what's called DIY RAC, do-it-yourself RAC, right? Which is okay. But the problem is when you have 20 or 40 different...
6:27when you're a large enterprise and you have 20 or 40 different teams all building their own DIY rag in a different way, you have a number of problems. First, it's hard to control the costs across all of them. Second, if one of the engineers that built this DIY thing left, this other engineer that built this other DIY, I have no idea how they did it. They cannot maintain it because it's not standard. And then third, the central IT team cannot govern them. It cannot make sure they're all living up to the same accuracy standards security standards explainability standards etc etc bias and so on so one of our key benefits as well is we usually say the aspirin the medicine to cure the rags pro is now we can have a central platform that your it team can govern properly that your developers can learn one and once and keep building on over and over again whether that be on premise or in the cloud.
7:22We actually support both modernities. So that's another very key aspect as well. Yeah. This battle against hallucinations has gone through several iterations. I mean, originally it was going to be solved through reinforcement learning with human feedback and guardrails. And then it moved on to RAG, Retrieval Augmented Generation. And then people seemed to be dissatisfied with RAG for various reasons, and they started fine-tuning models, supervised fine-tuning. And now there's reinforcement learning fine-tuning, and there's test time or inference time compute with something like the reasoning models.
8:18What are the drawbacks of RAG? And do you cover that whole gamut now in addition to RAG? Yes. So first, in my humble opinion, there is no drawbacks to RAG. You absolutely want to have rag if you want to have an accurate response coming out of your system you need to continuously ground the model in the truth in the facts yeah because by definition with all of these techniques supervised fine-tuning or or otherwise you will always have a probabilistic model at the end of the day and when you have something that's probabilistic that means there's always a probability of error and you can see that so if any of you yourself or any of your listeners go on Google and just search for hallucination leaderboard.
9:04The number one result you will get is a leaderboard that actually Victara maintains, which ranks all of the models out there by their what we call grounded hallucination rates. What is that? Let me explain. Grounded hallucination rates is when I already find the facts, I find the most relevant documents for your question, and I give the LLM, here are the facts, here is my question. What is the probability that the response is going to live up to the facts and there's not going to be anything extra in it. And you'll find out that even when you do that, even when you ground them in the knowledge, they will still inject things in the response that could be wrong.
9:43So the best in the world right now is actually Google and OpenAI. They're about 0.8 % and 0.9%, which is okay for consumer applications. And that's which models at Google and OpenAI? That's the O3 model from OpenAI and Google Gemini 2.0. Yeah. So they, and again, that's under grounding when you ground them. If you don't ground them and you ask them to give a response out of their parametric knowledge, meaning the knowledge stored in the large language model itself that was fine-tuned into it, the hallucination rates are way higher, way, way, way higher. They're like 10 % and above, right? And that's why you still always want to ground them in the truth so that it's kind of like what we do with humans, right?
10:22Before we go to our high school exam tomorrow, we refresh our memories, right? We refresh our memories so we give better answers. And you as a reporter, when you finish writing your article, if you want to make sure, extra sure that it's accurate, what do you do? Who do you call up? You call up a fact checker. Yeah, or I am the fact checker. Yeah, or you double read it or you double check it. Usually you want to get somebody else because if you hallucinated something and you read it, you're still going to think it's correct because you made it up yourself. So most organizations that have news organizations, when they have an article written by a reporter, they will have an independent fact checker that will read that article and make sure it does not deviate from the facts to ensure truthfulness.
11:03So that's exactly what we do in our system as well. So in real time, we have a model that's called the Hughes Hallucination Evaluation Model that in real time, in a matter of milliseconds, reads the response that came back and correlates that response back with the needles in the haystack that were retrieved from your ground truth and gives you a factual consistency score. 100 % means you're perfect. Everything was exactly like the facts said. And then if you go down to 0%, it means throw this away. This is horrible right now. right and we'd mark in the like with a highlighter just like a fact checker with which check your article for you they would mark oh this sentence looks off you might want to remove the sentence from the response to the human leveraging the response and this way you avoid all these embarrassing things that we see in the media like you must have heard about the lawyer that leveraged chat gpt and then it made up a number of completely made up citations of prior cases and that that gentleman get this barred because of that or for example, Air Canada, they were using customer support and then it offered the ticket for $1, right?
12:06Because it hallucinated that as, that's an offer I can give you right now. And that went all the way to the Supreme Court. And the Supreme Court, of course, sided with the plaintiff, meaning the customer and said, yeah, you have to honor that. It's not his mistake that your systems gave that ticket. So that's why you cannot take that chance. You need to have a real-time fact checker that is continuously checking the responses. And that's how we keep the accuracy under fault. So all the techniques that you mentioned, the supervised fine tuning, the reinforcement learning, the reasoning, the real-time reasoning with inference computing to think more about the responses, that's all in the LLM domain.
12:42That's in the LLM itself. But even when you do all these things with the LLM, you're still going to have a lacination rate of 1%, right? Like we see with O3 or like we see with Google Gemini 2.0. So that's why around that LLM, you want to have these two things. On the input to the LLM, you want to have a very, very, very, very, very, very, I'm stressing the importance of it, good retrieval system that finds the most relevant of all the documents and all the knowledge that you have. What's the most relevant needless in the haystack that are needed to give the right response? So that's number one.
13:16You need to have that. You give that to the model. You tell the model, now, okay, now think about the response. Don't use the knowledge in your head. Use the knowledge I'm giving you right now in the context window and answer this question or do this task for me And now you're gonna get a response, but that response might still have a chance of 1 % being wrong That's one of you. That's where you want to bring in the fact checker and tell the fact checker check this response against the needles and give us a score And that's exactly what humans do by the way like that's exactly a legal a lawyer before preparing a legal draft will study the prior art first, will prepare the legal draft, and then the paralegal will review the draft to make sure it's correct.
13:51So it's the same mechanisms. In other words, what I'm trying to say is rag will never go away. You will always need to have rag there, and you will always need to have explainability around it that sees when it goes off so you can fix the data. Because the next problem where this, sorry for digressing here, but the next area where problems can occur is bad data. And in fact, that's the biggest problem for all of us right now, by the way, is bad data. And one of the funniest examples I always like to share is when Google launched AI in Google Search last year. One of the most embarrassing examples they had was somebody was searching for, I'm trying to cook a pizza.
14:30And every time I cook it, the cheese is falling off the pizza. How can I make the cheese stay on the pizza while I'm cooking it? And then the AI from Google replied back and say, put some super glue between the cheese and the pizza. And we all thought at the beginning that that was a hallucination. that the AI just made that up. It wasn't. That was a data quality issue. What happened there is there was somebody who posted that question on Reddit a number of years ago, and a very evil human being replied back sarcastically and said, did you try putting some superglue in the middle? But lots of Reddit users found that answer so funny, they gave it thumbs up.
15:03So that answer got like 30 ,000 thumbs up. So the AI looked at that and said, that must be the right answer, right? So if you don't have the lineage of where did that response come from, that's what we we call the explainability, then you won't be able to go and fix that document that was wrong, fix that error so it doesn't happen again. So you need to have very good observability on the output as well, not just fact checking, because sometimes you do the fact checking and it will tell you 100%, but guess what? The fact was wrong. You had the wrong fact and a human would spot that. So now you need to go back and correct that fact.
15:32So that's why you need to have all of these subsystems integrated holistically to get the best quality of response. And this is on a platform that, so it's a SaaS offering. People go on and - It's on-premise as well. So you're right. It's a SaaS-like platform. So it's cloud native. In the cloud, it's fully SaaS. So in the cloud, you upload your data to us and it's completely SaaS, just like Snowflake as a database. But we also have an on-premise option as well that you can deploy on-premise. Because some of them, especially in the government and financial sector, they're very cautious about doing AI in the cloud.
16:06Yeah, sure. Yeah, it's like a double whammy for them. It's like, what if the AI figures our business for us and then opening AI now opens a bank? You see where this is going? So many of them, they wanted to have it run under their control, inside of their premises. Yeah. You know, I've spoken to a lot of companies doing things that sound very similar to what Vectera's doing. Yes. Is it a crowded space? I mean, I had Edo Liberty on a couple of years ago from Pinecone talking about vector databases and RAG. So we are not a vector database. I want to highlight that. So we are not. We use, inside of our system, we use two vector databases, open source ones.
16:48There is one from Facebook called FIAS, the FIAS library. And then the other one, oh, man, it's spacing out in their name right now. Starts with a W. Come back to me. Yeah. Yeah, it's not Pinecone, though, because we needed something open source that can also be deployed on premise. But we are not about the vector database. We're not about that. We are about the algorithms around the vector database that make sure you have the best embeddings inside of it, that you do the vector matching in the vector database, but you also do keyword matching at the same time. That's called hybrid search. And then once vector databases are very good at what's called retrieval, meaning finding the best needles in the haystack, but they're not good at ranking them.
17:26Exactly. I was going to ask that. All of the vectors around the given direction will be equidistant, so it doesn't know which one is more important. And that's where you need to have a re-ranking model. So we have a very specialized re-ranking model that we built ourselves that ranks these facts correctly. And then we have on top of that what's called chained re-rankers, where we allow the customers to define their own rankers. Because sometimes they might say, oh, I really trust information coming from VentureBeat more than I trust information coming from TechCrunch. or I really want to trust information coming from the last month over information from the last year or I want to trust information that had more than 20 thumbs up.
18:03So we don't know what kind of metadata things they want to do that on. So we allow them to plug in their own re-rankers as part of that pipeline. So sorry for the long spiel here. I want to highlight that we are not about the vector database. Yeah, but you answered my question because there's a lot of dealing, building a vector database and applying RAG to a vector database is an art. Yes, exactly. Retrieval, yes. So that's why, like you were saying, there's many other companies doing like what you do. Yes. But what differentiates us are these two things. Retrieval, amazing retrieval. Because the quality of the result you're going to get is going to directly be proportional to did you retrieve the most important knowledge to be able to answer this question?
18:51That is directly correlated and really, really good at that. And then the hallucination detection on the other side. Again, as I said, we are the benchmark in the industry for that. If you search on Google, like all the big models, when they come out, they call us up, please add us to your dashboard. And then our open source model is called the Hughes Hallucination Evaluation Model. It's open source, actually. It has more than 3.5 million downloads, way more than any of the other vendors out there. So that's kind of how, sorry for bragging a bit here, but that's how we differentiate ourselves, is by the accuracy of the results that we provide.
19:21Yeah. And then you do fine-tuning and you do test-time compute. You also do those things. And how much does that add to your accuracy? Yes, that can help a lot. So it depends. It depends on the use case. It depends on the use case. First, I remember the name of the vector database that came to me. Melvis. Okay. And so it doesn't start with W. It starts with M and it's Milvus. I was going to confuse it with VV8, which we also evaluated, but we lined up with Milvus. It's very, very scalable. It's very high performance. We're very happy with it. And it's open source, so anybody can grab it and use it.
19:57But going back to your question, again, we are not about the generative model. The generative model, it's pluggable for us. So rank is retrieval and then generation and then detection on the other end to detect whether you did something right or wrong. We are all about the retrieval and the detection. This is what we focused on. For the generation, we take the best models out there. So we actually plug into all of them. We plug in with OpenAI. We plug in with Google. We plug in with Anthropic. We plug in with Mistral. We have a version of DeepSeq that we're running ourselves. We have a version of QAN.
20:30QAN, in my opinion, is the best open source system out there. QAN is amazing. It's way better than DeepSeq. I think it's unfair how much press DeepSeq got compared to QAN. and we took actually a version of QAN and we fine-tuned it to your point, does fine-tuning help? We took a version of QAN, the QAN model and we fine-tuned it with a different cost function so usually when these models are being built for consumers, they are encouraged to be know-it-alls they are encouraged to always answer always make the user happy and answer, that's what they're optimized for and that can lead to the error rate being higher because the cost function is about giving an answer always.
21:10It's like when you go to a multiple choice exam and there is no penalty for getting wrong choices, right? You're going to answer everything and you're going to get some answers in the middle because they want to maximize your grade. So that's the cost function that most consumer models have built around. But for the enterprise, you want to penalize it for the wrong answers. So that's exactly Mockingbird. Mockingbird is a derivative model that we built. It's based on QAN. That fine-tuned QAN with a cost function based on DPO algorithm that says is every time you get an answer wrong, I'm going to slap you in your hands.
21:41So when you get an answer right, I'm going to give you a cookie. But when you get it wrong, I'm going to slap you in your hand. So that now makes the QAN model more aligned to only answer when it's confident in its response. And if it sees something where it thinks it's going to be guessing a lot, says, sorry, I'm not going to answer this. I'm afraid to be whacked on my hand. That's interesting. Make sense? Yeah. Build the future of multi-agent software with Agency. A-G-N-T-C-Y. The agency is an open source collective building the internet of agents. It's a collaboration layer where AI agents can discover, connect, and work across frameworks.
22:17For developers, this means standardized agent discovery tools, seamless protocols for interagent communication, and modular components to compose and scale multi-agent workflows. Join Crew AI, Langchain, Llama Index, Browserbase, Cisco, and dozens more. The agency is dropping code, specs, and services. No strings attached. Build with other engineers who care about high-quality multi-agent software. Visit agency.org and add your support. That's A-G-N-T-C-Y dot O-R-G. Join other engineers who care about building the Internet of agents. And that's a reinforcement learning fine-tuning. Yeah, using a technique called DPO.
23:15Yeah, and you do that before you hit the rag? Yes, that's during the training phase of the model. Right. So that's in compile time. I usually make this analogy. There is compile time and there is runtime, right? Model training is compile time. And model inference is runtime, right? And by the way, there's going to be way more inference than training in the long term. Today, we have lots of training going on because we're still at the beginning of the market and people are still building the models. But inference is going to be way bigger in the future. And that's why you see new companies now jumping into the inference space like Grok, not Elon Musk Grok, the other Grok.
23:53And Cerberus and Sambanova, you're right. Absolutely. So we do this technique of fine tuning and penalizing for giving the wrong responses. We do that during training time of the model. And once it's trained, we use it over and over again for all of our customers. Now, some customers might come to us and say, we built our own generative model and we fine tuned it on our own data set. which makes it more accurate. And that's correct. We'll take it. We'll plug it in. It's fine. Like, we'll leverage that. So we are pluggable on the generative model side. Yeah. I also spoke to...
24:31Take your time. I'm going to forget the name now. Give me a clue. Guy who spoke this morning. Glean? I've met too many people. Yeah, yeah. Take your time. Anyway, but they're... oh databricks i'm sorry oh yeah yeah and he was talking about how uh you know applying their technology to your uh data and and uh you know cleaning and rationalizing and improves the quality of responses yeah do you guys are they complementary to you they're actually they're partners they're actually investors in us i see so databricks is one of our investors and we are partnering with a number of other organizations as well on the data ingest side, like Confluent, Kafka, if you're familiar with the Kafka framework.
25:22We're working on other similar partnerships with other vendors. Airbyte is another one that we're partners with. And absolutely, as you bring data in, if you can clean it first before you load in, that's great. But in our case, even if you don't clean it, if you just give us good metadata around the document so that we have the metadata of who was the author, TechCrunch versus VentureBeat, what was the timestamp, which group within the company, how many thumbs up or thumbs down, any knowledge, supply chain metrics that you have around this document coming in, previous legal cases that you won, any metadata coming in can be leveraged to do that cleaning on the fly.
26:02Because what will happen then is the re-ranking algorithm will push the dirty data out, down, way down, and will keep the good data higher in the ranking. and the generative model is incentivized, give more attention to the stuff at the beginning of the prompt and not the stuff at the end of the prompt. So this way it provides a good response. It's a good hack to still give good responses without having to clean all of your data a priori. I'm not saying that to say that you should not clean your data. You absolutely should clean your data and in fact that's the biggest problem in the industry today, by far.
26:38In fact, I love the tagline from Informatica who I'm talking to them hopefully to be a partner as well. We're not partners yet, but hopefully they will be. I'm giving them free advertising right now. So Informatica tagline is lovely. I love it. It goes like this. Everybody is ready for AI except your data. And that's absolutely right. Yeah. But where do you think, so vector databases you think are here to stay, RAG is here to stay, oh no. No, vector databases. So it's funny. So actually, I'm an investor in a company called QDrent, which is a vector database company. I should disclose that. But I did that three, four years ago when the market was still at the beginning.
Read the full transcript
27:19What became very apparent is vector matching is an essential thing that any database should be doing. It's just like date matching, string matching, geolocation matching, address matching, name matching, vector matching. It doesn't need to be its own beast, right? And that very quickly came to be true. Like Postgres now has vector matching. Oracle has vector matching. MongoDB has vector matching. Elastic has vector matching. BigQuery and Google has vector matching. SQL Server, Microsoft, and CosmoDB have vector matching. So vector matching now became a feature that all platforms have. So I think that I could be wrong here, but my prediction is the vector matching companies will either have to evolve beyond vector matching, maybe jump into the RAC space, into our space, but good luck to them doing that, or get acquired.
28:14They're not going to have another option. These are going to be only two options in front of them because the hyperscalers win when it's a feature. If it's a full platform by itself that's unique, you have a chance, you have a fighting chance. But when it's just one feature of the many capabilities they offer in their core platform, it's very hard to compete with that. So you're saying that your RAC service doesn't necessarily have to retrieve from a vector database. It could retrieve from SQL servers. So anybody that has vector matching. Yes, exactly. Is that right? That's correct. And from SQL servers, from SQL or from microservice APIs.
28:53So part of our, like this is a bit deeper in our framework, but part of our offering is called Egentic RAG. So what is Egentic RAG? You're still doing retrieval, augmented generation, but the retrieval now is more agentic in nature. What does that mean? You will ask a high-level question like this, for example. What are the NVIDIA and Intel revenues for last year? And please also compare the risks they disclosed in their 10K filing. A question like this needs to be broken down to multiple questions. First, a SQL lookup that will look up the revenue. to select from revenue from table where a stock symbol is Intel or stock symbol is NVIDIA, and you get that.
29:35And then another lookup, which will do the semantic matching on the documents to retrieve the risks from the SE filings. And then all of that will be given to a generative model now. These are the needles in the haystack that we fetched from some from microservices, some from SQL, some from semantic systems. And then we give all of that to a generative model, and we say, okay, now please put all of this together into a nice document that gives you the answer to this final question, and they assemble everything together. Make sense? Yeah. So you have multiple sources. Graph systems as well. Like graph systems also can be an input into our system.
30:05I see. Where do you think this is going? I love that question. Yeah. I mean, it's an open-ended question. I was waiting for it. But I just... No, I love that question. You'll see why in a second now. Okay. Because, you know, you've got these multi-agent systems. I just was writing about Manus from this company, Monica, in China. Yeah. But you know, it turned out to be a hoax. I mean, not a hoax. Well, it's built on Anthropics cloud, right? Yeah, it's just a layer around Anthropics. Yeah, but it's a pretty impressive layer. Yeah, the UI is impressive. Yeah. Or do you think that it's not? I mean, nobody else has come up with an orchestration.
30:50No, no, that's not true. That's not true. So, Anthropic has a Tropic computer, which does that. And OpenAI has operator. I've been using operator for a few months. It does that. Just as good, if not better, more reliably, with many, many websites as well. It does not use the computer, though. It only uses the browser. And then Google has something called Project Mariner, which does that. And Microsoft is working on a product that's going to be part of Windows, that does that directly inside of Windows. The only company that's behind that bit is Apple, to be honest, in terms of providing it natively.
31:20So the question is going to be for Manos is, again, the same question is, will they have the muscle to keep funding the cost of this during the competitive phase, which is commodity, until the maturity phase where there's going to be lots of money being made but at the very thin margins? And that's usually, it's very hard to compete in that environment, right? So I'm not saying I am impressed by their marketing in the same way I'm impressed by the DeepSeek marketing. but if you look at the fundamentals, QAN, way better than DeepSeek. Anthropic, with what they're doing with computer use, is way more powerful than Manus, way more powerful.
31:56I could be wrong. That's my humble opinion. But going to your question, though, where is this going? The future that we're all working towards, I like to use an analogy from movies. I'm a big movies fan, especially sci-fi. First, hopefully the answer is yes, The Matrix. Is The Matrix one of your favorite movies? I've certainly seen it and enjoyed it. Yeah, the first Matrix specifically. The second one and third one are crap. But the first one and the fourth one was like a complete hilarity. But the first one, there is a very key line. In fact, the most famous line from that movie describes exactly where this is going.
32:31It's exactly what we're all building towards. So what is that most favorite line? I'm not going to say it yet. I'm going to tee it up. There is a scene where Neo, Neo who's played by Keanu Reeves, is plugged into the AI. And unfortunately in the movie they have these big things you stick in your head. Hopefully we don't need that. Hopefully we just wear a helmet or something without sticking a big nail in the back of our heads. But he spends eight hours with the AI and then he wakes up after these hours, eight hours, and he says this. I know Kung Fu. I know Kung Fu. And he knew nothing about Kung Fu before, right?
33:08So he went from being a novice, a beginner who knows nothing about that domain to being an expert in that domain. And this is really what this is about. This movement is all about that. How can I do that over and over and over again for any task that humans are doing? One of the examples, like many people will tell me sometimes in my talks, oh, I would never use AI. I'm afraid of AI. I would never use it. I would never touch it. I would never get near it. I would never listen to it. And I ask them one question. Have you used Google Maps? Yeah. It's funny you say that. I do the same thing. People say, I'm never going to use it.
33:43I say, when was the last time you opened up a paper map? Yeah, yeah, exactly. I've used Google Maps in the last month, and all of them were raised in hand. I mean, who do you think is routing the map for you in Google Maps, figuring out the weather conditions, the traffic conditions, the optimum routes, the shortcuts, the road closures, the construction, and giving you the best route? Do you think it's a human being at Google sitting and doing that? It's the AI. Are you listening to it when it says go right? Yes. Are you listening to it when you said go left? Yes. But the difference with Google Maps is Google Maps was custom-built AI.
34:10It was AI that Google had to build just for that task. The beauty of this new world of large language models is it works across any task. It's literally any task you could think of. Transformers are amazing. We don't understand how they do that, all of us. And that's the genius of it. So I'll give you, and hopefully this will be the conclusion, a couple of examples from our customer base. We have a customer called Sonosem. At Sonosim, what they do is they help radiologists use ultrasound machines. When you're using an ultrasound machine, you have to calibrate it correctly. And the calibration is a function of so many parameters.
34:48It's a function of the model and manufacture of the machine, the patient's gender, race, age, and the condition, pregnancy, heartache, muscle ache, whatever, right? The expert, expert radiologists, they know it's solid. They can just do it with their eyes closed. The average, they struggle with it all the time. The normal human layman has no idea how to do it. So what they did is, Sonosim, they took, with our platform, they took all of the manuals of the machines. They took all of the best practices from the experts, from the guys that know Kung Fu, how to calibrate this machine, loaded all of that up in our platform.
35:22And what they have right now is an AI agent that gives you the answer in real time. You just describe the condition. Here you go. This is the perfect calibration comes out. So now that average radiologist became expert, just like that. Another quick example is in manufacturing. This is one of our largest customers in the Middle East. They make bottles, like water bottles and drinks bottles and so on. They have hundreds of thousands of workers in their factories, and sometimes the machines in the factory will break down. The workers are not technicians. They don't know how to fix the machine. They don't know Kung-Kru.
35:55So they just sit there waiting for the technician to come. The technician can take two days to show up to fix the machine. That's downtime, that's lost of wasted energy, money, and productivity, and so on. So what they did with our platform, they took all of the manuals, again, of these machines, all of the troubleshooting tickets, customer support tickets, maintenance tickets, loaded all of these up into our platform, and now the workers have this app. They take a picture, they ask the question, and it gives them the step-by-step how to fix the machines themselves. They now know Kung Fu. That's less downtime in the factory.
36:26That is cheaper cost. You don't have to bring in the expensive technician. And last but not least, the workers actually get upskilled because they learn how to do it themselves. The same way we are learning from Chad GPT how to do things right now as well. So that's where this is going. We want to enable the I know Kung Fu era for all of us. Okay, Amr, this is fascinating. Thank you. I really enjoyed it. I look forward to following it and talking to you again. And I'll go back and take a second look at Manas with a little more critical eye. Yes. By the way, the rabbit did this as well. The rabbit had the task thing a year ago.
37:04A year ago, I could tell rabbit, go to Amazon and do all these steps for me. Go to LinkedIn, find all the people that are second connections from this company and add them and will do that for me. So tell me, I've never really understood why you need a separate device. Why isn't it enough? You're absolutely right. This is going to fail because of that. I agree with you. So I think the cool thing about it, it was only$200 and it's AI for life. You don't have to pay for the AI. There's no monthly subscription and it's multimodal, has voice and everything. That's the cool thing about it. But it's AI for the lifetime of the company because the company will fail, just like Humane failed.
37:42I don't know if you saw what happened with Humane, the pin guys. And the reason why is, yes, we don't need one more device. I already have my phone. I already have my watch, which has an assistant in it. And I have my Ray-Ban glasses in my pocket. Oh, you do. I would, yeah. So, yeah. So, I don't need to carry one more device. So, that's why this device is not going to succeed. Exactly because of that. Yeah. How do you, do you use them? Yeah, these are amazing. For translation, I could be, if I'm traveling in Japan, the street signs, like, read this for me and it will read it for me in English.
38:11If I'm having a conversation, do live translation for me right now. It will do live translation. I will hear my ear. You can be speaking any language and I will hear it in my ear in English. Wow. So, so, so useful. I can look at a legal contract and say, check this contract for me and tell me what things I should be cautious about. I can look at the menu. I don't eat pork. And I can look at the menu and there could be like 100 choices. And I just tell me the three choices I should go for. I would say go three. Or look at my clothes in the mirror. Are my clothes matching right now? Tell me, no, that jacket is awful.
38:38Change it right now. I'm going to get a pair. Yeah, yeah. I highly recommend it. It's really fun. Yeah. Okay. Okay, great. You're a wonderful interviewer. Oh, thank you. Thank you. I hope we meet again. Yes, me too. I really enjoyed this conversation.
From the publisher
AGNTCY - Unlock agents at scale with an open Internet of Agents. Visit https://agntcy.org/ and add your support.
What's stopping large language models from being truly enterprise-ready?
In this episode, Vectara CEO and co-founder Amr Awadallah breaks down how his team is solving one of AI's biggest problems: hallucinations.
From his early work at Yahoo and Cloudera to building Vectara, Amr shares his mission to make AI accurate, secure, and explainable. He dives deep into why RAG (Retrieval-Augmented Generation) is essential, how Vectara detects hallucinations in real-time, and why trust and transparency are non-negotiable for AI in business.
Whether you're a developer, founder, or enterprise leader, this conversation sheds light on the future of safe, reliable, and production-ready AI.
Don't miss this if you want to understand how AI will really be used at scale.
Stay Updated:
Craig Smith on X: https://x.com/craigss
Eye on A.I. on X: https://x.com/EyeOn_AI




