In short
Podcast Episode Summary: How Glean CEO Arvind Jain Solved the Enterprise Search Problem – and What It Means for AI at Work
Episode Overview In this episode of "Training Data," hosts Sonya Huang and Pat Grady engage in a detailed conversation with Arvind Jain, CEO of Glean, a company that is redefining enterprise search and knowledge management through AI. With a background as an early Google employee involved in search algorithms, Jain shares insights on the complexity of enterprise search and how Glean aims to revolutionize knowledge work using generative AI.
Key Themes
- Challenges of Enterprise Search
- Variability in permissions and access for employees complicates search functionality.
- The scattered nature of enterprise data across various systems makes it difficult to index and retrieve information effectively.
- Glean's Unique Approach
- Glean functions like "Google or ChatGPT for your enterprise," allowing employees to query company knowledge seamlessly.
- The platform integrates deeply with common enterprise applications (e.g., Salesforce, Google Drive) to create a cohesive knowledge base.
- The Role of AI in Enhancing Knowledge Work
- Generative AI is utilized to synthesize information and provide context, thereby improving the search experience.
- The concept of Retrieval-Augmented Generation (RAG) is emphasized as a way to combine generative AI with structured data retrieval.
- Future of Knowledge Work with AI Assistants
- Jain envisions a future where AI assistants will handle a significant portion of knowledge work, making employees more efficient and informed.
- The role of AI will shift from reactive (answering questions) to proactive (providing suggestions and insights).
Episode Structure
- 00:00 - Introduction
- 08:35 - Search Rankings
- Importance of ranking relevant documents based on user context and recent activity.
- 11:30 - Retrieval-Augmented Generation (RAG)
- Explanation and significance of RAG in enhancing AI’s capabilities.
- 15:52 - Where Enterprise Search Meets RAG
- Discussion on how Glean integrates RAG into its platform.
- 19:13 - How Glean is Changing Work
- Real-world applications and user experiences with Glean.
- 26:08 - Agentic Reasoning
- Exploration of how AI can reason and infer in enterprise contexts.
- 31:18 - Act 2: Application Platform
- Transition from search functionalities to developing applications on Glean's platform.
- 33:36 - Developers Building on Glean
- Features that make Glean a desirable platform for developers.
- 35:54 - Five Years into the Future
- Jain’s vision for the future of work and AI integration.
- 38:48 - Advice for Founders
- Jain shares lessons learned and strategies for success in AI ventures.
Key Takeaways
- Enterprise Search is Complex
- The intricacies of permissions and data organization make traditional search methods inadequate.
- AI as an Enabler for Knowledge Workers
- Glean aims to empower employees by streamlining access to information and driving productivity through AI.
- Proactive Assistance in the Workplace
- Future AI assistants will not only respond to queries but will also anticipate needs and suggest actions.
- Importance of User-Centric Development
- Glean emphasizes user feedback to refine functionalities and ensure that the platform meets real-world needs.
- Building Trust for Data Usage
- Establishing credibility with users is essential for companies to entrust their data to AI solutions.
Conclusion Arvind Jain's insights provide a glimpse into how Glean is positioning itself at the forefront of enterprise AI and search technology. The episode underscores the significant transformation that AI can bring to knowledge work, with the potential to redefine how individuals and organizations operate in a data-driven world.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00majority of the work that we do today is not going to be done by us anymore in five years from now. And that applies to me, that applies to you, like, you know, we both do very different things, but still, like, I think, you know, we are knowledge workers. And I think a lot of our work is actually going to be done by these amazing AI assistants that are actually in many ways, you know, more powerful than us, like, you know, the, they, they have access to all of our companies data or knowledge. They have all the context from all the past conversations and meetings, they don't forget anything. And they can really sort of, on top of that, they have the reasoning capabilities that allow them to be super helpful to you and any task that you do.
0:41So that's our core belief that what if our work is actually going to be done by these AI companions or assistants, and we want Gleene to be that assistant in the workplace.
0:56Please.
1:08Join us today as Arvin Jane, co -founder and CEO of Clean. Earlier in his career, Arvin was instrumental in building Google search and co -founder and CTO of Rubrik. Glean began life as an enterprise search company and today has evolved into a general purpose work assistant. Bringing AI into an enterprise context is notoriously difficult because of the integrations, the permissions, the ranking, the parsing, all the other magic that needs to happen to make AI work on your company data. Arvin joins us today to share how Glean is solving this problem where other companies have failed and what he's learned as one of the first successful AI -native application companies.
1:47Arvin, thank you so much for joining us. We have a lot of questions about Ragn, agents and knowledge graphs and all of that. But before we do that, can you give us one or two minutes on what is Glean and what are you building? Yeah, first of all, thank you for having me. Glean, think of it as the Google or Chat GPD, but inside your enterprise. It's a place where your employees go and ask questions and Glean answers all of those questions using your company knowledge, regardless of where that knowledge is, being said all back to you. So that's what Gleene does. Gleene is also an AI platform. So if you want to actually build AI applications inside your company, you can use the Gleene Rack platform to build those applications quickly.
2:32Wonderful. And since you make the analogy of a Google for Work, Google for Work I think is something that every CIO has described as their Holy Grail and we have two decades of fell attempts at building it. You were actually a star search engineer at Google before, and even Google never managed to crack this category before. Maybe, can you just say, why is this such a hard problem? And how did you do it? What I mean, search is hard because it's actually magic, in some sense, like you can come and ask any question that you have, and you expect the system to actually give you back the right answer.
3:08So expectations are always high. And it's a difficult problem, especially in the enterprises, because there's so much information inside the enterprises spread across so many different systems. It's both hard to actually even get hold of that information, but then even harder to actually make sense of what information is actually good, what has become out of date. So there is lots and lots of challenges around building that system. And in the past, I would say that there were no good attempts made. Because the problem was so hard, it requires so much R &D, so much investment, it was not really start -up friendly in many ways.
3:53And you couldn't even build a product. Just connecting with all of enterprise data meant that you had to spend like an year sitting with an enterprise, trying to actually bring that data into your search system, and then actually solve the real problem, which is make that information searchable. Arvin, one of the things that I think is so interesting about Glean is you are probably one of the first and best examples of what an enterprise AI application company can or should look like. And we're going to focus most of this conversation on the AI aspects of Glean. However, I know there are a lot of layers to stack.
4:32You've got the infrastructure, you've got the connectors, you've got the governance engine, you've got the knowledge graph. Can you say a couple words about all the stuff you had to build before you even got to the AI part to make the AI work? Absolutely. So as you said, search starts first with the data and the knowledge that you actually make searchable. So the first part of the green tech stack is these deep integrations that we've built with most common enterprise systems. So think of systems like Salesforce, or confluence, Gira, Google Drive, SharePoint, ServiceNow. Your enterprise data typically lives in all of these different systems.
5:09I need to bring it all together in one place. So that's the first part of our technology stack is these integrations. But then if you think about enterprise data, and this is one of the most unique things about enterprise search versus if you think about Google Search on the web, is most of an enterprise information is actually private in nature. When you author a document in Google Drive, you know, like this document may actually be private to you or you may share it with a few other people. And so you can't build a search engine which, you know, where you just dump all the company knowledge and make it accessible to everyone.
5:43You have to actually understand permissions of each content. So when you go and search, the system should understand who you are and only retrieve information that you actually have access to. So that's our governance layer. Understanding governance across all of these hundreds of different systems, which is quite complicated. And then the third part, and this is where really most parts I've filled in the past, is search is not about just putting like a whole bunch of documents in an index. And then like, you know, and somebody comes and asks a question, take those words or take that question, and just match it up semantically or with keywords, with the right content.
6:22You got to actually also understand who's the person who's asking a question. You know, I can come in and ask for an onboarding guide because I'm new employee. But then which onboarding guide should be actually given to me, it depends on whether I'm in the marketing team or I'm in the engineering team. So understanding people and understanding knowledge and relationships between them, that's a big part of actually making search or a question on service work inside an enterprise. So we do that. So we actually build a deep knowledge graph, where we look at all the employees, understand what roles do they play in the company, look at all the documents, then they sort of try to understand what documents are meant for what departments, what documents actually are popular.
7:07Is there like what are the relationships between a particular individual and a particular document? And that is what we used as a core foundation that sort of then governs, like when somebody comes and asks a question, what are the most relevant pieces of knowledge for them? So we had to do all of that work. And actually, interestingly, you mentioned what came before AI even became relevant. For us, the AI was actually part of the course first technology from day one. We were actually working with LLMs in 2019, or at least language models, like these were bird -based language models that you could use.
7:45And so when Gleen - MLMs. That's right. Is that the name for them now? But yeah, in the search, like engineering, community, we just call them language models at the time. And so the language models are actually part of the course search experience from day one, because it really allowed us to understand content at a semantic level. So that we already baked in our course search experience on day one when we were actually trying to look at a question from a user. We were never sort of limited by the actual exact keywords that you know users use. We were able to actually understand the meaning behind the question and actually match it up with the right documents.
8:27But still, that's sort of like all the work that you have to do before you can actually even do anything with LLMs. Can you say where the rankings? I think part of what makes Google work so well is I always get the answer I want at the top of the page. In the case of the public internet, you have so much web data and links and all that to make the rankings really work. How much, to what extent is that part of the magic for Glean and how do you guys do it? Yeah, so that's of course the core of the product is all the effort that we put in to actually build a really good ranking system for our search.
9:05I'll give you some examples of the kind of things that go into that mining, what documents are the best ones to rank for a given question. So of course, if you imagine that there's a document that people in the company are constantly looking at. So that gives you obviously a signal that there must be something about it. It's actually important. People like to spend time on it. If there's a document that was actually written in the last one or two weeks and there's some engagement around it, you know that this again is information that people care about. It's not actually become obsolete yet. Then if you think about a particular document that we see is not popular when you think look at a company level, but you look at one individual team inside the company, we see this heavy usage for that document inside that particular team.
10:02So that sort of tells us more about that this document may actually be relevant for this particular set of people. Or the last thing, the last example, imagine that somebody had a question and they're not like bothering to a search, they go and Slack and ask a question and then somebody else posts a link to a document as a response to that and the person who asked the question gave you know gave a thumbs up to it. Just you know imagine this interaction like what it means it actually means that that particular document was actually a really good answer for that question that the user had asked and so if you keep that association in mind like it's going to help you later when somebody else came and asked a similar question.
10:45So those are some of the signals. We have to sort of constantly look for all these signals. You know, they have to collect them differently in the enterprise setting. Then on the web, like, you know, Google only has to look at all the activity that's happening, right on Google itself, because that's the gateway to like, you know, any sort of knowledge quest. But if you look at in the enterprise, you know, not all the things are happening, you know, through the search paradigm. So you have to sort of go and look at all the activity around all the knowledge in different systems like your communication systems, your document systems, and just try to learn from that human behavior.
11:20Because ultimately that's how you learn. You learn from what people are doing inside the company. The more you can actually collect that information, the better your ranking systems are going to be. Can we spend a minute on Rang? As Pat mentioned, you were kind of in the right place at the right time. You could put together all the hard stuff so that when the LLM has gotten really good. You kind of had all the infrastructure in place. I think you've been one of the experts in using Rags to make these LLMs useful on your corporate content. Can you explain Rags to me like I'm five years old and what are the secrets to making it work?
11:59What are things that people don't talk about? What are examples of things that you can do thanks to Rags that you can't in a generic chat interface? Yeah, well, first of all, I think since you were talking about a five -year -old, let's first talk about what right. 10 -year -old. 10 -year -old. I need five. First of all, let's start with five. If you think about all these amazing models in GPT and Gemini and Cloud, these models are all trained on the world's public knowledge and data. And so if you were to actually go into GPD and ask a question like, hey, how many days do I get off, you know, you know, with my ptopasy, it has no idea.
12:41I can't answer that question because, you know, that's my company's private knowledge, you know, the answer is somewhere there and the model is not trained on it. So how do you bring your private enterprise data, you know, to, you know, to these models so that you can actually have a, I create that magic for you. That's what you, it lag based AI application architecture allows you to do. So the way it works is that you come and ask a question and you have a search engine or a retrieval engine, whatever you want to call it. And given the question, this retrieval engine finds potentially relevant documents, that could actually answer your question.
13:24And then you're going to take those documents or those sort of content fragments and make the model work on it. You'll tell the model like GPT that, hey, I have this question and I have this company knowledge that I think is relevant in terms of answering that question. Now, you answer that question using this knowledge. So, this is how like most AI applications today are being built in the enterprise. the only way to actually connect your private enterprise data to the power of these language models is basically a search engine that's sort of sitting in the middle. So, we like in our given that at Gleen also like we of course built a search engine over all of our enterprise content in the last five years, it actually allowed us to actually become like the best rag systems that allows you to now.
14:21Not only like of course, we deliver our own and user application, which is a green assistant using this rag -based application architecture, but we're also allowing companies to actually build more and more applications using rag. Now, I think in terms of while this is the architecture that is emerging as a canonical architecture for building AI applications. I think it is still full of challenges. It's actually really hard to build great AI applications using RAAG. Because one of the things, you know, that how models themselves are sort of wild, very powerful, they're also still an emerging technology.
15:02Models hallucinate, they make things up. And what you're doing now is you're actually adding one more complex layer of technology in this application architecture. So think of it like you're chaining two things, which are both not perfect. So oftentimes, you will see a rack -based AI application not perform well, because you asked a question and the failure actually happens at the racks, at the retrieval stage, where you didn't even actually, you know, were able to find the right pieces of knowledge, or maybe you found like, you know, stale information that then you're actually giving to the LLM to work on.
15:41And then of course, it's going to give you bad So it actually, while it's only way to bring knowledge together, it creates these interesting challenges for you as well. Let me see a question. In just a paraphrase a bit, what you were seeing at the start of this conversation, Act 1, Enterprise Search, Act 2, Application Platform, for that Act 1, which is Enterprise Search, how do the concepts of Enterprise Search and Ragn relate? Is one a superset or a subset of the other? Are they similar but distinct? Are they the same thing? How does enterprise search and Ragn? How do those concepts relate? So I think of search and Ragn as being in some sense, they're one of the same thing.
16:26The real code technology is taking all of your knowledge, enterprise knowledge, and putting them into this search system, where now you can actually ask questions and the system is able to actually give you relevant pieces of information back. So that's sort of the core technology. Now, you can actually use this technology either as a standalone product. So that's what the Gleene Search product is, for example, where people come in, they ask questions and we can give them the relevant documents. That potentially are useful to them based on that question. or you could actually use this as an API layer in your overall AI application.
17:14So where the search system, the search module, now is only one component of your overall AI application architecture. And the, so I think in that sense, it's sort of similar. But the industry on the other hand, I think like what we've seen, And most of these rack -based applications in the enterprise today, they actually use a much more simpler version of a retrieval system in their rack application, typically like a vector search -based system, which doesn't really have full enterprise context. And so that's how I would say that's the key difference. So for us, our approach always has has been that really think about how to build a standalone search system.
18:04Something that is as good that you can actually put them in front of the users as a standalone product. That's really the real test for how good the search is. That then actually when you put it behind the scenes in a rack -based application, it's actually going to actually create better AI experiences. So is it fair to say that kind of the magic that you've done in terms of getting a good ranking of search results? That is exactly like you've made that ranking good for people. It turns out making that ranking good for people is also what you need to make the ranking good for machines in order to get the best possible results.
18:39And that's why like, you know, what you've built is very different from somebody that's just DIY -ing at data pipeline and their own little retrieval system. Yeah, that's correct. I mean, I think it's really hard to build these systems, you know, like yourself and build them in matter of weeks, I think, like, I think like you can build a great AI demo in one day, maybe like in two hours now, but I think like to actually build, you know, like, you know, we're sort of like, you know, it's robust, it's stable, it actually adds value to your, you know, with a near enterprise, you know, like, you know, it's a hard problem.
19:13So we've talked a little bit about how you've built what you've built, and we know that it's working. We know that the company is quadrupling, year over year and we use it here internally and there are a lot of happy people out there who are customers of yours. The real measure of success in some ways is how your product is changing the lives of its customers. And so I'm curious to hear from you. When you look at your customers and sort of how they operate day -to -day, pre -glean versus post -glean, what are some of the changes you notice? How does this help people do their jobs? So, Glean is actually a product that is used quite heavily by people.
19:55There's like many, many different types of things. We're often surprised by what people are using Glean for. But I would give you a few examples. So, for engineering teams, they find Glean super useful in terms of trouble shooting. Whenever you run into any kind of roadblock or an issue, like sometimes an error, like, you know, just, you know, programs are not working properly. And so, Green serves as a really good troubleshooting tool for them. Like, you know, it's a place where you go and debug because, you know, you post the issue more often than not, you're not the first one, like, who's going to experience an issue, like somebody else has experienced all these issues before.
20:34So just getting the context from all of those, you know, like all the other people and and how they solve that problem before, like it sort of helps you solve that issue for yourself. So that's a big use case for engineering.
20:54For some roles, like for support, their life, day in, day out is about resolving, answering people's questions. And I think the tool likely, actually, fundamentally, change, change, you know, like how they work now, because, you know, by default, now they don't think about, like, given a question, like, trying to go and look for, like, answers in different knowledge spaces and whatnot, like, you know, instead, like, the first, you know, reaction, you know, that they have now is that, you know, there's a question that's coming in from a customer and then Gleen on the side is actually already answering those questions for them.
21:30So, so there's sort of, like, model of working change, you know, changes is from trying to find things to actually trying to validate what AI is telling them, as is the right answer, and then just share that back with users. Some teams are actually really changed their behavior. Stills people, for example, they use Glean as a way to prep for meetings. So before customer call is coming up, they will just ask Glean, they can be lazy, they can ask Glean, help me prep for this meeting, and Glean is actually going to bring that that 360 view of all of the data from that customer, like what happened in the last meeting, who's in like what opportunities in our open, with them and things like that.
22:13So it sort of really helps them, prefer a meeting, then actually a running meeting well, because customers always have lots of questions. And so it saves people feel more confident. Like running that meeting, because if somebody throws a curve ball at them, they can just ask clean right in that meeting, get down, so we're quickly sort of get those responses back. So in fact, in our company, we don't allow sales people to actually bring in sales engineers in the call. They have to answer those questions themselves. So that's one change in behavior that we drive in the first few calls. Yeah, those are some of the things.
22:54But overall, the use cases are unbounded. The one that I think is the one that's universal across everybody inside the company is finding other people who can help you, that is one of the things that Gleen makes it really easy for people. We help you connect with the right subject matter experts based on what questions you have. That's one thing that we see everybody in the company make use of a lot. Is there a North Star Metric e -track? I think these are wonderful stories of customer impact. I guess how do you benchmark yourself objectively? Yeah, so our key metric is how many questions people ask on a daily basis and actually get successful like you know, we were successful in answering those questions correctly for them.
23:40So similar to like Google's like search set metric then like okay. Yeah, can you can you share anything about those numbers or are you prefer the qubit private? Well, we have so yeah, so we have this technical metric. I don't know how much sense is going to make, but they would be tend to actually keep that number at 80%. So I think it's a proxy for that the 80 % of the sessions that users had with us, they were actually successful in getting what they needed. And is that, do you measure that success? Is that explicitly they thumbs up this was good or is that implicitly they take action on the basis of the results that you served and you can see that action taking?
24:16How do you actually measure the success? It's actually implicit. So we will track that actions. For example, in search, when you come and ask a question, and then you click on one of the top two or three results, and go to the destination, and then stay there for a long time. So that sort of gives us an indication that you were happy. You didn't come back and ask another question quickly or to find your search. So that's how we track, whether somebody's successful or not. What are some of the top things that are not in the product yet that you think will make people more successful? I think the, like I was, I referred to this, you know, when we started that building a product like chat GPT or clean, it's sort of like magic, like, you know, the expectations are in finite and because, you know, like, supposed to basically not just answer any questions that people have, but also, like, you know, perform like any tasks that they actually ask you to, you know, to do.
25:13And so, so, so for us is not so much about what features are missing, like the big thing that we have to actually keep working on is actually be successful at this core feature, which is like answer people's questions correctly, and answer questions of like higher and higher complexity, correctly for them over time. So we feel like us or anybody else, like out there today, we're all very, very far away from that true vision for our product, which is that we want Gleen to be that AI assistant that can actually answer any questions that you have using your company knowledge that can actually do half of your work for you in the future.
25:56And so I think like I would say maybe 2 % of the way there, AI for all said and done like we're still in very, very early stages of making that impact. So we're only 2 % of the way there. I'd love to ask you actually about Agentec reasoning. It's something that's been on our mind a lot as a partnership at Sequoia And I know it's been on your mind as well as a founder And one of the results that I was really impressed by in the coding space was that you know with rag You I think these coding agents can get to three four percent Completion rates, but if you give them more agentec reasoning capabilities they can get to 14 15 percent So like a multi -fold improvement and you know as simple as you know go reflect on what you just said or best of end or you know what are the techniques are.
Read the full transcript
26:47I'd love to understand how you guys are thinking you know, incorporating more agentic reasoning into your products and anything else to kind of get us from that 2 % where you said we are today to say what you hope to build one day. Yeah. And I want to clarify that 2 % is something that I was making up like as you know it's not not a measured number. Yeah. I just wanted to sort of express how early things are today and how much amazing things we're going to see in the future. I was just basically trying to talk more about that. But in terms of like, an agentic sort of like behavior, One of the things that we are doing on that front is first, try to actually get a lot of input from our users.
27:45So we have a concept of building a workflow inside clean to actually answer a complicated question. And today we actually seek a lot of help from our users in sort of, you know, competing, you know, that workflow. We'll actually, like, say, for example, if you come and ask a question, like, help me write a weekly status report of all the work that my team did. So this is your question. Now, if you think about this question, like, you know, it's complicated. Like, you know, there are a few things you need to do to actually go and really figure out the answer to it. The first thing is you have to understand, what do you mean by your team?
28:25Like who's your team? You do maybe go in your HR system, try to figure out who are the people who report to you. Then we are talking about work. So where does work happen for each one of these team members? You do sort of build an understanding of that, and then go sort of pull a bunch of knowledge from all these different systems. So I think right now what we're doing is we're actually, trying to actually get help from our users, and we will sort of create a plan for a complicated question, try to actually get the user to actually input and tell us, like, you know, we're getting it right. Sometimes, you know, users can actually, like, they can completely ignore what we do and just build a workflow on their own.
29:06And I think that's going to be essential for us to, like, build that, you know, fully -agentic behavior for the future. I think some, you know, you can build agentic behavior for specific narrow set of problems. But in green, since our footprint is so wide, there is the range of questions that people can have. The range of tasks that they want to perform is so broad that we feel like first we have to learn from, like workflows that people are going to actually create manually. And then build these models, which can then sort of take complicated questions in the future and automatically build those, convert them into these sort of like agent -like loop or a complicated workflow.
29:52So that's the approach that we are taking on it. So you're saying since you have such a broad surface area, you can't build agentic reasoning for every single possible task. And so instead you're exposing workflow engine for your users to individually be able to build different automations and different agents. Yeah, and then you learn from it. And then you learn from it. So once you see people building these workflows, that sort of then fits as it goes into a training data set to allow you to actually automatically build new workflows based on complicated questions that people have. So those agent capabilities are coming.
30:33But again, when it is hard for you to answer simple questions, then if you want to do complicated tasks, like it's equally hard, because you're gonna make mistakes. And imagine an agent that actually breaks down a complex task into a series of 10 individual tasks, then your error rate is gonna compound, like you have each step is 90 % accurate. So there is, it is incredible, but I think it is still something that we are, I feel like the human assistance is actually critical like building these complicated workflows. It's also, Arvin, it might be worth saying a word. Maybe this is obvious to people who are listening, but just to say it explicitly, how Act One, which is the enterprise search business, gives you the moral authority or the unfair advantage to get into Act Two, which is the application platform or the platform for agenteic behavior.
31:38It may not be totally obvious to people how Act One leads to Act Two. Can you just say a couple words about that? By building the search product, which immediately adds value to our customers, to our users, we are able to actually solve a bunch of complicated problems that you will typically run into an enterprise.
32:01The first part of that is security. So if you think about the green product, we are actually telling our customers that, hey, give us all of your data. And hopefully do something useful for you after you give that data to us. And that's a big demand. Like, no, it's not easy for companies is to actually trust a new product company and start up and with all of their data, and they're not actually getting any immediate value from it. And so that is one of the things that we've seen to be super helpful to us, because we actually have this search product that people understand that people want, and they want to deploy.
32:46So it's already deployed now, and so Glean is running and it's already connected with all of the enterprise data inside the company, And so then like helping, you know, like, you know, us going to our customers and then saying them that like look, you know Use that as your core AI data platform is it much easier to tell because we don't we not actually have to convince them again to To actually like you know to you know give all of the data to us. It's already there I say this this might not be a perfect analogy, but hopefully it's not a terrible analogy that you know Tesla had an advantage in self -driving because they're already selling cars You guys have an advantage in delivering AI agents because you're already selling a data platform that organizes all the enterprise information, makes it accessible, makes it secure, kind of puts people in a position where they're already asking questions of it, it's kind of a logical next step to ask it to start taking actions.
33:33Absolutely. I think you also announced this out of APIs that let developers build on Glean? Maybe say we're them that, like I think that was in response to customer demands. But what makes developers want to build on Glean versus directly access their own data? I think it's probably a similar effect to what you just talked about. Yeah, so like a lot of like AI applications that our customers are wanting to build They they need to actually tap into data that lives in like multiple different like SAS, you know cloud -based SAS systems and And like I think it's quite tedious for them to like first go and actually bring that data Like you know in one place, you know build a build a search or a reveal there using that that the integrations are hard, understanding permissions and governance is really, really hard on that.
34:24And I think when people, like as these models actually became accessible and developers started to actually develop AI applications, they realized that the, well, they were really excited about building these new cool AI apps, but basically they realized that building an app, the 90 % of the work was actually, this sort of boring infrastructure work that they didn't want to do. Like, you know, bringing data from all these different systems, running this ETL and data pipelines, and then sort of building a good search over it. And like, you know, so you'll spend like so much time before you actually even get to play with AI.
35:03And so that's the thing that, you know, that they find very useful with clean because, you know, we're actually solving all the problems around ETL, building a great search properly, will be in governance within your company. All of the stuff is done for you. You just have a search API, and you can sort of focus all of your attention on the business problem that you're working on and how I can help you sort of achieve that automation that you're looking for. In some ways, all the hard work you've done to ETL and put all the data together with data government governance reminds me a lot of snowflake.
35:39And you know, you're really doing it with like text data and unstructured data. but kind of just that central data platform that companies can build around, build apps on top of reminds me a lot of the snow -like story. Yeah, Arvin, can we ask you a question about future state? If you lost a dream for a few minutes, five or 10 years from now, how do you think Glean is showing up inside of a business? And maybe more importantly, if you're the typical knowledge worker five or 10 years from now, and you are equipped with Gleene. What is your life like then? That's a great question. I think let's keep it five years, instead of 10.
36:20And I think the, well, I mean, one belief that I have is that majority of the work that we do today is not going to be done by us anymore in five years from now. And that applies to me, that applies to you. Like, you know, we both do very different things. But still, like, I think we are knowledge workers. and I think a lot of our work is actually going to be done by these amazing AI assistants. That actually in many ways, you know, more powerful than us. Like, you know, they have access to all of our companies' data or knowledge. They have all the context from all the past conversations and meetings.
36:56They don't forget anything. And they can really sort of, and they have the, you know, on top of that, you know, they have the reasoning capabilities, you know, that allow them to be super helpful to you and like any task that you do. So that's our core belief that, like, you know, majority of our work is actually going to be done by these AI companions or assistants, and we want Gleene to be, you know, that assistant in the workplace. We want Gleene to be, you know, the place, you know, where most of your work happens. One of the things that we also think is going to change is today, A lot of AI is about, you go and seek help from these AI agents.
37:44For example, you go and ask questions, you get answers back. But the future is where this assistance is going to be proactive. If you think about, if you have an executive assistant, they actually help you a lot. A lot of their help is when you go and ask them for help, but a lot of their help is actually proactive in nature. They tell you what to do next. They manage your day, everything about your work life and the guide you to be effective throughout the day. And I think AI is going to actually allow that luxury regardless of who you are. Today, some executives in the company have that luxury, but in the future, like everybody's going to have these really powerful AI -based assistance that are going to actually help them do their work.
38:35So we're really excited about bringing that change to the workplace and hope that Gleene can be that the world's most successful AI assistant. Love it. Arvin, can we change gears a little bit? And I'd love to step back and hear your advice for other founders. You are one of the most successful application level AI companies. I think probably number two behind co -pilot in scale and you did it as a startup as an independent startup. I think you've also had to navigate some unique challenges, right? Like OpenAI, for example, is one of your providers and also one of your competitors, one of your top competitors.
39:11Maybe just tell us what that dynamic is like. Well, first of all, from a point of view of building a start up. In fact, I've been actually coding you guys in many places, like PAD I remember, the slide where you talk about the overall software market being $600 billion, but then AI is expanding that market to 15 trillion or 12 trillion, something massive. That's actually the reality of where we are today, which is that everything that we do is going to actually change, fundamentally change. AI is going to be a key component that's going to drive that change. So first thing that I like as a founder, I don't actually worry about what other people are actually working on.
40:02Because even if all of us are working on a lot of great things, it will still not be enough. It will still not be enough to actually solve all the problems that need to get solved. And so that's the first mindset. So like I think from a advice to other founders, that's the thing I want to tell them. Like if you found a problem, just go work on it and don't worry about if somebody else is solving it. Because the chances are that other people are not, and they won't solve it the same way as you will. But for us, coming back to Gleen specifically, the dynamics for us, we felt the same way. For the first four years of our existence, we were working on a problem where we had no competition.
40:44nobody was actually interested in solving the problem that we were solving. It was a dead market and we had to create a category which I should generate interest, be evangelical. But we knew that we were working on an important problem. But then certainly, Chad GPD happened and search has become hard and now in fact every company that you go and talk to wants to build a is that bad news for us, like how do you think about it? You know, from our perspective, like it doesn't matter, like you know, either way we feel is that, it's actually great news for us. Like now everybody is interested, everybody wants to buy our product, and yes, we have to compete with many, many other vendors, but that's the place where we think, you know, we'll win, because you know, we will, you know, like we have the desire to actually solve this problem, and you know, stay focused on this problem, keep working on it, and there's no reason for us to not do a better job than others.
41:44Part of what I heard in there was that building an AI company is just building a company, find an important problem, and solve it in a compelling way. I'm curious, particularly because this is not your first rodeo. Rubric was obviously wildly successful, and of course you were pretty core to some of the early days of Google. How much of building an AI company is just building a company versus things that are AI specific in some way? So good question. I think of AI mostly as a tool in your arsenal and one of the tools. I don't think you certainly become a different company because you are doing something with AI.
42:30In reality, there's going to be no new company that is not going to use AI technologies in some shape or form. So my point of view is that look, you have to actually find a business problem that you're planning to solve. And hopefully, you can actually solve that problem in a much better way because of the technology that AI is actually providing to you now. So I don't think it actually changes. I don't think it actually feels different. like whether you are, like we don't think of ourselves as an AI company for that matter either. Would you ever train your own models? I guess maybe more broadly, how do you think about where Gleens core competencies start and stop and you know, if you're the 100 R in D chips, where do you want to place them?
43:17We don't have plans to train super large models. But at the same time, we do train models which are smaller in size, for every individual customer of ours. These language models that we train for an individual customer, it sort of goes through all of their own enterprise corpus and sort of starts to understand, the link of the speak, the acronyms, the code names, all of that stuff within their corpus. So model training is actually a core part of the clean core technology, but not in the sense of training a model like GPT -4. We don't do that, we don't have plans for it. We plan to partner with a lot of other great companies that build models of that scale.
44:09Wonderful. Arvind, thank you so much for joining us today. This was a wonderful conversation. We really appreciate it. Thank you for having me.
44:21machine.
From the publisher
Years before co-founding Glean, Arvind was an early Google employee who helped design the search algorithm. Today, Glean is building search and work assistants inside the enterprise, which is arguably an even harder problem. One of the reasons enterprise search is so difficult is that each individual at the company has different permissions and access to different documents and information, meaning that every search needs to be fully personalized. Solving this difficult ingestion and ranking problem also unlocks a key problem for AI: feeding the right context into LLMs to make them useful for your enterprise context. Arvind and his team are harnessing generative AI to synthesize, make connections, and turbo-change knowledge work. Hear Arvind’s vision for what kind of work we’ll do when work AI assistants reach their potential.
Hosted by: Sonya Huang and Pat Grady, Sequoia Capital
00:00 - Introduction
08:35 - Search rankings
11:30 - Retrieval-Augmented Generation
15:52 - Where enterprise search meets RAG
19:13 - How is Glean changing work?
26:08 - Agentic reasoning
31:18 - Act 2: application platform
33:36 - Developers building on Glean
35:54 - 5 years into the future
38:48 - Advice for founders




