Aravind Srinivas - Building An Answer Engine - [Invest Like the Best, EP.363]

5 Mar 2024 · 1 h 9 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Invest Like the Best - Episode 363 Summary

Episode Overview Guest: Aravind Srinivas, Founder and CEO of Perplexity Host: Patrick O'Shaughnessy Duration: 1 hour 12 minutes Air Date: [Insert Date]

In this episode of *Invest Like the Best*, Patrick O'Shaughnessy interviews Aravind Srinivas, the founder and CEO of Perplexity, an innovative startup that positions itself as an “answer engine” leveraging AI technology. They discuss the intricacies of building a competitive search engine, the evolution of AI technology, and future aspirations for Perplexity to revolutionize the search landscape.

---

Key Topics Covered

  1. Defining "Great" in Search
  2. Redefining Search Experience: Aravind envisions search should provide direct answers rather than just links, contrasting with traditional search engines.
  3. "Insanely Great" Search: Inspired by Steve Jobs, Aravind emphasizes the need to create a seamless user experience that anticipates user queries and simplifies the process of asking questions.
  1. Foundations of Perplexity
  2. Initial Development: Started as a project to improve search capabilities on social media, particularly Twitter and LinkedIn.
  3. Transition to Comprehensive Search: The ambition grew to create a powerful answer engine capable of searching the entire internet efficiently.
  1. Mechanics of Perplexity
  2. Behind the Scenes:
  3. User queries are reformulated and relevant web links are retrieved.
  4. An AI model processes this information to provide concise answers with proper citations.
  5. Speed and Latency: Perplexity aimed to deliver faster response times compared to traditional search engines.
  1. Business Model Strategy
  2. Initial Growth Focus: Limited focus on monetization during early growth phases.
  3. Subscription-Based Model: After validating user interest, they adopted a subscription model similar to ChatGPT+, aiming for long-term revenue stability.
  1. Index Building and Data Quality
  2. Importance of Quality Over Size: A focus on building a high-quality index by prioritizing reliable sources rather than sheer quantity.
  3. Challenges in Data Retrieval: The process involves fetching relevant chunks of information and ranking them based on relevance.
  1. Addressing the AI Landscape
  2. Competition Awareness: While Google remains a dominant player, Aravind focuses on differentiating through speed, accuracy, and user experience.
  3. Talent Acquisition: Competition for AI talent is fierce; attracting the right expertise has become a significant challenge for startups like Perplexity.
  1. Future Aspirations
  2. Personalization in Search: Plans to enhance user experience by integrating more personalized features based on user demographics.
  3. Agentic AI Development: Exploring ways to develop AI that can automate decision-making and take actions on behalf of users.
  1. Key Takeaways for Entrepreneurs
  2. Execution Over Strategy: Emphasis on the importance of building a product and iterating based on user feedback before strategizing.
  3. Focus on Unique Value: Startups should carve out unique value propositions instead of trying to compete directly with established giants.
  1. The Future of AI and Search
  2. Vision for Disruption: Aravind aims for Perplexity to revolutionize consumer search categories like shopping and travel, which are currently dominated by traditional search engines.
  3. Anticipating Innovations: He believes that future advancements will come from focusing on reasoning capabilities in AI models and potentially using synthetic data for continuous improvement.

---

Conclusion Aravind Srinivas shares insights into the challenges and opportunities within the AI and search engine landscape. By focusing on user experience, continuous innovation, and leveraging AI's capabilities, Perplexity aims to emerge as a significant player in the future of search technology.

Recommended Actions

  • Explore Perplexity: Engage with Perplexity to experience the answer engine firsthand.
  • Follow Industry Trends: Stay updated on the evolving landscape of AI and search technologies to understand emerging opportunities.

For more details on this episode, check out the full show notes, transcript, and additional resources at [joincolossus.com](https://www.joincolossus.com).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Something I speak about frequently on Invest like the best is the idea of life's work. A more fun way to think about it is that I'm looking for maniacs on a mission. This is the basis for our investment firm, Positive Sum, and it's the reason why I'm so enthusiastic about our presenting sponsor, Ramp. Not only are the founders, Kareem and Eric, life's work level founders, certainly maniacs on a mission. They have created a product that is effectively an unlock for founders and finance team to do more of their life's work by streamlining financial operations, saving everyone their most precious resource, time.

0:30Ramp has built a command and control system for corporate cards and expense management. You can issue cards, manage approvals, make vendor payments of all kinds, and even automate closing your books all in one place. Speaking from my own experience using Ramp for my business, the product is wildly intuitive, simplistic, and makes life so much easier that you'll feel bad for any company who hasn't yet made the switch. The Ramp team is relentless, and the product continues to evolve to save you time that you would never have dreamed of getting back. To me, there is nothing more interesting than technologies that reduce friction for other entrepreneurs to be able to build the thing that they want to.

1:05So much attention has gone to cloud computing, APIs, and other ways of making life easy for founders. What Ramp has done and is doing is build yet another set of tools in this category. To get started, go to ramp.com. Cards issued by Celtic Bank and Sutton Bank, member FDIC. Terms and conditions apply.

1:29Hello and welcome, everyone. I'm Patrick O'Shaughnessy, and this is Invest Like the Best. This show is an open-ended exploration of markets, ideas, stories, and strategies that will help you better invest both your time and your money. Invest Like the Best is part of the Colossus family of podcasts, and you can access all our podcasts, including edited transcripts, show notes, and other resources to keep learning at joincolossus.com. Patrick O'Shaughnessy is the CEO of Positive Sum. All opinions expressed by Patrick and podcast guests are solely their own opinions and do not reflect the opinion of Positive Sum.

2:05This podcast is for informational purposes only and should not be relied upon as a basis for investment decisions. Clients of Positive Sum may maintain positions in the securities discussed in this podcast. To learn more, visit p-s-u-m.vc. My guest today is Erebin Srinivas. He is the founder and CEO of Perplexity, a startup that he describes as an answer engine built from scratch with AI, has set out for Perplexity to become the most powerful answer engine backed by up-to-date sources. He helps me pick apart the technology, describing the behind the scenes of what it takes to build Perplexity to reach its potential and compete along the likes of Google and OpenAI.

2:46Our conversation goes deep into programming this kind of infrastructure, the competition around latency, and constructing a business model around deep learning. There is so much on the horizon, so please enjoy my conversation with Eravind Srinivas. I thought a fun place to begin would be a game I like to play with people, which is I call it the one-minute bio. I'd love to hear the one-minute summary of your life up until the founding of Perplexity, just to set the context for all that we'll talk about. Yeah. I grew up in India from a pretty humble background, was focused a lot on engineering and programming right from the beginning, studied in one of the IITs, got really excited about AI because I happened to do a course on machine learning.

3:33DeepMind published a paper on training AIs to play Atari games. That got me to be resourceful. I got consumer gaming GPU cards from other people on my lab, traded my lab desk for them so that I could stay at the hostel and work on training these neural nets. And in general, when I look back, I've always realized I've been good at making deals, but traditionally you would consider me a nerd just excited about research and programming. My research there got me into Berkeley. I got to do more exciting AI research in Berkeley. that got noticed by institutions like OpenAI and DeepMind, which got me more exposure to the cutting edge.

4:14And at DeepMind, I had a pretty crappy apartment when I was an intern. So I just mostly stayed in the office. And I used to stumble upon books like How Google Works There. And in that, I got really excited about entrepreneurship. And that got me into thinking a lot of starting a company to work on a hard problem. Little did I realize I would literally go on to work on search itself, but that was faith laws irony sort of thing where the people here are excited about you taking on them. I worked at OpenAI and entrepreneurial ambition exceeded what I could do as an individual contributor there. So I left and started Perplexity initially as a small project to just work on searching over Twitter and LinkedIn and things like that.

4:57And as we kept seeing the progress, we just expanded our ambition to just searching over the entire internet, providing a much better experience by giving answers rather than links. What is your theory of good deal-making? Create a win-win situation and don't be greedy. There's this show called Succession. Yeah, great show. There, Logan Roy advises, it's not a deal unless you really screw over the other person. But in Silicon Valley, at least, long-term deal-making is the best deal. When I think about you, I picture that scene in Star Wars where Luke Skywalker is flying the little tiny plane into the Death Star and you're the plane and the Death Star is Google or something.

5:40Google's been considered one of the most unassailable moats in all of business and the product's been ubiquitous forever. We've all used it many times a day for decades now. Tell me about how you conceive of the components of a great search product. To go up against Google, you obviously have a theory of this is what great search should be. And maybe Google has gotten away from that. Just talk us through what great is to you when you think about search. When I think about the word great, I'm reminded about Steve Jobs. Insanely great. You don't want to be just great. You want to be insanely great.

6:14Look, search has always been a hack. 10 blue links was always a hack to get us information. but not needed anymore when we can more or less answer your question directly. So that is the great experience. And what is insanely great then? Insanely great is the AI doesn't even let you struggle to articulate a good question. This is the next part. Of course, you still have to make the first part work reliably, accurately, address a long tail of mistakes. but assume that that's going to get solved with a good amount of engineering. The next part actually to get into insanely great territory is make it so easy to even ask a question.

6:59Why are very few people good podcasters? Why are very few people good interviewers? Because asking good questions is not a skill that most people have. Everybody in the world is curious. Curiosity is unlimited, unbounded. But not all curiosity in every individual can be precisely articulated into good, interesting questions that elicit the most from an interesting mind. And AI is an interesting mind. It is a very knowledgeable mind. It has access to basically all the world's knowledge in an instant. But it's up to you to harness the power. Now, why do we require humans to be great prompt engineers?

7:40Why do you want to sell that vision? The open AI vision of, oh, AIs are amazing. You guys figure out how to be good at using them. But what would Steve Jobs do? Steve Jobs would be like, bring the power of these amazing AIs to the mere mortal. In fact, he's used the word mere mortal in many of these emails. Our job is to bring the joy of personal computing to mere mortals. Those who can see it before others, but democratize the joy. That's what we want to do for knowledge. Bring the joy of learning to everybody. And so that's what we want to address. Either helping people ask questions or start with some dumb version of the question and help the AI refine it for you and profile you enough that suggests interesting questions to ask and have a knowledge feed of questions that you just see because you asked questions about those topics before.

8:38and just every single day make you some X percent smarter. That's what we want to do. It's really cool to think about the sequencing to get there. We've had search engines. Like you said, it's a hack to get to answers. You're building what I think of today as an answer engine. I typed something in, you're just giving the answer directly with great citation and all this other stuff we'll talk about. And the vision you're articulating is this question engine. Anticipate the things that I want to learn about and give them to me beforehand. And I'd love to build up towards that. So maybe starting with the answer engine, explain to us how it works.

9:10Maybe you could do this via the timeline of how you've built the product or something, but what are the components? What is happening behind the scenes when I type something into perplexity, either a question or a search query or whatever? Walk us through in some detail the actual goings on behind the scenes in terms of how the product works itself. Yeah. So when you type in a question into perplexity, the first thing that happens is it first reformulates the question. It tries to understand the question better, expands the question in terms of adding more suffixes or prefixes to it to make it more well formatted.

9:45This speaks to the question engine part. And then after that, it goes and pulls so many links from the web that are relevant to this reformulated question. There are so many paragraphs in each of those links. It takes only the relevant paragraphs from each of those links. And then an AI model, we typically call it large language model. It's basically a model that's been trained to predict the next word on the internet and fine-tuned for being good at summarization and chats. That AI model looks at all these chunks of knowledge snippets that you surface from important or relevant links and takes only those parts that are relevant to answering your query and gives you a very concise four or five sentence answer, but also with references.

10:33Every sentence has a reference to which webpage or which chunk of knowledge it took from which webpage and puts it at the top in terms of sources. That gets you the nicely formatted rendered answer, sometimes in markdown bullets or sometimes just generic paragraphs. Sometimes it has images in it, but a great answer with references or citation so that if you want to dig deeper, you can go and visit the link. If you don't want to and just read the answer and ask a follow-up. You can engage in a conversation. Both modes of usage are encouraged and allowed. What percent of users end up clicking beneath the summarized answer into a source web page?

11:12At least 10%. So 90 % of the time, they're just satisfied with what you give them? Depends on how you look at it. Do you want it to be 100 % of the time? People always click on a link. That's the traditional Google. And do you want it to be 100 % of the time where people never click on links? That's ChatGPT. We think a sweet spot is somewhere in the middle. People should click on links sometimes to go do their work there. Let's say you're just booking a ticket. You might actually want to go away to Expedia or something. Let's say you're deciding where to go first. You don't need to go away and read all these SEO blogs and get confused on what you want to do.

11:45You first make your decision independently with this research buddy that's helping you decide. And once you've finished your research and you have decided, then that's when you actually have to go out and do your actual action of booking your ticket. That way, I believe there is a nice sweet spot of one product providing you both the navigational search experience as well as the answer engine experience together. And that's what we strive to be doing. Can you talk about the relevant pieces of third-party technology that makes something like this possible and how you blend them with your own technology?

12:19So we could talk about retrieval augmented generation here. We could talk about the various lineup of OpenAI versus Cloud versus Google and how you think about the underlying LLMs that you can use in TAP. I'd love to talk about how you think about that as a business person, too. I hear a lot of chat GPT wrapper or something like this. Tell us about the stack that you use. And don't be afraid to get technical, super smart audience. I'm just curious how these things all tie together and work in combination. First of all, we started off with$2 million in funding. When you have$2 million in funding, you have no job trying to build infrastructure yourself.

12:56Yep. Your only goal is to validate if you have a product that people want to use on a day-to-day basis, or at least a weekly basis, and get enough traction and awareness among users. And then think about building infrastructure that allows you to scale from where you got to 10x, 100x more. Now, that is the level to which we were ambitioning at the start. And we decided to be a wrapper. We want to do things that allow you to get the product out as quickly as possible. So we decided to be a wrapper. We connected Bing API with GPT 3.5 API and launched Perplexity. Anybody could have done that, except once you do it, that's when the game begins.

13:41The game begins after you got some excitement and some users are using your product, even after initial hype, and you get the sustained usage. Now that's when you say, hey, look, this thing I wrapped together is not going to scale. Sure, OpenAI is going to continue making 3.5 more scalable. And Bing has already done decades of work to make the search engine API scalable. But the orchestration layer that takes both of these together and handles so many queries and can handle any outage in any of these APIs requires you to build infrastructure yourself. And that's when we started building infrastructure ourselves.

14:18When 10 ,000 people were on the site at once, when Jack Dorsey tweeted about us, even though it's a wrapper, it went down because we don't have the right rate limits with OpenAI. We don't have the right rate limits with Bing or ChantGPD goes down frequently. There are all sorts of issues. AWS servers go down. That's when you start to actually build all the foundation layer of your infrastructure and backend to support a more scalable product. And when you start doing that, you keep encountering newer issues every time. And every time you solve the newer issues, your infrastructure keeps getting more and more robust.

14:54And after a year, if you look back, you're like, oh damn, this is impossible how far we've come. And we would have never imagined we could have built such a sophisticated backend when we started off versus two or three people. Can you explain from an insider's perspective and someone building an application on top of these incredible new technologies, what you think the future might look like, or even what you think the ideal future would be for how many different LLM providers there are, how specialized they get, scale the primary answer. So there's only going to be a few of them. How do you think about all this and where you think it might go?

15:29It really depends on who you're building for. If you're building for consumers, you do want to build a scalable infrastructure because you do want to ask many consumers to use your product. If you're building for the enterprise, you still want a scalable infrastructure. Now it really depends. Are you building for the people within that company who are using your product? Let's say you're building an internal search engine. You only need to scale to the size of the largest organization, which is like maybe a hundred thousand people. And not all of them will be using your thing at one moment. You're decentralizing it.

16:01You're going to keep different servers for different companies and you can elastically decide what's the level of throughput you need to offer. But then if you're solving another enterprise's problem where that enterprise is serving consumers and you're helping them do that, you need to build scalable infrastructure indirectly, at least. For example, OpenAI, their APIs are used by us, other people to serve a lot of consumers. So unless they solve that problem themselves, they're unable to help other people solve that problem. Same thing with AWS. So that's one advantage you have of actually having a first party product that your infrastructure is helping you serve.

16:42And by doing that, by forcing yourself to solve that hard problem, whatever you build can be used by others as well. Amazon built AWS first for Amazon. And because amazon.com requires very robust infrastructure, that can be used by so many other people. And so many other companies emerged by building on top of AWS. Same thing happened with OpenAI. They needed robust infrastructure to serve the GPT-3 developer API and chat GPT as a product that once they got it all right, then they can now support other companies that are building on top of them. So it really depends on what's your end goal and who you're trying to serve and what's the scale of your ambition.

17:22Say a click more about how you view the relative merits of this, what seems like just a pure arms race that's happening. This context window gets longer, this latency gets lower, the sophistication goes up. The generations of models are dizzying. How do you build the business in a way that wins no matter what happens at all these companies or however many LLMs there are to make it forward compatible with a rapidly changing environment? I think there is no solution. You obviously need to have a good engineering team and people who are very nimble and can use and learn new things pretty quickly.

17:56One thing I would say very inspired by Jeff Bezos is the end user doesn't care what models you're using or what indexes you're using. All they want is a great product experience. None of your users are going to say, hey, Arvind, one year from now, I want your product to be slower. Or hey, one year from now, I want your product to be less accurate. One year from now, I want your product to be rendering the answer in these huge paragraphs. I don't want a better format. They're not going to say that they only want these three things to keep getting better and they don't care how you achieve it. So the way we think about this particular question is whatever helps us get there.

18:37Ideally, it's something we build ourselves because usually for speed, you have to build stuff in-house. If you rely on others, at some point you'll hit the limits of how much you can speed up. On the other hand, if you're building the backend in-house, you can go even further in terms of making speed a priority for you. Let me give you an example. There are people who build on the Flutter or React native stack so that there's one single code base for both mobile platforms, iOS and Android. But because you choose to do that to minimize engineering overhead for you, the apps might get slower for the end user, more unusable, more large in terms of memory that it consumes on the device, things like that.

19:21And it just makes for a worse user experience. On the other hand, if you build natively, if you directly go to SwiftUI and build your app, the apps can feel a lot faster, snappier, consume less memory, and allow you to take more advantage of the native components, render more natively. Earlier, we used to render the answer on perplexity by using a render on the web and using the web view to render on the app. But then if you render more natively on SwiftUI and build components for that, it feels even snappier, even better. Small, small optimizations like these, if you keep shipping them every few weeks, the user loves it.

19:58And everyone loves a fast, snappy, accurate, reliable app. And then that increases your retention, your word of mouth and grows. You're now a better company than before. So I genuinely think backwards from what the user wants and try to do whatever it takes to get there. When I think about the history of the product, which I was a pretty early user of, the first thing that pops to my mind is that it solves this hallucination problem, which has become less of a problem. But early on, everyone just didn't know how to trust these things. And you solved that. You gave citations. You can click through the underlying web pages, et cetera.

20:31I'd love you to walk through what you view the major timeline product milestones have been of perplexity dating back to its start? The one I just gave could be one example. There was this possibility, but there was a problem and you solved it. At least that was my perception as a user. What have been the major milestones as you think back on the product and how it's gotten better? I would say the first major thing we did is really make the product a lot faster. When we first launched, the latency for every query was seven seconds that we actually had to speed up the demo video to put it on Twitter so that it doesn't look embarrassing.

21:09And one of our early friendly investors, Daniel Gross, who co-invests a lot with Nat Friedman, he was one of our first testers before we even released the product. And he said, you guys should call it a submit button for a query. It's almost like you're submitting a job and waiting on the cluster to get back. It's that slow. And now we are widely regarded as the fastest chatbot out there. Some people even come and ask me, why are you only as fast as ChatGPT? Why are you not faster? And literally they realized that ChatGPT doesn't even use the web by default. It only uses it on the browsing mode on Bing.

21:45So for us to be as fast as ChatGPT already tells you that in spite of doing more work to go pull up links from the web, read the chunks, pick the relevant ones, and use that to give you the answer with sources and a lot more work on the rendering. Despite doing all the additional work, if you're managing an end-to-end latency as good as ChatGPT, that shows we have even a superior backend to them. So I'm most proud about the speed at which we can do things today compared to when we launched. The accuracy has been constantly going up, primarily due to two things. One is we keep expanding our index and keep improving the quality of the index.

22:23From the beginning, we knew all the mistakes that previous Google competitors did, which is obsess about the size of your index and focus less on the quality. So we decided from the beginning, we would not obsess about the size. Size doesn't matter in index, actually. What matters is the quality of your index. What kind of domains are important for AI chatbots and question answering and knowledge workers? That is what we cared about. So that decision ended up being right. The other thing that has helped us improve the accuracy was training these models to be focused on hallucinations. When you don't have enough information in the search snippets, try to just say, I don't know, instead of making up things.

23:02LLMs are conditioned to always be helpful. Always try to serve the user's query despite what it has access to. May not be even sufficient to answer the query. So that part took some reprogramming, rewiring. You've got to go and change the weights. You can't just solve this with prompt engineering. So we have spent a lot of work on that. The other thing I'm really proud about is getting our own inference infrastructure. So when you have to move outside the OpenAI models to serve your product, everybody thinks, oh, you just train a model to be as good as GPT and you're done. But reality is OpenAI's mode is not just in the fact that they have trained the best models, but also that they have the most cost-efficient, scalable infrastructure for serving this on a large-scale consumer product like ChantGPT.

23:48That is itself a separate layer of mode you can build, tech mode you can build. And so we are very proud of our inference team, how fast, high throughput, low latency infrastructure we built for serving our own LLMs. We took advantage of the open source revolution, Lama and Mistral, and took all these models, trained them to be very good at being great answer bots and serve them ourselves on GPU so that we get better margins on our product. So all these three layers, both in terms of speed through actual product backend orchestration, accuracy of the AI models, and serving our own AI models. We've done a lot of work on all these things.

24:26Can you explain the early discussions that you and your team had about how to set up the business model? Because while the 10 blue links is fundamentally broken, user experience may be from this point forward. It did play very nicely with an amazing ad-based business model. This is a different business model like ChatGPT and you can use a version for free or you can pay for a better version. But I'm really curious how you consider the various different business models. What were the business models that you almost did but didn't do? How do you think about the trade-offs of the business model that you chose?

24:57This is all happening in real time. We're trying to figure out how to build companies and products around this new technology. And I'd love to hear the early story of how you made those decisions. Honestly, I had no idea of business models when we were initially growing. Of course, investors were all, you're growing, keep focused on the growth. Don't worry about making money. when you're growing, people are willing to support you to keep growing bigger and see how far it can grow before deciding what is the business model to stick to. So until we got to late hundreds of thousands of queries a day, we were not even concerned about making money.

Read the full transcript

25:30But at one point, we were really interested to know, hey, are you guys getting all this usage? Because people want to use some chat GPT alternative when it's down or some free GPT-4 usage? We had like 10 queries a day of GPT-4 or something at one point. Are you guys just getting all this usage because of people not wanting to pay for OpenAI or when OpenAI is down, but you don't actually have real product market fit? So that was a question that we were asking ourselves. And it was a valid question. So how to best answer this question? You create a subscription version of your product that has the same pricing as ChatGPT+.

26:14You charge it the exact same thing,$20 a month, and see how many people convert to paying users. They cannot be just paying for GPT-4 because they're getting that in OpenAI as well. And they cannot just be paying for browsing alone because they're also getting that on OpenAI. So despite that, why are they coming and paying for you? They come and pay for you because they like your product experience. They like what you offer. And so that needed to be known or else there's no point raising another round to scale this thing further up and build a business. Because when you don't really have true product market fit and you're just a subsidy, then you shouldn't be raising more money.

26:51There's no real PMF there. And that was what we validated quickly. We put in a subscription plan. We saw how many people converted to paying users. We hardly even tried to convert them. So despite that fact, a lot of people chose to convert and started paying and the business was growing really fast. We just said, the subscription model works. It's not the ultimate model. I don't think that's the final piece in the profits that are going to be generated in the AI chatbot sector, but it's a good start. Everyone is doing that. OpenAI started it. We are doing it. Google is also trying that. Microsoft is trying that.

27:29So it's a good start. It was like something that potentially could be a billion dollars in revenue for us. It's already billions of dollars in revenue for OpenAI. Let's start there. And also when you have sufficient scale of usage, try to think of what advertisements do. What does advertisement in this new medium looks like? How would it even work? How would you not compromise the quality of the answer? How would you not compromise the quality of the citations? Despite that, how can you help creators of content reach more consumers? So that's an interesting challenge to figure out. So we will also work on that and expect the others to also keep thinking about these things.

28:07And the other business is like APIs. We have our own online LLM APIs, which is basically an LLM that has no knowledge cut off. So it's always live, up-to-date, real-time information, unlike the GPT APIs. And we are slowly expanding access to it and letting other people build on it. For example, Rabbit devices are using those APIs. We're also partnering with other devices. Browsers like Arc are using those APIs. So it's small steps towards also building a developer or enterprise focused version of the business, which can make use of all the infrastructure we built, similar to how we use OpenAI's infrastructure.

28:47How does that work? So if you think about it in simple terms, OpenAI is training this huge model up to a date and it's using information available up to that date to train the thing that the knowledge cut off. How do you continue to have something that valuable that is up to date, including today's dates to web pages? How does that actually get built? How do you build that? It's the same thing as a product. You ask a query and it goes and pulls pages from our index and then uses signals from the web to rank it and then gets you back to answer that uses knowledge from these snippets that it pulled up in the form of a concise paragraph.

29:22So it's whatever happens in the product, it's the same thing, except it's been black boxed to you as an API. And then you just send in your request and you get a completion. And you can make it a chat assistant too, so that you can create products that enable this conversational answer engine experience. And people have built WhatsApp assistants using that. You ask a question, grab it, it can hit the API and give you the answer. So we can do a lot more. The reason our APIs are valuable is because nobody else offers this level of speed and accuracy for an end-to-end search plus LLM experience. Can you expand on index?

29:58You've referenced that a few times for those that haven't built one or haven't thought about this. Just explain that whole concept and the decisions that you've made. And you already mentioned quality versus size. But just explain what it means to build an index, why it's so important, etc. Yeah. So what does an index mean? It's basically a copy of the web. The web has so many links and you want a cache. You want a copy of all those links in a database. So a URL and the contents in that URL. Now, the challenge here is new links are being created every day on the web. And also existing links keep getting updated on the web as well.

30:37New sites keep getting updated. So you got to periodically refresh them. The URL needs to be updated in the cache with a different version of it. Similarly, you got to keep adding new URLs to your index, which means you got to build a crawler. And then how you store a URL, the contents in that URL also matters. Not every page is native HTML anymore. The web is upgraded a lot, rendered in JavaScript a lot. And every domain has custom ways to render the JavaScript. So you got to build parsers. So you got to build a crawler, indexer, parser, and that together makes up for a great index. Now the next step comes to retrieval, which is now that you have this index, every time you hit a query, which links do you use?

31:23And which paragraphs in those links do you use? Now that is the ranking problem. How do you figure out what is relevance? Relevance and ranking. And once you retrieve those chunks, like the top few chunks relevant to a query that the user is asking, that's when the AI model comes in. So this is the retrieve part. Now the generate part. That's why it's called Rack, Retrieve and Generate. So once you retrieve the relevant chunks from the huge index that you have, the AI model will come and read those chunks and then give you the answer. Doing this ensures that you don't have to keep training the AI model to be up to date.

31:57What you want the AI model to do is to be intelligent, to be a good reasoning model. Think about this as when you were a student, I'm sure you would have written an open book exam, open notes exam in school or high school or college. What do those exams test you for? They don't test you for rote learning. So it doesn't give an advantage to the person who has the best memory power. It gives advantage to the person who has read the concepts, can immediately query the right part of the notes, but the questions require you to think on the fly as well. That's what we want to design systems. It's a very different philosophy from OpenAI, where OpenAI wants this one model that's so intelligent, so smart.

32:39You can just ask it anything and it's going to tell you, we'd rather want to build a small, efficient model that's smart, capable, can reason on facts that it's given on the fly and disambiguate different individuals with different names or say if there's not sufficient information, not get confused about dates. When you're asking something about the future, say that it's not yet happened. These sort of corner cases handle all of this with good reasoning capabilities, yet have access to all of the world's knowledge in an instant through a great index. And if you can do both of these together end-to-end orchestrated with great latency and user experience, you're creating something extremely valuable.

33:15So that's what we want to build. If you think about the history of the business so far and every episode of what you've had to build, what stands out in your memory as the most difficult period or thing that was built? What would you least want to go back and live through again in terms of its difficulty and stress? Well, I started working in deep learning in 2014, and we were not even doing deep learning in Python at the time. Everything was done with C++ and CUDA. I was using this framework in deep learning called CAFE that literally, if you had to build a different architecture outside of the traditional CNNs, you had to go and write those layers in C++ and CUDA, recompile the library again, because everything's in C++, it has to be compiled.

34:00And after you get a compiled object, you write the new neural net using those layers, create a protobuf file, and just rewrite all the data layers again, and then launch shops. It was a nightmare. Most of the CUDA drivers would have to be reinstalled again for different GPU cards. There was no standardization. So you would probably spend hundreds of hours just installing CUDA and installing these libraries, changing layers, changing the libraries. That the amount of patience and willpower you needed to still do all this to succeed in your research was just crazy. But it's good. It's a good proxy to test if somebody is a good engineer or not, because usually people give up very fast.

34:44I didn't give up. And of course, life got a lot easier once Python-based symbolic languages came, like Tiano from Montreal and then TensorFlow from Google. Then TensorFlow was a pain in the ass too, because debugging it was really hard. Every time you got something wrong, you had to actually change the graph and not be able to print any intermediate things. And then PyTorch came. This is the core AI deep learning stuff that was such a pain when we used to work with. And I would not want to go back to those days, honestly. How would you explain the feeling of the transformer coming online and what that was like to experience as an engineer?

35:21How would you explain what a transformer unlocked to a person that's less technical? Yeah. So what the transformer primarily did is it just made the description length of the architecture of a neural net so minimal. It's very homogenous architecture. Until the transformer, you would have a recurrent layer, a convolutional layer, and a bunch of hidden layers. You would have to be sophisticated to know the right combinations of them. It's almost like you're cooking a meal, but you have to get the right mixtures of so many different parts that any mistakes anywhere could just cost you so much. What the transformer did is one simple model, which is two layers, attention, matmos, attention, and Matmos alternating each other, that it's the same layer repeated again and again.

36:09You just had to decide three or four hyperparameters and that's it. So it just lowered the barrier to entry. You don't have to be a sophisticated neural network expert to design architectures anymore. Instead, the work went more into getting the data right. The architecture problem was solved. you just got to literally scale it up in terms of layers a number of hidden dimensions but that's it more work was spent on getting the tokenizer right the data right how the word is converted into the vocabulary what parts of the internet you're scraping quality of the data which do you leave out which do you train on how do you ablate for what do you evaluate on in terms of how do you know the model is good so that created the different set of skill sets who are more like physics PhDs who had that rigor and experimentation, less background in ML to come and have a big advantage right now.

37:09And that is the core skillset of the Anthropic team. There's a company called Anthropic. They used to work at OpenAI. Their CEO, Dario Amodi, he's actually a physics guy, physics PhD. But he became incredibly skillful for leading teams like this because of his background. And he hired people like that. He hired people who were having physics background to come work with him. And they built all GPT-3 testing for new capabilities, ablating clearly at the smaller scale, forecasting, scaling loss. They brought in this new discipline there. That's what has led to most of the breakthroughs that we see in ChatGPT and all the stuff.

37:46Nobody launches a$100 million run YOLO. You cannot do that. It's most likely going to fail. It's not like how people on Twitter talk, well, why is Google not doing this? Why are they not taking all the data that they have and launching a huge model and just destroying OpenAI? Because you cannot do that. If you just put a lot of data in, model is going to be confused. It's going to look at so much that it's not going to learn any one thing properly. So there is a science towards figuring out the right data mixes at smaller scale, forecasting what will happen if you scale it up, and then rigorously launching larger and larger runs.

38:23And I believe that it's less about being a great transformer designer and more about being a great data expert and experimenter. Do you think that the transformer architecture is here to stay and will remain the dominant tool or architecture for a long time? This is a question that everybody asks in the last six years or seven years since Sederent's transformer came. Honestly, nothing has changed. The only thing that has changed is the transformer became a mixture of experts model, where there are multiple models and not just a single model. But the core self-attention model architecture has not changed.

39:01And people say there are shortcomings, the quadratic attention complexity is there. But any solution to that incurs costs somewhere else. to most of the people not aware that majority of the computation in a large transformer like GPT-3 or 4 is not even spent on the attention layer. It's actually spent on the matrix multiplies. So if you're trying to focus more on the quadratic part, you're incurring cost in the matrix multiplies, and that's actually the bottleneck in the larger scaling. So honestly, it's very hard to make an innovation on the transformer that can have a material impact at the level of GPT-4, complex costs of training those models.

39:42So I would bet more on innovations, auxiliary layers, like retrieval augmented generation. Why do you want to train a really large model when you don't have to memorize all the facts on the internet, when you literally have to just be a good reasoning model? Nobody is going to value Patrick for knowing all facts. They're going to value you for being an intelligent person, fluid intelligence. If I give you something very new that nobody else has an experience in, are you well positioned to learn that skill fast and start doing it really well? When you hire a new employee, what do you care about?

40:14Do you care about how much they know about something or do you care about whether you can give them any task and they would still get up to speed and do it? Which employee would you value more? So that's the sort of intelligence that we should bake into these models. And that requires you to think more on the data. What are these models training on? Can we make them train on something else and just memorizing all the words on the internet? Can we make reasoning emerge in these models through a different way? And that might not need innovation on the transformer, that might need innovation more on what data you're throwing at these models.

40:44Similarly, another layer of innovation that's waiting to happen is the architecture, like sparse versus dense models. Clearly, a mixture of experts is working. GPT-4 is a mixture of experts. Mixtral is a mixture of experts. Gemini 1.5 is a mixture of experts. So even there, It's not one model for coding, one model for reasoning and math, one model for history. That depending on your input, it's getting routed to the right model. It's not that sparse. Every individual token is routed to a different model, but it's happening every layer. So it's still spending a lot of compute. How can we create something that's actually 100 humans in one company?

41:22So the company itself as an aggregate is so much smarter. We not created the equivalent item model layer. more experimentation on the sparsity and more experimentation on how we can make reasoning emerge in a different way is likely to have a lot more impact than thinking about what is the next transformer. I'm curious then to think about bottlenecks in two ways. So bottlenecks specific to perplexity and what it wants to build and your perception of what the bottlenecks are in AI writ large. If you had to answer for both, what do you think the number one bottleneck to progress is in both those cases?

41:56I would say for us perplexity, the main bottlenecks today is just getting reasoning to emerge in smaller models. If that happens, the cost per query is just going to go down tremendously. If you don't need GPT-4 for being accurate, let me give you a rough statistic. It's not actually rigorous. Let's say a model like GPT-3.5 or a mixed trial fine-tuned version of that that matches 3.5 gets 8 out of 10 queries or something like 7 out of 10 queries, no hallucinations. And GPT-4 will get 99 out of 100. The accuracy rate is so much better in the long tail. Now, if I can get a 3.5 or mixed trial model to be as good as 4 at hallucinations, which is basically connected to reasoning capabilities, when you don't have enough information, just say no, deduce it, then that makes a tremendous impact.

42:50I no longer need to serve a large model anymore. And the service can run way more profitably. That sort of a skill is lacking in smaller models. And that connects to the first point I made about making, how do you train models in a different way? So that the most reasoning capable model shouldn't necessarily be the largest model. That hasn't happened yet. And if that correlation breaks, I think it'll have a huge impact. The other impactful scenario in general for the field, not just specific to perplexity, is synthetic data. What happens when all the data on the internet is saturated? You've trained on all of it, that every new data set that's being created on the internet doesn't add a lot of value to it.

43:30It's very marginal. How can you make these models create the next generation of the data for themselves and recursively improve? I'm not talking about scenarios where these models are going to go rogue and start thinking for themselves and take over humanity or something. Very simple experiment where GPT-5 was designed by GPT-4 largely instead of human annotators. Now this will have an impact because I spend a lot on human annotation for hallucinations. I don't have to, I can have the model do it. I can have a smarter model do it for me. So I can spend less and get data annotated faster. I can make improvements on my core models that are sold on production much faster because an AI can look at a million queries, figure out what's wrong, annotate what is wrong, and tell my smaller AI model to train on them.

44:21And I can finish the training run in a week instead of doing it over a month. That way my users get to feel the product get better much faster, accumulate more users that way. And I get more data. The improvements on the product can be tremendously faster. So I think both of these will have a lot of impact, not just on any other startup. synthetic data and reasoning in smaller models. You talked about how in search, speed, latency, and accuracy are obviously two things that you focused on a ton. I'd love to talk about when you think they'll become customized, more context aware of who I am relative to the next perplexity user, how that happens.

44:59And also when they become more agentic, when they can actually start doing stuff for me. Because if you think about the broken 10 links, you could argue the answer engine's broken too. At the end of the day, I just want to have an idea and have an action happen. The end of that idea that I don't have to do. How do you think those next two components of context awareness and agent behavior might start to find their way into models like yours? First, let's start with context awareness. Maybe we call it personalization, more hyper-personalized versions of perplexity. Let's achieve very simple things.

45:32Location, gender, age will already solve a lot of personalization for you for what it's worth. people think you need to literally put all of your activity on the prompt. For what it's worth, these models get confused when you throw a lot of information at them. People advertise long context a lot, but the more you throw at these models in the context, the more confused they get in terms of what to focus on. So personalization can be done when you know exactly what to retrieve from your past and focus more on the highest order bits like location, gender, and demographics, and create a much better experience that's more catered to you than the average user.

46:15I think we can already do this this year, and we will focus on doing that. The second part, agentic versions of perplexity. We believe that's likely to happen very fast. Let's say there's three parts towards taking an action. You do your research, You make your decision and then you take your action. We are doing the first part pretty well, allowing you to do your research. The decision you're still exercising, you want to decide. And the final part, the action, you just want to task the AI as if it was your executive assistant. That's what you want to do today. Now you can go a step further and say, I don't even want to do this decision.

46:51Let the AI decide everything for me. Let the AI do the research for me. Let the AI act for me. I just want to have no agency. I just want to sit and chill, watch Netflix all the time. And AI is like, will work for me. Have meals show up for me. So I think the second part seems more dystopian. I don't want that to happen. Though if people want that to happen, it likely happen. I think the first part we can work towards that once we have models that are better reasoners than GPT-4. Today, I can confidently claim that GPT-4 is not there yet to be a good action bot. That is the biggest reason why the GPT plugin store failed, because it cannot handle all these different APIs calls together at once, and therefore it didn't work.

47:35When you want an action bot to work, it needs to chain a lot of decisions together and handle corner cases. Now, why do we still need executive assistants? The reason you need them is because sometimes you're scheduling something with somebody and they might not have availability for the availability you have. And you might want to move some things around because you might want that to happen that week itself. These kinds of thinking and corner cases, you want to be able to handle. And you don't want the AI to keep coming. Hey, Patrick, that guy's busy. What about this? You want your assistant to think and act on your behalf so that you're able to focus on other things.

48:13Now, we don't have the AIs that can do this today. GPT-4 cannot do this today. Maybe 4.5 can, maybe 5 can, I don't know. When that model is available, definitely we will also add more agent experiences in our product. We are not capable of training those models today. We don't have the budget. We don't have the compute and the talent to do that. I think somebody has to show the proof of existence of such models, and then we can figure out how to get there. And until then, we have our jobs right in front of us to just reduce hallucinations and improve the research part. You mentioned the word talent there, which feels like such an important topic to cover.

48:49What is the talent in AI right now as a leader of a business that obviously wants to and needs to recruit awesome talent? This is certainly the most actively fast moving, exciting area in technology today. It's attracting lots of really talented people, but I'm sure you're supply constrained. There's not enough great AI talent out there. So yeah, just describe the talent. What's it like? How do you participate in it? Anything you can share would be really interesting. Yeah, I would say that if you want to compete on pre-training, a large model like GPT-4, Anthropic Cloud, Mistral, Llama, Gemini, that's it.

49:29Maybe Elon's XAI. A lot of potential, but not done anything major there yet. But that's it. That's over. Everybody who has some chops at doing large-scale training, a lot of rigorous scaling law analysis, data experimentation, are working in these six companies today, competing for building a competitor and a number seven. You just not only have to raise a lot of money and give these people a big cluster, you also have to hire the talent away from them, zero sum at this point. And people don't want to leave because when you don't have anything, when they have peers to work with. And when they already have a great experimentation stack and existing models to bootstrap from, for somebody to leave, it's a lot of work.

50:16You have to offer such amazing incentives and immediate availability of compute. And we're not talking of small compute clusters here. I tried to hire somebody from Meta, very senior researcher. And you know what they said? They said, come back to me when you have 10 ,000 H100s. And you know what? 10 ,000 H100s, billions of dollars over a five to 10 year period. Why would I have it? And then how do I create it? Also, by the way, it's not just about having the money to buy these clusters. You need to make it available today. And there is a supply chain problem. Most of the GPUs are getting booked out one year in advance or two years in advance.

50:54So even if you have the money, you have to wait. By the time you waited and got the money and booked the cluster and got it, the guys that are working here have already made the next generation model. And they're like, look, the world has changed. I'm already in the next generation. I'll come when the next version of the model is finished training. This time you'll come back to me when you have 20 ,800. So it's a game that is being played by seven people right now, six people, largely for a fire, I would say. My hope is that it gets commoditized. It's not a hope in vain. There is some good reasons to why it could happen.

51:33AWS, Azure, GCP, all have incentive to commoditize this and be the biggest winners of all these models, actually, more than OpenAI or Anthropic or Mistral. So the cloud service providers and NVIDIA, of course, all of them have incentive to commoditize these models and have so many other businesses make use of these models and deliver a lot of profits than just a few people eating all the profits. That way they get to win the most. So my hope is that since the cloud service providers are the ones bankrolling all these companies directly or indirectly, they will put it on their clouds for enterprises and then enterprises can take them and post-train them.

52:19I'm talking about a different kind of training. The first type of training I've talked about is pre-training. But pre-training alone is not enough. You cannot take a model and just put it into a product and do nothing. it'll still fail at a lot of consumer use cases, customer use cases. You have to post-train them and address the long tail of issues you get on serving a product. Now there, you actually have a huge advantage if you have a lot of users because you have the data flywheel. So if you have a lot of users and establish the data flywheel and establish all the tooling and evals to constantly make use of your existing data sets to improve your product and build that machine, that flywheel machine, and accumulate enough users and a brand, you have tremendous advantage to create a lot of value.

53:04And we are focused on that. And that talent doesn't need to be as sophisticated or scarce as the pre-training talent. This talent, you can hire people who want to get into AI from other industries, like crypto or like e-commerce, and teach them. And there are so many resources out there. They are fast learners. They even can learn themselves. and they can add a lot of value that way. Forplexity is one of the companies doing that. I'm sure there'll be many more companies there. If you were to oversimplify it and ask the investor community, what is stopping them from funding more AI application companies, not infrastructure, not LLMs, companies more like yours, I think they would say, well, we're worried about defensibility.

53:49We don't know how these companies control their margins and control churn, et cetera, over time, if they're very reliant on underlying infrastructure. I'm not saying this is my opinion, I'm just saying a common take. What would you say to the people that want to build AI applications as they think about their business model, their defensibility, their sustainable differentiation? What advice would you give to other entrepreneurs that want to be successful building apps using AI? I would say that this is going to be an issue until you're a monopoly in your sector. Honestly, I have thought about this too and I've expressed my frustration to a lot of people.

54:27I'm getting asked the same question that I was asked when I had 10X fewer users. When this is ever going to stop? Will it stop when I have another 10X more? And the answer that guy gave a pretty successful billionaire entrepreneur was, no, you will still be asked. You'll continue to be asked. So get used to this. And you'll only stop getting asked when you are the number one. And then what you'll be asked is, oh, look at these smaller guys trying to do the same thing. When are you going to crush them? So it's always going to be a thing. I would say the best answer is always, this is a hard market.

55:04There are big players, but what we are offering is this. Look at what users are saying. Trust the users. That's a signal. It is not unfair for investors to feel like they might be making a mistake. Look at what happened in the Sora text-to-video release that OpenAI did. until then RunwayML and Pico were every investor wanted to piece in those companies. Now like, oh, what should I do? Because until then their mindset was perplexity is such a dumb investment to make because they are directly in the text interface. And OpenAI is not focused on alternate modalities like voice or video. And therefore I'll go and invest and start doing that.

55:46And now does that reasoning apply anymore? I would rather do enterprise chat GPT as an investment because OpenAI doesn't care about enterprise. And then they come and release chat GPT for enterprise. Everything has competition. And I think at one point, you just got to realize that for one company to be doing so many projects at once, like OpenAI is doing today, definitely not all of them are going to succeed. Even if they do succeed, their success doesn't mean you're failure. they are trying to create as much value for themselves. Same thing with Microsoft, same thing with Google. For you, the only thing you can focus on is fast execution and be the best in your sector.

56:25If those two things are not true anymore, if someone else is kicking your ass there, then you have to be worried. But that is true anything you do. You're always going to have competition on anything that can create billions of dollars in revenue. Because for these guys, billion dollars in revenue is what they care about. If you're focused on building something that is only going to be hundreds of millions of dollars in revenue, you're just focusing on building a billion-dollar company. Then you don't take as much venture funding. Do it yourself. Try to be bootstrapped like all the mid-journey guys doing it.

56:53There's a great quote that I've seen you post, which is, the successful warrior is the average man with laser-like focus, which is a great Bruce Lee quote. It sums up everything you just said. There's one more quote, by the way. I don't fear the man who practiced 10 ,000 kicks once. I fear the man who practiced one kick 10 ,000 times. And that's what you're trying to do. If you have only one thing to protect, you go out of your way to be the best at it. If you apply that quote to your own experience, the laser-like focus piece, what have you not done in order to stay focused? What have been some things that you maybe would love to have tried or gone and done, and maybe you'll do them in the future, but things that you've actively said no to in order to stay on that laser-like focus?

57:36Well, we never did image generation. We have image generation as part of the answer. if somebody wants to enrich the answer, but not as a chat experience where someone can just type in a prompt and get an image. And then we've never done free form chat. Everybody said, why don't you just support both modes like Bing does? Dude, I'm just trying to build a search product here. Everyone's like, AI is not meant for information, man. Hallucination is the feature. You should build products where hallucination is a feature, not a bug. When you're building AI chatbots for search, hallucination is a bug.

58:13So you're doomed. You should take advantage of the hallucination being a feature. You should build something like character AI. You should build a GPT store. You should have a travel perplexity like health or like shopping. You should have a store, how it looks on character AI. You should go click on one agent and talk to it. And you should allow people to create their own agents too. And you should work on agents. You should allow people to book restaurants. You should go enterprise. You should allow internal search. Honestly, I give a lot of credit to one of my co-founders, Johnny Ho. He's actually running our product division, basically.

58:48And he was a competitive programmer and a very successful quantitative algorithmic trader too, before starting Perplexity. I would give him even more credit at saying no to things than even myself, because sometimes I'm also tempted to have that founder experimentation energy. And, but the good thing is I'm not very arrogant. So when somebody that's done a lot more thinking about it than me says, no, we should not do this, tend to trust their advice. And of course, there are some times I've not trust, despite the saying, no, we have to go and do this thing. And it's happened. But most of the time I do listen to the Steve Jobs quote of, I'm as proud of saying the things I said no to as I am of the things that I chose to do.

59:37If you think about the future now and where this all might go, I'd love you to paint the biggest possible picture for us in search specifically, not for AI at large, could go any direction, but for you and what you want to build. We're in 2030 or whatever, pick your date. What gets you the most excited? What potential future states get you the most excited? I would say that disrupting search categories where we are currently clicking on a lot of links and having a lot of commercial intent there would be insanely amazing. And I think it's possible to do that. We will be working hard on that. So right now, perplexity is associated with this amazing knowledge assistant that you use for fact checks, trivia, learning about things, digging deeper, but more like a research buddy.

1:00:27That experience needs to expand to the average consumer search categories, shopping and travel and insurance and legal and medicine. There's huge amounts, dollars of advertising thrown at Google for these categories that Google has zero incentive to make these categories good on Gemini, even if they wanted to. And that's what we want to go and disrupt. What gets you the most worried? I would be lying if I said I wasn't worried of OpenAI. I was trying to do similar things to us. Bravado and all that is great, but let's be honest, I'm definitely worried about ChatGPT trying to go more in the direction of search.

1:01:06And the only thing that we have going for us is our speed and accuracy in our UX that is very much better than what they have, at least regarded by many users that way, even if it's not my opinion. If they chose to focus more, it's going to be more like competition there In that case, the differentiation is going to come from us executing even faster and better. Because unlike a big company, if you call them a big company, they're actually pretty fast. So we have to be even faster. So that is one thing I do think about. Unlike what most people say, I'm not really worried about Google. Not because they cannot execute on this.

1:01:41They actually are way better engineers and researchers than us. It is their own business model. Yeah, it's a counter-positioning thing. Yeah. You've just raised a round from some well-known investors and individuals. So I'm sure you have talked to lots of investors that invest in this sort of stuff. If you think about all of those experiences, what do you think investors understand best and least about your kind of company right now? I think what they don't understand, at least a lot of them, is how hard it is to actually create a product like this. their mental models are all in the vertical SaaS era where they found companies that went more verticalized and found customer lock-in effects and then succeeded as a business.

1:02:29What they failed to understand is verticalization might actually be the wrong strategy in AI because one generic bot that can do many things is a lot more valuable to the end user. When people think of AIs, they think of the most generic way of interacting, which is natural language powered. Whereas when you are going verticalized, there are always going to be certain query categories you cannot handle. But you cannot instruct a human user to only interact in a constrained way. You have to design the product that way. It takes a lot of product design and verticalization on the product layer to still let humans interact in the most free-form way, yet have a constrained experience.

1:03:11Nobody has succeeded at this. And further, people don't realize that incumbents in the verticals can just sprinkle an AI chatbot within that app and have all the other layers like product depth and other support like databases, existing vendors on that one single platform already going for them. Most of the things are handled. So that is one thing that took me a lot of time to explain to people in the beginning. Only one particular investor said this to me without even me having to explain this, Mark Andreessen. Mark Andreessen talked to me in January, 2023. And he said, all I'll tell you is when Google came out, there was so much of investor frenzy in investing in Google for a vertical.

1:03:59You don't know any of those companies today because they all shut down, but they got a lot of funding. So don't do that mistake. Everybody's going to tell you to make perplexity a vertical so that in their mind, they're investing in something safe. But if you do that, you're doomed. You better go all the way and go all in, or you just don't do this company. I was like, damn, finally one guy said exactly what I was thinking and happens to be the guy who pioneered the browser. So I got a lot of courage from his advice today. From the outside, it's so interesting to watch a company like yours just as a user, and it's been a blast to use it and see it progress.

1:04:41Is there anything else that we haven't talked about that you think would be the most surprising to the non-builders out there, whether that's investors or users or people that aren't actively building these products themselves. Is there anything surprising about the state of things today that you think we should talk about that we haven't? This is not surprising if you put in more thought into this, but a lot of people are worried about competition and modes, ensuring they don't get destroyed by open AI. Some amount of thought there is very useful. I'm not discarding that at all. But if that is the only way you make your decisions on what to do for your company or your product, you're likely to fail because these guys will do everything under the sun if it's actually valuable worth doing.

1:05:29So as a startup, your only job is to figure out if there's something you can do that has not been done yet already, and that can deliver value to the world. That is truth. Startups are all about finding a truth vector. If it is truth and if it can generate value, there's no reason the existing players don't want to do the same thing too. They will also try. The way you establish your modes, continue to execute faster, make sure that it's not easy for the other companies to do, because it takes a lot of non-trivial work and keep going. On the other hand, if you're like, hey, I want to build a sales AI co-pilot because OpenAI is not going to go stat vertical and Google doesn't care.

1:06:11Microsoft doesn't care. The Salesforce would do it. It's not hard. These are things that HubSpot would do it. So just don't be so naive in the way you think about strategy and modes. And don't spend so much time in the beginning of the company thinking about strategy. Try to iterate. We made a lot of mistakes. We built Texas SQL. It's the dumbest idea you can work on, honestly, because nobody even writes SQL. 80 % of the SQL that actually makes money for Snowflake or Databricks is not even being written. It's just Power BI generated or already written queries that are constantly periodically running.

1:06:45When you start thinking about these, charging someone based on consumption rather than writing new queries, TextoSQL is such a bad idea because once the same SQL is being written, you're not making any money out of it. So usually when you try to think on white paper, a great idea or a whiteboard, it doesn't happen. Very few people have built companies of that nature. And I would say probably all the existing big players were all built with iteration and trying out things and stumbling upon something awesome, and then building the strategy around it. Build the execution muscle first. Don't try to be a great strategist right away.

1:07:23Build something, make sure it has some traction, get the muscle that you can keep iterating, and then you deserve the right to strategize. This has actually been said by this other guy called Frank Slootman as well, the Snowflake CEO. He has this whole line in his book, Amp It Up, where you only deserve the right to strategize once you have earned the track record of execution. And I strongly believe in that. Really, really interesting closing thought. I'm so appreciative of you letting us behind the scenes here into building one of these things very actively. I'm sure it's been stressful and all-consuming and very fun and very interesting.

1:08:01I ask everyone that I interview the same traditional closing question. What is the kindest thing that anyone's ever done for you? We were going through the SVB incident. You remember the bank collapse? Yes. I was supposed to be on vacation that weekend and I was not. Yeah, we were going through that and I was very stressed. And a lot of people checked on me during the time. And Nat Friedman just came and said, I'll give you the money. Don't worry. I'll take care of your payroll. That was very nice of him to do that at the time. Him and Daniel obviously have done an amazing job of supporting this ecosystem.

1:08:34Pretty cool. Yeah, exactly. I had no idea because I was actually in Redmond at the time doing some company visit and this whole thing was going on. I was in the airport before I could get on the flight. I tried to wire the money out to my personal account, in fact, because I didn't even have another account. And I checked with my lawyer. This is okay. He said, this is very emergency. Just get the money out somewhere, man. It doesn't matter. And Nat and Daniel were like, don't worry. Even if that money goes out, we'll fund you. Amazing. Well, thank you so much for your time and for a great conversation.

1:09:02Thank you, Pat.

1:09:32Thank you.

From the publisher

My guest today is Aravind Srinivas. He is the founder and CEO of Perplexity, a startup that he describes as an “answer engine” built from scratch with AI. Aravind has set out for perplexity to become the most powerful answer engine backed by up-to-date sources. He helps me pick the technology apart, describing the behind the scenes of what it takes to build Perplexity to reach its potential and compete alongside the likes of Google and OpenAI. Our conversation goes deep into programming this type of infrastructure, the competition around latency, and constructing a business model around deep learning. There is so much on this horizon, so please enjoy my conversation with Aravind Srinivas. 

Listen to Founders Podcast
For the full show notes, transcript, and links to mentioned content, check out the episode page here.
-----
This episode is brought to you by Tegus, the only investment research platform built for fundamental investors. How hard do you work to get the insights you need to make a great investment decision? How many hours do you spend digging through public records and expert transcripts, or manually updating complex models? Investors should compete on their ability to analyze investments, not how well they aggregate data. That’s why Tegus offers a unified, end-to-end research platform that combines robust qualitative content sets, up-to-date financial data, management and culture checks, and more — all in the same easy-to-use, streamlined user experience. 95% of the top 20 global private equity firms use Tegus. Shouldn’t you? Learn more and get your free trial at tegus.com/patrick.
-----
Invest Like the Best is a property of Colossus, LLC. For more episodes of Invest Like the Best, visit joincolossus.com/episodes. 
Past guests include Tobi Lutke, Kevin Systrom, Mike Krieger, John Collison, Kat Cole, Marc Andreessen, Matthew Ball, Bill Gurley, Anu Hariharan, Ben Thompson, and many more.
Stay up to date on all our podcasts by signing up to Colossus Weekly, our quick dive every Sunday highlighting the top business and investing concepts from our podcasts and the best of what we read that week. Sign up here.
Follow us on Twitter: @patrick_oshag | @JoinColossus
Editing and post-production work for this episode was provided by The Podcast Consultant (https://thepodcastconsultant.com).

Show Notes:
(00:00:00) Welcome to Invest Like the Best 
(00:04:08) First Question - Redefining 'Great' in Search
(00:07:16) The Mechanics of Perplexity
(00:10:59) The Evolution of Perplexity From Wrapper to Robust Infrastructure
(00:13:19) The Future of Large Language Models (LLMs)
(00:22:39) Aravind's Strategy For Constructing The Business Model
(00:28:16) The Process of Building an Index for a Search Engine
(00:35:55) The Impact of Scaling and Data Quality on AI Models
(00:40:00) Bottlenecks in AI Development
(00:47:06) The Talent Landscape in AI
(00:57:59) The Vision for the Future of Search
(01:00:19) Importance of Execution in AI Startups

More from Invest Like the Best with Patrick O'Shaughnessy

All 200 episodes
Aravind Srinivas - Building An Answer Engine - [Invest Like the Best, EP.363]Invest Like the Best with Patrick O'Shaughnessy · 1 h 9 min
Listen in VO