In short
Parallel Web Systems is rebuilding web search for agentic AI, arguing that human click data is a “bug.” Instead, agents should provide feedback to improve indexing/ranking, using model research to compress information into fast, high-signal excerpts for LLMs. Parallel launched first as a “search agent” that crawls after a query (trading crawl latency for inference-time compute) to incrementally build an index, then later shipped lower-latency versions (TurboNow; 200ms). It also proposes a new content monetization model for agents using incentive alignment and Shapley values to estimate each publisher’s incremental value.
Guest backgrounds
Parag Agrawal—ex-Twitter CEO (sold to Elon), now founder/CEO of Parallel Web Systems; previously led at a post-product-market-fit scale with large feedback loops.
Key claims
agent feedback replaces click/rating data; agents need different query interfaces than humans; agent search avoids “pre-AI slop” by extracting authoritative excerpts instead of forcing clicks through fast-loading SEO pages; web search will shift from pull to push via background agents.
Notable examples
insurance underwriting/claims, sales data enrichment, finance workflows replacing overnight human curation; SEC filings for authoritative financial numbers; meeting-prep agents that run tens/hundreds of web searches per meeting.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOIntroduction to Parallel and Agentic Search
0:00 to 0:58
Learn about Parallel's approach to search and the importance of agent feedback.
“Our view at Parallel is that human click data is a bug and agent doing work with search should rely on agent feedback, not human feedback.”
The Future of Search and Web Technology
1:23 to 3:08
Explore how Parallel aims to reshape search technology for agents.
“I'm really excited about this conversation.”
Understanding Web Search Challenges
3:08 to 5:00
Discuss the fundamental issues web search faces for humans and agents.
“Let's talk about, let's talk about parallel.”
Reinventing Search with Agent Feedback
5:00 to 7:44
Discover how agent feedback changes the landscape of search technology.
“One way to think about it is it's a billion to billion matching problem, right?”
Building Search Agents and Use Cases
7:44 to 9:46
Learn how Parallel developed search agents for various industries.
“Because I imagine this is one of those things where customers want full coverage day one.”
Optimizing Search Performance
9:46 to 12:34
Understand how Parallel optimizes search results for efficiency and accuracy.
“outsourced human work on top of web data in order essentially to collect evals to run agents to figure out what search for agents should look like in the first place with empirical use cases instead of theoretical evils.”
The Role of Research in Agentic Search
12:34 to 14:00
Examine the research aspects that contribute to enhancing agentic search capabilities.
“So you will do, we're trying to organize information in memory across the memory hierarchy in a way that we can access it fast.”
Optimization in AI Models
14:00 to 15:00
Learn how optimization in AI models can enhance performance and reduce costs.
“And so you can think of it as, but every time, if you use only half the tokens, if your model is context limited or memory limited, you can now do more problems.”
Human vs. Agent Search Behavior
15:00 to 16:18
Discover the differences between human and agent search behaviors and their implications.
“you now give the model the ability to do more.”
Challenges of Web Search
16:18 to 18:32
Understand the challenges presented by low-quality web content and how agents can address them.
“Yeah, I'm like, I feel like I have this incredible, I can just see it coming a mile away.”
Show all 24 chapters
Handling Search Queries in Agents
18:32 to 20:54
Explore the process of how search queries are crafted and retrieved in agent systems.
“And so you can call it pre-AI human slop, or you can call it catering to a lazy human and being successful at SEO, right?”
Quality Control in Search Systems
20:54 to 23:15
Learn about quality control measures and the importance of optimizing search results.
“running at each stage with different architectures, pulling more features to ultimately get down to here are the thousand tokens I want to bring back to this AI, which has the highest signal.”
Future of AI and Search Infrastructure
23:15 to 25:39
Discuss the future of AI in search infrastructure and the evolving relationships between companies.
“So in the super AGI-pilled worldview you, it feels like if you can push on quality across web search and in the model layer, why wouldn't you?”
Search Agents Dynamics
25:39 to 28:03
Examine how search agents operate and their impact on search volume and efficiency.
“and we will see who builds and who buys and who partners and how things evolve.”
Exploring the Role of Agents in Search
28:03 to 29:11
Understand how agents alter search dynamics and multiply search results.
“You want to talk about that dynamic a little bit?”
Building Custom Agents for Efficiency
29:11 to 31:30
Learn about creating custom agents to enhance productivity and automate tasks.
“Now can we, and it relies on a bunch of web data.”
The Evolution of Search Quality and Speed
31:30 to 32:42
Discover how search quality and latency affect user experience with agents.
“So there'll be some rationalization in all of these agents.”
The Future of Search: Human vs. Agent Queries
32:42 to 34:48
Discuss the balance between human and agent-generated queries on the web.
“Do you think there are more agentic queries than human queries on the web now?”
Challenges in Internet Economics and Content Monetization
34:48 to 37:51
Explore the implications of agent-driven interactions on internet economics.
“Yeah, no, I think this was perhaps part of the thing that totally convinced me that this needed to be done by me and build parallel.”
Aligning Incentives in the New AI-Driven Landscape
37:51 to 42:00
Learn about the importance of incentive alignment for content creators in AI.
“and you currently put it on the web, your one available business model pre-parallel was to be in the head and be able to transact with the lab on some fixed fee contract.”
Understanding Shapley Values in ML
42:00 to 44:34
Learn about Shapley values and their implications for collaborative work in machine learning.
“It might be a different result, but as far as the end customer value is concerned, it's the same quality.”
Incentive Alignment in Content Publishing
44:34 to 47:25
Explore how incentive alignment could transform content publishing and owner participation.
“except you can really well estimate them if you build the right kind of data and models around it.”
The Evolution and Naming of Parallel
47:25 to 49:32
Discover the journey of naming the company Parallel and its vision for a dual audience.
“That's very exciting, especially at a time when you see the traffic data, the stack overflow plummeting.”
The Future of Agent-Driven Web Interactions
49:32 to 53:33
Understand the anticipated evolution of web interactions driven by AI agents.
“So there are these very high value portions of the web that have flipped already to being agent first.”
Transcript
Automatic transcript. May contain errors.0:00Our view at Parallel is that human click data is a bug and agent doing work with search should rely on agent feedback, not human feedback. We believe that these models are really good at compressing information and we can benefit from a lot of the research that have gone into building models and apply it to search indexing and ranking. and so you can now make many, many arguments and that's the arguments we made back then that actually now it's way more tractable as a problem because of the existence of agents, not just as in technology, but as a distinct customer.
0:57Parag, thank you so much for joining us today. We're delighted to have you on the show. For those who don't know, Parag of Twitter CEO fame was CEO of Twitter before selling it to Elon and is now back on the founder arc. You founded a company called Parallel Web Systems, which is scaling up agentic search for the agentic web. We're very excited to have you here today to talk about the future of search and the future of the Internet. So let's get started. What's Parallel? Thank you, Sonia, for having me. Thanks, Andrew, for joining us. I'm really excited about this conversation. At Parallel, we're building a bunch of technology in order to allow agents to search and use the web.
1:39So just like humans forever have figured out how to use browsers and search engines and clicking and browsing around the web to access information, agents need to do the same things. We started parallel with the bet that agents would do it a thousand X more than humans ever have. And as a result, we need to reinvent the technology that can power search for agents, but also the business models that go alongside it. And that's what we're doing at parallel. Okay, I want to go back into what makes human and agent search so different. But before we get there, you've told us that, you know, you're unlearning a lot of the lessons that you learned from running Twitter as you build Parallel.
2:20Why is that? Listen, when I was at Twitter in leadership roles, Twitter was a post-product market fit, extraordinarily scaled business where your feedback loops were from the hundreds of millions of customers using the product for 30 plus minutes every day. right? In that world, you operate differently than a pre-product market fit company based on the premise that in a few years, a new customer is going to show up on the internet and let's go build technology for the not yet here customer that we are all learning every day and every week. Fantastic. Okay. Let's talk about, let's talk about parallel.
3:10I have a question, actually. Can I jump in? No, you cannot, Andrew. I'm Andrew. I didn't get an introduction, but I'm also happy to be an inaugural guest on the Training Data podcast. Andrew is making his podcast debut on Training Data. We are honored to have you finally here. Thank you. I feel like all of us in Silicon Valley have a very surface level idea of search. Like, what is search? Like, the job to be done is getting the answer makes sense. but we all know there's crawling and there's an index and there's ranking. But maybe let's take a big step back and explain what is the problem of web search, be it for a human or an agent, and then we can dive into the differences.
3:49The problem of web search, and we all know it and experience it, when we want to find something and we do not know where it is on the web, we go to a search engine like Google, and the search engine then hopefully surfaces the answer to us in the most convenient of locations. That's the base problem. Now in order to do this, what is the search engine doing? The search engine is going and crawling the web, which is finding every URL that's out there, trying to read it, trying to organize that information in what might be called an index. So that by having all of this stuff in memory in one location, you don't have to, once the query comes in, you don't have to spend time loading up pages because you already have them, you've already done pre-processing to organize them.
4:37And then when the query does come in, you quickly understand the query, you find the most relevant results. And then there are many, many, because you're essentially taking hundreds of billions of pages and narrowing it down to what, five or 10 or ideally one in terms of what someone is looking for right now. So you go through these multiple stages of retrieval and ranking in order to surface the most relevant result. So that's the broad problem. One way to think about it is it's a billion to billion matching problem, right? So you have hundreds of billions of pages and hundreds of billions of queries over time.
5:14And you need to figure out how to match make across these two. That sounds like an enormously expensive infrastructure challenge. And I think for the longest time, basically only Google and Bing had, you know, done full web scale crawls and indices. Why did you think about, you know, you as a young company could go off and tackle that problem? It seems like a problem with the giants. So it is expensive in the long term. So what's actually interesting is that when I was starting the company three years ago, you could imagine a world where one, some of the reasons it was difficult for others to compete which was not having access to great feedback in terms of is this a better search result than this?
6:00To collect that feedback at scale, there was a problem around human ratings. There's a problem around click data. You need to access those. Now it is of course expensive to crawl the entire web and index it. But as you think about what agents and the large models that we now have access to have enabled, it's the ratings data can be now created by experts way more cheaply. Our view at Parallel is that human click data is a bug, and agent doing work with search should rely on agent feedback, not human feedback. We believe that these models are really good at compressing information and we can benefit from a lot of the research that have gone into building models and apply it to search indexing and ranking.
6:53and so you can now make many, many arguments and that's the arguments we made back then that actually now it's way more tractable as a problem because of the existence of agents, not just as in technology, but as a distinct customer. And then we figured out a way to go about building this business, which did not require us to go spend all of the money on infrastructure up front before we could service a customer, right? So if you can paint a path of incrementally being able to build an increasingly larger and more sophisticated index over time, as you solve problems for more and more customers, that's the insight that actually convinced me that this was a viable problem for us to take on.
7:44And how do you make that happen? Because I imagine this is one of those things where customers want full coverage day one. So how do you go about making that happen? When we first launched the product, right, we launched and we did not launch a search product first. We launched a search agent product first. Our search agent could go essentially crawl the web after a query arrived. So if you're doing deep research, you have patience to the extent of a minute. And we've had products which sometimes take 10 minutes of research. That's a lot of time to be able to crawl a lot of pages. if only you have enough of a map to know how to prioritize crawling, right?
8:24So you can make up for shortcomings. Like an index is oftentimes, you can think of it as a latency optimization. So if you give up on that dimension, if you're competing with humans, that's why our search agents were competing with the alternative being outsourcing to humans to curate amazing data, right? So we said, it seems like humans sitting on search engines are far way easier to compete with than a search engine on day zero. So by building a product that was a search agent to do real work on top of web data, we were able to incrementally go build our index. Well, there's some examples of what people built with your search agents.
9:07In the very early days, there were search agents being built with us for doing some kinds of insurance underwriting workflows and claims processing workflows. people in sales were doing all kinds of sales data enrichment people in finance who would previously and historically go to build a model collect data by sending it overnight to a set of humans who would curate data that would feed into a modeling exercise they would do would start using us to do that instead. And so we were going initially after replacing where there was outsourced human work on top of web data in order essentially to collect evals to run agents to figure out what search for agents should look like in the first place with empirical use cases instead of theoretical evils.
10:11I see. You're trading off the crawl for inference time compute effectively. Yeah. You guys were started before the term like NeoLab came into existence and you have, you know, web systems product, but you also have like a bunch of AI people. From the outside, before we met for the first time, like it wasn't obvious to me how much of a web systems company or an AI company that you've built. The more we spent together, more time we spent together, the more clear it's like, by some definitions, this is a neolab. Do you want to talk about like where the research side of what you're doing comes in, not just the like infrastructure and system side of what you're doing?
10:48So we don't categorize ourselves as a neolab. Well, of course you would not. We, no, because I think, I don't think our output is a model. Like I somehow, maybe my definition is broken. My definition of a neolab is an output is a model. I think our output is a complement to a model. What we build is something that multiplies on top of a model in order to give either, you can call it the model gets better or the agent built with the model gets better and has superpowers, right? So we've always want to be in a place where whenever someone ships a better model, a Neolab, for example, if somebody else ships a better model, they have a higher hill to climb.
11:31For us, somebody ships a better model, they have now unlocked four more use cases where we can be valuable. Now, whether we need to do work that might be framed as research, that remains. So in that sense, we do have to do research, but we're also not pre-training extraordinary large models. In fact, our job is to compress things down into tiny ranking models, right? Like I think if you go back to my framing around like this 100 billion to 100 billion matching problem. Every query is essentially, give me a thousand tokens from a trillion web pages on the web and make sure they're the right thousand tokens.
12:17Like that's the query that we are getting effectively from to our search engine. And so now what do you have to do? And until last week, we would give ourselves three seconds to throw a bunch of compute at read time to do it. This week, last week, we shipped a product which now does it in 200 milliseconds. So you have now that much time to go figure out how to pick those. So you will do, we're trying to organize information in memory across the memory hierarchy in a way that we can access it fast. We're trying to figure out how to train what model to select the best thousand tokens. And so there's a large amount of research that goes into figuring out how to take a pre-trained model of some kind, adapt its architecture for this new problem and allocate effectively compute in a way that produces the best output with a limited compute slash latency budget.
13:14So a lot of these use cases seem like the deep research kind of shaped use case. And when I think deep research, at least in the earlier incarnations, it was effectively like an agentic loop with the model reasoning and then basically just calling a search engine calling Google or some proxy to Google. Yeah. Right? Why is that insufficient in your eyes versus what I'm hearing from you as agentic search is kind of a net new capability? So calling Google for every query in a deep research thing, it gets you somewhere, right? Okay. If you use parallel search, you will, for the most part, use under half the tokens in your agent.
13:59It will become more accurate and be faster into it. And so you can think of it as, but every time, if you use only half the tokens, if your model is context limited or memory limited, you can now do more problems. You can do the same problems cheaper or faster. So everything to me, when you have a infinite appetite for information and relevant information for all kinds of work, it's at its core an optimization problem around quality, cost, and latency. Every model advancement is about how do you squeeze out more intelligence and then how do you distill it down into keep most of it at a tenth of the cost, right?
14:49And that applies to search as well. And every time you can produce all of the signal with less noise in your search results to give to a model, you now give the model the ability to do more. And so a lot of the bet here is intelligent compute allocation across the model layer and the agent layer down to the search layer. The other interesting thing to observe is the interface actually changes when you think about serving agents versus humans. Humans rely on keyword search. With agents, we've had to innovate quite a lot on what does an agent tell our search engine. The more we know, the better we can do.
15:33Interesting. So humans can only hold a few words in memory, basically, versus an agent. No, I think we can. We're just lazy. Yeah. Right? Like, we can have a conversation. We can hold a lot of words. But when you start typing, we want to tell Google, like, two incomplete words with a typo in there and hope for the best or rely on some form of a drop-down autocomplete thing to avoid typing those three words in the first place. So we're fundamentally lazy. Turns out, like, models less so. fewer typos better specified queries perhaps longer queries less for the search engine to guess what the agent might want uh so you get to solve a different class of problems one other thing about humans is i feel like we know how to decipher the you know pre-ai slop that populates many of the top ranking google results you know for best product for x y and z and you You get all the affiliate advertising things.
16:39And yet, obviously, they rank highly. I'm very susceptible to it. Are you? Yeah, I'm like, I feel like I have this incredible, I can just see it coming a mile away. How do agents deal with things like that, right? Where these are, you know, well-trafficked pages. They have seemingly good answers, but you just know they're not real. Perhaps with parallel, agents have to deal a little bit less with that. Let me tell you why those pages exist in the first place, right? Let's work through a simple example. If you ask for a public company's most recent financials, like just the headline revenue number, we can sit here and know that there exists an authoritative filing with the SEC, which has that number, perhaps on page 73 of a PDF.
17:25But that is the authoritative number, right? Now imagine Google decided, like, I care about authority. and whenever you ask that query that's the first result you see right and then you click that this pdf takes what like three and a half seconds to load instead of the second that another page does you're already frustrated and then you see a earnings page that's the alternative but here you have to now grab your way to page 77 to find your answer we're lazy we're not going to do the work right? So now there exists an entire class of content on the web, which is like, okay, this information is needed by a lot of people.
18:11It is worth putting it on a page that loads fast, where this information is above the fold. All of you can go there. It's 99.99 % right. So you're not like so skeptical when you go there that this is 100 % wrong. And it's added real value in the process because it identified out of the 300 page earnings report to empty bits of information that should be above the fold. And so you can call it pre-AI human slop, or you can call it catering to a lazy human and being successful at SEO, right? The good news is with agents, we're not making the agent click around and fumble around and crap a PDF, right?
19:00We're taking an excerpt from the most authoritative place on the web and trying to bring it to the agent's context window. And so we aren't forged into this weird trade-off. And this trade-off existed with humans in the first place because with standard browsers and protocols and everything, we just didn't figure out how to have us point exactly consistently across everywhere on the web to the exact right highlighted tiny paragraph. With agents, we get to bring that to the agent's context and then let it figure out what's next. Can you walk us through what actually happens when one of your developers sends a query off to your agentic search API?
19:45It's in some sense pretty standard. So we've done some models to figure out what this query is and enrich it to figure out how it will flow into the rest of our system. We have a bunch of indexes which organize different subsets of the web in different ways. And so the first layer will essentially craft queries for each of these different systems. Each one of these systems perhaps is our sort of big index. One of these is perhaps our fresh index. One of these is perhaps like you can, some people will describe it like a knowledge graph. some people will describe it like a structured index. There are a bunch of these, right?
20:23So you're now deciding which ones this query needs to go to. You're figuring out what is the query rewrite for each of these. Then each of these has like a big retrieval layer and then a ranking layer and then more ranking layers. So you're going to try to boil down tens, hundreds of billions of URLs or documents down into thousands, tens of thousands, down into specific excerpts and paragraphs in those tens of thousands with more and more bigger models running at each stage with different architectures, pulling more features to ultimately get down to here are the thousand tokens I want to bring back to this AI, which has the highest signal.
21:10And now if you look at are various versions of our search API. They just throw for different latency and cost constraints, different amounts of compute at various points in this journey to hit those limits. So in the abstract, it's simple, right? What's interesting is what each of these models that I described, how you curate and collect the training data for them, right? How you build those models and optimize them. The index itself, how you use the memory hierarchy to store it to be able to hit certain cost, quality, latency thresholds. What's the North Star from a quality perspective? Like, you know, on Google, there's that, you know, did you get this result you wanted in the first three answers or something?
22:01What is the equivalent North Star for you? So I think of, I don't know if I'm right on this, but like my take is that a billion to billion matching problem is a forever problem. And so the real question is at what point incremental optimization isn't worth the squeeze? So I don't think there is a thing as like, okay, we're done on improving this thing. The question is at some point, it's going to get harder and harder to improve this thing. and it just won't be worth it, right? But I'm hoping that we don't get there actually because like if you think of what we're doing with AI, we will have more intelligence that gets cheaper every few months.
22:49And as a result of it, it will come down to having great models which are cheap, having great information, matchmaking across a need and all of the information available to you whether it's your own or on the web and doing something unique and differentiated with it to produce more knowledge, right? And anytime you can do something 20 % better than somebody else, that might give you an edge. So why wouldn't you? Right? So in the super AGI-pilled worldview you, it feels like if you can push on quality across web search and in the model layer, why wouldn't you? I've heard a point of view that this is so fundamental to the model companies that they're just going to own it.
23:39And in part, because as they're collecting data for pre-training, that's a very, very expensive infrastructure exercise. That kind of is your source of truth for the crawl. What do you think of that? I don't see empiric data on the ground to support that view. To build a fresh web index, I don't think that crawl is particularly useful. And let me frame why. So if you think of what we are building, we're building a complement to models. We like to crawl things that people don't like to crawl for pre-training. because like if the model already was trained on it like it's not useful model companies for the training aren't patient enough to go in a completionist way try to wait for really slow random javascript to load because the number of tokens you get per amount of compute you throw at it is like one order of magnitude so two order of magnitude too bad and so like is it worth the extra effort to get these tokens?
24:49For us it is because we're completionist. For them it's like I'll take X trillion tokens. So that's one. Now I do think this is a core part of every agent. My world view is that if you're buying LLMs for doing work for 9 out of 10 use cases you will want them to have access to the web and create search infra optimizing for agents. So it is a real adjacency for all kinds of LLM inference. And that supports your view that model companies could, should have the best in class infra for it. So now they can build it or they can buy it. And that's the conversation. and we will see who builds and who buys and who partners and how things evolve.
25:48Are you partnering with any of the model companies that you can share? We want to, we will. I can't share anything on that. We did announce, and I don't know your definitions of model companies, we announced today actually that we are working with Google Cloud to be a search and grounding provider for their enterprise agent APIs. So if you think of grounding Gemini models or other models available on GCP, when you build agents on GCP or chat apps on GCP or do any other inference with LLMs on GCP, when you want to attach web search to it, your options are Google search or parallel search. And parallel search is product integrated.
26:38The integration is optimized We've spent time with technical teams and training teams and product teams and commercial teams to make sure that when people use Gemini models with parallel, they get exceptional and great results. So yes, there are these partnerships now emerging. I bet that there'll be several of these. They will all look somewhat unique. That's a big deal from the search king. Congratulations. Google's the original NeoLab. I don't know how to frame Google as a model lab versus a hyperscaler. And I don't know what precise lessons to learn from this one, whether it applies to other labs or to hyperscalers.
27:27And so we will see. The way that I use voice agents and I need to get a reservation for dinner at night, what restaurants should I go to? And then that request gets fulfilled by the agent. When we talk to a lot of the parallel customers, there's this background agent, whether the monitor product, these agents that sort of are always watching the world or watching the web. And when something happens, then they go off and take actions and do something with it. It might be worth because if we think about what does a thousand X more mean, there's like the depth of research. And then there's just like the that, you know, what is actually initiating the tasks?
28:02Is it a human initiating the search or is it an agent itself? You want to talk about that dynamic a little bit? Yeah. So there's a bunch of dimensions here. Let's go back to for a moment on search agents, right? If you run a typical search agent, and even without doing deep research, it'll do somewhere between 5 to 20 searches, even if it answers within a few seconds, which is why not, right? So already, if you transition from using a chat GPT medium or all the categories, but like somewhere not on the high tier, like instant low medium, every time you write a prompt to it in chat GPT, it will do five to 10 searches.
28:48As you dial it up, it'll do hundreds and thousands of searches. So one interesting thing to observe is like a human action to a multiplier on number of searches that happened, right? So just by using an AI app, you're kind of multiplying your way to perhaps one order of magnitude, more searches. now a lot of our initial takes on the product and the market were to go after bigger multipliers than even that right so we were much more interested when you said i have a portfolio of 10 000 small businesses where i have given out credit to for all of them every month I have this human process that runs to feel out how my risk is going up or down.
29:44Now can we, and it relies on a bunch of web data. We're trying to use agents for doing this. So here a developer is effectively, the multiplier there is hundreds of thousands or a million in terms of the number of web searches that happen because a human goes and programs that instead of now doing this process every month, we can do it every week, right? So doing a lot of searches, when this agent runs every week to create a dashboard on the portfolio and a collection of action items that somebody needs to look at. Right. Now you can go to another example, which is even more interesting. I know, do you use something to do meeting prep documents for you all that's an agent?
30:28There's a Sequoia one. There's a Sequoia agent. Yeah. I was going to say James Flynn, but he's one of our great young guys. So I use notions agent and you can build custom agents, which look into all the internal data that we have at parallel, plus all of the web data using parallels APIs to create meeting prep docs. One time I went and created one prompt to build this custom agent. Now it does tens and hundreds of web searches for every meeting I have. Every time I build a new agent for a new use case, that keeps multiplying. So I think the path to these background agents doing more and more and more work all the time for us, it's only going to be bounded by value versus spend.
31:22You know, like I don't think it's rational right now. Like I'm probably spending more on it than I should be. but it's not too much, so I don't care yet. So there'll be some rationalization in all of these agents. But I think we will deploy background agents to the extent that there is incremental value in doing that compute. And the same thing applies to searches. All of these background agents would do a bunch of searches. And so our first set of products were really obsessively focused on these. In part because we were building, growing the index and we decided that our company was based on three dimensions.
Read the full transcript
31:58quality, cost, latency. And for the first couple of years, we said, let's focus, let's ignore latency, and let's just nail the other two. Because optimizing systems, distilling to smaller models, is much more known art than unknown research. So once we achieved the best quality search and search agent products, at every price point, we started working on latency. And that's what we shipped with a product we call TurboNow. It is the fastest, highest quality agentic web search on the market by lot. Do you think there are more agentic queries than human queries on the web now? I don't think yet. Just given some of the multipliers you mentioned and given the background agents, it seems...
32:52Yeah. I think, I don't have to remind you, but we are early. we are very early in agent adoption like you go step outside of our bubble people haven't heard the word fable like so we are very very early i think there are now people like me who are probably operating at the thousand x i don't know what do you think how many google searches a day did you do three years ago before Chargibity. 20, 30? Yeah, I would have guessed something like 20, 30. I think today, if you just look across at all of my agents, I bet they're doing 1 ,000x more than that. Maybe 100 to 1 ,000. If you count some of the things that happen at my company, which isn't assigned to a human, it might easily be more than a thousand x but i think we are the outliers rather than the norm so i think we're very very early on this journey i i do think i recently saw i think it's cloudflare that said that in their monitoring of web traffic the ai traffic is about the same as human traffic in terms of page reads, which is slightly different from searches because it includes perhaps all crawlers that are out there and a bunch of other stuff.
34:23But I think it's going to happen. So maybe there's a good segue to talk about a topic that I know you are passionate about, which is the economics of the internet as we know them. Some of the kind of fundamental assumptions there, you know, human eyeballs, scarcity of attention seem to be falling right in front of us right now. Are the economics of the internet broken now and what's going to happen? Yeah, no, I think this was perhaps part of the thing that totally convinced me that this needed to be done by me and build parallel. I know most people hate ads. I used to do ads and build systems for ads.
35:02I love ads. Oh, wow. I love shopping. okay yeah so you get good ads i don't know if i love ads but i intellectually love ads because ads is the reason that so much amazing content and technology is available for free to all of us right like google search wouldn't be free without ads twitter wouldn't be free without ads and these are truly useful pieces of technology a lot of content on the out on the web out there that we can access for free wouldn't be free if not for ads. And so ads is a very efficient monetization scheme. It's a very efficient monetization scheme because it is exceptional at differential pricing, right?
35:51So most queries, Google loses money on. Some of them make it up and it's an extraordinary business with extraordinary margins. Same with Twitter. Most users... Exactly. You're welcome. I'm subsidizing the free information you guys are getting. You are subsidizing all of us. Thank you. But I think ads is extraordinarily efficient at differential pricing and monetization on the web, which is why it has been a dominant business model. Now, the core assumptions, as you noted, going into it around limited human attention to translate into outcomes. if humans don't show up and their agents show up on the web?
36:32Like, what does this mean? How does the business work? And yeah, I think this was the, if we don't figure out a new business model, we're seeing it already, right? Like people are going to say, okay, I don't want my content to be accessed by an agent because I have a business model for humans. So I actually want to go pay someone to SEO, optimize myself so more humans show up but then if their agent shows up I'm going to cut it off right which seems confusing and discontroting like okay ultimately this agent is acting on behalf of a human but we haven't found business model alignment yeah in a way you can't monetize that visit you can't monetize that visit yeah so let's say you are in the business subscriptions right so you get a bunch of you get a thousand humans and you convert 20 of them into monthly subscriptions, right?
37:29You don't know how that loop works yet with agents. You can't distinguish. You don't have statistics. You don't know if these will lead to subscriptions or they'll just keep stealing your content as nameless agents. And so there are real challenges around the old business models breaking and us not figuring out real scalable new business models, right? So if you own high quality content and you currently put it on the web, your one available business model pre-parallel was to be in the head and be able to transact with the lab on some fixed fee contract. That's literally, which includes some amount of training and liability and then inference time access, right?
38:15That was your one option. That option is not available to most content on the web. It's a very head phenomena. And even for the head, it is a broken business model because when AIs or inference grows, let's say 7x this year and another 7x the next year. On this 50x, their deal size is not growing 50x. Like none of them after signing a two-year deal believes that their share isn't going to decline materially at renewal. And so these are like fixed price constructs in the world of AI inference. And so which doesn't drive for sustainability for all of these businesses. Now, our solution is trying to learn all the lessons from my work at Ads on having transacted.
39:14Like I sold Twitter data to OpenAI. Having transacted on that side to figure out what actually might work and be incentive aligned. So what might work? Efficient differential pricing. Paying for differentially a lot for extraordinarily high value content. Paying a lot for high value work, accessing the same content. So it's differential in both dimensions, quality and value of work being done with it. And a way of doing this scalably and not just with bespoke deals. So those are the properties needed for any reasonable solution. And the biggest property of it all is incentive alignment, right?
40:01Like at what point do people want to collaborate into this enterprise? Why do the model companies need to pay anything at all? You should ask them. My understanding is, one, you want fresh data during inference time to be able to display it in products like chat GPT or CLOT. Two, you want training data. And three, you want some liability protection for training that you already did. And so the payments are some combination of these three things. And I can't be sure of like how they value each of these three. You said, you know, even a very simple query will go do 10 searches. And so how do you do attribution between all the different sources that, you know, boil down to one paragraph response?
40:58Yeah. At parallel, we like to build models. I heard. So, no, I think it goes back to my point around incentive alignment. So what are you, before we try to build a model, let's try to figure out, like, if you were going to try to do this intellectually, theoretically, how do you go do it? So you have to ask the question of, okay, how much incremental value did somebody's content add? So you can run all of these simulation exercises, you take one piece of content out of the corpus, and then you say, let's run the agent, let's see if the quality of our outputs declined, how much, to claw back that quality.
41:39Perhaps if I threw a little bit more compute or a better model in some way, could I claw back that quality? Oh, I could. It cost me a cent. huh, I could get the quality I lost by not having this source. My alternative was to throw more compute. A cent worth of compute to get that same quality. It might be a different result, but as far as the end customer value is concerned, it's the same quality. And so you're like, okay, this source is worth close to a cent, feels like, right? That's intuitive. Now, the formalization of this kind of an intuition is the core framework we use. It's called Shapley Values.
42:23Yes, let's go. What's a Shapley Value? Music's my ears. It's a game theoretic. You're a game theorist? A bunch of game theory, yeah. Yes. Shapley Values is this very theoretical mathematical concept. It is used in Shap Values and feature importances for those who are ML folks here. Let's say the three of us collaborate on something and the whole is bigger than the sum of parts in that moment. The theoretical question is, okay, how do I divide up this sort of bigger pie that we created by collaboration? So all three of us have incentive to collaborate, right? And Shapley values is a mathematical way of effectively answering this question, right?
43:06Now, that sounds amazing, right? Like you could, if you're co-founders, you could figure out how to divide equity. unfortunately it's not that useful it's not that useful because in order to compute Shapley value you need to simulate all words where some subset of the two of us collaborated but the third one doesn't and play out those realities to then impute back to today in terms of how we should divide the pi which in practice you can't do most places in ML models when you do feature importances you kind of can you can hold a feature back run your model and see how well it did. In web search, we can run simulations of if we did not have access to this URL or this domain or this collection of them, how would the agent perform?
43:54We can, if you're good at evals, if you're good at assessing quality, you can build that data by running various scenarios. Collect a bunch of this data and then you can train models to... Your favorite thing. Your favorite thing. So the challenge with Shapley values is computing Shapley values in our context is way more expensive than the amount of dollars we spend on the agent. Forget the amount of dollars we want to pay a publisher, right? So to compute a content owner, to compute that a content owner gets a dollar, if I decide to do the full Shapley value computation, that might take several dollars.
44:32So it doesn't make any sense. except you can really well estimate them if you build the right kind of data and models around it. But we have confidence that our estimations are good and that this is sound theoretically that if there was perfect information symmetry, people would want to collaborate. So the same way in the old ads days, people did second price auctions and believed for better or worse that people would reveal their true bids in an ad auction and end up paying less than that. Once feedback loops get established in a market, like today, everything is like an auto bid, right? Like most people are measuring ROI when they're doing advertising and running on auto bid instead of making up bids.
45:22I think a solid foundation based on incentive alignment that Chaplin MAD drives will ultimately maximize participation of content owners into this, as well as content seekers via AI in an optimized system. And we have some positive evidence to support it. Like we've had some interesting partnerships that we've been able to do and announce. And yeah, like imagine sitting with content owners and explaining Shapley math. Takes a moment. but ultimately once you pull out the properties that you participate in the value if you have unique differentiated data you get paid more if a banker in an expensive job reads your data versus my retired dad reads your data the banker ends up paying more for that read because it's a part of high value work and the macro math also seems to work.
46:31Like if you're going to spend a lot of money on inference on LLMs for knowledge work if we allocated like 2 to 10 % of it to data on the web that's way bigger than all web data business models today. Outside of like walled gardens like Facebook and LinkedIn. So the macro math supports it. It is a scalable approach. And as agents on the web grow an order of magnitude year on year, it's like by my calculations, like we're 12 to 24 months from this math being able to give meaningful dollars for a very wide range of content owners on the web. That's very exciting, especially at a time when you see the traffic data, the stack overflow plummeting.
47:34You see a lot of the human internet as we know it going away because of incentives. That's very exciting to see how you're thinking about incentive alignment for people to keep publishing. Yeah, it's why we started the company. Why is this company called Parallel? The company's original name was Shapley Inc. Really? it was when do you know this i knew that yeah when i incorporated i'm telling you i was obsessing about everything to do with the problem space so while knowing that the first set of also one shapley inc is a terrible name for a b2b product it was not going to be the long-term name so we we but i incorporated as shapley inc for lack of a better word and shapley.ai happened to be available.
48:25Shapley.com is a parked pond domain. So it was not going to be a name. I went around not talking about my company as Shapley and my badges at all events used to call it Nuko or Stealth Company. So, and it took us almost six, eight months to figure out what the real name of the company will be. we end up at parallel in part because it's like one at the time we're doing a lot more in parallel and two we started visualizing this sort of a parallel web built for ais and how its properties are different and this metaphor that when you publish you're now thinking of as we all are now of two audiences okay i'm going to create a page i know humans will read it what should it look like to them and then how should i make sure that agents can read it too so it feels like you're dual publishing to two audiences and so we had this like this parallel web for agents will emerge and so that's why we started liking parallel love it i actually was thinking about this recently with the um you know things like earning transcripts right like i feel very confident that more people are consuming earnings transcripts through agents than are actually certainly listening to the audio and honestly probably been reading the transcript itself.
49:52So there are these very high value portions of the web that have flipped already to being agent first. And obviously the communication has not yet flipped. But I think if I were a public company CEO today and I was doing an earnings report, I would make very clear that the thing that I'm saying will be transcribed and interpreted correctly by the agents, not just by people listening theoretically. Same for us when we're publishing docs for our APIs. It's like our customers are building AI solutions. They are using AI to do it. It's their agents reading our docs and code in our FDKs. It's not humans fumbling around docs pages for the most part.
50:42In fact, for us, the primary audience is an agent and that's how we test our docs. A parallel web for agents. That's very cool. Maybe close us out. Tell us, just give us a snapshot of where Parallel is today and if everything that you hope to build comes true, what does Parallel look like? What does the world look like? What's your role in it? I think of three levels in this journey. Level one is people are building simple agents that use the web more as like a tool, like a web search tool because it's familiar because the first set of agents we built to your point was just like model and give it the same tools because these models have been trained to use all the tools humans have been used to and let them do work by and large if you think of like most work it's there today of our influence it's there today we have now started seeing some subset of customers who are in this world where they're building more sophisticated multi-agent systems which use sub-agents and have agents wake each other up or orchestrate in interesting ways right people are seeing that as sub-agents most familiar one is sub-agents with encoding agent harnesses but we see a lot more of that like if you're building an ai scientist some of those systems are very interesting very sophisticated very long-running and like just throw large amounts of compute and data at a really hard problem.
52:16And I think the third layer for the web specifically is the web goes from pull to push. So today by and large, across the first two modalities I described, either an agent is calling a tool or a sub-agent, but it is telling it a request, saying, go find this for me right now. And I think where we will end up in a couple of years is a variety of use cases will be the web parallel. Call me if this happens so that my agent can do some work or a human can do some work. So I use this line with my team all the time, which is like, imagine like agents are everywhere and they can do a lot of things. Right.
53:06if there is something you can do today and it's worth doing and we still have GPUs available we'll just go do it right like we won't like we won't be like oh let's just do this tomorrow for the sake of it if it can be done today do it so what will we do tomorrow we will do tomorrow work in response to something that circles through either another agent's work or something changes in the world as visible in satellite imagery or some customer commentary that happens, some agent finishing some compute, a human having a new insight to trigger work. But there are going to be a few feeds like this which will drive new agentic work tomorrow.
53:52And one of those feeds is going to be everything that changed on the web, which is really exciting for me because then you're framing not a point in time need but a long term here is what is actionable for me if something like this happens as evidenced by all of the information on the web that is changing all the time call me and then I'll run my agent on it and so we get to then allocate compute onto the entire web all the time on behalf of all the customers and that's really exciting. That's awesome. Really ambitious vision. You're clearly extremely passionate about this and exciting to see you building.
54:40Thank you so much for joining us today, Parag. And thank you for joining us, Andrew. You're very welcome, Sunya. Thanks, Parag. Thank you, Andrew, for your debut.
55:00Thank you.
From the publisher
Parag Agrawal is making a bet that goes against two decades of web search: agents will query the web a thousand times more than humans ever have, and the infrastructure built around human clicks is wrong for them. The former Twitter CEO, now founder and CEO of Parallel Web Systems, explains why Parallel treats human click data as a bug and trains on agent feedback instead. He unpacks the counterintuitive choice to ship a search agent before a search engine, building an index incrementally, and how the new Turbo product cut agentic search to 200 milliseconds. But the problem Parag keeps returning to is economic: the ad-supported internet collapses when agents show up instead of people. His fix draws on Shapley values to pay content owners for the value their pages provide agents, with real dollars reaching publishers, he predicts, within 12 to 24 months.
Hosted by Sonya Huang and Andrew Reed, Sequoia Capital




