Harvey Co-Founder Gabe Pereyra on the Token Pricing Reckoning Coming for AI

18 Jun 2026 · 45 min · 15 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Token pricing reckoning for AI agents, and how Harvey’s open Legal Agent Benchmark (Lab) measures agent performance and cost tradeoffs for real legal tasks.

Guests

Gabe Pereyra, co-founder of Harvey; previously at Google Brain, DeepMind, and Meta. He focuses on scaling vertical AI and agent infrastructure. Nico (interviewer/guest on the later segment) discusses benchmark design and data creation; he has an academic research background and works on Lab.

Key claims

AI usage is exploding, making token-based serving costs rise sharply; users will face “what did my agent cost me $10B?” bills. Law firms can’t rely on one model provider due to conflict and platform risk, so enterprises need routing across models. Benchmarking enables model selection by quality-per-dollar and practice-area performance, not just overall scores.

Notable examples

Lab tasks mimic Big Law work like diligence—e.g., analyzing change-of-control provisions from a data room to generate a diligence memo. Pricing discussion compares Opus 4.7 vs 5.5 and Gemini 3.5 Flash speed/cost tradeoffs. Data set created via agent-generated synthetic legal documents reviewed by former big-law attorneys.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Episode Discussion

0:00 to 14:03
“The models are getting more expensive and they're getting better.”

Transitioning from Seat-Based to Consumption Pricing

14:03 to 17:08

Explore the shift from traditional models to data-driven consumption pricing in AI.

“Like one company is this traditional enterprise seat-based business.”

Transitioning from Seat-Based to Consumption Pricing

17:15 to 18:11

Explore the shift from traditional models to data-driven consumption pricing in AI.

“Nearly 40 % of startups fail because they run out of cash.”

Transitioning from Seat-Based to Consumption Pricing

18:13 to 18:34

Explore the shift from traditional models to data-driven consumption pricing in AI.

“That's why companies like NVIDIA Anthropic, Salesforce, and Gemini partner with Turing.”

Building the Agentic Layer in AI

18:39 to 23:05

Discuss the rapid architectural changes in AI and agent-based solutions.

“I mean, you might have missed the whole SaaS wave.”

Understanding Token Costs and Pricing Challenges

23:05 to 27:02

Examine the complexities of token-based pricing and its implications for users.

“couple things we're thinking through i think the most obvious is just doing better and better routing of models.”

Navigating the Future of AI Pricing

27:02 to 28:00

Analyze future pricing dynamics and the evolution of consumption-based models.

“I've not heard someone kind of break it down that it'll...”

Navigating AI Pricing Pressures

28:00 to 30:04

Learn how market dynamics affect AI token pricing and model competition.

“My agent's going to use a ton of tokens and you're kind of stuck.”

The Importance of Research Diets in AI

30:04 to 31:14

Discover how AI researchers keep track of advancements in the field.

“And yeah, I think that's kind of like the main way.”

Lessons from Influential Figures in AI

32:14 to 34:13

Explore key influences and lessons learned from leading figures in AI.

“I think Winston's definitely one of them.”
Show all 15 chapters

Benchmarking Legal AI Agents

34:13 to 36:39

Understand the significance of benchmarking in legal AI applications.

“So about two weeks ago, we launched Lab, which is our legal agent benchmark.”

Creating Innovative Legal Datasets

36:39 to 38:46

Learn about the unique challenges of generating legal datasets for AI.

“Gabe had emphasized a lot was there was no existing data set for this type of information.”

Evaluating AI Performance Metrics

38:46 to 41:08

Discover how to evaluate AI performance considering cost and quality.

“But I don't think that these like aggregate average performance measures tell the whole story for two reasons.”

The Future of AI in Legal Applications

41:08 to 42:05

Explore the evolving landscape of AI in legal tech and its future potential.

“I also think that the benchmark is a great kind of way to ground research directions that we want to invest in, right?”

The Role of Specialization in AI Development

42:05 to 44:24

Learn how specialization and domain expertise enhance AI capabilities in legal contexts.

“And the harness is basically the way that you define all of these tools, skills, how agents can delegate to other agents to complete tasks.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Gabe Pereyra:The models are getting more expensive and they're getting better. We're just seeing this huge explosion of usage, cost, tokens, everything. We need to talk about tokens. Winston mentioned you guys got up to$13 trillion. We are consuming a ton of tokens for just serving all this stuff. The big misconception right now is I don't think people realize how expensive this is going to get. And they're going to be like, what did my agent do that cost me$10 billion? We launched Lab, which is our legal agent benchmark. It's a benchmark for measuring the performance of agents on real-world legal tasks. We posted our first sort of initial results of how closed-source frontier models like those from OpenAI and Tropic Deepline perform on the benchmark.

0:49Gabe, welcome to Sorcery.

0:51Gabe Pereyra:Thanks so much for having me. Thanks for having us here. We're in the secret speakeasy back room. Exactly. Yeah, I see. So today we're going to walk through the latest release of your benchmarks, which is such a fun topic. Winston was so excited to talk about it, too. So I guess to start, let's break it down. What are you measuring with your benchmarks and where did we get to today? Yeah, so the thing we open sourced was our legal agent benchmark. and I think the thing that we're super excited about is we have done a bunch of benchmarking in the past so before Legal Agent Bench we had Big Law Bench which was more of a chat-based benchmark which was kind of a QA data set and we shared some of that we worked with kind of labs and the providers but at the time we're a smaller company and we didn't have the bandwidth to like fully open source it and support it and so what we wanted to do with Legal Agent Bench was fully open source this so anyone working on agents could run their agents and kind of work with us in the community to evaluate these things.

1:56Gabe Pereyra:And the way I would think of this data set is we tried to really mimic what the coding agent benchmarks did. And the way to think of the coding agent benchmarks is with things like SWE bench, terminal bench, you have a GitHub repo, which is kind of all the code for a project. You have an issue, which is someone said, hey, this doesn't work. Can you fix it? And then you have a bunch of unit tests that will tell you, you know, if you write a bunch of code and run it, the unit test will tell you, you know, you broke this thing or this thing passed. And the way the agent gets scored is, did it write all this code that passed all the unit tests?

2:34Gabe Pereyra:And I think there's a bit of a misconception that, oh, legal is subjective, so you can't do this. And I think the thing for especially big law is a lot of the work you can actually quantify and this is what senior associates and partners are doing when their junior associates do a bunch of this work and so as maybe an example of a task in the legal agent bench it would be can you do diligence which is you get a data room so you're trying to buy a company and you'll get all of that company's contracts and what the law firm needs to do is go through all the contracts and make sure that there's nothing you know questionable or that might blow up the deal in there and so the way we set up this data set like one of these tasks is you have a data room that's all the contracts you have a partner request which is a very high level instead of instructions you would get from a partner if you're an associate and usually for something like diligence it would be hey we got the data room can you look at it right and you would need to infer I know what I need to do based on that.

3:39And then I think part of the novelty here is we use both of those to

3:47Gabe Pereyra:create essentially legal unit tests. And so what are all the things that a partner would be looking for when they're trying to assess an associate's work product? And so you would take that data room and you would generate a diligence memo, which is what the law firm gives to their client to basically say here's all our findings on this diligence and the things you would look for. So one section is going to analyze all the change of control provisions, which is basically if I'm buying a company, there's going to be a change in control of the company. And many contracts have things that can be triggered.

4:21Gabe Pereyra:And most of them don't matter, but there can be ones where this important vendor contract will change in a material way and you need to flag this. And so that would be like a unit test of like you want to check that. And then the challenge with scaling this benchmark is that's just for diligence. And then as you scale this to every practice area that these large law firms help their clients in, you have to think about how to do this for fund formation and IP litigation and capital markets. And so a lot of what we're building is essentially this taxonomy of here's all the tasks that a law firm would do.

4:56Gabe Pereyra:And here's all the subtasks that associates would do and how partners or senior associates would grade that work. And I think what's nice is that framing is kind of the same thing you need to evaluate these agents to do these tasks. Benchmarking makes things a lot more competitive and people love that. I mean, AI is like the most competitive world ever right now. And so I'm really curious on that standpoint, like how are you then like, where does this lead to next? Yeah. So I think the interesting thing and kind of to your point is, okay, we released this benchmark and now we're going to have all of the labs and everyone else building models competing on it.

5:38Gabe Pereyra:And I think there's obviously some risk of, oh, what if some lab providers models are the best? Then why do they need Harvey to do this is maybe one of the potential risks. But what we've seen is actually every different model is good at something different. And so with the initial results we saw, Anthropics models are quite strong, but there's areas where 5.5 is better. There's some areas where open source is better. And increasingly, it's not just which model is the best, it's which model can solve the task at the lowest price point. Because as these models get better, there's kind of intelligent saturation where it's like, okay, for some simple task, I actually want to know that this open source model is good enough.

6:20Gabe Pereyra:And so a lot of the value that we're trying to provide to law firms is you can't just use one model to solve all these tasks because it's getting too expensive. So how do you think about all the tasks that your law firm does and which model you should use from which provider to solve those tasks? I know you've answered this before, but your research partners are your biggest competitors. OpenAI, Anthropic, NVIDIA, we have Base 10. There's like a handful of across categories, right? Right. So why open source it? Doesn't that pose a risk to you? Yeah, I think we think of all of these providers as there's obviously overlap with the labs and, you know, Anthropic and OpenAI have done kind of like Claude for Legal, Codex for Legal.

7:09Gabe Pereyra:But I think the problem that we're solving is different. I think a lot of the products that they're building are individual productivity and the products we're trying to build are organizational productivity for law firms or large in-house departments. And so when we think about open sourcing, I think there's kind of the initial point that I made of most of these law firms and enterprises can't rely on a single lab provider. So like there's a big risk if you're a law firm that if you just use Anthropic or you just use OpenAI, you run into conflict risk. So imagine you're using only Anthropics models as a law firm and you want to represent OpenAI.

7:52Gabe Pereyra:OpenAI is not going to let you send their sensitive legal data to Anthropics models. And so as a law firm, you need to support all the different models if you want to represent OpenAI and Google and Microsoft and all these providers. So there's the conflict risk, there's like platform risk. Like what happens if you pick one provider and they run out of compute, their models fall behind. And so you need to kind of abstract all these things away. And so we think of the providers similar to cloud. And I think you see the same thing in cloud with like Snowflake and Databricks and Datadog, like all these companies built on top of the cloud, but also compete with the cloud.

8:29Gabe Pereyra:So I think it'll look like that. And then I think the bigger point is we don't see the general legal intelligence as where the value of Harvey comes from because most of the valuable legal data can't go into the general purpose models. And so if you think of these large law firms, the work that they're doing is a very sensitive internal investigation, like mergers and acquisitions that would move the market and they can't go into public models. And so we think of our strategy as we want to open source the general stuff and work with all of the providers to make these models as good as possible at general legal.

9:09Gabe Pereyra:And then we want to build infrastructure for these law firms and enterprises that help them own their own models and build their own systems on their unique data. And that's where we think kind of our unique value comes in. And how are you training the models? You can't pull or pool data from your clients. So how are you learning off of them? Yeah, so the way I would think of the big challenge and where law firms are so different than most enterprises, like at Harvey, all of the data in our company for the most part is ours. At a law firm, a lot of the data is all of their clients and you can't mix this.

9:45Gabe Pereyra:And so a lot of what we're starting to work on with our clients is we have a product called Shared Spaces, which essentially lets a law firm and their client collaborate on a legal project. and so you can imagine a world where you have a large private equity firm and their law firm and they're doing all of their fund formation and all of their kind of M &A together in these shared workspaces. Where you can start training these models is how do you use that relationship to build these unique models and so a lot of what we want to figure out is given all of the like ethical wall constraints, regulatory constraints, client data constraints, how can we help law firms use their data in a way that respects kind of all of their client obligations, but still be able to improve their systems based on that.

10:32So speaking on your background, you were at Google Brain, DeepMind, and Meta. What were the biggest fundamental lessons that you learned from each of them that are driving the decisions you're making?

10:43Gabe Pereyra:Yeah, I think especially Google Brain and DeepMind when I was there, this was kind of in 2016, 2017. And so right when deep learning was starting to take off. And I think it was a really cool experience seeing both of those labs because they had kind of opposite strategies of how they approached research, which I think both were correct, but led to kind of different ways of executing on kind of the vision of where they thought AI was going. And so when I was at Brain, it was very bottoms up. It was, I mean, both had some of the smartest AI researchers in the world, but Brain, the approach was, you know, let's get all these smart people, give them a bunch of compute and then kind of let them do their own projects.

11:25Gabe Pereyra:And I think part of the outcome there is you had someone like Noam Shazir and Ashish Viswani and kind of the rest of the folks that invented transformers. And so that approach worked. And then DeepMind was much more top down where Demis just had this vision of, okay, we're going to create AGI. Here's all the things that I think are required. And so they kind of had this tech tree of how they were going to do it. And then let's have all these teams working on research projects that would solve all these milestones to get us to AGI. And I think our approach is a bit more inspired by the DeepMind one, just because I think because we're in a vertical, the end goal is very clear.

12:06Gabe Pereyra:Like we kind of know here's all the legal work that needs to get done. And so you can think about that North Star as, you know, here's all the practice areas that these law firms do. Here's how we would build systems that help them, and do large parts of those with AI. And so it feels much more defined and it's more of an applied problem than it is kind of a pure research problem. But I think seeing that approach really inspired like how I think about these things. How has your role technically evolved over time? I remember I was like, I was just listening to this interview or I forget what it was, but they were talking about Andre Carpathia now.

12:47Like even back at Tesla, like you'd think like these roles back then were like super glamorous and stuff, but he was data labeling all day. But you know, now he's the almighty God of like AI. So like, well, what is the biggest difference over time?

13:02Gabe Pereyra:Yeah, I think it's definitely been interesting at Harvey where it's like when we started the company, I was more focused on research than we should have been at that stage of the company where I kind of had a strong sense that what is happening now was going to happen. And so I wanted to figure out, you know, how do we train models? How do we scale this? But it became very clear pretty quickly, like, okay, first we need to build an enterprise business because, you know, we were a bit ahead of the curve of like, like when I first, when we first started the company, I was pitching people agents.

13:40Gabe Pereyra:And at that, like four years ago, people were like, what are you talking about? And so we quickly had to kind of like think about what is the right product now for customers that we can make for them that's useful, get revenue and scale up the GTM motion, kind of a traditional enterprise SaaS motion. And Winston and I always had kind of this vision of we're going to need to build two companies in parallel. Like one company is this traditional enterprise seat-based business. And then the second will be kind of the transition you're starting to see now as these things as these models get better things move to consumption and i think you'll need both of these because it takes a while to educate buyers on how to buy this like buying off tokens is insanely complicated um it'll probably at some point go back to something like subscription because it's hard to like forecast these things um but i think my role kind of shifted from when i started the company we're trying i was trying to do research and i was like okay we need to do traditional like build a great product build a great enterprise business and then now it's gone back to research where we've put all those things in place.

14:44Gabe Pereyra:I think a lot of the infrastructure has caught up, like de-agent stuff is happening now, that now I'm full-time. And a lot of it is like data labeling and how do we create good data sets, how do we work with all these partners, but I think it's kind of back to doing the things that I really enjoyed when I was at Brain, DeepMind, and Meta. And so I think now is kind of the time where it seems like all these application layer companies, that you saw Composer 2 with Cursor. Like everyone's starting to make this stuff work. And I think that's like super exciting. The inference layer has gotten really hot right now.

15:17What is your general sentiment about what's going on? How is that going to evolve?

15:21Gabe Pereyra:Yeah, I mean, I think it's tied to kind of what I was just saying is like, as this is starting to work in terms of open source models are catching up, I think a lot of the like performance from the models is coming, not just from, you know, doing this large pre-training, but all of this test time compute of, you know, how do you make these models better at reasoning? How do you build good kind of RL environments for them and good harnesses? And so the more that works, I think the more that kind of shifts the ability to application layer companies to do interesting things. And then for example, for us, we work with Base 10 and Fireworks and Together AI, Applied Compute, Trajectory, kind of a bunch of these companies, Ngram, that are doing either inference, RL as a service.

16:06And I think it just feels like the research is starting again, where there's just so many different avenues.

16:15Gabe Pereyra:All of these NeoLabs that we're working with are trying different ways. We're also working with the traditional model providers because they're also making a bunch of progress. And so I think it's just interesting as we're able to build more of these data sets. And even like we have a bunch of customers that want us to figure out better ways to train on their data. I think kind of all of these things are starting to take off and then I think the inference providers you're seeing them do very well where there's customers like us where we do need to move some of our traffic to open source models because it's just becoming super expensive to serve like the largest closed frontier model.

16:52Gabe Pereyra:And so I think it's like a mix of like performance. We have a lot of customers that want to own their own models, own their own data. And like, I think that's kind of also the infrastructure providers will serve that. And so, yeah, I think they're in a good spot. Sorcery is brought to you by Brex, the financial stack trusted by more than 30 ,000 companies, including one in three venture-backed startups in the U.S. Nearly 40 % of startups fail because they run out of cash. Rex is literally built to help founders avoid that. Unlike traditional banks that let your money sit idle, chipping away at it with fees, Rex is designed to help you spend smarter and move faster.

17:32Their all-in-one solution combines checking, treasury, and FDIC protection into one powerful account. You can send and receive money globally at lightning speeds, get 20 times the standard FDIC coverage through their partner banks, and even high yield from day one, with same day and even same hour liquidity. Access your funds anytime. Companies like Scale AI, DoorDash, Service Titan, HIMSS, Anthropic, Flexport, Robinhood, and Plaid trust and use Brex. Start today at brex.com slash sorcery. That's B-R-E-X dot com slash sorcery. Turing is training the next generation of AI with tasks that require real expertise and real world judgment.

18:16That's why companies like NVIDIA Anthropic, Salesforce, and Gemini partner with Turing. Turing builds realistic reinforcement learning environments and data systems based on real operational traces, the kind of infrastructure frontier labs need to train superintelligence. Visit turing.com slash S-O-U-R-C-E-R-Y. I guess to dig even deeper, architecture is changing really fast. I mean, you might have missed the whole SaaS wave. But starting off in AI, it was all chat based. Now we've completely moved to the agentic era, which is like a thousand times more cost competitive. You need compute, power, memory, all these sorts of things that are reorganizing how companies are being built.

19:03And of course, the stock market too. So how are you thinking through the agentic layer and building that out?

19:09Gabe Pereyra:Yeah, so I think this was like a huge transition we've gone through in like the past six months where exactly like you said, the original product we built was kind of this chat based co-pilot for lawyers. And then we built all this product around it that let it work at a law firm on client matters and all these things. And then maybe six months ago when these coding models started getting really good and you could do kind of things like cloud code and codex and that started working really well in the command line. We started building a bunch of the infrastructure of, you know, how do you get these agents that were running in your command line and could execute all these tools and work really well and how to move that into the cloud.

19:52Gabe Pereyra:And I think now you're seeing this with kind of these managed agent solutions. And to your point, it's like now you need sandboxes. the models use way more tokens. We have queries that like we have simple like assistant type queries where you say, you know, draft me a document that a single query can cost$20. We have like a review product where you can upload a hundred thousand contracts and ask the models to review them. And some of those can cost$20 ,000. And so it is just getting incredibly expensive. And I think in parallel, what we're seeing is, I think even two years ago, the models were so good with coding that most programmers would just use them and be like this is super useful I think we're just starting to see in the past six months that inflection for legal where the models can now generate entire documents and they're starting to work in a way where most lawyers who aren't using this technology just for fun they're like this needs to be so much better than the way that I'm used to doing things for me to change my routine to do it that like we're starting to see that absorption.

20:59And so a combination of like the models are getting more expensive

21:03Gabe Pereyra:and they're getting better. They're using more tokens because they're running for longer and they work so much better that users are using them more. We're just seeing kind of this huge explosion of like usage, cost, tokens, everything. We need to talk about tokens. I mean, there's like so many different ways to take this question because there's obviously the margin sentiment on it. And then there's also the internal usage and efficiency on tokens. And I think Winston mentioned you guys got up to$13 trillion. So how are you thinking through that and the cost optimization? And then we can also talk about efficiency internally.

21:42How do you...

21:43Gabe Pereyra:Yeah, so I think there's just kind of that inflection point I was talking about where up until recently, we've always been capability constrained. And so we always wanted to use the largest model because for the most of our tasks and like if you look on the like legal agent benchmark, even the best models weren't good enough to complete those tasks. And so and then the usage, you know, and we weren't at a scale that the usage, you know, it wasn't something like cursor where it was so expensive to serve these things and we didn't kind of have the latency requirements. but that has changed in the past six months where now we are consuming like a huge number of tokens I think for some of the labs we are like the largest consumer of embeddings we are consuming a ton of tokens for just serving all this stuff and so I think it's a good problem to have in the sense of like the usage is so high we obviously need to think through you know how do we serve these customers because if we kind of look at our usage uh like plots right now it's like this is just the start and it's like there's still so much room where it's like if everyone starts using this long horizon agent stuff the way some of our powers are power users are using it um there is just going to be a lot of like cost optimization you need to do and so i think there's like a couple things we're thinking through i think the most obvious is just doing better and better routing of models.

23:11Gabe Pereyra:So like as the largest models get better, not every task needs to be sent to the largest models. And so that's like a great use case of the legal agent benchmark. Increasingly like serving more traffic from open source models, post training models, like a lot of these like very large frontier models are large because they're good at everything. And so I think a lot of the opportunity for these like specific verticals will be, okay, I probably don't need a trillion parameters if I just need the system to be good at diligence. And so how do you build products and models in a way that you can get that frontier performance without kind of all of the frontier costs?

23:48What do you think the biggest question is people are not asking?

Read the full transcript

23:51Gabe Pereyra:I feel like probably the big misconception right now is I don't think people realize how expensive this is going to get. And I don't think people realize how difficult it is going to be for customers to deal with that. Like I think when I talk with most people, they think, oh, just move to consumption pricing and that will solve all the problems. But I think there's going to be this very interesting dynamic where there's kind of a couple ways to price these things. And I think most VCs, what they want to see is like, can you price the work, right? Like, can you sell the value of the work you're selling?

24:30Gabe Pereyra:And then like, that's one end of the extreme. And then the other end extreme is just like price by tokens. And I think the problem you run into of pricing the work is actually the same problem that law firms run into when they try to do fixed fee pricing, right? They're trying to say, hey, here's the fixed rate of this cost. And I think there's going to be this great irony where when we started the company, one of the questions we got asked the most was, you know, what's going to happen with the billable hour? And I think most people's assumption is just, oh, everything's going to move to fixed fee.

25:00Gabe Pereyra:So then these law firms can protect their margins. And I think something people don't appreciate about the billable hour and why it's such a good mechanism is it lets you price incredibly complex work at massive scale in a way that the entire industry can agree on, right? Because all these law firms we've talked to about pricing changes, like whenever we talk to them about fixed fee, they're just like, we just have 10 ,000 clients. We can't negotiate every engagement and price this, and everyone's different. And I actually think something similar is going to happen with this token based pricing where you've seen a bunch of headlines of like the Uber CTO where he's like, we just ran through all of our coding tokens for the year, right, in three months.

25:48Gabe Pereyra:And so all these customers are going to start getting these consumption bills of like$10 million. And they're going to be like, what did my agent do that cost me$10 billion? And if you think about how these law firms solved it, they're like, they got a bill for$10 million from their law firm. And they're like, tell me what every associate did for every six minute increment on the entire project. And they wrote that. And I think you're going to start seeing things like this for token billing, where there's going to become this whole ecosystem of like, how do you optimize around this? Because you kind of have these weird misaligned incentives from the model providers, right?

26:24Gabe Pereyra:Because they're selling you consumption. like they're somewhat incentivized to have their agents use as many tokens as possible and then how do you solve that it's like well we can help you benchmark that and route all these and say oh actually for this task you don't need to use all these tokens and so i think there's going to be like i think this is going to be much more complicated than people realize and it's the same as like it's why it's so hard to price legal work and all of this work where it's just so hard to quantify. Like if I think of our token usage for our programming team, it's like what did they use all these tokens on?

26:57Gabe Pereyra:I'm like they're definitely more productive but it is very hard to like quantify these things. And so I think that will be like a big challenge. Wow. I've not heard someone kind of break it down that it'll... I mean obviously you guys are biased but that is a very rational way to think about monetization for them. Yeah. Yeah. And I think the same way. So I think it's more complicated than that in the sense of like when most people look at law firm billable hours, they're like, oh, the incentives here are completely misaligned. This must lead to bad pricing. But I think what people don't think about is like, OK, the billable hour creates some pricing misalignment.

27:37Gabe Pereyra:But the fact that these law firms need to compete with every other law firm means that that's kept in check. And so it's like this actually converges to roughly the right pricing and so I think the same thing will happen with the model providers where it's like if you look at Opus 4.7 and 5.5 Opus 4.7 is three times more expensive than 5.5 but it's you know 10 or 20 percent more performant but like these are gonna cause pressure on each other and then if open source catches up you're also gonna have pricing pressure and so even though there is like if you just had one model provider and they owned all the models, then they would just be like, here's how much tokens cost.

28:15Gabe Pereyra:My agent's going to use a ton of tokens and you're kind of stuck. But I think this is exactly why you're seeing the inference providers and all the other model providers be successful because you just are going to need a lot of options the same way. That's how you deal with kind of like pricing with professional service providers. But yeah, it is a weird, like there's a bunch of different levels of like the pricing there. So how are you personally keeping track of everything that's going on? Like what is your research diet? Has it changed over time? Yeah I think it's changed a lot. I mean I think when I first started doing research the thing that was so exciting about deep learning is everything was open source and everyone published everything and so I could just go on archive for six hours a day and I would just sit there and read papers all day and you just You could read every breakthrough real time.

29:06Gabe Pereyra:That was the thing that initially attracted me to the field because I was like, this is so interesting and it's so easy to talk to everyone about it because everyone's so open about it. I think that has somewhat changed now, or not somewhat, that has completely changed in the sense like no one really publishes what they're doing anymore. And so it is harder to keep track of the research breakthroughs and what's going on. I think the way that I do it now is we work closely with the labs. A bunch of the people I worked with at Brain, DeepMind, and Meta are at these labs. And so, you know, keeping up with everyone that way.

29:41Gabe Pereyra:And then I think the thing that I feel like I'm most excited about, for example, with this benchmark, is we are going to start publishing some of our research that we're doing with the labs or the other providers and open sourcing more models, more of the work we're doing. And so I think there's hopefully, as open source catches up, there's going to be more and more of this. And yeah, I think that's kind of like the main way.

30:32an idea like AI-powered supply chain companies with positive free cash flow or defense tech companies growing revenue over 25 % year over year. Publix AI then dispatches a swarm of agents that scan every single US stock, evaluates them, and instantly builds a custom index around your thesis. What really stands out is how clearly it explains why each stock is included. And before you invest, you can even backtest your idea against the S &P 500 so you're making decisions with real context, not just guessing. And beyond generated assets, Public lets you invest in stocks, bonds, options, crypto, all in one place.

31:06They'll even give you an uncapped 1 % match when you transfer your investments over from another platform. If you want to build a portfolio that actually reflects your thesis, visit public.com slash sorcery. Paid for by Public Investing. Full disclosures in the description. Enterprise AI runs on Merge, the AI infra platform for integrations, agent tooling, and model orchestration, so your teams ship product, not plumbing. Mistral, Dropbox, and Drada already trust Merge in production. Start building at merge.dev. Founders scale faster on Deel. Set up payroll for any country in minutes, hire anyone anywhere, get visas handled fast, and get back to building.

31:45Visit deal.com slash sorcery. That's D-E-E-L dot com slash sorcery. Well, speaking of performance, I have to do a little ad slot. So Brex, who I think you're a customer of, they're all about performance, spending smarter, moving faster. And so with that question, I usually ask, maybe this is a layer deeper. Who are the people that you most look up to that you've learned from and what are the biggest lessons that you've taken from them? Yeah, that's a good question.

32:16Gabe Pereyra:I think Winston's definitely one of them. I think he is in terms of like his ability to like scale the company, keep track of everything. I think that's something that just watching him grow with the company has been kind of super inspiring. And I think there's like a lot of things that I learned from him in terms of like when I did research, I think what I was good at is how do I focus on just one thing? and I think scaling a company is being able to keep that one thing in your head but then solve thousands of these other problems in parallel and I think as a researcher like I was not good at that I was like good at like okay I'm going to solve like you know I'm going to try to figure out AGI and I'm going to ignore everything else and like you just can't do that when you're like scaling a company and watching Winston's ability to like do that and also kind of build a team and all these things while keeping that in mind I think I've like learned a lot from him doing that um let's see who else I think my like old roommate uh when I was at Brain Barrett's off is someone that like I've learned a lot from he was kind of the best AI researcher that I like worked with when I was at Brain and he's now at OpenAI like ran their post-training team and I think kind of similar ability of like just being able to keep all of the context in his head always kind of know like the right like research directions and things like that i think that was like another person i was super impressed by um and then i think like the obvious ones of like i think jensen is super impressive uh like dennis i think a lot of these people kind of leading the like like satya leading the top like labs cloud providers all the like large players in the i space i think there's like a ton you can learn from from all of these and then um i think kind of the other like top application layer companies like the cursor folks and Brett I think all of these folks are kind of doing super impressive stuff given learning from all of them what are the key traits that you look for in new hires yeah key traits in new hires I think the biggest is like are you just obsessed with the topic like the best hires that we've made it's usually really easy to talk to them about what we're doing because they know so much already and so I remember like with Nico Daniel Spencer like a lot of our like early hires when I would just pitch them here's what we're doing it just they immediately were like that totally makes sense here's a bunch of spin-off idea and we could just like talk forever about this and so I think that's something that I always look for of like usually if I can have these conversations with someone where we're just building off each other like that's always been a strong indicator like I think I think it's somewhat unique but I think Winston and I's like hit rate on like executive hires is like insanely high and it's usually like if we can both have a conversation like that with someone and they're you know background everything's a good fit we're like okay this person is it's gonna work out and I think that's been like pretty accurate and you can usually just feel someone who's like insanely passionate or like obsessed with a topic versus like someone who's like doesn't know what's going on and so i think that's kind of one of the biggest like indicators what are you most excited for this year training some models great answer nika welcome thanks for having me congrats on the recent breakthroughs thank you yeah we're very excited about it can you talk through them yeah Yeah, absolutely.

35:58So about two weeks ago, we launched Lab, which is our legal agent benchmark. It's a benchmark for measuring the performance of agents on real world legal tasks. So we designed the benchmark to basically mirror how legal work is done at law firms. And then today we posted our first sort of initial results of how closed source frontier models like those from OpenAI, bi-anthropic defined perform on the benchmark. So we're excited about what the benchmark offers, not only in terms of making model and agent performance accessible to lawyers, but also a number of different research directions that will spawn from here.

36:38So something that Winston and Gabe had emphasized a lot was there was no existing data set for this type of information. So how did you guys create that and how did you train off of it? Yeah, so this was actually, I think, one of the most interesting and innovative parts of the project. So obviously, legal data is some of the most sensitive in the world. Firms can't just freely share that for something like an open source benchmark. Crazy. Yeah, so we had to get creative with how we constructed the data set. And the way that we did it was actually agent-led with lawyer review. So coding agents, these agentic systems are actually so good at generating synthetic data now that even for non-public documents like certain contract types, etc., they can generate a pretty good first draft.

37:32Obviously not good enough to be lawyer passing. So the way that we approached this was we have a team internally at Harvey that we refer to as Applied Legal Research. They're all former big law attorneys. They come from a variety of different practice areas and subdomains of law. we basically mapped out 24 practice areas that law firms have on their websites that they have lawyers associates kind of assigned to within the firm and they went through all of the the sort of tasks that the firms do or the associates that the firms do within those practice areas and they have the priors right they're lawyers they've done this work before to write out okay for this type of work we use these types of documents and they look like this and this is what a realistic deal kind of looks like.

38:16We had agents generate the data from there and then we sent it to sort of like our larger network of lawyers to review the outputs, the rubrics, document quality, all those sorts of things. But it ends up being a pretty like human in the loop but scalable way to generate data. So for someone reading this how should they be evaluating the benchmarks? So I think there's two things here. One is just making agent performance legible, right? Right. So there's kind of overall performance we have and we'll continue to maintain a leaderboard of just here's how this model, this agent performs on the benchmark.

38:52Obviously, higher score is better. But I don't think that these like aggregate average performance measures tell the whole story for two reasons. One, people aren't really thinking about performance in terms of just like quality maxing anymore. Like I need to quality maxing. Oh, my God. I mean, that's been the story of like the last two years, right? It's just like I'm going to achieve the highest quality possible at whatever cost, whatever latency, whatever burden on my users. But people are thinking about it a lot more now in terms of what is the quality I get for some amount of money spent or some amount of time spent.

39:26And so we're able to sort of dissect model performance in these sorts of ways, right? So you can see what the trade-offs are in quality for a model that costs one-third of the price of our leader, which is Opus 4.7 right now. Or like Gemini 3.5 Flash, it's like seven times faster at completing work than some of the frontier models. So you can look at it that way to make real kind of production agent decisions. And where are you getting, outside of this, where are you getting all of your research information diet? Yeah, so I think there's sort of three. there's three kind of avenues here. One is I did come from research so I still have kind of a fondness for academic research.

40:07Like we obviously have partnerships with the labs. They're not publishing as much but a lot of my old collaborators are still publishing archive papers and NURPS, ICLR, all these academic conferences. Even new conferences are spinning up around agents. Twitter is just like the real time. Really? Twitter? Feed it directly. Yeah I think you have to have a great filter. You have to understand the signal to noise ratio of Twitter. But if you're following folks like Karpathy, Noam Brown at OpenAI, Shalto at Anthropic, some of these kind of leading voices within the labs, as well as some of these academics who have a presence on Twitter as well, it can be helpful.

40:51You talked about the biggest takeaways from the benchmarks, but what are you most excited about? Yeah, so what I'm most excited about is, so we get this cost, we get this latency, we get this practice area breakdown, right? That basically tells us where do we need to invest to improve performance of our product for our customers. I also think that the benchmark is a great kind of way to ground research directions that we want to invest in, right? And so we've seen an insane amount of activity from the open source community, both AI researchers and legal tech, which believe it or not has a bustling sort of open source community, to investigate a number of areas, right?

41:30So like post-training open weight models, I think we're starting to see early results that show that open weight models when post-trained kind of close the gap with the closed source frontier. The agent harness is like the buzziest topic right now. Yeah, tell me more about that. So the harness is essentially the term of art that Twitter has adopted for the infrastructure and scaffolding that goes around the model, right? So you have some model, it's making decisions. You give it some number of tools, like I can read a document, I can write a document, I can do a web search, I can write code. And the harness is basically the way that you define all of these tools, skills, how agents can delegate to other agents to complete tasks.

42:14It's basically the infrastructure and scaffolding around the models. It's a really interesting area of research for us, though, because the thing that we're seeing over and over again is that specialization matters and domain expertise matters. Right. And so we can have our lawyers, the open source community, our AI researchers all working on legal specific skills and tools that make the agents better at these tasks that lab and our customers are sort of experiencing. So you're even more bullish on Harvey now. Yeah, I've been bullish since day one. I think it's only increased. I know. I think I saw you have a bobblehead behind you.

42:51I know they have a fancy term, but yeah, seems like you've been here for a bit. That is one of the most unique three-year anniversary gifts I think I've heard of from a company. But yeah, three years in about a month. Exciting. Okay, so as we close out, how many times are you going to beat this benchmark? I do think the benchmark is now a target. You created your own competition internally. Is this good? We did. My perspective on the entire kind of like AI benchmarking and hill climbing exercise is the extent to which everybody is investing in making models and agents more capable for legal. That is purely to our benefit.

43:29Because actually, like the innovative work that we want to do is not just base model capabilities. And so you asked, what is the thing I'm most excited about? I think this year, one, I do think the benchmark gets saturated within a year at least. But two, I think this year is the year that we see intelligence at an individual level kind of brought into intelligence at an organizational level, right? And so for Harvey, what that means is moving one layer of abstraction up in decision making, like how do lawyers collaborate with lawyers on our platform? How do human agent teams collaborate? And I think there's a lot of really interesting product infrastructure, but also AI problems to explore there that, frankly, we haven't been able to explore yet because the models are not good enough, right?

44:15So if we make the models good enough, then we can do really innovative research, in my opinion. So the collective IQ will go up 50 points. At least, yeah. Okay. Well, Nico, thank you so much. Thanks for having me. Hey, it's Molly. If you enjoy our interviews, check out our newsletter, Sorcery.VC, where we deliver a once-a-week top deals and tech headlines email and also go deeper on our podcast interviews. Subscribe to Sorcery today. And don't forget to subscribe to the podcast on YouTube, Spotify, Apple, or wherever you listen. Link in description to sign up.

From the publisher

Gabe Pereyra is the co-founder and President of Harvey. Before Harvey, he was an AI researcher at Google Brain, DeepMind, and Meta, working on deep learning at both Brain and DeepMind in 2016 and 2017 as the field was taking off. Valued at $11B, Harvey has passed $300M ARR, 960 employees, 2,000 customers, and roughly 13 trillion tokens processed this month.

Harvey has raised over $1.2B to date from Sequoia, Kleiner Perkins, GV, Coatue, Elad Gil, the OpenAI Startup Fund, and GIC, with Sequoia and GIC co-leading the most recent $200M round at $11B.

Niko Grupen is Harvey's Head of Applied Research. Both he and Gabe led the open-sourcing of LAB, the Legal Agent Benchmark, the first open-source benchmark for measuring AI agent performance on real-world legal tasks. LAB covers 1,200+ tasks across 24 practice areas, built with agent-led data generation reviewed by Big Law attorneys. Initial frontier-model results from OpenAI, Anthropic, and DeepMind posted today, with Claude Opus 4.7 leading the leaderboard.

In this episode, Gabe and Niko sit down with Molly in Harvey's San Francisco speakeasy to break down what LAB measures, why Harvey gave the rubric to its biggest competitors, and what the early results show about long-horizon legal agents. Gabe goes deep on the token economics now hitting application-layer AI: single queries that cost $20, contract reviews that cost $20,000, and why he says Harvey is the largest embeddings consumer for some of the labs. He also covers the multi-model strategy, why open-source matters for law firm conflict risk, & the architectural shift from chat-based products to cloud agents.

Niko closes with the methodology behind LAB, how Harvey generates synthetic legal data with agent-led generation and lawyer review, and what's next for the open-source legal research community.

This is the second episode in Sourcery's Harvey series, following the last conversation with co-founder and CEO Winston Weinberg.


𝐄𝐏𝐈𝐒𝐎𝐃𝐄 𝐋𝐈𝐍𝐊





𝐒𝐏𝐎𝐍𝐒𝐎𝐑𝐒

• Brex—The modern finance platform, combining the world’s smartest corporate card with integrated expense management, banking, bill pay, & travel. https://brex.com/sourcery 

• Turing—Turing delivers top-tier talent, data, and tools to help AI labs improve model performance—and enables enterprises to turn those models into powerful, production-ready systems. https://turing.com/sourcery 

• VCX—VCX is the public ticker for private tech, allowing investors of all sizes to invest in venture capital. View The Portfolio at http://GetVCX.com  

• Deel—Deel is the global people platform that helps startups hire, manage, pay, and equip anyone, anywhere. Trusted by more than 35,000 fast-growing companies, Deel is the people platform that just works, so teams can scale without the chaos. Visit: https://www.deel.com/sourcery

• Public–Investing platform Public just launched Generated Assets, which lets you turn any idea into an investable index with AI. With Generated Assets, you can build, backtest, refine, and invest in any thesis with AI. Gone are the days of one-size-fits-all ETFs. https://public.com/sourcery 

• Merge—The leading provider of customer-facing integrations and agentic tools for frontier LLMs, Fortune 500 organizations, and B2B SaaS companies. Visit https://merge.dev  

Follow Sourcery for the latest updates!

https://www.sourcery.vc


Disclosure

Paid Endorsement. Brokerage services by Open to the Public Investing Inc, member FINRA & SIPC. Advisory services by Public Advisors LLC, SEC-registered adviser. Crypto trading provided by Zero Hash LLC, licensed by the NYSDFS. Generated Assets is an interactive analysis tool by Public Advisors. Output is for informational purposes only and is not an investment recommendation or advice. See disclosures at public.com/disclosures/ga. Matched funds must remain in your account for at least 5 years. Match rate and other terms are subject to change at any time.

More from Sourcery

All 190 episodes
Harvey Co-Founder Gabe Pereyra on the Token Pricing Reckoning Coming for AISourcery · 45 min
Listen in VO