In short
Arena CEO Anastasios Angelopoulos explains why AI model “leaderboards” should be grounded in real-world utility, how Arena evaluates models using user-driven signals, and what the rise of open-source models means for the AI market, competition, and regulation.
Guest backgrounds
Anastasios Angelopoulos is Arena co-founder/CEO; previously a Berkeley statistician and PhD researcher in theoretical ML, later building Chatbot Arena as a research project. Host Juven (Kleiner Perkins partner) and guest host Mamoun Hamid co-host the interview.
Key claims
API “lock-in” is overstated; open-source models could deliver ~90% of the work at ~10% of the cost, pressuring closed vendors. Arena’s public leaderboard is free/marketing, while labs pay for evaluation dashboards on their model checkpoints. Arena’s rankings move beyond human preference to task completion, steerability, hallucination rate, and factuality via automated claim extraction and web evidence checks.
Notable examples
“Pelican riding a bicycle” benchmark satire; Kimi topping web development leaderboards; GPT Image 2 vs 1.5 nearly 100% win rate; “Leaderboard Illusion” paper alleging bias.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe AI Landscape and Revenue Growth
0:00 to 1:16
Discussion of the rapid growth in AI revenue and the potential for vendor switching.
“Anthropic revenue has been booming, booming, booming on the API.”
Origins of Arena and Its Growth
1:50 to 3:11
Anastasios shares the early days of Arena and its evolution from a project to a significant company.
“I didn't think I was going to be in tech.”
Evaluating AI Models and User Feedback
3:11 to 4:35
Insights on how Arena evaluates AI models using user feedback and real-world performance.
“And then I remember what happened was that at the beginning was just open AI.”
The Importance of Real User Value
4:35 to 9:06
Exploration of how user interactions shape the evaluation of AI models and the benchmarks used.
“Yeah, and we had like, you know, 20 models.”
Understanding Arena's User Base
9:06 to 11:24
Discussion on the demographics of Arena's users and their primary use cases.
“The point is that the user incentive and the platform incentive are perfectly aligned.”
Shifts in AI Development and Leaderboards
11:24 to 13:11
Anastasios discusses the rapid pace of AI development and changes in leaderboard dynamics.
“It tends to be more knowledge worker professionals that are using RENA.”
Emerging Trends in AI Competition
13:11 to 14:01
Anastasios shares insights on the competition between American and Chinese AI models.
“You know, the easy answer would be that the pace of development is just extraordinary.”
The Economic Value of Web Development
14:01 to 15:10
Explore the significance of web development in the current tech landscape.
“In a very important category, by the way, most software developers are doing web development.”
The AI Leaderboard: A Shift in Dominance
15:10 to 16:28
Discuss the emerging dominance of Chinese AI models over American ones.
“there is room for a trillion dollar American company based on open source.”
Challenges Facing American AI Companies
16:28 to 18:52
Analyze the struggles of American AI firms in competing globally.
“They completely imploded because of some issues that they were having internally and restarted.”
Show all 28 chapters
Navigating Human Preferences in AI Evaluation
18:52 to 22:34
Understand the criticisms and improvements in AI evaluation methods.
“Um, but, but like a company like OpenAI Anthropic has built such a mature pipeline around hardware, around data.”
The Data-Driven Future of AI Companies
22:34 to 25:08
Learn about the future revenue models and data utilization in AI.
“And so now all the rankings that we have on Arena go far beyond human preference.”
The Journey from Academia to CEO
25:08 to 28:00
Hear about the transition from PhD student to leading a successful AI company.
“that are being generated within your company.”
Navigating People Problems as a CEO
28:00 to 29:50
Learn how a CEO balances people management and technical work.
“And that's my job as CEOs, deal with people's problems.”
AI vs. Electricity: Transformational Technology Debate
29:50 to 31:48
Discover the arguments regarding AI's potential to surpass the impact of electricity.
“Like, I'd be very curious your argument for why.”
The Decoupling of Labor and Capital
31:48 to 34:29
Explore the implications of labor becoming detached from capital in capitalism.
“I'm curious, what's your reaction to that argument?”
Assigning Value in a Changing Economy
34:29 to 37:52
Understand how the value assignment may change with AI and autonomous agents.
“And so I think that is what we all have to sort of grapple with as a society is how do we live in this world where more and more of the value of the world will accrue to potentially fewer companies.”
The Future of AI Model Improvements
37:52 to 40:08
Examine the rapid advancements in AI models and their implications for the industry.
“Not necessarily even at the rate that they have.”
The Economics of AI Applications
40:08 to 42:00
Discuss the potential economic shifts in AI applications and their value creation.
“and that certain businesses might have all their capabilities be able to be done by the mini light version of whatever the next Kimmy is?”
The Rise of App Layer Companies
42:00 to 43:49
Discussing the increasing demand for app layer companies in various industries.
“And we already see that the army of lawyers that AI lawyers, or whatever we're going to call it, at Harvey is the use just continues to go through the roof or take any app layer company.”
Regulation and Open Models
43:50 to 45:02
Exploring the potential regulation of open-source AI models and its implications.
“you know it especially as open source unless like we all just lock down open source from china right Which, I don't know.”
Navigating Rapid Growth
45:03 to 47:35
Insights on managing a rapidly growing company in the AI space.
“Um, yeah, I think we could probably do with three or four.”
Personal Challenges and Company Responsibilities
47:36 to 52:58
Sharing personal struggles and the impact of family during business stresses.
“I would think about it more like, how do we put the right people in the right place at the right time and make great products?”
Future Plans for Arena
52:59 to 55:20
Overview of upcoming priorities and innovations for Arena.
“Can I ask you about, can I ask another question?”
Agents and Ecosystem Expansion
55:21 to 56:00
Discussing the development of agents and their role in Arena's ecosystem.
“of course um what does the rest of this year look like for you and the company yeah a lot of big things coming for Arena.”
Building AI Routers and Agent Products
56:00 to 57:29
Learn about Arena's approach to AI and its focus on developing routers and agent products.
“It includes a a bunch of sub-models and sub-agents.”
Hiring Insights and Company Vision
57:29 to 58:42
Discover the roles Arena is hiring for and the vision behind their AI ecosystem.
“Infrastructure engineers, AI engineers, those are probably the most valuable.”
Defining Grit and Personal Resilience
58:42 to 59:32
Explore the concept of grit and how it relates to perseverance in the face of challenges.
“When you hear the word grit, what do you think of?”
Transcript
Automatic transcript. May contain errors.0:00Anthropic revenue has been booming, booming, booming on the API. But to some extent, it's also easy come, easy go. If you are spending that much money that quickly, there's no chance that you're not able to switch vendors. And so if an open source model comes by and it is able to do 90 % of the work at 10 % of the cost, which is totally feasible that that might happen over the next couple of years, I think that's a pretty structural problem. People are assuming a degree of lock-in to APIs that I don't think is necessarily going to exist. We work with labs to help them improve their models. The product that we give to them is basically an analysis and insights product called an evaluation.
0:33In the process of developing models, Laz will come to us and say, hey, we have like 50 checkpoints this week. How good or bad are these at different types of tasks? Is this good at front end? Is this good at full stack? Is it bad on tool calls? We take those checkpoints, we evaluate them on our data, which represents real usage, and then we hand those results back to them as an insights dashboard. And so this is the product that scales us to over$100 million in revenue. This has proven so useful to the labs. How can we take an arena and put it within every business in America, in every business globally?
1:15Welcome to Grit. I'm Juven, partner at Kleiner Perkins, a show where we go beyond the highlight reel and explore the personal and professional challenges of building history-making companies. Today on Grit, I'm talking with Anastasios Angelopoulos, co-founder and CEO of Arena, the platform the entire AI industry watches to answer one question. Which model is actually best? Anastasios went from a Berkeley statistician to building Chatbot Arena as a research side project to running a company that hit$100 million in annualized revenue this June at a$1.7 billion valuation. My partner, Mamoun Hamid, joins me as guest host.
1:50You were born here. I was born here. Yeah. So, L.A. kid. L.A. kid. LA's the best place to grow up. You know, it's like... That was awesome. A lot of art, a lot of music. Yeah. I didn't think I was going to be in tech. Well, the first time we met was actually on the Berkeley campus. That's right. In a... I don't know if I should call it charming or depressing little rectangular box. Yes. With five people. The, like, yeah. We were, like, in the basement of Soda Hall. And, yeah, that's what we, like, pitched you guys. That's where you pitched us. It was awesome. Me, Yanwei Lin, and then you and us.
2:26Lee Marie and Ilya, I think. Was it right? Bucky too. Bucky, yeah. Buck. Why'd you meet at Berkeley, not Stanford? Well, he was finishing off his PhD at Berkeley. Dude, this was never meant to be a company. We started this off as a student project three years ago at Berkeley because at the time we didn't know how to evaluate out. It was kind of a side project of a side project. I, meanwhile, was doing my PhD in theoretical ML. I was like proving theorems in a basement at the time. When really I was happy doing that. And that's also how my wife, because we were both theorists and that's how we both also got our jobs at Stanford.
3:03And that's, um, and then we, you know, in the meantime, we started Arena and it just grew and grew and grew into this project. And then I remember what happened was that at the beginning was just open AI. And then Google came in the mix and Anthropik came in the mix and it just exploded because people wanted to compare. People wanted to understand the differences. and they were kind of surprising because you would test them on like MMLU or whatever multiple choice questions. And it would be like, the model that does well on the test is not the same model that I like to use. And that's doing well in the real world.
3:35That's because they all trained at the test. And so we had this philosophy that it's really about reality. Reality is the only judge that you can actually trust. So we built Arena, which is this platform where people go to use AI in the real world. And by using it, they're giving us the feedback that we need to make these leaderboards, which you'll see online. And it just exploded from there. But we met very early in the days back when we were still basement dwellers. Yeah. And do you think it was at that time, you know, this is a year and a half ago or so now, you had OpenAI, Anthropic, Google, and DeepSeek was relatively new.
4:17was there a moment around deep seek that made you think hey like i need to go start a company because there'll be not just three models but there'll be you know not infinite but lots of models i wouldn't say it's deep speaks it's deep deep seek specifically but i there you know throughout the history of arena there have been all these times when people kind of being like do you think that they're going to consolidate do you think that the models are going to consolidate people still ask me that question like once a week i get oh but what if the models consolidate and it's always the opposite people are just not understanding that ai because it transforms every industry is going to lead to like many flowers blooming um and so yeah for sure that that was a trend that we didn't fully appreciate at the time but it has proven to be really helpful for our business yeah because at the time it was still the leaderboard for those the obvious company.
5:12Yeah, and we had like, you know, 20 models. You know, we started with like eight models. Yeah. How many models do you have today? 500. And when you say have, can you explain what have means? Oh, we just like rank them on the leaderboard. Yep. Really, we've started deprecating some of the older models. Yep. But you know what's happening now is that every week there'll be like 10 releases. That's insane. It's insane. Today, new release from Gwen. Gwen Max. And there's more coming from Quinn. The dragon has awoken. Yeah. Yeah. You've become somewhat of an expert on some of these open source models.
5:54For sure. And like, I think you guys are sort of leading the charge on, you know, almost like, not actually anointing, but like really understanding deeply, like what's happening with Kimmy and Quinn and others. Yeah. where does that expertise within the company come from well it all comes from the users the point is that like the point of arena is that real users and the utility that these models provide to real users should be at the center of the ai conversation there shouldn't be some random benchmark that's made by some dude that thinks they're smart some researcher out there that's like, oh, how does AI do on this weird prediction task?
6:37And I think that long-running finance workflows of XYZ type are the thing that determines AI progress. It's like, dude, get your head out of your ass. You're the calcet for AI. We're the calcet for AI. There you go. I'm serious, though. People make all these f***ed up benchmarks. And they're so weird. Why are they making those benchmarks? Because benchmarks attract attention. And people don't know how AI models, people don't know how to rank them. People don't know what performs best to their task. And so they're creating all of these like wacky ways of measuring them. We were all like for the last year on Twitter looking at pelicans riding bicycles.
7:24Because people were making benchmarks around SVGs. You know what I'm talking about, right? The pelican riding the bicycle? No. Well, there was this common benchmark that was made, you know, as a particular individual who likes benchmarking things, and made a pelican riding a bicycle. And then everyone on Twitter is looking at this. Every time the new model comes out, oh, where's the pelican? It's like, let's show you the pelican. Who the f*** cares about the pelican, dog? I'm not kidding you. I mean, it sounds ridiculous because you guys haven't heard about it, but it's like... Yeah, I have not.
7:53No, you should be. I'm so happy to hear that. Yeah, it actually reminds me of when you were in the graphics industry in circa 2000, everyone had the spinning video of Venus de Milo. Oh, yeah. Totally. That was a test. Yeah. So this is the pelican riding the bicycle. This is how far we've gotten is from mythology to... Yeah, 100%. Now I think that benchmark has been saturated because it's made its way into the training data now. people know how to generate SVGs and pelicans riding bicycles but people will create all sorts of like random little tests and toys examples but what it doesn't speak to is real user value what happens when you go try to use them all for your actual workflow you're trying to code with it you're trying to do your finance workflows on it spreadsheets about which companies to fund and which companies not to fund which you guys do every day that kind of stuff there should be benchmarks that measure it and it can't come from just somebody writing down a set of questions It needs to come from you.
8:52You need to be the user and you need to tell me, hey, this is good, this is bad. And it's not because you are trying to benchmark, but because you're in the natural flow of work and you're getting your job done. That's what Arena is all about. It's measuring utility by providing utility. The point is that the user incentive and the platform incentive are perfectly aligned. Because we're trying to measure utility and the only way to measure utility is to give utility to people. so it's like calci meets reddit where you can do some hierarchy upvoting to make sure that the users are the ones that are actually deciding where what's good or not yeah and i think um a different way of thinking about it would be like a uh model agnostic chat gpt um or a model agnostic cloud co-work because uh the insight is that the user of this product should never feel like they have to give feedback.
9:47They shouldn't feel like they have to upvote, downvote. And a lot of times the feedback that they give is not of that nature. It would be implicit conversational feedback. You know, I like this. Can you keep going? Or, you know, what a terrible response. You know, undo it and give me something else. Anytime you're doing that, that's explicit natural language textual feedback that we can use for our rankings. And then there's implicit natural language feedback. You ask a question. You try to get your response. It doesn't work. So you try to ask it another way. Almost like in the days of search, you have query reformulation.
10:26There's query reformulation also for co-work type products. Where you're just trying to get like, how many times have you had the experience where it's like, I'm just trying to do X with my AI. And I keep prompting it to do X and it's doing it wrong. Or it's doing X plus Y and the Y I didn't ask for it. It's like, I've been trying to like, you know, figure out where to put a couch in my living room. And it's like, I want the painting here. And I want the couch there. Can you do that? And it's like, sure. And then it's like, I'll do that. And I'll also modify the window. What are you doing? And so I'll try it again.
10:59No, do it again, but don't mod. And so that is a piece of feedback. Every interaction is a piece of feedback. Download button. That's a piece of feedback. And so if you take all this together, you can learn about which models provide utility to users. you know which ones are they actually getting their job done where they're not having to reformulate their query 50 times which ones are they downloading the files and then actually using them which which prs are they merging into their code base all this stuff we measure and it becomes the highest quality signal of model performance on the market what kind of people use arena on a regular basis so this is very interesting um what would you guess actually you don't know our business, what would you guess about the organic users?
11:42Engineers. Why do you guess engineers?
11:48Like assuming maybe that if the breakaway application is coding and having spent a bunch of time with engineers over the last year, they seem to care absolutely more than anyone being on the frontier, depending on what coding model they want to use. Yep. That's why I guess. Totally. And you're right. It tends to be more knowledge worker professionals that are using RENA. 28 % of them are software engineers. By the way, at the scale of tens of billions of users, this is a lot of freaking software engineers. Yeah. So 28 % software engineers, 17 % scientists, 15 % finance, 6 % legal, 6 % medical.
12:26It just follows the categories that are, like you should expect as these new categories. Yeah. Like if Harvey and Sierra ever catch up to cursor, you would expect those numbers to probably start going up in arena. For sure. Yeah. It's no joke. It's no joke. No joke. So it's not just a leaderboard. It's not just a leaderboard. There's this massive community consumer following that powers all the data and analytics that we show. And I'm curious, like, okay, let's just assume that you have more ground truth on model quality and progress than maybe anybody in the world. What has surprised you over the last year on how the leaderboards have changed?
13:11You know, the easy answer would be that the pace of development is just extraordinary. I think it's faster than any of us could have predicted. But I think the number one thing that's starting to happen now is there's a phase shift between closed and open models. And I think that the implications of this are not well understood by our industry. The fact that Kimi was number one on our web development arena was huge national, international news. I don't know if you saw it online. But the fact that Kimi is number one on web development, it's the first time, really, that we're seeing a violation of the distillation narrative.
13:50The distillation narrative being that the only reason why China is keeping up is because they're distilling American intelligence. But how can that be true if they're number one and America's number two? In a very important category, by the way, most software developers are doing web development. So it's not a nothing category. It's where most of the economic value is coming from. And so what that means is that, yes, of course, they may be distilling, but distillation is not the only part of the story. There's something else they're doing actually to push the frontier beyond American labs. And so this is what we're getting into.
14:24you know, I don't think Americans are fully appreciating the intelligence and creativity that's coming from Chinese scientists. But that is being shown, it's being proven as we speak. And Quen is coming out, you know, Quen just came out with their MAX model. It's a frontier model. And there's going to be many more of these. And so I think that we're starting to get into the flippening, the flippening of the American and Chinese. And so what happens when this actually comes to fruition? If we actually get a Chinese model that's as dominant as Fable has been over the past couple months, in the sense that it's topping every leaderboard, you know, you should tell me what's happening to the capital markets.
15:05I think it's not going to be good for us. But the other aspect of it is that there is a silver lining, which is that I think that there is room for a trillion dollar American company based on open source.
15:22I think everybody would probably agree with that. I don't know. Would they? Well, I mean, I think everyone here would like to agree with that. Like, I'm curious if the leaderboard suggests that there's anything even close. Nope. Nothing. Even close? No. Thinking machines, the closest. If you look at open source only models, they would be number 10. Wow. And then there's nine Chinese above them. Wow. Nine. And actually, there's probably more now. They're probably not even top 10 because of Quenna. And so it's very, very, very competitive. And we are way, way the hell behind. Mm-hmm. So, What gives?
16:07Like, what are we missing here as, you know, the East versus the West? Is it not putting enough resources behind distillation, but also other novel techniques that allow Kimi and Kuen to produce these frontier models? Yeah, I think, you know, there's some aspects of this that are incidental because Lama was a pretty good model at the beginning, and then they killed their program. Yeah. They completely imploded because of some issues that they were having internally and restarted. Maybe if they get back into the open source game, things will work out. But we seem far from that. My take is that there's business model issues that have not yet been resolved.
16:56We're starting to see a path for those companies monetizing now, but it wasn't clear earlier. Because what happens, I'm going to spend$5 billion training a frontier model, and then I'm just going to give it away. It sounds stupid, right? It sounds like a stupid idea. And so not a lot of American companies did this. And the Chinese can do this because they have different corporate dynamics in China, which you guys are very familiar with. Um, but in the U S it's like, well, how am I going to make the money back that I used to train this model? And, you know, there's a couple of ways that people are approaching it now.
17:29One way is to have licenses that are kind of open source and also kind of not, that have like revenue thresholds. Hey, if you use this in a product that's over$20 million in revenue, we're going to do a rev share. If you're an inference provider and you want to serve this model, then we're going to do a rev share. That's one way of going about it. And then the other way of going about it, which I think is becoming more and more common with like Thinky and Reflection and Mistral and others, is to take the open source and then use it as a way to insert yourself into companies and then build basically an AI modernization motion around that.
18:04But to Mamoun's question, like let's take Thinky, for example, you're saying it's barely scratching the top 10, but it seems to have extremely competent... Sorry, barely scratching top 10... Open source. Open source. Open source. It seems to have raised enough money to be dangerous and have all of the right people around the company. Like why, why hasn't it been able to break through the top 10 and open source in the benchmarks? I mean, listen, the talent game is very, very, very competitive. The model game is also competitive. They're starting from behind, let's be honest. I mean, you saw they restructured their team like seven months ago.
18:42So they like blew everything up and then restarted. And then seven months later, they're at a model that's top 10 open source. I would say that's actually not bad. Yeah. It's like good, strong momentum. Mm-hmm. Um, but, but like a company like OpenAI Anthropic has built such a mature pipeline around hardware, around data. Um, you know, they, they've accrued advantages that, that need to be overcome. There's no question that thinking machine is the underdog. Mm-hmm. this might be a silly question but do you get pressure from these labs like uh you've become very influential and let's call it engineers perception of their product like pretty much more than anybody no they care they must care a lot they care a ton how do they show you they care well they care about the results and so they're coming to us and asking hey can you co-release with us or, you know, what's my score going to be and all that stuff.
19:45So we work very carefully with them. The trust is a good asset for our company. But you have to recognize these people are smart, so they understand incentives. Their incentive is the same as our incentive, which is for the platform to be trusted and neutral. If the platform is not trusted and neutral, it doesn't have value to anybody. Mm-hmm. And so we've never really had a model for a to come to us and try to say, hey, will you do this under the table? They have too much on the line, and they know that if they do that, that damages the platform forever. Yeah. So we don't have issues like that.
20:23Perhaps surprisingly. A lot of big personalities. You would think that they'd be like trying to jockey for... Sure. Yeah, but people want to play on a fair playing field right now. What's the biggest criticism of the platform? I'll tell you historically. ARENA started with battle mode. So you would get one, you'd input one prompt, you get two responses, and then a user would choose which one is better or worse, according to their preference. And we built an incentive so that they would vote their real preference. The reason being that the answer they vote for, they continue with. Which means that you shouldn't vote for a bad answer because it'll pull your context.
21:03I would say the biggest criticism of the platform historically has been human preference is really an incomplete signal of model quality. Like, why should I care if humans prefer it? I care if it's verifiable, I care if the code's high quality, I care if it's factual. And humans aren't like checking all the facts when they vote for answers. They might just vote for what they like. You know, you might vote for something that confirms your pre-existing biases. And so, what we've been doing is we've been going way, way beyond the battle mode to address these concerns and to take advantage of, like, richer signals.
21:37So, we've moved largely away from battles as the main signal on Arena for this reason. We incorporate human preference still, but the number one signal that we have now is task completion. Are you getting your job done? And you don't necessarily need to tell us. We can also measure it straight from the data. So, it also fixes the gap between the stated and revealed preferences. you know, that we have. Task completion, steerability, hallucination rates, factuality. We now rank factuality by using pipelines that we build that are automated, extracting all these factual claims and then, you know, shoveling them into a bunch of search models that go search the web to see if we can find any evidence to support or to disprove any claims that are made by these models.
22:18And you'll be surprised, they're not all factual. Despite the way that they're trained or maybe because of the way that they're trained, and memorizing a bunch of lies on the internet, which you can find. And so we've been trying to neutralize the concerns around human preference by going way beyond it. And so now all the rankings that we have on Arena go far beyond human preference. Yeah. So you're not a leaderboard. You're not just, for a consumer, you are a model agnostic, chat GPT or cloud. But as a company, as a business, you're a data company. Yeah. Absolutely. And by a data company, we mean that the value in the product comes from an accrual of a massive amount of data, high quality data that we can collect through our various platforms.
23:10You want to talk more about what you guys sell? Yeah, sure. So we work these days primarily with labs and we work with labs to help them improve their models. The product that we give to them is basically an analysis and insights product called an evaluation. So the first thing to note is that the numbers you see in the leaderboard, like the leaderboard that you see publicly, never touches money. It's 100%. We basically run it as a charity so that the world can see. Of course, it's great marketing for everyone involved, ourselves and others. But we evaluate all the models on there for free. Now, where we do charge money is in the process of developing models, Laz will come to us and say, hey, we have like 50 checkpoints this week.
23:54how good or bad are these at different types of tasks? Is this good at front end? Is this good at full stack? You know, is it bad on tool calls? What can you tell us about where we need to improve? And we take those checkpoints, we evaluate them on our data, you know, which represents real usage. And then we hand those results back to them as an insights dashboard. And so this is the product that scales to over a hundred million in revenue. And a lot of the future of the business is saying, okay, this has proven so useful to labs How can we take an arena and put it within every business in America, every business globally?
24:30How can we have every business take advantage of their data in the same way we are to know which models are best for them, which models are best for their users, which models are best for their employees, and then how to save money? Because that's a big part of it is the trade-offs. People don't know how to make the trade-off between performance and costs. They don't even know how to define performance, right? Because it's not like you just, it's just speed, you know, it's not like just a, you know, a CPU or something that you can just, hardware benchmarking. It's very, very mushy. Like, and you need to know whether people are getting their jobs done better, uh, and whether they're getting more satisfaction, like CSAT type measures, from the actual, you know, organic traces that are being generated within your company.
25:13Uh, this is sort of representing the future of the business. Mm-hmm. How quickly did you get to that$100 million? Eight months. How big is the company now, people-wise? About 75, 80. And how big was it eight months ago? Eight months ago, it was probably like 15 to 20. It's insane. It's completely insane. Absolutely insane. Yeah, imagine what I felt like coming out of the basement where I met. That's what I was just thinking. Yeah, as a PhD student, just proving theorems. And then now I'm here running this company. It feels completely surreal. Surreal fast. Surreal fast is surreal, just like weird.
25:57Do you enjoy it? Yeah, I enjoy it. Which part of it? You know, I mean, obviously like it's so exciting to be able to participate in the creation of this technology. That's gotta be the number one. It's all of us. I mean, you guys included, have to be just so excited to be like in the nexus of the most important thing that humanity has ever invented. It's like, you're gonna be bigger than electricity. It's just gonna touch everything. It could even result in like a new species, guys. I'm not kidding. Like, I feel weird, like even with those words coming out of my mouth. Yeah. It's like, yeah, you're like chanting the name three times and it appears.
Read the full transcript
26:38Do you know something that I don't know? Yeah, I do. I do. You do? I know what's happening in Southern West. They're making aliens. We've had contact with RSI. RSI. RSI is here. No, you know more than me. I know more than... But yeah, I think that's number one. It's just so cool to be part of this. Mm-hmm. And then there's things about being a CEO that I love. I would say I love working with people, a diverse array of people from all sorts of walks of life. You know, as an academic or as a PhD student, you interact with technical people, people who are experts at their craft, really honest people, sweet people that have no idea about how necessarily to productionize something, but that are brilliant at what they do.
27:28And I love those folks, but it's only like a tiny, tiny slice of the world. And now I work with BD people who everything is a negotiation. And I work with, you know, uh, like lawyers, and I work with investors, and I work with, uh, you know, engineers who are like great at building things and taking something from an idea into a product that you actually love to use. And I work with marketing people. And all of the people that we have at Arena are absolutely best in class. They've built generational companies before. Yeah. And so it's really cool to be just, you know, working with so many amazing people.
28:03And then people have problems. And that's my job as CEOs, deal with people's problems. I'm sure you guys know. You know, you're really good at it. Actually, you're exceptional at it. From the basement as a PhD student to a year later, running a company at this scale, I think just from point A to point B, I don't think I've seen a slope higher than yours. That's so sweet. I really appreciate that. You've seen a lot of slopes. So that's a high compliment. But I'll tell you, the people problems, Of course, they take up more and more of the time as time goes on, because everything ultimately is a people problem.
28:39A company, you design the company, the company builds the products, right? So, I mean, there's some truth to that. I love being part of the details, so I'm always gonna be part of the details, but at the end of the day, you need to make sure the right people are in the right place at the right time, and are pointed in the right direction to make things happen, and that they're not caught up with, like, dumb crap. But I kind of love dealing with people's problems, too. And you just gotta love people. You're like, oh, this person's having this problem because they are emotionally feeling this. And oh, how can I help them fix this insecurity that they have?
29:12Or this interpersonal, this argument that's happening because it's something stupid. How can we help resolve that? And I kind of relish that. Now there's parts I don't like. Parts that I don't like about being a CEO, primarily that my calendar is just filled with meetings. I just have like 10 hours of meetings a day. I do 10 hours of meetings i go home i do whatever work i couldn't do in the meetings and then i rinse and repeat meetings are not always pleasant sometimes you just want to be a little like i love doing technical work too um so i try to carve out five to ten percent of my time to do that but you know that's always last last on my priority can you steel man why you think this is more important than electricity?
29:58Like, I'd be very curious your argument for why. Okay. So why is it the most transformational technology that humans will ever invent? Obviously, like electricity - Ever invent or ever have invented? Ever have invented. Yeah. Ever have invented. Obviously, electricity is foundational, right? We can't have AI without electricity. So in some sense, like if you think about it hierarchically that way, electricity would be more important. Now, the counter argument to this would be just in terms of the realized impact, that sort of the incremental value that AI would add would be greater than that of electricity.
30:34And so why would that be? AI has the potential of literally creating intelligence on demand, which is something that humans have never seen before. and could be influential insofar as it actually even deprecates humans. They could literally make humans obsolete in the sense that we're able to create an organism that can pursue any goal faster than we can, that can create productivity at a greater degree than we can, that can explore the universe more efficiently. And the impact of that, I think, sort of if you were to try to decouple everything else electricity can do versus everything AI can do, that impact may be greater than electricity, because it will just supersede everything that humanity has ever accomplished.
31:35That would be the argument. So to be clear about the thesis, it would be, if you take everything that electricity has done and will do, outside of AI and compare it to the impact of AI, then the latter will be a greater. I'm curious, what's your reaction to that argument? It's beautifully said, and I fully agree with you. There's nothing to disagree on. I think, you know, we've talked about this for the last few years, that this is the coming together of all the pieces of technology that have been created over the last 60 years from the transistor in silicon to firmware software operating systems the internet mobile cloud everything pulled together allows this moment to exist yeah right so it is the it is the coming together the culmination of everything's happened in our industry ever since Our firm was started 55 years ago, and it is greater than all those combined.
32:47And I actually love the framing of if you took electricity and took to one side what electricity has done for humanity and put to the other side what AI will do for humanity, and the way you said it, the latter will be greater than the former. Yeah, no debate. no debate no debate do you have a different opinion no i mean i hope you're right because something's got to justify these these valuations so so you know i think the real question is what happens to money yeah will money be so it'll it's interesting because money there's there's something that has made capitalism work so far which is that capital and labor are coupled like you can earn money by through labor and that is no longer really going to happen right so So that's where this capitalism stops working for people.
33:39And then you have to start thinking, what's going to happen to people? I think that's the fundamental question is the decoupling of labor from capital. Yeah. Very fundamental. That is really the crux of the matter for humanity is the decoupling of the two. Historically, the stock market revenue for companies, which is tied to the value of the stock market, was tied to the labor. Yeah. Input was labor. And as labor becomes detached from the creation of value, the stock market capital value diverges from labor input. Totally. Right. And that's really what we have to all grapple with as a society over the next decade or so, which I think we have to flush this stuff out.
34:26um there's just one one stat that really jumps out um is if you look at the overall market cap of um the stock market it's like i think the u.s stock market's like worth um 100 trillion dollars something like that in that order roughly and uh it's it's sort of in line with the the gdp of the world so in like 20 30 years ago uh the overall value the stock market was way lower than the overall gdp of the world so capital value has appreciated a faster rate than gdp yeah and it's going to accelerate in that way um and it's already actually if you look at the u.s stock market um a lot more of the value accrues to um the delta is like the u.s gdp is like 22 trillion the u.s stock market's like 60 trillion so it's like a 3x yeah and the overall worldwide is like a 1x so the u.s is already at a 3x capital value to overall gdp totally and that's going to accelerate that's going to crazy degree right and so it's the separation of labor from capital.
35:42And so I think that is what we all have to sort of grapple with as a society is how do we live in this world where more and more of the value of the world will accrue to potentially fewer companies. Totally. And also how do I think the idea of how we assign value is another thing that we are going to need to think about. because the assignment of value comes from scarcity and scarcity comes from the coupling of capital and labor the fact that you can only make x dollars because you own a salary and therefore you have to spend that money carefully and therefore value is a signal like price is a signal value right the whole point of like this these economic theories is that of capitalism is that price is information.
36:33The price is the information about how valuable something is. And that information is driven by scarcity because the sort of like decentralized market of consumers of whatever that product is, they go and they pay whatever it's worth to them as opposed to other goods. But what happens when those two things are decoupled? And I think there's two ways of seeing it. One way is what happens to humans? How do humans assign value in a world where, let's be honest, there's going to be a big fraction of people that just can't make money? And then there's another side of things, which is agents. What if agents can assign value?
37:11Do they have the right to assign value? Is value just something within the construct of a human? If you're making life better for a human, that's valuable. Or is there a construct under which you're making life better for an agent? and therefore an agent with a wallet can pay for that, we call that value, and that the money that the agent is spending doesn't need to come from a human. So can agents autonomously be earning value? And that's something we're going to need to decide as a society is whether or not we're even going to allow the possibility for autonomous agents to directly assign value themselves.
37:51Is all of this assuming that the models will continue to get better at the rate that they have? Yeah, models are harnesses. Not necessarily even at the rate that they have. This might happen in 20 years. It might happen in two. My guess is it's going to happen closer to two than 20. So I think we better start thinking about it now. Why is that your guess? It's just the rate of improvement hasn't decreased. If you look at the arena, we're seeing, at times, bigger gaps between successive generations of models than we have in history. still to this day. You're saying that the newer models from Anthropic and OpenAI are taking bigger step function leaps and improvements in your leaderboard than like GPT 3.5 to 4.
38:40Yeah, it's happened before. I mean, GPT 3.5 to 4 was another one of the biggest leaps in history. But we are still seeing some of the biggest leaps in history to this day. So for example, GPT Image 2 versus 1.5 was, I think, the biggest gap that we have ever seen in history on ARENA. It was nearly a 100 % win rate. And by the way, it's almost impossible to get people to agree on anything. Like, it's almost impossible to get a large number of humans to vote the same way on any political issue, on, you know, basically anything at all. But 100 % of them are voting GPT Image 2 better than every other model that we had ever tested at the time that it was released and we're seeing phenomena like that still happening so yeah i don't think and rsi you know it's the buzzword of the day if this happens um we'll absolutely be seeing acceleration question for you what are you long and what are you short i think that the api layer people are assuming a degree of lock in the apis that i don't think is necessarily going to exist.
39:46Anthropic revenue has been booming, booming, booming on the API. But to some extent, it's also easy come, easy go. It's like if you are spending that much money that quickly, there's no chance that you're not able to switch vendors. And so if an open source model comes by and it is able to do 90 % of the work at 10 % of the cost, which is totally feasible that that might happen over the next couple of years, and that certain businesses might have all their capabilities be able to be done by the mini light version of whatever the next Kimmy is? Well, I think that's a pretty structural problem. And so what are people doing?
40:21What are labs doing to do this? You know, OpenAI Anthropic, OpenAI has this huge consumer business, which is kind of like going to be able to keep it afloat anyway. By the way, one thing I'm long is the consumer ad marketplace on OpenAI, which has obviously not been fully utilized to this day, but I think it remains one of the biggest opportunities in AI. Because the question is, how are you going to generate hundreds of billions of dollars in revenue for them? Right? like annualized. And there's like only a few businesses that know how to do this. And it's like ads and great hardware businesses.
40:54And I actually don't know anything else. You probably know better than I. No, those are your, you know, take Google. You take Amazon. You take NVIDIA. Exactly. Right? Tesla. Tesla. You know, consumer or great hardware. And like ads or hardware. If he's right, let's just imagine like if open source... delivers, breaks out. Do you think that value goes up the stack? Like, let's just say right now, so many of our application layer companies have very expensive costs to serve. And if you got rid of the gross margin problem, these would be some of the most incredible businesses of all time, like ever.
41:37So I guess if open source can alleviate the gross margin problem because their cost to serve goes down by 80 to 90%, then naturally, does the logic follow that value continues to go up the stack? I think that's been our fundamental belief here, that we're at maxis here because we believe that the models will continue to get better, the open source models will be even better, and the cost will come down. And we already see that the army of lawyers that AI lawyers, or whatever we're going to call it, at Harvey is the use just continues to go through the roof or take any app layer company. There's an insatiable demand for the apps that get the job done.
42:24And you don't want like the half job done. You want like all the particular needs of the legal industry met or the medical industry or the coding. So, yeah, fundamentally, like, all the things we're talking about benefit the appellator companies. And that's always been sort of the, that's how it's always played out in every sort of cycle of the industry, is that companies, big companies especially, buy from other companies that package up a really nice solution, you know, show up with the, the person who's going to help them implement it, get it up and running, getting it, getting the value delivered to the end user.
43:11And that usually comes in the form of an app layer company doing that for them. And so I think the tailwinds for these app layer companies are totally there. I know they were called rappers for a very long period of time. And I think that's, it's always the, you know, the very lazy way of, you know, identifying these companies or uh describing these companies uh but but yeah i'm i'm in the the camp of there's just a lot of value that will accrue to these app layer companies well i think that this is a really interesting point so then how do you view the lapse it's to your point like you know it especially as open source unless like we all just lock down open source from china right Which, I don't know.
43:59Regulatory capture. Yeah, do you think that happens? I think that there will be serious regulation on open models, unfortunately. Yeah. Yeah, and I don't know exactly what shape that'll take. I think in part, it's being shaped right now, and that people like you guys probably have the power to make sure that's done responsibly. Right. But I would predict that we regulate those. There will be regulatory capture. Yeah. Maybe the silver lining is then we get a good open source model here. Yeah, I think that that's the saving grace. And to be clear, I think the commercial incentives favor this anyway, because we are going to need American models so that people can build on a reliable supply chain and also one that doesn't have a tax factor, frankly, a foreign adversary.
44:43So we're going to absolutely need American open source. Maybe this will accelerate it. But I just hope the regulation doesn't slow things down, too, because it could also slow things down. Would that be the worst thing for your business? Oh, for sure. Our business sucks if we only have one model provider, you know? Right. And if we have two, we're barely making it.
45:06We want 10. We want 50. Um, yeah, I think we could probably do with three or four. Um, but you know, regardless of our incentives, it's bad for everyone. I mean, who, who wants an oligopoly? Like, the whole point of this, like, capitalist system is that we should try to encourage competition. And so now, if we have the winners going to Washington and convincing politicians that we should create these regulations that are going to be harder for, you know, that basically only incumbents are going to be able to satisfy and make it harder for other upstarts to be able to increase adoption and get their models out there, I mean, come on.
45:46Like, this is the most anti-competitive possible behavior. yeah you know um i was listening to um brett taylor we're at some thing and he was talking and like i would consider brett to be more in the know than like anybody and like about as competent as competent like whatever competent is that's him like he is i have so much admiration for him and obviously he's on the board of like he knows more than most of us and whenever he talks about anything AI related, he's like, listen, I have no idea. Like when Brett talks, he's like, I genuinely have no clue. Cause he's like, I've never seen anything like this.
46:28And he's like, I know more than anybody in this room and I have no idea. And I'm like, you know, I kind of respect that, you know, I'm just like, it's just, we've never seen anything like this. Yeah. We've never seen it. It's insane. Yeah. Completely. Going back to the, like running your company for a second. Sure. How do you think about, like if you did 100 million in eight months, I don't know what the goal is in a year from now, but - We don't either because we keep beating it. A lot more. We keep setting it, we keep beating it. How do you support the company? Like how do you, I don't, how do you support revenue growth that is going parabolic like that?
47:11Like from a company perspective, Like, how do you even build the scaffolding around a company that's growing that quickly? Yeah, I mean, the good thing is that I think AI has made it easier to run a very high revenue company with fewer people resources. So we don't need to scale people linearly with revenue. Fortunately, we've avoided doing that entirely. I would think about it more like, how do we put the right people in the right place at the right time and make great products? and how do we attract the best people in the world towards a small team that can support huge amounts of revenue. And of course we need to fill in to certain G &A things as we go.
47:53You know, we probably need more lawyers because we have more contracts. The number of contracts and number of lawyers kind of scale together. But it's not true of the raw revenue. Yeah, or the product. You know, the product runs on its own. It's arena.ai. It's this website. It has all these users. You know, of course we want to keep making the product better for them. We want to keep refining the user base and, you know, making the users better, giving them more tools, making it safer, decreasing abuse of the platform and so on. So we want to hire for all these things. We have our enterprise play.
48:24We want to hire great infrastructure engineers for. But it's not like that's a revenue scaling problem. It's more of a product scaling problem. What's your worst day at the company so far? Man, worst day at the company. Such a positive guy. I know. Well, it's easy to be positive when you recruit 100 million in eight months. No, there's some low periods. Yeah. You know, at the beginning, in the early days, there, um, it's so hard for me to talk about. And there was this moment. Uh, we were just, you know, in the early days of trying to start the company, trying to get it off of the ground. And we were still supporting labs and open source and other academics.
49:11sort of on an ad hoc basis, giving people data, helping people do their research. And we were about to raise our seed, which was, of course, stressful for me. I had no idea how to do it. Yawn was helpful. You, of course, supported us in that time. And a very stressful period for me. And then my father-in-law passed away suddenly. He passed away of just a heart attack, I guess, in Serbia, my wife's servant. And so I kind of had to drop everything and go to Serbia. He was also her person. They were so, so close. And they're similar people. Like, when I talk to her, I see him. And I was close to him, too.
49:57I really loved him. And so he died literally while we were in the, hey, term sheet's coming. How are they? Stage of our seed. And so I had to call people like Moon and say, hey, guys, I need a few days. And then we went through the funeral. And then two days after the funeral, I was on calls again. Like negotiating the deal. Right? Because I needed to, because the company needed me to do that. And maybe I didn't 100 % need to at the time, but I thought I did at the time. That people wouldn't wait. And that the deal was hot. And I didn't know what to do, but I knew that I needed to do what's right for the business, because there's people's livelihood relying on me.
50:37And then after that, you know, I think it took a long time for my wife to forgive me for the fact that I had to go be there during that period. And, you know, she was mourning so deeply. And I would say that she cried every day for many months. And then eventually, it was every two days. Maybe six months later, every two days. And then three months after that, it was every four days. And now she cries every so often, she says, I remember my dad. And I'm starting to forget my dad. I forget the sound of his voice. And that's just so difficult. Mamoun knows, because he's been through a similar experience.
51:25Yeah. And then, on top of that, a couple weeks later, there was this awful moment where there was this paper that was released called Leaderboard Illusion. People don't know this about Leaderboard Illusion. You guys might have seen it. Not only is it filled with a bunch of lies with data, basically incorrect analysis about ARENA, that sort of is, you know, it's this coalition, MIT, Stanford, you know, whatever. So it looks like it's a reputable source. But actually, when you dig into the analyses that I had about Arena at the time, it was about, oh, this company is favoring closed-source models over open-source models and blah, blah, blah from Cohere.
52:08Of course, Cohere was like number 70 at the time. So, you know, you say what you want about the incentives. We've always been fair to everybody. And that was very difficult because it felt like a stab in the back from some people that could have been my future colleagues if I ever decided to go to Stanford. Yeah. And... was another difficult experience on top of the sort of personal challenges that we had at the time. It was very hard for me and my wife, and it happened a few months before our wedding. And so we thought that our wedding was gonna be a sad time. We thought that we were not gonna have a good day.
52:44But, you know, we ended up pulling, she ended up kind of like rallying for that day, and we both had an awesome time at our wedding, thank God, because it's only something that you get once. but there have been some very low points in arena. Startup-like, the lows are low and the highs are high. Hi, Goosebumps. Thank you for sharing that. Of course, yeah. Can I ask you about, can I ask another question? Of course, yeah, yeah, yeah, yeah. I'm an open book. When you, I thought the framing of how you said it took a while for her to forgive you for that. Was it her forgiving you or was it you forgiving yourself?
53:20It was probably both, but she definitely needed to forgive me. Yeah, I think it's for her - Prioritizing the company? Yeah. Right. Right. I think it felt like a betrayal to her that I would prioritize in the way that I prioritized. And probably she also had more clarity of thought and was a little further away from the fundraising details for me to say, hey, maybe I get why you would do this meeting, but I don't get why you would do this other meeting. and like not all the decisions i think made sense to her probably because she was right and they didn't all make sense to be honest i don't know for what it's worth um i got to see a very small maybe like 0.01 of you during that time in terms of what came through to me but the level of care and love that you showed towards her and your father-in-law was actually unique to hear from a founder CEO who's building something in AI there's a level of humanity and empathy and compassion that you had towards others that you actually don't see from founder there's so much egocentricity.
54:39So it made me, and we'd just gotten to know each other, respect you so much as a human being, but like someone who is more than just a PhD, Stanford professor, founder in AI. It was just like, and maybe that's why from point A to point B, you have this slope because you may be the full package. i appreciate it no i think we really resonated together on that personal on that personal note that we've been through some similar difficult experiences and um yeah for for better for worse i have a big heart um so i you know i appreciate that you value that man thank you for sharing that of course um what does the rest of this year look like for you and the company yeah a lot of big things coming for Arena.
55:34Two big priorities, I would say. The first is agents, and the second is building out the enterprise side of our business. So on agents, you know, we've been historically known as a platform for evaluating models, but when you get into agents, it's just such a more heterogeneous ecosystem. It's not just the model, it's also the harness. And the harness can be very, very difficult to, you know, evaluate. It can be a multi-component system. It includes a a bunch of sub-models and sub-agents. And it's something people don't know. You can't go to somebody today and say, hey, which model, which harness combination is the best?
56:14But we can actually do that at Arena. And so what we do, if you go to arena.ai slash agent, we have agent mode, where you can do coworking-type tasks, like cloud coworking-type tasks on Arena in browser. You can connect your code base, you can do everything. And when you type something that will basically give the agent a computer and it will go in and it will allow you to do, you know, open-world tasks. And as part of that, we can randomize both the orchestrator and the harness for the agent, so we can collect all of those interaction effects that make it challenging. The goal being that people shouldn't really have to think about what harness to use or what orchestrator to use.
56:50It should just be an automated system. We can even do that via routing, which we've also built. We've been building routers for the last year. Routers are now hot. But we were in the routing game. So hot. I know. So hot. We've been in the routing game for quite a while. And then the second would be, how can we take these ideas, the sort of like self-improving data flywheel that we built on Arena that allows us to build the best routers in the world and then copy paste them into every business, help businesses take better advantage of their organically occurring data so that they can build a better routers, avoid vendor lock and get AI sovereignty and so on.
57:21And so those are the two big priorities. Are you hiring towards both of those priorities if people are listening? What are you hiring for? Infrastructure engineers, AI engineers, those are probably the most valuable. If you can come help us build out our agent product and our enterprise product, which requires really high availability, high uptime, you know, this routing functionality, data engineering, that's all really important. Machine learning research. You know, I come from a research background myself, so I know a lot of ML researchers, but if there's people out there that are interested in these fundamental problems around, how do we optimize, basically?
57:58How do we measure? And then how do we optimize? Because you need to know what target... Basically, you need to know the target in order to optimize it. And that's what evaluations are all about. I view Arena as being the system of value. That's where I see us in the ecosystem, the AI ecosystem. We can help developers, businesses, and the whole community define value through the real world. and then once you have a value system, then you can hill climb it. And the whole game is about picking the right value system. And so our value system is a utility to real people. And so anybody who's passionate about that mission should come join us.
58:44Thank you for doing this. Of course, Will. Anything else for you? Last one for me then. When you hear the word grit, what do you think of? Oh, grit. I think of people that never give up. I think of people that day in and day out, they are coming in and they have their nose to the grindstone. And there might be terrible things happening around them, or they might be pulled in other directions, but they're never going to lose focus. They're always going to pursue the goal, and it's never going to be about themselves. They're never going to have any ego about it. Grit means somebody who um is always putting the greater mission ahead of themselves and is willing to do anything to get the job done my um i was going through hell and back on the company over the last couple of weeks on some stuff and my fiance on the mirror of our bathroom wrote the man in the arena quote yeah the full teddy roosevelt man in the arena quote just literally in marker on the entire on the entire mirror.
59:53And yeah, it reminds me of what you said. Like, sometimes you just got to be in it. Totally. Appreciate you doing this. Thank you so much. Amazing. Yeah, thanks for being here, guys. That's it for now. If you liked the episode, please leave us a review or go back into the archives where we've done more than 200 episodes with some fantastic folks. This podcast is a Kleiner Perkins production and I'm Juven. Thanks for listening. I'll see you in the next video.
From the publisher
AI could be bigger than electricity.
Anastasios Angelopoulos is seeing that shift firsthand, with Arena ranking 500+ models and tracking roughly 10 new releases every week.
On Grit, he explains why there may never be one dominant AI model and why the API layer may be less sticky than people think.
Guest: Anastasios Angelopoulos, Co-Founder and CEO of Arena
Connect with Anastasios Angelopoulos
Connect with Mamoon
Connect with Joubin:
Email: grit@kleinerperkins.com
Follow Grit:




