In short
Podcast Episode Notes: Latent Space - [State of Evals] LMArena's $1.7B Vision
Episode Summary In this episode, host Anastasios Angelopoulos discusses the journey and future of LMArena, a platform that has quickly become a major player in AI evaluation. The conversation dives into the company's recent funding, its vision, and the controversies surrounding AI model leaderboard integrity. Angelopoulos emphasizes the importance of community feedback and the company's commitment to building a transparent evaluation platform.
Key Takeaways
- Funding and Valuation
- LMArena raised $150 million at a $1.7 billion valuation, with $30 million annual revenue following the launch of their evaluation product.
- Origin Story
- The company began as an academic project within a Berkeley basement, incubated by Anjney Midha at a16z. The decision to transition from academia to a for-profit entity was driven by the necessity to scale up their operations.
- Platform Usage
- LMArena has facilitated 250 million conversations, with 5+ million users. Approximately 25% of users are professionals in software development, indicating a strong technical user base.
- Leaderboard Integrity
- The episode addresses the "Leaderboard Delusion" controversy, where a paper criticized LMArena's testing methods. Angelopoulos refutes these claims, stating the paper contained factual inaccuracies and highlighting the transparency of their testing approach.
- Community Engagement
- The platform encourages user participation through preview testing, allowing users to test unreleased models and contribute feedback, which helps in maintaining a fresh and relevant evaluation system.
- Future Plans
- LMArena aims to expand into specialized fields such as medicine, legal, and finance, as well as explore multimodal capabilities like video evaluations.
Detailed Discussion Points
- Funding and Financial Strategy
- The primary use of the $100 million raised will be for:
- Covering inference costs for free user access.
- Transitioning from Gradio to React for better platform flexibility and developer hiring.
- Hiring top talent across machine learning, product development, and marketing.
- User Base Analysis
- LMArena's active engagement metrics:
- 250 million conversations
- Diverse demographics with 50% of users logged in, enhancing user insights and feedback.
- Addressing the Leaderboard Delusion
- The paper critiquing LMArena claimed inequities in model testing due to undisclosed private testing:
- LMArena's response highlighted factual errors such as misrepresentation of open vs. closed-source model testing.
- The community appreciates the preview testing experience, which contributes to user retention.
- Commitment to Platform Integrity
- Public leaderboard is viewed as a charitable initiative:
- Models cannot pay to be included or removed, ensuring unbiased rankings based on real user interactions.
- Future Developments and Categories
- Upcoming expansions may include:
- Occupational and Expert Arenas covering various professional fields.
- Introduction of video evaluation capabilities.
- Potential API offerings to enhance platform accessibility.
- Community Management Strategies
- Key community retention tactics:
- Implementing persistent history features for logged-in users.
- Regularly seeking feedback to improve user experience and engagement.
Conclusion Anastasios Angelopoulos shares his vision for LMArena as a central evaluation platform for AI, focusing on transparency and user engagement. The discussion highlights the platform's rapid growth and the significant role it plays in the AI ecosystem, as well as its approach to community management and model evaluation integrity.
For more information, listeners are encouraged to visit [Latent Space](https://latent.space).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Evolution of LMArena
0:45 to 2:42
Discussion on the transition of LMArena from LMSIS and the impact of branding.
“Out of the LMSIS sort of, you know, conglomerate at Berkeley.”
Incubation and Investment
2:42 to 3:38
Insights on how LMArena received support and funding from investors.
“It became clear that the only way to scale what we were building was to build a company out of it.”
Scaling and the Need for a Company
3:38 to 5:32
The necessity of building a company for LMArena's growth and reach.
“First of all, we don't necessarily need to spend all that money, right?”
Funding Strategies and Usage
5:32 to 8:48
How LMArena plans to utilize its $100 million funding and its operational costs.
“It is, but about half of our users now are logged in.”
Understanding User Demographics
8:48 to 11:01
The composition of LMArena's user base and insights from data gathered.
“Did you tell the Hugging Face folks that?”
Comparative Analysis of Competitors
11:01 to 12:20
Comparison between LMArena and other AI platforms, focusing on their strategies.
“and we'll expose them and so on and so forth.”
The Leaderboard Delusion Paper
12:20 to 13:33
Discussion around the criticisms from the 'leaderboard delusion' paper and LMArena's response.
“Yeah, I mean, to us, it just looked like sort of a random thing.”
Impact of Image Generation in AI
13:33 to 14:00
Exploration of the role and significance of image generation in AI applications.
“I don't know about that, but yes, the information reported this.”
The Evolution of AI Models
14:00 to 14:59
Discussion on the evolution and impact of multimodal AI models.
“you know Gemini used to get a lot of complaints for generating racist images or whatever.”
Core Principles of LMArena
15:00 to 16:13
Exploration of LMArena's core principles and community engagement.
“We're just going to be, our paper figures are going to be made by Amazon.”
Show all 15 chapters
Maintaining Platform Integrity
16:14 to 17:30
Insights into maintaining integrity on the LMArena public leaderboard.
“new use cases of AI, and gives the whole world sort of ground truth for how real users are using these models and how good they are on those use cases.”
Future Directions for LMArena
17:31 to 18:59
Overview of future plans and categories for LMArena's growth.
“play systems never going to be like that.”
Strategies for Community Growth
19:00 to 21:22
Discussion on strategies for growing and retaining AI community members.
“We're soon to launch on the site at some point, you know, later this year or early next.”
Seeking Partnerships for Growth
21:23 to 22:32
Exploring potential partnerships and collaboration opportunities.
“And how do I build in all the retention mechanisms so that they stay and then they're also bringing their friends?”
Highlighting AI Capabilities
22:33 to 23:37
Discussion on showcasing AI capabilities and ongoing projects.
“So I think the most straightforward thing would be for someone like Cognition.”
Transcript
Automatic transcript. May contain errors.0:09all right we're here with anastasius from arena i actually don't actually know your Star 7A, it's like very... Angelopolis. Yeah. There you go. Congrats on all the success. You got the arena handle. Yeah, we did. We got the arena handle. Thank you. Big branding moment. I mean, I think X is like being more commercial, so obviously you bought it, but at least you like have a place to go to where you can be like, hey, like, we really like this. But I do think like dropping LM has changed the feel of it. I don't know. I don't know. I mean, the reason we kept the LM at the beginning is because we were like, we started as LMSIS, right?
0:50Out of the LMSIS sort of, you know, conglomerate at Berkeley. So we decided to... Language models, yeah. Exactly. So we wanted to maybe broaden a little bit. Yeah. And we were the first Serena, so we feel like let's kind of try to own that. Yeah. Last time we had you guys on, you hadn't really spun out yet. And I did a call with Alessio and I was like, these guys are going to start a company. and i didn't know i think you actually were already started at the time i don't i don't remember uh maybe because anj i had a i chatted with anj and he said he was your founding ceo he was indeed uh which like people don't know yeah the anj is a very interesting character we have a podcast schedule with him yeah yeah he does a lot more than normal vcs he does he's been incredible to us you want to shout out some stuff yeah absolutely so you know the way the company started was as an incubation by onge yeah so what he did is he kind of like found us at berkeley and picked us out of the basement and was like hey these guys seem like they're onto something and started working with us really early gave you know gave us some grants he was not you know a16 was not the only one to do this we also had a great grant from sequoia but um onge was in particular quite quite supportive of us and you know gave us some resources in order to continue building out arena before we even were committed to starting a business and in that capacity he sort of like, you know, formed an entity for us and this and that.
2:12And, you know, he was like, hey, you guys can walk away at any time if you guys don't want to start a business. I mean, it was really incredible, very aggressive investment move by him, right? Because of course, any money that he spent at the end of the day, like we could walk away and leave him with nothing. But I think he knew wisely that the right thing for Arena was to start a company out of it. It was the only way that we could scale and that, you know, Waylon Yon and I would ultimately see that and be excited about doing it ourselves, which ended up being the case. Was there a moment for you where you were, I'm sure you were debating it yourself, you had other opportunities.
2:47What was the deciding factor for you? It became clear that the only way to scale what we were building was to build a company out of it. That the world really needed something like Arena. Arena being really a place to sort of measure, understand, and advance the frontier AI capabilities on real-world users, on real-world usage, based on organic feedback. And that in order to achieve the scale and distribution necessary and the quality, of course, of the platform necessary to do this effectively, we would need to start a company out of it. We considered other options. Are we going to keep doing this as an academic project?
3:27Are we going to do it as a nonprofit? Blah, blah, blah. But ultimately, under those constructs, we didn't feel like we'd have the resources necessary to accomplish our mission. So you raised$100 million or$80? $100. $100. $100 million. That's a lot of resources. It's great. Yeah, what's it for? Well, you know, obviously... That's on behalf of, like, everyone. Yeah, everybody. Dude, it's an arena. How are you going to spend your money? Yeah. It's great. Yeah, please tell. First of all, we don't necessarily need to spend all that money, right? The purpose of money at a company is to give you cards to flip.
3:58It's to say, hey, you have enough resources necessary so that if your first bet fails, you can make another bet and another bet. Of course. So that's not to say that we're going to spend all of it. Of course, you want to spend things responsibly. Having said that, the platform is actually quite expensive to run. We fund all of the inference on the platform. You know, the way that it works is the platform... Like you pay market rates? They don't give you discounts? No, no, we get discounts, but they are standard enterprise discounts. The same that would be given to any other customer. And if you disclose any, I don't know any numbers.
4:30I see numbers of votes. But like, what's that in like monthly tokens or I don't know. I don't know about tokens. You know, I'd have to back that out. But we have like, you know, let's say, again, this is off the cuff, but I can safely say we have more than 5 million now. 5, 6 million. We have probably 250 million conversations that happen over the course of the platform. We're on the order of, you know, mid-tens of millions of conversations every month that are happening on the platform. It's actually quite a large, you know, it's one of the largest consumer platforms for LLMs. Of course, nothing is really comparing to ChatGPT, but, you know.
5:04But no, still, I mean, I think the largest skilled ones. The benefit of this is that, like, it's actually quite a diverse population. So 25 % of the people on our platform, for example, do software for a living. still at this scale. How do you know? Because we like do all sorts of, we either survey them or we'll like analyze the prompt distribution that's coming to the platform. I'm happy also to share more. We've done something called Expert Arena, which is like trying to understand the distribution of experts that are coming to the platform. A lot of that can be like unauthenticated or whatever usage.
5:34It is, but about half of our users now are logged in. Yeah. And so we have some ability to understand them. And we also have surveys that we run on the platform that tell us a bit more about who are the actual users. Of course, there's always response bias in surveys, so you have to take it with a grain of salt. But nonetheless... And if you don't know Anastasia's background, he's the guy to correct for response bias. There's a lot of guys like that and girls. You're not the only player. There are this artificial analysis starting at the arena. Yup. AI. It's like some crypto people that started this.
6:09I don't know if you have a conversation with them on like, well, hey, this is our thing or let's work together on something i i don't know if you've uh what the no you know so i've talked to the artificial i've actually talked to both groups yeah both seem you know am i missing any major players it's just no i think those i don't know yeah i don't think so okay i think those are like some of course large depending on how you define the term now artificial analysis obviously has like huge like market mind share um around the analysis of of different ai systems yeah they told me they're They're just like, we are going to be Gartner of AI.
6:44Yeah, that's kind of their goal. And I think they're going after that consulting market and so on. Artificial analysis, from what I understand, is like a team of consultants that are doing this. They seem like really, really nice guys going after that particular market. It's different from our platform in the sense that their analysis is based on sort of like aggregating public benchmarks and turning those into analytics. Independently rerunning. And independently rerunning. Which matters. Yeah, yeah, yeah. Which matters, yes. And using those in order to compile sort of reports and so on that educate the field on the performance of all these different models.
7:22But they also have arenas. They have arenas, but the arenas are not based on organic usage. Like the thing that distinguishes our platform versus theirs is that the users are actually inputting their own use case. They're actually asking their own question. And that gives a level of realism that platform doesn't have. Of course, they specialize in a slightly different thing. But I see those platforms kind of diverging in that sense. Yeah. And sometimes it's like it's the only way to do this. So, like, for example, for AA, their video arena is pre-generated videos. You can't enter in your own video process.
7:55That's correct. But we're doing it organically. Yeah, exactly. And so... As a voter, it does help in terms of I don't have to wait. It does. But also, why would you go? Do you actually care about, like, other people's videos? To form your own intuition. Maybe, yeah. Maybe you're, like, interested in comparing. Like, I'm a shitty prompter, right? Are you? I don't believe that. I'm a terrible prompter. I learned by example. Don't denigrate yourself. Don't denigrate yourself. There are many prompters much better than me that say that, right? That's a fact. People have all sorts of cool ways of prompting LLMs.
8:26Yeah, it's educational to say. The only way to learn is by, like, well, look at other people's prompts and see the results and go, like, oh, I didn't know you could do that. And that's hard. Yeah, yeah, yeah. Okay, so let's come back to Arena. Oh, one thing I do want to say is the number one use of funds is getting off Gradio. Oh, yeah. Yes. Well, listen, Gradio, incredible platform. Gradio scaled us to a million miles. Yeah. That's incredible. And of course. Did you tell the Hugging Face folks that? Of course. Yeah. We're really, really grateful to Gradio for taking this so far. Eventually, you know, it became time for us to move off of that and be able to react.
9:01I'm sure Hugging Face would have loved you to stay on. They would have, I'm sure. was there a technical reason you just couldn't get the performance? Yeah, it just became hard to develop, and there were all these tools that we wanted in React, and to do all the fancy things that you can do in React became kind of... One example, I don't know. What's a feature that you really wanted? Let's say we wanted to create our own custom loading icons for video with notifications. How are we going to do that in React? It's hard. Yeah, we'll make a custom component. I'm sure the Higginfee guys are going to come in and say, hey, you can do that in Gradio, which maybe you can.
9:38Of course they're going to say that. Also, fewer developers know. How are we going to hire for that? Do we have to re-skill them? Do people are less familiar with that stack? Yeah. So it's full React, Next.js, all this. Yeah, yeah, yeah, all that. Okay, cool. Other use of funds that might be interesting? Like, you know, basically... No, that's basically it. How you can deploy the resources. Okay. Yeah, it's on primarily inference. So it funds the free usage of the platform. and then also hiring, of course, headcount. Yeah. We have an office, you know. That's us. That's an SF. I'll tackle one of the major things this year, which I'm sure you're tired of thinking about, but for people who are not in the loop, this is going to be news to them, the leaderboard delusion, the whole thing with Cohere.
10:19Let's summarize what they said and then your response. So the leaderboard delusion is a paper that critiques Alamarina. pretty like rudely well you know I would say unscientifically and let's be curious Cohere wasn't doing that well on Cohere was like 74 it's all good you know it's actually not it's a respectable place that they had on the leaderboard I don't even think it was really Cohere people like the Cohere model developers doing this it was more their research side but in any case what does the leaderboard illusion say it says that Alamarina was the claim is that Alamarina was doing this undisclosed, quote-unquote, private testing on our platform, that model providers will send us pre-release models and we'll expose them and so on and so forth.
11:04And that this creates so-called inequities in the leaderboard, you know, due to that pre-release testing. For example, they cited that Meta at some point tested some amount of models with us. Yeah. Of course, we can't disclose all of the details of how all that was done. But that is the main claim of the paper. Now, our response to that paper, it's online. you can find it on response to leaderboard illusion. And our response to that paper is essentially pointing out a series of factual mistakes in the paper that, that question the validity of the claims. So you can go look at the first version of the paper on archive yourself and you'll see the claims.
11:40I think most scientists would view that as, Oh, they've corrected it. They've correct, of course, because we, I mean, but they didn't correct everything. They just correct some aspects that were just blatantly unscientific and false. But, you know, there, for example, um, said that we were, that we only sampled like 9 % open source models and like, you know, 60 % like closed source models and just created a gap between open and closed source. But in reality, we're actually really supportive of open source models. And it was more like 60-40. And so that was, you know, one of the examples of an error in the claims.
12:17Another example is that they were claiming that there was some sort of bias introduced by this pre-release testing and that it was undisclosed. In reality, as you probably know, we've been doing this pre-release testing for a long time our community loves it they love basically getting like secret code names yeah the secret code like Nano Banana yeah all that so Nano Banana by the way started on you started on us right people loved it went like global sensation like non-zero fraction of the global population using Nano Banana did you talk to Naina about naming it Banana or was that no it was their decision I believe but it was sort of this randomly generated thing and it just went no no so apparently Naina who's a PM yeah yeah is named after her because her nickname is Nano Oh, that's sweet.
12:54I didn't know that. I was doing it. I put banana on it. Oh, that's sweet. I didn't know the origin story. Yeah, I mean, to us, it just looked like sort of a random thing. But it was clearly heads and shoulders above. Huge. Which, like, before that, there was Reeve Image, remember? Yeah, I do. Of course. Yeah. And also BFL and all those. I mean, all those models are also great. And I think those teams are also improving quite quickly. But Nano Banana was a sensation. I mean, that moment alone changed Google's like... Roadmap. Yeah, market share. Seriously. I mean, Google stock, billions of dollars are moving because of Nanobl.
13:31And now there's like an opening I code red and everything. I don't know about that, but yes, the information reported this. I would say like the image generation, I would say has been like this weird part of AI overall because it's not strictly AGI critical. Like it's not reasoning. it's like feeding more context into the model it is the model generating a visual representation. So it's basically like I always think like well you know Gemini used to get a lot of complaints for generating racist images or whatever. That was a hilarious moment. And Chagipiti also had it in the past. And I'm like well can we just get rid of this?
14:12Do we have to do image generation? Because let's just focus the positive reputation of AI in general on language models and coding and the other stuff. But I'm wrong. I'm such a huge Nanam Banana pro show. Yeah, I totally agree. I was also kind of wrong about this. I didn't see the positive benefits. But actually, I think that these multimodal models are going to become some of the most economically valuable aspects of AI, both in consumer and also in enterprise. Because one of the fastest growing market segments in AI adoption is marketing. and marketing and design. Yeah, ads. So I'm a content creator, right?
14:52Yeah, of course. I'm sure you're using it all the time. Infinite supplies of diagrams and explainers and infographics. Yeah. Soon we're not going to be even making the papers. YouTube thumbnails, yeah. We're just going to be, our paper figures are going to be made by Amazon. Yes, yes. I do think that actually one shot. So DeepSeek came out with V3.2 recently. I took their explanations, which are very wordy. They like very concise papers. It's 23 pages long, but it's very dense. Yeah. And so I just took their explanations of the RL environment stuff and I fed it into Nano Banana Pro and I fed an image that I used to understand the paper better.
15:25And the fact that I can just casually generate like a paper quality diagram that would usually take a PhD student like a month in Photoshop or something to do is incredible. Yeah, it's incredible. It is amazing. I want to ask about your principles running Arena. I think you manage a giant community, 5 million Mao. What have you decided are the core principles, I guess, before becoming a company and now that you're a company? I don't know if there's anything that's changed for you. I don't think anything has really changed. We want to provide the North Star of the industry and center the use cases of real users, foreground those so that people know what to target.
16:03The goal is to create a benchmark that is constantly fresh, that does not suffer overfitting because of the fact that we constantly have new data points coming in that tracks all the different new models, all the different new use cases of AI, and gives the whole world sort of ground truth for how real users are using these models and how good they are on those use cases. We continue to do quite a few open source data releases. We've probably released more data than basically anybody on the real world use cases of AI. millions and millions of conversations, real-world conversations from real users that the community is using to study and improve on.
16:46Yeah. And then I think in terms of what you will build versus what will not build, I guess I'm not necessarily caught up on everything that you've launched. I know recently you've done the dev or code arena. Yeah, code arena. That's the most part. Code arena, expert arena. Yeah, expert arena. So basically, what is in the critical path for you, let's say for next year and what what have you decided you'll never do so let me first talk about things that i'll never do the platform integrity comes first to the platform the basically the public leaderboard that we show on alan marina i think of as a charity it's a loss leader for us we don't really make money on the public leaderboard you can't pay to get on the public leaderboard it's not like a gartner in that sense it's not like any of these like uh you know pay play systems never going to be like that.
17:35Models are going to be listed on the leaderboard, whether or not the providers pay and whether or not they're getting a good score. They can't pay to take it off either. And so what that means is that the leaderboard has a certain integrity that will never be compromised. But not all preview models will make it onto the... No, but that's okay. Those preview models have never been released. Who cares about putting unreleased models on the leaderboard. The point is that for every released model, the score that you see on the leaderboard is statistically sound. It reflects the real world capabilities of the model.
18:10Yeah. Why? Because millions of people from around the world have voted for it. And that's where that number comes from. All we do to compute that number is millions of people are voting. We take those votes, we turn them into a number. That's always going to remain a transparent and fair reflection of model performance. Where are we going? Lots of different new categories. I don't know if you recently saw, we expressed, we, um, we exposed, uh, occupational and expert categories. So now single digit percentage of our user base, we're millions, millions to tens of millions of users, right? So single digit percentages means a lot.
18:42Single digit percentage of our user base are in medicine, in legal, in business, you know, finance, accounting, creative marketing, stuff like this. And we're able to show the performance of these models in all these different verticals because we have all these users in our, in our user base. Uh, and we're, we're working more towards multimodal, you know, video. We're soon to launch on the site at some point, you know, later this year or early next. So lots of things in the pipeline. Amazing. Would you expose an API? We've thought about it. Yeah. I think it's a, it's a possibility. Yeah. Well, what are the counter arguments?
19:17Why not? Well, there's obviously a need for an API. The question is more of focus of our company. Just because we're a startup. And so we really should be doing one thing well. Arenas. Yeah, arenas. So I'm not sure how far we want to sort of splay out and on what timeline we would want to do that. Yeah. Any other sort of like community management tips? You know, more broadly, like every AI company like really wants to grow their community. You're obviously one of the strongest in the world. What's really, really worked? Well, so first of all, I want to give a shout out to our community manager, Greg, who is doing an awesome job managing our community, whether that's on Discord or on Elamarina.
20:00He's really incredible. So I would say hire Greg. But don't hire Greg. Don't hire Greg. Don't hire Greg. He's ours. Find a Greg. Find a Greg. But in general, you know, the question of how do you get to so many users, that is a tough question. And keep and retain them. That is a tough question because consumer is one of the hardest markets in the world. There's a lot of websites in the world that people can go to. You know, why should they go to yours? And the reality is if you want to create a really dominant product, you have to provide people value. And to be frank, I don't think we're all the way there yet.
20:31It's not like I have the solution and answer for how to build a great consumer product. If I did, we wouldn't be at tens of millions of users. We'd be at hundreds or, you know, we'd be at a billion users. Is there a world you like are bigger than Chachapiti? I don't know. I don't know that we need to be. And I don't know that we ever will be because that's an extraordinary generational product that they built, right? And it took a lot of time. And then to some extent, it also involved luck. There's a lot of lightning in the bottom moments like NanoBanel was for us where our user base just like goes up by a lot.
21:04But when those users come, they can just as easily leave. So the way I think about it is every user is earned. You have to earn them every single day. They can leave at any moment. They're fickle. And so all the time you have to be thinking about how do I provide this person value? Learning, how are they using my website? What more could I give them? And how do I build in all the retention mechanisms so that they stay and then they're also bringing their friends? Is there one that's working in terms of retention? Like you said, a lot of people are signing in now. Yeah, sign-in was a big driver of retention.
21:35No, no, but what did you give them in order to encourage them? Oh, like history. Persistent history. That's it. That's enough. Yeah, that's one thing that has had a big impact. Okay. Yeah, cool. What do you want from people? What are you looking for help on, like any calls to action? Yeah, we are always looking for people to come and join us. If you are one of the best people in the world in your area, whether that's consumer product, whether that is machine learning, whether that is, you know, B2B, go to market, marketing, all these things, we need you at Arena. We're building like a high performance team of real experts in everything that they do.
22:08And, you know, I'm always looking for excellent people to work with. You need like, what are our partnerships, right? Let's say I'm at Cognition. I want to partner with Elimarina or just Arena. What works for you? What existing partnerships do you already have that's really fruitful? Yeah. So, I mean, we, of course, partner with all of the major model labs. Yeah. And that's just straightforward. Like, hey, we have a new model. Here you go. Exactly. So I think the most straightforward thing would be for someone like Cognition. It's like, let's evaluate that. But that's an agent. Yeah. But we should be continuing to shape.
22:41Well, Coderine is an agent. Yeah, that's true. And it's more focused on, like all these arenas tend to focus on the model rather than the harness. But that maybe should change. Maybe we should be evolving towards that direction. And I think the code arena is a good example of an arena that would support a full featured harness like a Devin. And so in my view, if I'm talking at Cognition, I'm saying, hey, let's get Devin on the arena and figure out how to loop together that harness. just so that we can. I'm sure that there's something that could be really valuable there, especially given Devin.
23:15Last week, people were talking about Devin dead. Did you see that? Yeah. People were saying Devin's gone. Devin's not gone. Devin's everywhere. It's doing very well. But people, so can we highlight that for people and show them, hey, Devin is actually the best or one of the best in the world at doing what it does. Almarina can actually do that. And our place as a central evaluation platform allows that to happen. Yeah, love it. All right. Thank you for owning the State of V-Vals. Thanks so much. Congrats on a wonderful year. Appreciate it. Congrats to you too. Congrats on all the growing momentum in your podcast and in your career.
23:47Thank you. It's really impressive to see.
From the publisher
We are reupping this episode after LMArena announced their fresh Series A (https://www.theinformation.com/articles/ai-evaluation-startup-lmarena-valued-1-7-billion-new-funding-round?rc=luxwz4), raising $150m at a $1.7B valuation, with $30M annualized consumption revenue (aka $2.5m MRR) after their September evals product launch.
—-
From building LMArena in a Berkeley basement to raising $100M and becoming the de facto leaderboard for frontier AI, Anastasios Angelopoulos returns to Latent Space to recap 2025 in one of the most influential platforms in AI—trusted by millions of users, every major lab, and the entire industry to answer one question: which model is actually best for real-world use cases? We caught up with Anastasios live at NeurIPS 2025 to dig into the origin story (spoiler: it started as an academic project incubated by Anjney Midha at a16z, who formed an entity and gave grants before they even committed to starting a company), why they decided to spin out instead of staying academic or nonprofit (the only way to scale was to build a company), how they’re spending that $100M (inference costs, React migration off Gradio, and hiring world-class talent across ML, product, and go-to-market), the leaderboard delusion controversy and why their response demolished the paper’s claims (factual errors, misrepresentation of open vs. closed source sampling, and ignoring the transparency of preview testing that the community loves), why platform integrity comes first (the public leaderboard is a charity, not a pay-to-play system—models can’t pay to get on, can’t pay to get off, and scores reflect millions of real votes), how they’re expanding into occupational verticals (medicine, legal, finance, creative marketing) and multimodal arenas (video coming soon), why consumer retention is earned every single day (sign-in and persistent history were the unlock, but users are fickle and can leave at any moment), and his vision for Arena as the central evaluation platform that provides the North Star for the industry—constantly fresh, immune to overfitting, and grounded in millions of real-world conversations from real users.
We discuss:
* The $100M raise: use of funds is primarily inference costs (funding free usage for tens of millions of monthly conversations), React migration off Gradio (custom loading icons, better developer hiring, more flexibility), and hiring world-class talent
* The scale: 250M+ conversations on the platform, tens of millions per month, 25% of users do software for a living, and half of users are now logged in
* The leaderboard illusion controversy: Cohere researchers claimed undisclosed private testing created inequities, but Arena’s response demolished the paper’s factual errors (misrepresented open vs. closed source sampling, ignored transparency of preview testing that the community loves)
* Why preview testing is loved by the community: secret codenames (Gemini Nano Banana, named after PM Naina’s nickname), early access to unreleased models, and the thrill of being first to vote on frontier capabilities
* The Nano Banana moment: changed Google’s market share overnight, billions of dollars in stock movement, and validated that multimodal models (image generation, video) are economically critical for marketing, design, and AI-for-science
* New categories: occupational and expert arenas (medicine, legal, finance, creative marketing), Code Arena, and video arena coming soon
Full Video Episode
Timestamps
00:00:00 Introduction: Anastasios from Arena and the LM Arena Journey00:01:36 The Anjney Midha Incubation: From Berkeley Basement to Startup00:02:47 The Decision to Start a Company: Scaling Beyond Academia00:03:38 The $100M Raise: Use of Funds and Platform Economics00:05:10 Arena's User Base: 5M+ Users and Diverse Demographics00:06:02 The Competitive Landscape: Artificial Analysis, AI.xyz, and Arena's Differentiation00:08:12 Educational Value and Learning from the Community00:08:41 Technical Migration: From Gradio to React and Platform Evolution00:10:18 Leaderboard Delusion Paper: Addressing Critiques and Maintaining Integrity00:12:29 Nano Banana Moment: How Preview Models Create Market Impact00:13:41 Multimodal AI and Image Generation: From Skepticism to Economic Value00:15:37 Core Principles: Platform Integrity and the Public Leaderboard as Charity00:18:29 Future Roadmap: Expert Categories, Multimodal, Video, and Occupational Verticals00:19:10 API Strategy and Focus: Doing One Thing Well00:19:51 Community Management and Retention: Sign-In, History, and Daily Value00:22:21 Partnerships and Agent Evaluation: From Devon to Full-Featured Harnesses00:21:49 Hiring and Building a High-Performance Team
Get full access to Latent.Space at www.latent.space/subscribe




