[State of Evals] LMArena's $100M Vision — Anastasios Angelopoulos, LMArena

31 Dec 2025 · 11 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

```markdown

Latent Space

The AI Engineer Podcast

Episode Summary

[State of Evals] LMArena's $100M Vision — Anastasios Angelopoulos, LMArena

Overview In this episode, host Anastasios Angelopoulos discusses the journey of LMArena, a platform that has evolved from a basement project to a $100 million venture aimed at becoming the leading evaluation platform for frontier AI models. The conversation took place live at NeurIPS 2025 and covers the history, challenges, and future aspirations of the platform.

Key Takeaways

Origin and Growth

  • LMArena began as an academic project under Anjney Midha at a16z.
  • Transitioned from an academic initiative to a startup model to scale effectively.
  • The company raised $100 million, primarily for inference costs, platform migration, and talent acquisition.

Platform Usage and Metrics

  • LMArena handles 250 million conversations with 5 million+ users, with millions of monthly interactions.
  • A diverse user base, with 25% consisting of software professionals.
  • Half of the users are now logged in, allowing for better data collection and user understanding.

Leaderboard Controversy

  • Addressed critiques from a paper by Cohere researchers alleging unfair practices in leaderboard evaluations.
  • LMArena refuted claims of bias and lack of transparency, asserting their commitment to an open leaderboard based on user votes.
  • The “Nano Banana” moment was highlighted as a pivotal event that positively impacted Google’s market share and validated multimodal AI applications.

Platform Integrity and Future Directions

  • Emphasized that LMArena operates as a charitable entity regarding its leaderboard, ensuring integrity and fairness.
  • Plans for expansion into occupational verticals (e.g., medicine, legal, finance) and multimodal domains (video capabilities).
  • Consumer retention is a primary focus, achieved through features like sign-in and persistent history.

Detailed Discussion Points

  1. The $100M Raise
  2. Funds are allocated for:
  3. Covering inference costs for user interactions.
  4. Migrating from Gradio to React for better performance.
  5. Hiring skilled talent across various domains.
  1. User Engagement and Demographics
  2. The platform's success is driven by a broad spectrum of user engagement.
  3. Engagement strategies include responding to user feedback and maintaining an active community presence.
  1. Technical Improvements
  2. Migrated from Gradio to React to leverage modern development capabilities and enhance user experience.
  3. The transition aimed to overcome limitations posed by previous infrastructure.
  1. Addressing Critique: Leaderboard Delusion
  2. Responded to accusations of inequities with evidence showcasing openness in model testing and representation.
  3. The importance of community involvement in preview testing was noted, enhancing user experience and engagement.
  1. Future Roadmap
  2. Plans to introduce new expert categories and enhance multimodal capabilities.
  3. Potential exploration of API availability to widen user interaction and facilitate partnerships with other platforms.
  1. Community Management Insights
  2. The importance of providing consistent value to users to retain them.
  3. Strategies such as creating a seamless sign-in process and offering persistent user history have proven effective.

Conclusion The episode encapsulates the evolution of LMArena, addressing its strategic decisions, community engagement, and commitment to integrity in AI evaluations. Anastasios Angelopoulos's vision for the platform emphasizes continuous improvement and adaptation in the rapidly evolving AI landscape.

Resources

  • LMArena Website: [lmarena.ai](https://lmarena.ai)
  • Follow them on X: [Arena on X](https://x.com/arena)

Chapters

  • 00:00 - Introduction: Anastasios from Arena and the LM Arena Journey
  • 00:01:36 - The Anjney Midha Incubation: From Berkeley Basement to Startup
  • 00:02:47 - The Decision to Start a Company: Scaling Beyond Academia
  • 00:03:38 - The $100M Raise: Use of Funds and Platform Economics
  • 00:05:10 - Arena's User Base: 5M+ Users and Diverse Demographics
  • 00:06:02 - The Competitive Landscape: Arena's Differentiation
  • 00:10:18 - Leaderboard Delusion Paper: Addressing Critiques
  • 00:12:29 - Nano Banana Moment: Market Impact
  • 00:19:10 - Consumer Retention: Effective Strategies
  • 00:21:49 - Hiring and Building a High-Performance Team

```

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Journey from Academia to Startups

0:46 to 2:39

Discussion on how Anastasios transitioned from academia to founding LMArena.

“Out of the LMSIS sort of, you know, conglomerate at Berkeley.”

Funding and Financial Strategy

2:40 to 4:14

Overview of LMArena's funding, including their $100 million raise and its intended use.

“Was there a moment for you where you were, I'm sure you were debating it yourself.”

Understanding Arena's User Base

4:15 to 5:51

Exploration of the user demographics on the Arena platform and insights gathered from them.

“No, no, we get discounts, but they are standard enterprise discounts.”

Competitive Landscape in AI Evaluation

5:52 to 7:59

Analysis of the competitive landscape including discussions about other AI analysis platforms.

“like he's the guy to correct for response bias.”

The Leaderboard Delusion Response

8:00 to 9:54

Anastasios responds to critiques from the Leaderboard Delusion paper regarding LMArena.

“it does help in terms of, I don't have to wait.”

The Impact of Multimodal Models

9:55 to 14:07

Discussion on the economic potential of multimodal AI models, including their use in marketing.

“So it funds the free usage of the platform.”

The Value of Multimodal Models in AI

14:07 to 15:37

Learn about the growing economic importance of multimodal AI models, especially in marketing.

“Infinite supplies of diagrams and explainers and infographics.”

Core Principles of LMArena

15:40 to 16:30

Discover the guiding principles behind LMArena's community-centric approach and its commitment to integrity.

“What have you decided are the core principles, I guess, before becoming a company and now that you're a company?”

Future Developments and User Expansion

16:38 to 19:16

Explore LMArena's plans for future features, including category expansions and potential API offerings.

“from real users that the community is using to study and improve on.”

Community Management and User Retention

19:17 to 21:20

Learn strategies for community management and retaining users in a competitive market.

“Well, there's obviously a need for an API.”
Show all 11 chapters

Partnerships and Collaborations in AI

21:21 to 23:54

Understand the importance of partnerships in AI development and how LMArena engages with other companies.

“And how do I build in all the retention mechanisms so that they stay and then they're also bringing their friends?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:12All right, we're here with Anastasias from Arena. I actually don't actually know your... Star 7A, it's like very... Angelopolis. Yeah. Yeah, there you go. Congrats on all the success. You got the arena handle. Yeah, we did. We got the arena handle. Thank you. Big branding moment. I mean, I think X is like being more commercial, so obviously you bought it. But at least you like have a place to go to where you can be like, hey, like we really like this. But I do think like dropping LM has changed the feel of it. I don't know how you feel. Yeah, I don't know. I mean, the reason we kept the LM at the beginning is because we were like, we started as LMSIS, right?

0:50Out of the LMSIS sort of, you know, conglomerate at Berkeley. So we decided to... Language models, yeah. Exactly. So we wanted to maybe broaden a little bit. Yeah. And we were the first Serena, so we feel like let's kind of try to own that. Yeah. Last time we had you guys on, you hadn't really spun out yet. And I did a call with Alessio and I was like, these guys are going to start a company. And I didn't know. I think you actually were already started at the time. I don't remember. Maybe. Anj, I chatted with Anj. And he said he was your founding CEO. He was, indeed. Which people don't know. Anj is a very interesting character.

1:28We have a podcast scheduled with him. He does a lot more than normal VCs. He does. He's been incredible to us. You want to shout out some stuff that you did? Yeah, absolutely. So, you know, the way the company started was as an incubation by Ange. So what he did is he kind of like found us at Berkeley and picked us out of the basement and was like, hey, these guys seem like they're onto something. And started working with us really early, gave us some grants. He was not, you know, A16 was not the only one to do this. We also had a great grant from Sequoia, but Ange was in particular quite supportive of us.

2:01and, you know, gave us some resources in order to continue building out Arena before we even were committed to starting a business. And in that capacity, he sort of like, you know, formed an entity for us and this and that. And, you know, he was like, hey, you guys can walk away at any time if you guys don't want to start a business. I mean, it was really incredible, very aggressive investment move by him, right? Because of course, any money that he spent at the end of the day, like we could walk away and leave him with nothing. But I think he knew wisely that the right thing for Arena was to start a company out of it.

2:32It was the only way that we could scale and that, you know, Waylon Yon and I would ultimately see that and be excited about doing it ourselves, which ended up being the case. Was there a moment for you where you were, I'm sure you were debating it yourself. You had other opportunities. What was the deciding factor for you? It became clear that the only way to scale what we were building was to build a company out of it, that the world really needed something like Arena, Arena being really a place to sort of measure, understand, and advance the frontier AI capabilities on real-world users, on real-world usage, based on organic feedback.

3:12And that in order to achieve the scale and distribution necessary, and the quality, of course, of the platform necessary to do this effectively, we would need to start a company out of it. You know, we considered other options. Are we going to keep doing this as an academic project? Are we going to do it as a nonprofit? Blah, blah, blah. But ultimately, under those constructs, we didn't feel like we'd have the resources necessary to accomplish our mission. So you raised$100 million or$80 million? $100 million. $100 million. That's a lot of resources. It's great. Yeah, what's it for? Well, you know, obviously...

3:45Ask on behalf of, like, everyone. Of course, yeah, everybody. Dude, it's an arena. How are you going to spend your money? Yeah. It's great. Yeah, please tell. So first of all, we don't necessarily need to spend all that money, right? The purpose of money at a company is to give you cards to flip. It's to say, hey, you have enough resources necessary so that if your first bet fails, you can make another bet and another bet, of course. So that's not to say that we're going to spend all of it. Of course, you want to spend things responsibly. Having said that, the platform is actually quite expensive to run.

4:12We fund all of the inference on the platform. You know, the way that it works is the platform... You pay market rates? They don't give you discounts? No, no, we get discounts, but they are standard enterprise discounts. The same that would be given to any other customer. And have you proposed any, I don't know any numbers. I see numbers of votes. But like, what's that in like monthly tokens or I don't know. I don't know about tokens, you know, I'd have to back that out. But we have like, you know, let's say, again, this is off the cuff, but I can safely say we have more than 5 million now, 5, 6 million.

4:45We have probably 250 million conversations that happen over the course of the platform. We're on the order of, you know, mid-tens of millions of conversations every month that are happening on the platform. It's actually quite a large, you know, it's one of the largest consumer platforms for LLMs. Of course, nothing is really comparing to ChatGPT, but, you know. But no, still, I mean, I think the largest skilled ones. And the benefit of this is that, like, it's actually quite a diverse population. So 25 % of the people on our platform, for example, do software for a living. still at this scale.

5:16How do you know? Because we like do all sorts of, we either survey them or we'll like analyze the prompt distribution that's coming to the platform. I'm happy also to share more. We've done something called Expert Arena, which is like trying to understand the distribution of experts that are coming to the platform. A lot of that can be like unauthenticated or whatever usage. It is, but about half of our users now are logged in. Yeah. And so we have some ability to understand them. And we also have surveys that we run on the platform that tell us a bit more about who are the actual users. Of course, there's always like response bias in surveys.

5:49So you have to take it with a grain of salt. But nonetheless. And if you don't know Anastasia's background, like he's the guy to correct for response bias. Yeah. Okay. There's a lot of guys like that. And girls. Guys and girls. You're not the only player. There are this artificial analysis started in the arena. Yup. AI. It's like some crypto people that started this. I don't know if you have a conversation with them on like, well, hey, this is our thing, or let's work together on something. I don't know if you've... No, you know, so I've talked to the artificial... I've actually talked to both groups.

6:20Both seem, you know, great people. Am I missing any major players? No, I think those... I don't know. Yeah, I don't think so. Okay. I think those are, like, some, of course, large depending on how you define the term. Artificial analysis obviously has, like, huge, like, market mindshare around the analysis of different AI systems. Yeah, they told me they were just like, we are going to be Gartner of AI. Yeah, that's kind of their goal. And I think they're going after that consulting market and so on. Artificial analysis, from what I understand, is like a team of consultants that are doing this.

6:51Yeah. Seem like really, really nice guys going after that particular market. It's different from our platform in the sense that their analysis is based on sort of like aggregating public benchmarks and turning those into analytics. Independently rerunning. And independently rerunning. Which matters. Which matters, yes. And using those in order to compile sort of reports and so on that educate the field on the performance of all these different models. But they also have arenas. They have arenas, but the arenas are not based on organic usage. Like, the thing that distinguishes our platform versus theirs is that the users are actually inputting their own use case.

7:34They're actually asking their own question. and that gives a level of realism that platform doesn't have. Of course, they specialize in a slightly different thing, but I see those platforms kind of diverging in that sense. Yeah, and sometimes it's like, it's the only way to do this. So like, for example, for AA, their video arena is pre-generated videos. You can't enter in your own video. That's correct, but we're doing it organically. Yeah, exactly. And so as a voter, it does help in terms of, I don't have to wait. It does, but also why would you go? Do you actually care about other people's videos?

8:09To form your own intuition. Maybe, yeah. Maybe you're interested in comparing. I'm a shitty prompter, right? Are you? I don't believe that. I'm a terrible prompter. I learn by example. Don't denigrate yourself. Don't denigrate yourself. There are many prompters much better than me that say that. That's a fact. People have all sorts of cool ways of prompting LLMs. Yeah, it's educational to say. The only way to learn is by, well, look at other people's prompts and see the results and go like, oh, I didn't know you could do that. Totally. Yeah, yeah, yeah. Okay, so let's come back to Arena. Oh, one thing I do want to say is the number one use of funds is getting off Gradio.

8:42Oh, yeah. Yes. Well, listen, Gradio, incredible platform. Gradio scaled us to a million miles. Yeah. That's incredible. And, of course. Did you tell the Hugging Face folks that? Of course. Yeah, we're really, really grateful to Gradio for taking us so far. Eventually, you know, it became time for us to move off of that and be able to react. I'm sure Hugging Face would have loved you to stay on. They would have, I'm sure. was there a technical reason you just couldn't get the performance? Yeah, it just became hard to develop, and there were all these tools that we wanted in React, and to do all the fancy things that you can do in React became kind of...

9:18One example, I don't know. What's a feature that you really wanted? Let's say we wanted to create our own custom loading icons for video with notifications. Okay. How are we going to do that in React? It's hard. Yeah, we'll make a custom component. I'm sure the Hagenfage guys are going to come in and say, like, hey, you can do that in Gradio, which maybe you can. But also fewer developers know. How are we going to hire for that? Do we have to reskill them? Do people are less familiar with that stack? Yeah. You know. So it's full React, Next.js, all this. Yeah, yeah, yeah, all that. Okay, cool.

9:49Other use of funds that might be interesting? Like, you know, basically. No, that's basically it. How you can deploy the resources. Okay. Yeah. It's on primarily inference. So it funds the free usage of the platform. And then also hiring, of course, headcount. Yeah. We have an office, you know. That's us. That's an SF. I'll tackle one of the major things this year, which I'm sure you're tired of thinking about. But for people who are not in the loop, this is going to be news to them. The leaderboard delusion, the whole thing with Cohere. Let's summarize what they said and then your response. So the leaderboard delusion is a paper that critiques Alamarina.

10:27pretty like rudely well you know I would say unscientifically and let's be curious Cohere wasn't doing that well on Cohere was like 74 it's all good you know it's actually not it's a respectable place that they had on the leaderboard I don't even think it was really Cohere people like the Cohere model developers doing this it was more their research side but in any case what does the leaderboard illusion say it says that Alamarina was the claim is that Alamarina was doing this undisclosed, quote-unquote, private testing on our platform, that model providers will send us pre-release models and we'll expose them and so on and so forth.

11:04And that this creates so-called inequities in the leaderboard, you know, due to that pre-release testing. For example, they cited that Meta at some point tested some amount of models with us. Yeah. Of course, we can't disclose all of the details of how all that was done. But that is the main claim of the paper. Now, our response to that paper, it's online. you can find it on response to leaderboard illusion. And our response to that paper is essentially pointing out a series of factual mistakes in the paper that question the validity of the claims. So you can go look at the first version of the paper on Archive yourself, and you'll see the claims.

11:40I think most scientists would view that as... Oh, they've corrected it. They've corrected, of course, because we, I mean, but they didn't correct everything. They just corrected some aspects that were just blatantly unscientific and false. But, you know, they, for example, said that we only sampled like 9 % open source models and like 60 % closed source models and just created a gap between open and closed source. But in reality, we're actually really supportive of open source models. And it was more like 60-40. And so that was one of the examples of an error in the claims. Another example is that they were claiming that there was some sort of bias introduced by this pre-release testing and that it was undisclosed.

12:22In reality, as you probably know, we've been doing this pre-release testing for a long time. Our community loves it. They love basically getting like... The secret code names. Yeah, the secret code, like Nano Banana, all that. So Nano Banana, by the way... It started on you. It started on us. Right, and people loved it. It went like global sensation, like non-zero fraction of the global population using Nano Banana. Did you talk to Lena about naming it Banana or was that their decision? No, it was their decision, I believe, but it was sort of this randomly generated thing. And it just went... No, no, so apparently, uh nana who's a pm yeah yeah is named after her because her nickname is nano oh oh that's sweet i didn't know that i was the banana put banana on it yeah oh that's sweet yeah i didn't know that didn't know the origin story yeah i mean to us it just looked like sort of a random thing but it was like clearly heads and shoulders of both huge which like before that there was reeve image remember yeah i do of course yeah and also bfl and all those i mean all those models are also great And I think those teams are also improving quite quickly.

13:17But Nano Banana was a sensation. I mean, that moment alone changed Google's like... Roadmap. Yeah, market share. Seriously. I mean, Google stock, billions of dollars are moving because of Nano Banana. And now there's like an opening, Code Red and everything. I don't know about that, but yes, the information reported this. I would say the image generation, I would say, has been this weird part of AI overall. Because it's not strictly AGI critical. It's not reasoning. It's not feeding more context into the model. It is the model generating a visual representation. So it's basically like, I always think like, well, you know, Gemini used to get a lot of complaints for generating racist images or whatever.

14:06yeah that was a hilarious moment and and chat gpt also had it had it in the past and i'm like well can we just get rid of this like do we have to do image generation because let's just focus the the positive reputation of ai in general on language models and coding and you know the other stuff yeah um but i i i'm wrong i'm i'm such a huge nanon banana pro show yeah i totally agree i was also kind of wrong about this i didn't see the positive benefits but actually i think that these multimodal models are going to become some of the most economically valuable aspects of AI, both in consumer and also in enterprise, because one of the fastest growing segments, market segments in AI adoption is marketing and marketing and design.

14:49Yeah, ads. So I'm a content creator, right? Yeah, of course. I'm sure you're using it all the time. Infinite supplies of diagrams and explainers and infographics. Yeah. Soon we're not going to be even making the papers. You can't make the emails. We're just going to be, our paper figures are going to be made by an animal. Actually, yes, yes. I do think that actually one shot. So DeepSeek came out with V3.2 recently. I took their explanations, which are very wordy. They like very concise papers. It's 23 pages long, but it's very dense. And so I just took their explanations of their our environment stuff and I fed it into Nano Banana Pro and I spent an image that I used to understand the paper better.

15:24And the fact that I can just casually generate like a paper quality diagram that would usually take a PhD student in like a month in Photoshop or something to do is incredible. Yeah, it's incredible. It is amazing. I want to ask about your principles running Arena. I think you manage a giant community, 5 Million Mal. What have you decided are the core principles, I guess, before becoming a company and now that you're a company? I don't know if there's anything that's changed for you. I don't think anything has really changed. We want to provide the North Star of the industry and center the use cases of real users, foreground those so that people know what to target.

16:03The goal is to create a benchmark that is constantly fresh, that does not suffer overfitting because of the fact that we constantly have new data points coming in, that tracks all the different new models, all the different new use cases of AI, and gives the whole world sort of ground truth for how real users are using these models and how good they are on those use cases. We continue to do quite a few open source data releases, we've probably released more data than basically anybody on the real world use cases of AI. Millions and millions of conversations, real world conversations from real users that the community is using to study and improve on.

16:46Yeah. And then I think like in terms of what you will build versus what will not build, I guess I'm not necessarily caught up on everything that you've launched. I know recently you've done the dev or code arena. Yeah, code arena. That's the most - Code arena, expert arena. Yeah, expert arena. So basically like what is in the critical path for you, let's say for next year, and what have you decided you'll never do? So let me first talk about things that I'll never do. The platform, integrity comes first to the platform. Basically the public leaderboard that we show on LM Arena, I think of as a charity.

17:21It's a loss leader for us. We don't really make money on the public leaderboard. You can't pay to get on the public leaderboard. It's not like a Gartner in that sense. It's not like any of these pay-to-play systems. Never going to be like that. Models are going to be listed on the leaderboard whether or not the providers pay and whether or not they're getting a good score. They can't pay to take it off either. And so what that means is that the leaderboard has a certain integrity that will never be compromised. But not all preview models will make it onto the... No, but that's okay. Those preview models have never been released.

17:57Yeah, yeah. Right? Who cares about putting unreleased models on the leaderboard? The point is that for every released model, the score that you see on the leaderboard is statistically sound. It reflects the real-world capabilities of the model. Yeah. Why? Because millions of people from around the world have voted for it. And that's where that number comes from. All we do to compute that number is millions of people are voting. We take those votes, we turn them into a number. that's always going to remain, you know, a transparent and fair reflection of model performance. Where are we going? Lots of different new categories.

18:30I don't know if you recently saw we expressed, we, um, we exposed, uh, occupational and expert categories. So now single digit percentage of our user base, we're millions, millions to tens of millions of users, right? So single digit percentages means a lot. Single digit percentage of our user base are in medicine, in legal, in business, finance, accounting, creative, marketing, stuff like this. And we're able to show the performance of these models in all these different verticals because we have all these users in our user base. And we're working more towards multimodal. Video, we're soon to launch on the site at some point later this year or early next.

19:06So lots of things in the pipeline. Amazing. Would you expose an API? We've thought about it, yeah. I think it's a possibility. What are the counter arguments? Why not? Well, there's obviously a need for an API. The question is more of focus of our company. Just because we're a startup, and so we really should be doing one thing well. Arenas. Yeah, arenas. So I'm not sure how far we want to sort of splay out and what timeline we would want to do that. Yeah. Any other sort of like community management tips? More broadly, every AI company really wants to grow their community. You're obviously one of the strongest in the world.

19:49What's really, really worked? Well, so first of all, I want to give a shout out to our community manager, Greg, who is doing an awesome job managing our community, whether that's on Discord or on Elamarina. He's really incredible. So I would say hire Greg, but don't hire Greg. Don't hire Greg. Don't hire Greg. He's ours. Find a Greg. Find a Greg. But in general, the question of how do you get to so many users and keep and retain them. That is a tough question because consumer is one of the hardest markets in the world. There's a lot of websites in the world that people can go to. Why should they go to yours?

20:23And the reality is if you want to create a really dominant product, you have to provide people value. And to be frank, I don't think we're all the way there yet. It's not like I have the solution and answer for how to build a great consumer product. If I did, we wouldn't be at tens of millions of users. We'd be at hundreds or, you know, we'd be at a billion users. Is there a world you like are bigger than Chachapiti? I don't know. I don't know that we need to be. And I don't know that we ever will be because that's an extraordinary generational product that they built, right? And it took a lot of time.

20:57And then to some extent, it also involved luck. There's a lot of lightning in the bottle moments like NanoBanail was for us where our user base just like goes up by a lot. But when those users come, they can just as easily leave. So the way I think about it is every user is earned. You have to earn them every single day. They can leave at any moment. They're fickle. And so all the time you have to be thinking about how do I provide this person value? Learning. How are they using my website? What more could I give them? And how do I build in all the retention mechanisms so that they stay and then they're also bringing their friends?

21:28Is there one that's working in terms of retention? Like you said, a lot of people are signing in now. Yeah, sign-in was a big driver of retention. No, no, but what did you give them in order to encourage? Oh, like history. Persistent history. That's it. That's enough. Yeah, that's one thing that has had a big impact. Okay. Yeah, cool. What do you want from people? What are you looking for help on, like any calls to action? Yeah, we are always looking for people to come and join us. If you are one of the best people in the world in your area, whether that's consumer product, whether that is machine learning, whether that is, you know, B2B, go-to-market, marketing, all these things we need you at arena we're building like a high performance team of real experts in everything that they do and you know i'm always looking for excellent people to work with you need like what are partnerships right like let's say i'm at cognition i want to partner with a little arena or just arena what works for you what existing partnerships do you already have that that's really fruitful yeah so i mean we of course partner with all the major model labs yeah um and that's that's just straightforward like hey we have a new model here exactly so i think the the most straightforward thing would be for someone like cognition it's like let's evaluate that agent yeah but we should be continuing to shape our well code arena is an agent evaluation and in fact i think focused on like all these arenas tend to focus on the model rather than the harness so but that maybe should change maybe we should be evolving towards that direction and i think the code arena is a good example of an arena that would support a full featured harness like a devin yeah and so in my view i'm if i'm talking at cognition i'm saying hey let's get devin on the arena and figure out how to you know loop together the so that we can i'm sure i'm sure that there's something that could be really valuable there especially given devin last week people talking about devin dead did you see that yeah people were saying devin's gone devin's not gone it's not gone devin's everywhere so very well but people So can we highlight that for people and show them, hey, Devin is actually the best or one of the best in the world at doing what it does.

23:30Almarina can actually do that. And our place as a central evaluation platform allows that to happen. Love it. All right. Thank you for owning the State of Evalues. Thanks so much. Congrats on a wonderful year. Appreciate it. Congrats to you, too. Congrats on all the growing momentum in your podcast and in your career. Thank you. It's really impressive to see.

23:54Thank you.

From the publisher

From building LMArena in a Berkeley basement to raising $100M and becoming the de facto leaderboard for frontier AI, Anastasios Angelopoulos returns to Latent Space to recap 2025 in one of the most influential platforms in AI—trusted by millions of users, every major lab, and the entire industry to answer one question: which model is actually best for real-world use cases? We caught up with Anastasios live at NeurIPS 2025 to dig into the origin story (spoiler: it started as an academic project incubated by Anjney Midha at a16z, who formed an entity and gave grants before they even committed to starting a company), why they decided to spin out instead of staying academic or nonprofit (the only way to scale was to build a company), how they're spending that $100M (inference costs, React migration off Gradio, and hiring world-class talent across ML, product, and go-to-market), the leaderboard delusion controversy and why their response demolished the paper's claims (factual errors, misrepresentation of open vs. closed source sampling, and ignoring the transparency of preview testing that the community loves), why platform integrity comes first (the public leaderboard is a charity, not a pay-to-play system—models can't pay to get on, can't pay to get off, and scores reflect millions of real votes), how they're expanding into occupational verticals (medicine, legal, finance, creative marketing) and multimodal arenas (video coming soon), why consumer retention is earned every single day (sign-in and persistent history were the unlock, but users are fickle and can leave at any moment), the Gemini Nano Banana moment that changed Google's market share overnight (and why multimodal models are becoming economically critical for marketing, design, and AI-for-science), how they're thinking about agents and harnesses (Code Arena evaluates models, but maybe it should evaluate full agents like Devin), and his vision for Arena as the central evaluation platform that provides the North Star for the industry—constantly fresh, immune to overfitting, and grounded in millions of real-world conversations from real users.

We discuss:

The $100M raise: use of funds is primarily inference costs (funding free usage for tens of millions of monthly conversations), React migration off Gradio (custom loading icons, better developer hiring, more flexibility), and hiring world-class talent

The scale: 250M+ conversations on the platform, tens of millions per month, 25% of users do software for a living, and half of users are now logged in

The leaderboard illusion controversy: Cohere researchers claimed undisclosed private testing created inequities, but Arena's response demolished the paper's factual errors (misrepresented open vs. closed source sampling, ignored transparency of preview testing that the community loves)

Why preview testing is loved by the community: secret codenames (Gemini Nano Banana, named after PM Naina's nickname), early access to unreleased models, and the thrill of being first to vote on frontier capabilities

The Nano Banana moment: changed Google's market share overnight, billions of dollars in stock movement, and validated that multimodal models (image generation, video) are economically critical for marketing, design, and AI-for-science

New categories: occupational and expert arenas (medicine, legal, finance, creative marketing), Code Arena, and video arena coming soon

Consumer retention: sign-in and persistent history were the unlock, but users are fickle and earned every single day—"every user is earned, they can leave at any moment"

—

Anastasios Angelopoulos

Arena: https://lmarena.ai

X: https://x.com/arena

Chapters

00:00:00 Introduction: Anastasios from Arena and the LM Arena Journey
00:01:36 The Anjney Midha Incubation: From Berkeley Basement to Startup
00:02:47 The Decision to Start a Company: Scaling Beyond Academia
00:03:38 The $100M Raise: Use of Funds and Platform Economics
00:05:10 Arena's User Base: 5M+ Users and Diverse Demographics
00:06:02 The Competitive Landscape: Artificial Analysis, AI.xyz, and Arena's Differentiation
00:08:12 Educational Value and Learning from the Community
00:08:41 Technical Migration: From Gradio to React and Platform Evolution
00:10:18 Leaderboard Delusion Paper: Addressing Critiques and Maintaining Integrity
00:12:29 Nano Banana Moment: How Preview Models Create Market Impact
00:13:41 Multimodal AI and Image Generation: From Skepticism to Economic Value
00:15:37 Core Principles: Platform Integrity and the Public Leaderboard as Charity
00:18:29 Future Roadmap: Expert Categories, Multimodal, Video, and Occupational Verticals
00:19:10 API Strategy and Focus: Doing One Thing Well
00:19:51 Community Management and Retention: Sign-In, History, and Daily Value
00:22:21 Partnerships and Agent Evaluation: From Devon to Full-Featured Harnesses
00:21:49 Hiring and Building a High-Performance Team

More from Latent Space: The AI Engineer Podcast

All 247 episodes
[State of Evals] LMArena's $100M Vision — Anastasios Angelopoulos, LMArenaLatent Space: The AI Engineer Podcast
Listen in VO