In short
Reuters-reported possibility that China stops releasing open-weight/open-source AI models globally, and how that could reshape the AI ecosystem, costs, security, and benchmarking. The episode also discusses “peak benchmark” concerns and why post-deployment evaluation matters, using Arena and Hippocratic AI examples.
Guests and backgrounds
- Anastasios Angelopoulos, founder of Arena (formerly LM Arena), which tracks comparative AI model performance via real-world traces; Arena raised about $150M in a Series A.
- Munjal Shah, founder of Hippocratic AI (clinical voice agents for healthcare). Hippocratic AI has raised over $400M (including $126M at ~$3.5B valuation) and released Polaris Err model version 5.
Key claims
- Open-weight models from China are “frontier quality” and enable multi-model architectures that are otherwise too expensive with closed models.
- If China pulls open-weight access, Western open-weight ecosystems (e.g., NVIDIA NeMo, Mistral, Google Gemma, etc.) would benefit, while enterprises dependent on Chinese open weights would lose.
- Benchmarks are insufficient without production feedback; benchmarks should be vertical- and outcome-specific.
Notable examples
- Hippocratic AI uses ~31 models in parallel for voice safety; GPT-4o-level models would cost ~$105/hour vs ~$9–$10/hour product pricing.
- Healthcare failure cases: drug-name disambiguation and background TV noise; also slurred speech post-stroke and planned “sobbing detection.”
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOOpen Source AI Models from China
0:00 to 0:33
Exploring the impact of Chinese open-source models on the AI landscape.
“The best open source models are coming from China.”
China's Potential End to AI Model Release
1:42 to 4:12
Discussion on the implications of China potentially stopping the release of AI models.
“I'm so glad to have you both here because AI news is dropping like you wouldn't believe.”
Impact on Startups and Costs
4:12 to 6:37
Exploring how the changes in AI model availability affect startups and their operational costs.
“We do massive post-training to those models.”
Performance Comparisons of AI Models
6:37 to 10:41
Comparing the performance of Chinese models against their Western counterparts.
“If you ran it all on Fable 5, how much would that cost you?”
Future of Open Source AI
10:41 to 14:00
Discussing the future prospects of open-source AI amid potential restrictions from China.
“we might lose overnight if these models were kind of taken off the market.”
Exploring Open-Source AI Drought
14:00 to 14:58
The discussion focuses on the implications of potential restrictions on open-source models from China.
“There are no voice-to-voice non-cascaded models that are good.”
Consequences of AI Dependency
14:58 to 16:19
Analyzing the risks of American businesses relying on Chinese open-source models.
“we would be in essentially a open-weight model drought.”
Monetization Challenges for Open Weights AI
16:19 to 18:18
Addressing the difficulties in monetizing open-source AI models and potential business strategies.
“I think today, about how Alibaba is really struggling to monetize its Quinn group.”
The Cost of Developing AI Models
18:18 to 22:30
Discussing the financial and logistical challenges of continuously developing AI models.
“Uh, Moonjil, if you had to cough up a chunk of your revenue to a open weight American or Western AI lab, would that dramatically change your economics?”
The Future of Open-Source AI
22:30 to 24:28
Exploring the necessity for a strong US presence in open-source AI against China's advancements.
“capitalize properly to do that unless it just becomes a frontier lab, effectively.”
Show all 32 chapters
Benchmark Evolution in AI
24:28 to 28:00
Discussing the saturation of benchmarks in AI and their relevance in practical applications.
“It's still an inferior model, technically speaking, to the Frontier models.”
Exploring Model Intelligence and Latency
28:00 to 28:50
Discuss the need for varying levels of intelligence and latency across different applications.
“Because how do you make a model smarter without deep thinking and planning and all the other things?”
Healthcare's Unique Needs and Benchmarks
28:50 to 30:02
Highlight the critical need for high intelligence and low latency in healthcare AI models.
“On the other hand, but in healthcare, we needed it.”
Performance, Cost, and Latency in Model Evaluation
30:02 to 31:36
Discuss how performance, cost, and latency are crucial factors in assessing AI models.
“By giving you your bank balance, I don't need it to be that smart.”
Innovations in AI Benchmarking and Model Comparison
31:36 to 33:26
Explore how the evaluation of AI models can be enhanced through innovative benchmarking.
“And that's happening, you know, millions and millions of times around the world.”
The Importance of Customized Benchmarks
33:26 to 34:54
Emphasize the need for tailored benchmarks to capture unique challenges in AI applications.
“Now, then we went to the cascaded model where you have an ASR that takes the voice and then converts it into text and put it in the LLM.”
Continuous Improvement in AI Safety Standards
34:54 to 36:32
Discuss the ongoing necessity of improving AI model safety and reliability in healthcare.
“that puts AI to use inside their product or service, regardless of the industry or vertical.”
Real-World AI Failures and Lessons Learned
36:32 to 37:36
Share insights on real-world AI failures and the learning outcomes from those experiences.
“Luckily, right now, we've done 200 million patient interactions without one significant safety incident.”
Evaluating Drug Name Stability in AI
37:36 to 39:28
Investigate the challenges AI faces with drug name recognition and its implications.
“Thanks to the generosity of a good friend.”
The Critical Role of Intellectual Property in AI
39:28 to 40:03
Examine how intellectual property influences benchmarking and competitive advantage in AI.
“spot of what benchmarks have we not created and what are the queries in the benchmarks.”
Challenges in AI Voice Technology
42:00 to 44:20
Explore the complexities in AI voice technology, especially in healthcare.
“is it's actually if somebody, a 14-year-old being called by the AI or a 15-year-old over the thing about their healthcare says it's not safe here.”
Data-Driven AI Performance
44:20 to 47:00
Understand how post-deployment data can enhance AI performance and benchmarking.
“back to the five nines point, five nines down to one nine, nine nine.”
Unique Features and Market Positioning
47:00 to 49:40
Learn about unique AI features that improve safety and performance in clinical settings.
“of a model's strengths and weaknesses and analysis that helps businesses build better models and understand which models to choose.”
The Future of AI Benchmarking
49:40 to 53:00
Discuss the limitations of traditional benchmarking and the need for data-driven approaches.
“It's like, well, how often do you have to do child or adult protective services?”
Personalization in AI Interactions
53:00 to 56:00
Examine how AI can learn from individual user interactions to enhance communication.
“Now, hopefully China can fix the Mandarin thing.”
AI Learning and Patient Interaction
56:00 to 1:01:11
Exploring how AI can personalize patient interactions based on past communications.
“And then you kind of learn that, oh, no, this is serious because you because there's something going wrong.”
Healthcare Revenue Milestones and IPO Discussion
1:01:11 to 1:05:23
Discussion on the recent revenue success of a healthcare company and thoughts on going public.
“a revenue milestone well done putting you on track to be public company size whenever you'd like to be.”
Investing in the Next Generation
1:05:23 to 1:10:01
Discussion on the importance of teaching financial literacy and investing to young people.
“Is that just the kind of the de facto view?”
Nationalization vs. Buying Assets
1:10:01 to 1:13:07
Exploration of the implications of nationalizing versus purchasing assets in business.
“Anyways, well, at least he's buying them instead of just saying, I take over all of your oil producing.”
AI's Impact on Healthcare
1:13:08 to 1:16:12
Discussion on how AI can revolutionize healthcare and improve patient outcomes by providing abundant resources.
“Well, yep, I think, well, I have a lot of thoughts about that.”
The Importance of Patient Monitoring
1:16:13 to 1:18:42
A personal anecdote emphasizing the need for proactive patient monitoring to prevent healthcare emergencies.
“What we do is we send them to the cooling center, which are set up all over the country.”
Future of AI and Politics
1:18:43 to 1:19:51
Discussion on the potential political implications of AI advancements and the importance of informed voting.
“like, okay, but her doctor knew she wasn't refilling for five years because she didn't ask for a refill.”
Transcript
Automatic transcript. May contain errors.0:00The best open source models are coming from China. It's basically going to hurt the AI strategy. Tons and tons of companies that are building with an open source first mindset. Mathematically untenable. One model GPT-4O's level from open AI is$6 an hour. So you can't use 31 because the math blows up on you unless you use open source. If you run our constellation all on open AI models, it'll cost you$105 an hour. That's more than the human. The real power of all of this AI is not replacing work you do today and making it cheaper. That's actually not working super well. Thanks to our friends at PayPal, the exclusive sponsor for This Week in AI.
0:36Try the payment and growth platform that's trusted by millions of customers worldwide. PayPal Open. Start growing today at PayPalOpen.com. Hey, everybody. Welcome back to This Week in AI. My name is Alex. And today we are talking about the possible end of open-way models from China, AI lab bearishness, benchmarks, the great AI data race, and more. To help us understand it all, we've brought two leading AI startup founders to the show today. In one corner, we have Anastasios Angelopoulos from the company Arena, now previously known as LM Arena. It's a brilliant data set for tracking comparative AI model performance.
1:10It's now a big darn startup, having raised$150 million in a blockbuster series A. Anastasios, welcome back to the show. Happy to be here. Glad that we have you back. We also have Munjal Shah from the hot shores of vertical AI. His company, Hippocratic AI, is working on all things clinical voice agents in the healthcare space. With more than$400 million raised, including$126 million at a$3.5 billion valuation recently, and version 5 of its Polaris Err model out, the company is shooting for Megascale. Munjal, welcome to the show. Thanks for having me. I'm excited to be here. I'm so glad to have you both here because AI news is dropping like you wouldn't believe.
1:45The latest thing that's going on that we all have to talk about is the possibility that Reuters reported this morning that China may end the release of its AI models to the world. Now, we all know that Chinese open-weight models have been incredibly clutch, super powerful, super cheap. A lot of startups build on them. And if that ends, it does seem to kind of reorient or reorder the broader AI landscape. Anastasios, starting with you, first reactions to this news, because I think we're all still digesting a bit. I mean, listen, a lot of people are going to be made really upset by this because open source models, the best open source models are coming from China.
2:20And so it's going to do a couple of things. The first thing it's going to do is it's basically going to, you know, hurt the AI strategy of tons and tons of companies that are building with an open source first mindset. The GLM is like basically a frontier quality model. It's comparable to a lot of the frontier offerings from open AI. GLM 5.2. Yep, that's right. GLM 5.2. And so, you know, if we can't get 5.3, 5.4 or GLM 6 or whatever, you know, what have you, what's going to happen is that people are going to have to reconsider what models to use, whether they're going to go with a closed source strategy or, you know, and this is the second thing that's going to happen.
2:55It's, of course, going to promote open source development inside the U.S. and Europe of models that, you know, we can actually rely on to be released within the confines of our actually sovereign nation. So do you think that this is a net benefit then to American and European AI labs or also American and European AI startups? Well, yeah, that's exactly the right question because there's a bunch of game theory you need to play out if you're China. If you're China, you should be thinking, okay, how do I make China win? And so if you shut off open source AI now, let's be honest, China is doing well there, but they're not exactly winning the AI race yet.
3:35So by shutting it off and not giving Chinese companies the ability to compete on the global stage, not giving them the ability to sort of access global data, not giving them the ability to access global customers, What they might actually be doing is sort of, you know, cutting the budding trees that they have in China related to AI and then promoting the development of Western open source AI. So to some extent, it might actually be good for our national security. Munjal, where do you stand on the winners and losers from this? I know your company makes its own models. So I presume you have a foot in all open and closed source camps.
4:11Yeah, actually, you know, so we only use open source as our base. We do massive post-training to those models. We put them in a configuration of 31 models together. And that's how we ensure clinical safety. We actually found one model kind of wasn't enough. And one model open source without significant modifications wasn't enough. Like we actually modify 100 % of the parameters in the open source, open-weight models that we have. And so, you know, we started originally by building on, you know, the meta open source products. And those were state-of-the-art at one point. But we have migrated to actually working and leveraging the Chinese open source models at this point because they are the best.
4:52And so, one, I think one of the big implications is open source today allows startups to compete with the big model companies. And so this will actually destroy their ability to do that, including ours. Like, it's not a great thing if this happens. Second, people don't realize the cost implications. So, you know, we sell our AI voice agents for about$9,$10 an hour. Okay. We're using 31 models in parallel because for voice, you need very low latency, right? So you got to fire it all at once. Now, just one model, GPT-4O's level from OpenAI is$6 an hour. So you can't use 31 because the math blows up on you unless you use open source.
5:38but actually using those 31 creates a level of safety. We've not been able to do with one model, no matter how smart it is, because you have a model checking, a model checking, a model kind of thing. And models staying very focused on just doing one thing, like looking for overdoses in a call. And so, you know, while the models can scale the context windows, you can't scale the attention span nearly as well. Right. You know, you put in a million tokens. Great. Which tends to focus on the last 10 ,000 or the first 10 ,000 or a random 10 ,000. You don't know. No one knows that. One million context window.
6:07So context window being good, no, not always, right? But I think that now you can't run multiple models if you don't have open source. And so there's an entire class of use cases that actually need a multi-model architecture to ensure safety, to accuracy at a high precision that become mathematically untenable. If you run our constellation all on open AI models, it'll cost you$105 an hour. That's more than the human. Yeah. If you ran it all on Fable 5, how much would that cost you? Because I'm curious. $10 billion an hour. $10 billion an hour. I mean, but so look, I think there's a cost implication.
6:52There's that. The other part is what Alex Garp was saying last week, which is that a lot of the businesses don't want to share all the queries back with some of the frontier model companies and open source hosted in your own environment on your own servers is which is what we do by the way we take the open source and run it in our own place we don't run it you know because we modify it so much we have a different um a different set of weights by the end but um that was is also so there's a security implication too uh which in healthcare you know one does care a lot about the security and the privacy of the data i just want to double click on something when you say you use 30 plus models in in harmony with one another.
7:29Does that mean you have post-trained each one of those from an open-weight base? Or is there one model that you trained heavily and then sharded into smaller components? Largely, most of them, we've taken an open-weight base of different sizes because we don't need all of them to do the same thing. One is checking and ensuring that HIPAA authentication is always done correctly. One is an escalation engine that actually transfers the call to a human if you're having chest pains and run something called the Schmidt Thompson Protocol, which is what nurse triage call centers run to ensure safety. I mean, we're one of the few LLMs that will literally hand off properly, will run an escalation, will hand off the call to human if needed.
8:13But we do not, not all of the 31 are post-trained or fine-tuned. There's, I think, a few of them that are just a good set of prompts on an open source model, but that's a very small number out of the total. Do any of them qualify as an SLM, a small language model? I'm just curious if there's any place in your multi-model setup. No. Okay. Yeah, no, this is a fallacy. I mean, it's like, I'm a small language model. I'm like, that's not a thing. This thing only worked because it was a large language model, folks. And so what we found in our system, for example, our entire constellation is now 5 trillion grammars.
8:57And so what we found was every time you use a big model, it catches more out of distribution things better, which in healthcare matters. We had the other day a scheduling call where everybody's like, scheduling? That's not hard. How difficult is that? Well, this guy called up. He said, I need an appointment with my neurologist. The AI was like, tell me why you need appointments so I can get you the right appointment type. He's like, I was struck by lightning yesterday at work. Huh? And the AI had to reason because the answer isn't send him to the ER 911. It's the next day. He's obviously okay.
9:28Not an imminent risk. But it's also not schedule him six weeks from now. It's like, you better get him in today or maybe tomorrow, a more high priority appointment. The AI, that's not in the rule set. If struck by lightning yesterday and still standing, do this. That's the power of generative AI. You need big models to do that. And everything else, they won't handle those. And so I think the smallest one we use is 7 billion parameters. The biggest one we're using is over a trillion in an MOE structure. A mixture of experts, which is a way to reduce the overall processing load of a larger model by looking for one slice of it that has a specific expertise.
10:11Anyways, Anastasios. So on the point of startups using a lot of open weight models to build their own thing, there's been a lot of conversation about how open models from China have been getting closer to what I'll just call the American frontier performance. I'm curious from the arena perspective how true that is because on one hand, it does seem that there has been a closing gap, if you will, but also it seems like a bit like that truck gif where the truck never actually hits the pole. They're never actually going to quite get there. So I'm just curious about how much performance we might lose overnight if these models were kind of taken off the market.
10:46Well, it's absolutely true that China has continued catching up to the US frontier. And so there's a lot of the reason why people are saying that is because of ARENA. They're looking at ARENA and saying, hey, GLM 5.2 is, yeah, exactly, is near the top of the leaderboard. Right now, if you look at agentic performance, we have agent ARENA. And Agent Arena is measuring the ability of models to do general purpose agentic tasks, like, you know, the quad code type tasks or quad code work type tasks in your browser. So we have millions and millions of traces that we collect every week that allows us to assess these capabilities.
11:20And what you'll see if you do that is that GLM 5.2 is about a GPT 5.5 level model. Which is impressive because that just came out. Which is super impressive. Yeah. And, you know, it's probably like tied, around tied with 5.5 high, which is pretty incredible performance. The X high version of GPT is probably a little bit higher. And what you'll see is that it doesn't quite have the same level of tax success rate. And yeah, you pulled it up on the screen here. It doesn't quite have the same level of steerability, but it's quite close. and especially with a little bit of fine-tuning can actually exceed the performance of these models on an enterprise's workload.
12:04So it's definitely true that China's catching up. And what happens if you lose these models? Well, the next best open-source model, I mean, let's look down the leaderboard. If you exclude China, it's like you're losing ZAI, you're losing Kimi, you're losing DeepSeek, you're losing Minimax, you're losing Quen. And then where are you? Like Gemma. Well, no. Well, kind of. I think you're more at NVIDIA's Nemetron models, which have been pretty... Oh, Munjal is nodding his head back and forth. Munjal, what's your opinion on the NVIDIA Nemetron model family? Because I think you have one. No, no. I mean, look, I think it's awesome that they're doing it.
12:45I really wish that they continue to progress and get to the same level that we're seeing with the Chinese models, right, in terms of performance. because we need it. But I think they have yet to release Nemetron Big or whatever their version of the largest one is, if I remember right, unless it came out already. Anastasia, you'd know, I guess, better than me. But we need these open source. And actually, there's a different reason we need them. I love these leaderboards because they allow us to rank, but I also believe the coverage of all the tasks you would need the model to do that's covered by the benchmarks is very small.
13:20And so, you know, there's this artificial jagged intelligence concept that kind of says, hey, you know, these models are good at some things, really bad at other things. As they've been improving, what we notice on the tasks that we need for healthcare activities and automatic conversations with patients is that the peaks get better, but the troughs don't move much. And so we have to do a lot of post-training in RL to move the trough up. And so, you know, Because if you think about all the possible outcomes of a model that are tested in terms of the benchmarks, it's still a very small coverage area.
13:55And we try to get benchmarks that are indicative of other areas, right? But it's not perfect. And so we have actually built a whole set of new benchmarks that we're using in the healthcare setting and then actually starting to benchmark all the different models on them, including the voice-to-voice models, by the way, which are really dumb. There are no voice-to-voice non-cascaded models that are good. everybody's like they sound amazing i'm like they do but they're dumb as heck and um and so i mean i think that uh we just have to realize that the other value of open source is that it allows you to improve these troughs and in in kind of the blind spots where the current benchmarks are not kind of measured we're gonna get to benchmarks in just a second so i have questions about that but i want to make sure that we're getting this right if the chinese government does decide to preclude global access to open-weight models from its myriad, high-performing, sometimes public, AI labs, there isn't an obvious immediate replacement from, broadly, the West.
14:57And so we would be in essentially a open-weight model drought. Is that fair, Anastasios? Yeah. I mean, listen, who are the losers? The losers are all of the enterprises that are building on those open-source models now and don't have a good alternative. And the winners are the current American open source ecosystem, which is not, you know, it's not caught up with China, but likely will if all of the traffic moves to them. So those players would be players like Nemotron. It'd be RC. It'd be Mistral. It'd be Google with Gemma. Reflection AI and theory. And reflection. Mistral is French, but I mean, we'll count it in our bucket.
15:35They're part of NATO, so it's all the same. I mean, one big happy family, right? There's no tensions whatsoever in that domain. the 51st state yes um i thought that was canada but you know we'll take them now um if there isn't a replacement for these and no company curly can match we're talking about seeing once again a lot of american businesses being overly dependent on china for a key piece of their operations and i feel like we kind of ended up in the same place worth manufacturing with ai and it's just disappointing but the reason why i'm a little bit skeptical of your claim that american open ai open-weight AI, just let's say it that way to be clear, will catch up as quickly as even in China, they're having a hard time monetizing these open-weight models.
16:18There's a story in the Times, I think today, about how Alibaba is really struggling to monetize its Quinn group. And if you look at the Minimax and Z.ai's earnings as public companies, very small compared to American closed source. So I wonder if there's even a financial incentive to do this, or if we're just going to see companies like Hippocratic kind of stuck without having the same ingredients to keep making improved soup with these open weights. Yeah, I think it's a great question. It's a question that I brought up actually many times in various podcasts, which is like, how do you think about the monetization strategy and the business model behind open source?
16:55So just to get deeper into the question, what is the reason why we're asking this? With software, the reason why you can create an open source business model, you can say, okay, I'm going to have this open source software, and then I'm going to become the best place to run this open source software. I'm going to, for example, have Spark, and then I'm going to build Databricks on top. Databricks is not just Spark. Databricks has like a huge and thick value layer on top of Spark. They're the best place to use Spark. They're the best place to store your data, blah, blah, blah, blah, blah. Apache Spark, the open source project.
17:27Exactly. And then that's how you build your value. Now, with open weights AI, it's a little different because you can just download the weights. Nobody else can contribute to the open source project, really, because nobody has access to the data and the compute and the blah, blah, blah to be able to train a big model. And then you just can run it on any, basically any cloud. Yeah. And so there's really, you know, it's unclear what the business model will be. Now, what I have heard that people are doing in order to try and create an actual business model around open weights AI is to do these licensing agreements where in the license of the model, they have a, basically, if you spend enough or if your company is large enough or earning enough revenue, that the fact that you're using this open weights model will entitle you, entitle the company to a rev share.
18:17So they'll start saying, okay, if you're like a$10 billion company with over a billion dollars in revenue and blah, blah, blah, then 20 % of your revenue goes to us because you're using an open-weight model. Uh, Moonjil, if you had to cough up a chunk of your revenue to a open weight American or Western AI lab, would that dramatically change your economics? Or could you afford a 20 % margin hit on your inference costs and still offer your product at an attractive rate? It wouldn't be a big deal. I mean, I think that it's so much cheaper than running it on the closed source frontier models. It's just so much cheaper.
18:58Because I mean, it's very simple. You as a company cannot sell on top of a frontier model because you can't double stack software margins. That's it. It's that simple. You can't take an 8 % margin, add another 8 % margin and then sell it. And you certainly can't put together lots of models around it. And so if there is a, I mean, I think Anastasia is right. You need, there needs to be a business model there because there isn't the same network effect that you get in traditional open source software where everybody's contributing to it and it's all getting better and everybody's gaining from it.
19:32But at the same time, you know, this is critical to the, to us not having a, you know, bipolar, unipolar world of frontier models and everybody having to pay the tax and everybody being worried like Figma that they're going to be stealing them. And now we serve pharma as well. And as you can see, Anthropic recently came out and said, oh, we're going to build drugs too. I'm like, oh, great. I'm sure the pharma guys are super excited to keep using Anthropic after that. Right. We're going to subsidize our own execution. I think I'll pass. Thank you. I think I'll pass. And so this is very critical.
20:07And there's a structural... I actually had this conversation a long time ago, actually, even with David Sachs when we were at the White House together. And I was like, David, the most important thing we got to do is ensure that US open source continues to exist and is on the leading edge of open source. And so I think, you know, maybe this gives us more impetus to get that initiative going. And maybe NVIDIA with obviously an infinite supply of GPUs is the guy to help pioneer that with Nemetron. That would be awesome. I just don't want NVIDIA to become the next open source AI giant because I don't trust them to not bias that in their own favor.
20:46They are an enormously profitable company. They're the most viable company in the world, Mitchell. I don't think they need more power. But I'm curious why, and I mean this with nothing but love and respect, your aspirations aren't higher here. Because you're already doing all this work to collect the data you need, to test it out in practice, to fill in the tross of the artificial jagged intelligence line, and you're doing all this AI work already. Why not just do it all yourself, soup to nuts? or form a consortium with some friends to share the cost if you had to? So two things. One is we have found the safest models are the biggest models.
21:21Biggest models are really expensive to train. And then I got to do it every year. Hey, look, my phone's ringing. David Sachs is calling with a check from his enormous pile of money. I don't know. Yeah, I mean, unless David's, I mean, it's probably like$400 million a year. And you have to build the team to do all of the pre-training and to gather all the data. And you have to figure out some of these tricks that are hard to figure out that somehow China's getting figuring out as well. Like not an easy place to be. And honestly, that asset you're making is valuable enough to other people who don't compete with you that it really should be a shared cost.
21:53Now, maybe there's some consortium of it, but I do think that the person best positioned, you know, also I had conversations there and NVIDIA is at least open to not only open sourcing the weights, but open sourcing the training data, open sourcing everything. So all of us can do continuous pre-training and other neat things that we want to do on it. So, but, but Nanette, it's a very expensive lift to have to keep making your own open source. Cause you got to do it like every single year, right? It's not a one-time thing. No, it's continuous. And you have to stay on that. I don't think one startup by itself could ever be kind of capitalize properly to do that unless it just becomes a frontier lab, effectively.
22:39And there's not an easy way to have a consortium around that. I think this is a critical piece of infrastructure that we need. And luckily, we had it in the beginning. It then the dominant players became Chinese in it. And now NVIDIA maybe can lead us back. But I don't see how startups continue to fight this war. even and verticalize and build unique products without it. And actually the companies that are built on the frontier models, what's happening to every single one of them? They're getting eaten by the frontier models. Yeah. Well, not only that, but right now their unit economics are all horrible or upside down even.
23:17So they're all like, oh, don't worry. I mean, recently I heard one of them raised, I won't say which one, a vertical AI company that's in many verticals. And it was like, yeah, we know we're totally upside down, but by the way, hey, we're going to build on open source later and swap out these models. And Cursor did it, and these guys did it, and it'll work. And I'm like, first of all, I mean, coding is probably very specifically trained on it, but they won't be able to do that swap if there is an open source to exist. So I think this is super critical infrastructure. I mean, look, I'm a capitalist, and I'm not a big fan of state capitalism, but if we're going to have a national AI project, Maybe this is the place we could be putting some more of our work.
24:02I'm not sure. China's advantage right now is that their product's being used as open source. I think they need to find a different answer than we're just going to block this. Because as soon as they block this, what influence do they have on the worldwide AI game? I mean, they're just going to build a frontier model company and say, hey, that generated a lot of money. So we're going to generate one. but ours isn't quite as good. So then what is it? Why am I buying it? Is it cheaper? It's got to be something. Otherwise, like, why am I using? It's still an inferior model, technically speaking, to the Frontier models.
24:39All right, let's go back to benchmarks, which we touched on just a little bit ago. I was going to bring this up in the context of the latest from our dear friends over at DoorDash, clearly a leading company in the broader AI conversation. That's sarcasm, if you didn't catch it. They recently dropped a thing called Dash Bench, which allowed them to train, essentially show pairs of AI models working together to find code to fix. And they found some interesting things about this. To me, this kind of felt like we've reached the possible apex of the number of benchmarks out there in the world. So Anastasios, from your end, clearly you have your own approach to this, but have we reached kind of like peak benchmark saturation today?
25:16And have they lost a lot of their bite? Because I don't care as much anymore about benchmarks versus actually using something and seeing if I like it myself. You know what? the benchmarks are only going to be accelerating. And here's the reason. It's because you cannot tell whether something is good until after you put it in production. It's the post-deployment evaluation of models that actually matters. And it's not like whether or not it's good at a multiple choice test. That's exactly the philosophy that we have at Arena, by the way, is that all of our benchmarks are based on the continuous usage of tens of millions of people of like actual AI products and realities.
Read the full transcript
25:49We just place them in people's hands, see what they do with them. It gives us the most diverse benchmark in the world because we can collect millions and millions of agentic traces every week from people and then measure how it's doing, not just, you know, on history questions, but how well it's doing for, you know, everybody in every country around the world in math and coding and instruction following and multi-turn tasks and legal and medical verticals and so on. And it needs to go beyond that, you know, our benchmark, you know, That's our philosophy that it needs to be about the post-deployment utility to real people.
26:24But in order to really continue to fulfill that mission, we really need to be inside every company in the world. We need to be helping every company define what good means, define outcomes, help them measure these outcomes and ensure that they are getting the reliability, the performance that they need at the cost that they want within their particular vertical. and I think that Munjal was talking about this earlier. That means that we need as many benchmarks as our businesses. So then they don't become benchmarks in a general sense, but they're more just company-specific quality checks, Munjal?
27:00So I'll draw an analogy for you. I think we've all gotten to obsess thinking everybody's building the same type of vehicle. But some people need pickup trucks and some people need station wagons and some people need SUVs and some people need convertibles and some people need caterpillar tractors that are super rugged. And so I'll give an example of where the benchmarks are actually coming in now that we're seeing. So if you do a two by two grid of latency fast enough for voice, we call that about 500, 600 milliseconds in the LLM because you need voice in, voice out, and the whole thing's got to be under about 1.8 seconds or less ideally.
27:37And or latency above that. And then high intelligence, low intelligence. You notice all the innovations come up here in the top left, which is high latency, high intelligence. So if you have all the time to wait in the world, the models have gotten really smart in the last year. If you go to the bottom right, which is you need it in 500 milliseconds, 600, have they gotten much better in the last year? Because how do you make a model smarter without deep thinking and planning and all the other things? We basically saturated the amount of tokens we had to kind of hyper trained them and with more and more data because we ran out of data quite a while ago.
28:16And oh, by the way, all this synthetic data stuff turned out to be false, which, you know, I'm not surprised by because it was basically inbreeding. I'm like, I knew inbreeding wasn't going to work, but okay, let's ignore that. But in the top - Attsbergs, just saying. Yeah, that's funny. So in the top right, you know, I think that very few people need high intelligence, low latency. If I'm taking orders at Taco Bell, I don't need it. Fine, I mess up a little and give you two tacos instead of three. I mean, that's what happened anyways when I go through the Taco Bell drive-thru with a human anyways.
28:50On the other hand, but in healthcare, we needed it. We need very high intelligence, very low latency, and we needed a way to build that. So, Hippocratic, we ended up, we've actually changed kernel implementations. We forked VLLM and created a VLLM that runs faster for MOEs where you hit a few experts more than the others, by the way, which we do. But it works best for large ones. And we've actually sped that up almost 10x. And we take all this latency surplus, use it to deploy bigger models because the bigger models have better reasoning and the better reasoning creates more clinical safety. And then we find the next tech innovation and we do it again and again.
29:32we've done flywheel after flywheel at the very low level to be able to do this. And I think that when you come back to this benchmark question of it, I would say there's not even just a question of all the benchmarks. There's also a question of all the different use cases drive it. We're in this low latency, high intelligence bucket that very few people need, but in healthcare we need. And there will be a whole new set of benchmarks to even talk about that. that. Low latency, low intelligence. There's a whole bunch of very simple... By giving you your bank balance, I don't need it to be that smart.
30:06It's fine. But most of the world right now is innovating only in the top left quadrant, which is infinite latency. Right now I'm like, write me this report, Claude. I'm going to go get a sandwich. You can take 10 minutes. I could care less. It's better than me riding it. And so I think we've gotten too obsessed on kind of thinking everybody wants the same car. They don't. They need different types of vehicles. And those will drive even different benchmarks. Yeah. No, I think this is a really interesting point. But on a success, when I look at, for example, the Text Arena leaderboard, it's super helpful for me to understand probably to Munjal's point, the most intelligent, but maybe not the fastest models.
30:49So how do you work in the latency point and the other things that Mujol just brought up into creating a reasonable public-facing metric for model quality? Yeah, absolutely. Well, I think there's three. There's the trifecta. There's performance, there's cost, there's latency. And right now, the hottest topic is performance cost. Because what people are finding is that the token maxing era happened, and people are just spending, spending, spending tokens. And you'll get crazy things. Because of the way the context works, you'll say, you know, thank you to Claude. And you'll basically tip it$3 because it had like all this big context that it's using to like read and say, what do I need to look at for this, you know, for this thank you?
31:34And it's like, you're welcome. $3, that would be$3. And that's happening, you know, millions and millions of times around the world. You know, it's happening today. Right now, somebody's tipping clock. And so you have to start asking the question, am I getting the outcome or am I just spending money for no reason? And so Arena is definitely trying to help with this. If you look at the TechSterina leaderboard and you go to the Pareto frontier plot, we are trying to help developers and businesses around the world who are making choices about their models. And so what you'll see exactly is that Fable 5 is on the extreme end where it's the best model on the text leaderboard, but it's also by far the most expensive.
32:22And the scale, by the way, is logarithmic. So as you go to the right, every line is in order of magnitude, cheaper and cheaper and cheaper models. And so what you'll see is that a lot of that Pareto Frontier is dominated by Google. If you look at it, one of the things we did in looking at it and creating a novel set of benchmarks is, look, here's some very specific things that are not in the normal benchmark. Drug name disambiguation. Can patients say drug names right? No. They leave entire syllables, right? They're like, what are you on? I'm like, something statin. I'm like, that's what I'm on.
32:59I'm on something statin. Okay. Well, which statin is that? Super statin? Simba statin? Which one is it? And so, but look at the things we compared to first. Did we compare to Gemini 3.3 and at that time we wrote this was GPT 5.0? No, we just said too slow for voice, too slow for voice, too slow for voice. Okay, we did. Now, then we looked at the speech-to-speech models that everybody's excited about, but it turns out they're all very slow parameter counts. They're just not very smart. Now, then we went to the cascaded model where you have an ASR that takes the voice and then converts it into text and put it in the LLM.
33:32And then you give it to a TT and it has a text-to-speech engine and speak it back out. And you notice there, we compared to all of those and it got better. But now if you look, then we looked at the Polaris model and said, all right, how good is that at it? And then we looked at the Polaris model where we actually, we looked at our 4.0, we looked at our 5.0 with just the main model, and then the whole constellation of the 31 working together. And we basically showed step by step by step that, A, these are the things you need to know. Do models know toxicity limits of OTC meds? Oh, can I take ibuprofen?
34:06Sure you can. Don't take more than 2000 milligrams. Does it nail it every single time? It needs to. Does it get, there's different mental health questions and muscle skeletal questions. There's wound and skin questions. Like we went really deep and this is the, there's payer questions on that. There's compliance questions on HIPAA and anti-kickback things for pharma calls. Like these are the detailed things and you can benchmark them and you can show So how much of the contribution is coming from the model? How much is it coming from the post-training we're doing? How much of it's coming from the constellation of the post-trained models altogether?
34:40This is why I think he's right. Literally, we've got to prove this point. Every business is not only one new rubric or one new eval, it's hundreds. So you think this type of rundown is going to become the norm for pretty much any company that puts AI to use inside their product or service, regardless of the industry or vertical. Yes. And now, but let's talk about our friends at, what's the legal one? Lagora, Harvey. Harvey, right? Harvey guys that, oh my God, we got beat on our own benchmark by the frontier models. Actually, part of that, I think is, and why didn't that happen to Hippocratic, is again, the latency.
35:21Their use case has infinite latency, right? Meaning you can wait however long to generate the document. And the big guys are gone after that in a big way and have improved everything they can to do that. So A, I do think this is exactly where it's going, but B, part of what's interesting is, like I said, the big guys are making a convertible and they were making a convertible and the big guys beat them making a convertible. I'm making a caterpillar tractor that needs reliability over everything else, even if it costs more. Okay, now on these benchmarks, I'm really curious because I looked through some of these before we jumped on, but I just scrolled through all of them.
36:01Went on for quite a long ways there. Quite a lot of your scores for the Polaris 5.0 Constellation, which is your newest and best model firing on full power, quite a lot of your scores are like 99.97. First of all, 10 points for getting yourself high marks, but how much work is there left to do? Like what's Polaris 6 going to bring if Pilaris 5 is already so damn close to perfect for these tasks that you have set as the critical work that it needs to do? Look, in healthcare safety, there's still patients that are hurt in the point three, right? Actually, you don't stop there. You stop at five nines.
36:39And so that's part of it. Luckily, right now, we've done 200 million patient interactions without one significant safety incident. So we're feeling like we caught everything, but I still think there's a need to get even better. The second part is actually it's the blind spot of what benchmarks are not there. That's the improvement because actually, how did we even find those? In fact, Anastasia and I, we were talking at that football game last year, remember? Yeah, we met at the Super Bowl. Yeah, yeah. I tried not to say which one, but now you just said it. But like the, but the - Wait, wait, wait, wait.
37:18Were you in the same box or just did you run into the bathrooms? We had the same friend who took us to the same box. Oh, okay. I understand. I understand. We did that. A friend. He took us to the Superbowl. Us entrepreneurs don't, don't go on our own. He runs a billion dollar company. I run a billion dollar company. He is a friend. I have a friend. No, that's how it works. We both went for free. Just officially. We went for free. Thanks to the generosity of a good friend. But the, because us poor entrepreneurs, you know, we, we just try to get into any game we can. But we were talking and he said, hey, put your benchmarks on our site and we can run them all.
37:54And in some ways, I'm excited to do that. But in some ways, I wasn't because actually a lot of the IP of the company is now seeing so many real world examples where we had failure cases and we had to address those failure cases and deal with that. I'll give you another example we found on the TTS side recently where, yeah, 11 labs is not drug name stable. So it's not only says the drug name wrong, it says the drug name differently each time. And it's actually, it's okay kind of to say it wrong because patients say it wrong. If you're consistent, you're actually okay. Although that's still not great because if they do, like people don't know how to say it, but they know when you said it wrong.
38:38and your clinician, your AI clinician loses a lot of credibility if you say it wrong. What kind of clinician are you? You can't even say the drug right. I think 100%. If I was on the phone with somebody and they mispronounced ibuprofen, I would be like, okay, click. I'm not going to - And 11 Labs recently released a new version that I think improves that. So good for them. They figured this out. But we found it was drug name stable 95 % of the time. It just wasn't drug drug name stable 99.9 % of the time. And so we ended up building our own TTS to fix that because it was such an important criteria.
39:11And by the way, our pharma customers were having none of it because it's their drug with their name. And so they're like, no, you can't say Manjaro, you got to say Munjaro the right way. And so I think that these are the different elements that come out of this. But a lot of the improvement is in the blind spot of what benchmarks have we not created and what are the queries in the benchmarks. And there is actually a tremendous amount of IP in what we're testing on that we had to learn the hard way. Yeah, I think that's exactly right. I would take it even one step further to say that the benchmarking evaluation is actually the fundamental and core IP of any business in the future.
39:50And the reason is because it encodes what you know about your domain. It encodes everything about what it means to achieve success. And so you should know that and nobody else should. Because if they have a verifiable method for hill climbing on what it means to do a great job in your domain, they can just replicate your business. So Anastasio, I want to double click on this. Your point is that if you as a company can come up with the correct benchmark for what you're serving through the medium of AI, you know what matters to your customers, what your system can do. And essentially it becomes kind of like a functional DNA of your in-market performance.
40:28Exactly. I mean, listen, like, Munjal runs Hippocratic AI. Imagine that Munjal releases all of the benchmarks, all of the data that tells exactly, that gives an exact roadmap to all of the AI companies on how to build the best voice assistant in medicine. And then they just can climb that and build the perfect voice assistant for medicine. You know, that's a pretty bad outcome for you. And that's not just true for your vertical. It's true in all verticals and all AI native companies. And the model labs are going after these. You know, the model labs are now competing with Harvey. And so Harvey better not release the benchmarks that they use and the definition that they have of, you know, what it means to be good.
41:13Because it's not that hard to climb once you have the North Star. Okay, but then if I'm a customer of Hippocratic AI and they say, we have a 99.59 for, I'm going to pick a thing here, child protective services for these agentic voice calls. And I say, cool, how'd you measure that? And then Munjal says, well, I'm not going to tell you, that's my secret sauce. Does that break customer trust? Or is that simply just protection of in-house IP in the same way we've seen IP protected in like patents and so forth? There's a very easy way. Go ahead. I mean, we would just show them. I mean, that customer under NDA, we'd be like, here's the queries we did and here's how.
41:49I'm not worried the customer's going to go hill climb it. It's highly unlikely. And so, happy to be very transparent and show them. But in that case, what that actual feature is, is it's actually if somebody, a 14-year-old being called by the AI or a 15-year-old over the thing about their healthcare says it's not safe here. By law in 50 states, you have to notify child protective services within a certain timeframe in a certain way. And you have to flag those and you have to do that. And if you could release an AI that is an AI clinician that calls patients and calls teenagers. And if it doesn't have this feature, you can't go live.
42:30Like you really can't. So there are some very detailed things, but there is IP in this. There is learning in this. I'll I'll give you another thing we learned that is on the periphery, not even in the core model. We used to fail on 25 % of all calls because of background TVs. Because when you're in healthcare, who are you calling? Typically older folks because they use healthcare more than younger folks. And what are they doing at home when you're calling them? They're watching TV. Is the TV loud? Oh, it's loud. They just go to my parents' house. My dad definitely struggles with his earring. and the TV is blasting.
43:09And getting a bit of background noise is easy. Background speech, though, is the thing you're listening for, right? And so we had to build an entire layer of this algorithm we call MRX that sits in front of all of our processing and listens for background TV and cleans it up. And now we still fail on about 1 % of all the calls, but it's better than 25, but we didn't even know that existed as a problem. Same thing for handling post-stroke victims with slurred speech. Same thing for, you know, there's numerous features both in the core clinical model, outside the core clinical model, in the TTS on both ends, on the input, on the output, in the middle.
43:53And these are all things you learn from just failure, but you can't, you know, you could replicate everything I've done and you'd still probably fail on 25 % of calls, which is unacceptable if you didn't get this background TV thing right. Wow. Well, you guys also just introduced cough detection, which I presume was easier than background TV removal, but still because it goes to the point of how many things you have to deal that are edge cases that aren't really, but instead come up quite often or at least often enough to take your, back to the five nines point, five nines down to one nine, nine nine.
44:25Yeah. I mean, the next thing we're actually working on, a future feature we haven't launched yet is sob detection. because think about this case. I'm talking to you and your words are saying something different than the rest of you. So you're like, yeah, I'm totally fine. I'm totally, totally fine. Could you imagine if the LLM response is, I'm so glad you're having a good day, Manjal. You'd be horrified, right? You'd be absolutely horrified. You'd be like, that is horrible that the LLM did that because you and I know they're not having a good day and they're misleading you with their words. But right now, the ASR just takes the words and says, here's the words.
45:10Please, LLM, give me an answer. And so these are critical features that no truly empathetic, responsible clinician would ever want. Like, they'd be like, no, I can't deploy this. Same thing with the post-stroke, people slur their speech. Well, that's not that often in the normal world. It's very often post-discharge calls in healthcare. Well, you've convinced me I'm not going to build a competing startup to Hippocratic AI because it sounds very tricky. Now, we've talked about data a couple of times. And, Majal, you've mentioned how many data points you have that you've put into the training of Polaris 5 and preceding versions of it.
45:47Anastasios, your company launched a product that lets people collect a certain type of data from your user base and turned that into a commercial product last year. And it's grown to 100 million run rate pretty quickly. How does that information you can bring impact the conversation about benchmarks and getting AI to this kind of five, nine levels of repeated performance? Well, Arena is all about the post-deployment utility of the AI to people. It's exactly the conversation that we've been having, which is if you actually take a model and you put it in the hands of individuals in the wild and they're using it for all of their super diverse use cases, how is it going to perform?
46:25It's data that we have access to that almost nobody has because we have one of the largest consumer apps in the entire industry. We have tens of millions of users on Arena that are coming all the time to use AI for whether it's their work, for their legal, for their medical, or for their software engineering, or for their personal tasks, for asking for research, for planning, for advice, for lookup. and so we're able to take that post-deployment data and turn it into very careful descriptions of a model's strengths and weaknesses and analysis that helps businesses build better models and understand which models to choose.
47:08I think that the future of the company is really around helping the rest of the world understand how to get the best out of their own data because it's not enough to just look at an external data set no matter how diverse. You know, I think Arena has the best coverage of any data set in the world, simply by virtue of the way that it's collected. It has all of those things that you didn't know you needed to measure because of the fact that it's in the wild. We don't determine distribution. The distribution is determined by this like super broad set of users. And so it inverts the problem from a, like, let me describe what my problems are to let me look at the data and see where the problems arise, which is a fundamentally more scalable and organic way of doing benchmarking.
47:57So can we take that template and help every business in America do better benchmarking? That is something that I would be excited to work with someone like Mujol on. Mujol, is that something that you could use at Hippocratic AX? I'm trying to figure out a little bit through the weeds here what arena is selling and how it grew from zero to 100 so quickly, because I was very impressed, Anastasios, by that rapid commercialization. Yeah, I mean, look, I think that there is this, we are actually always seeing this as an issue as well, which is that we've built all these detailed features. We know the product works better than anybody else's product on all these corner cases.
48:32It's not always easy to get that across to customers efficiently and fast. And, you know, in the early days of the company, it was like, you have voice AI, I've never heard voice AI. Then it's like, oh my God, you guys have a clinical voice AI and you've ensured it's safe and here's how you've tested it and here's how you've done it and the protocols. And so that was also differentiating. Correct. We're still the only ones that will escalate and transfer the call to a human, which is an important kind of safety element. And so you have unique safety. But, and we continue to have lots of people looking at different parts, but nobody wants to do clinical.
49:09Like they're all scared of it. And we're the only ones that have kind of gone into that. But at some point, somebody will. And it will be, there is a need to articulate some of these deep differences where otherwise it's just, and why they're important. Like if I just told you background TV, I figured out how to do it. You're like, that's great. But it was when I told you I failed on 25 % of all the calls because of it. You're like, oh shit. That's a really important feature. Sure. And the same things on all the different clinical elements that we have. It's like, well, how often do you have to do child or adult protective services?
49:46How often do you have a patient that has a case that you've got to handle better on the clinical side? How often are? And so I think that's where it's important to be able to articulate that. You do build a reputation over time of, hey, these are the safe guys. These are the ones taking the most things. But I think the The other part that's interesting is showing why the model architecture is actually bringing us many of these advantages. Why is the constellation better than a single model? Why is big models better than small models? By the way, one thing I forgot to mention when we were talking about the open source, I don't know if you realize this, but one of the reasons Kimi and all these guys have not gotten as much adoption of their open source is that the code isn't there to train them all.
50:32Like it's amazing how much code my team has found is missing to do some of the core trainings you want to do on them to optimize them. You run them on the inference engines. They're super unoptimized. So they actually don't run very fast. And especially the big ones. People have done a lot of optimizations on the small versions of them because they like to run them on their laptops. But the big ones have actually been largely ignored. We've had to do a lot of core work that we shouldn't have had to do, frankly. But now that this is, it would be great to create, I think over time, you'll end up with the benchmark for every vertical, for every type, or a series of benchmarks that'll be in a compounded benchmark and people will do it.
51:16But today we're still exploring, meaning we don't even know all the things that would need to go into that benchmark that actually matter and how to weight the benchmark. Like, let's say they're an aggregation of 500 different smaller evals. Okay. But your aggregate benchmark needs to be a weighting on those. And it's not an equal weighting because some problems occur. Well, you know, the other aspect is that 500 may not be enough. And it's probably not. I mean, it's probably 5 ,000 or 10 ,000 or 100 ,000. But the other part is you need that and the weighting. You would have to ask the question.
51:49Do you think that, you know, sort of at ridiculous, do you think that every person needs their own benchmark? No. Is the limit of X as we go to infinity, basically every single person has a set of benchmarks around it? I mean, I have that with my own codex instance, which I've trained to talk to me in a certain way. And I view its performance through the lens of what we can call Alex Bench, if you will. because you have your own set of tasks that you're doing and you have the model that you like to talk to you in a certain way because you are alex and i'm on a stasis and we have a different way of communicating and we might like different types of people and we might you know want different types of information we might have a different task distribution that we care about we might have our own different failings that we need the model to cover or help us understand and uh and to your point each language we've also found so for example we have made mandarin safe enough to use for clinical actions as long as you're not talking about drugs.
52:49No, we have not figured out how to get the drug part right. Exactly. Yeah, I mean, listen. We actually couldn't figure it out in Arabic either. And what's the reason? Well, the speech recognition system's error rates are so high in these low-resource languages. Now, hopefully China can fix the Mandarin thing. But we literally have not found an ASR that's good enough to be able to use for drug name identification for anything relating drug names. And so the error rate is so high, so we won't let our customers do Mandarin for anything related to drug names. It makes sense. And I think that the sort of conclusion that we're arriving to is that we do actually need that level of fine-grained detail to understand much more than we could ever possibly write down.
53:37And the implication of this is that benchmarking as an area, like building benchmarks, is actually not a scalable solution to this problem. Because you will never be able to write down all of the things that you care about. So you need to do something different. you actually need to invert the problem and you need to be looking at the data first it needs to be completely data driven and we probably need to be training models to do it too training models to help us understand on an individual basis you know even more finer grain than a business because you have customers it's not like your business is just one business serving one individual you have you know thousands hundreds of thousands millions of customers that you need to serve and they they they probably have their own customers that they need to deal with.
54:24And so, you know, when you're getting to that level of granularity, what needs to happen is you just need to be monitoring, you need to be looking at what is happening, and you need to have automatic systems for classifying where the errors are happening, what are the sort of principal components that are going into the failure rates, you know, for who it's working, for who it's not working. And those types of systems are likely to be the future of what we now call benchmarking and evaluation. Well, in that case, then I want to bring my own context from my own AI world to that conversation. And I wanted to be able to plug into Polaris 5 from our dear friends at Hippocratic AI.
55:00So it knows, oh, Alex tends to talk this way. He's hyperbolic and likes to kind of make some jokes. Don't take him too seriously. This is a medical call. You know, drive to the point. And for you, Anastasia, it could be he's a complete stoic or whatever. But I would want to bring that. And I don't think that a top down benchmark to your point would work. So then how do we create a system by which not only can I aggregate and collect my own personal or personal business AI context and bring that to bear in a setting like what Hippocratic AI pulls together? Do we need a new MCP for people? Very, very good question.
55:34And I think you're getting to something where there's a human performance and then there's a superhuman performance. If you were a human on the other end, the intelligent human would be talking to Alex on the phone. You know, Alex, you know, has is bleeding out or something like this. And he's funny. He's making jokes and blah, blah, blah. But you're kind of listening to who this is and you're trying to understand them as you are, you know, working with them from a medical setting. And then you kind of learn that, oh, no, this is serious because you because there's something going wrong. And he's telling, you know, I'm getting lightheaded.
56:10I'm dizzy. I'm about to faint. and you know things are going wrong but also lol you know i'm fine i guess like it's okay like dog with the fire meme okay so maybe that's you but the other person the person on the other end an intelligent emotionally intelligent and intellectual individual on the other end will learn that but where ai can go farther is it can potentially learn that across all of your interactions in every surface that you've ever seen we don't have the way to collect that yet But Munjal, if I was a patient of a hospital group that uses your technology and I had to talk to this voice agent, let's say twice a month, is it able to learn from me as Alex and therefore how I communicate and so forth and bring that to bear the next time?
56:58Or does my interactions flow into a kind of a shared bucket that shapes overall performance in a less individualized way? So two things, or three things. One, we actually do allow, we do store memories from the prior calls and bring them up. By the way, super hard to do. You think it's easy, I'll just take the whole prior call transcript and put it in this context window. Yeah, until you blow the context window. Because when you exceed 10 ,000 tokens in the context window, you slow down latency, even if you have space. So now you've got to compress the memories a bit, right? You've got a sandwich, all has two kids, these are their ages, blah, blah, blah, blah, blah.
57:35You can't just store everything. And so that's one issue. And it's a latency thing more than it's a context window size thing. The second element, yes, by the way, when you remember things about prior calls and bringing it up, the superhuman ability of that, patients love it. Like they just absolutely love it. Second, we've now built in a dynamic personality that we had to build because people were annoyed. I'm in a rush. It's not in a rush. If I'm in a rush, you're in a rush. If I'm joking, you're joking. so the third thing is i i want to argue a little bit the other direction for this kind of everything's got to be specifically personalized look there's a thing a really good thing we did in health care in america it's called the standard of care we standardized a lot of things and as long as you do the standard of care you're fine second do i need to personalize to every patient and be there well my first priority is is to get the things that have a safety risk so should i personalize everything for you?
58:31I mean, it'd be nice. It's not that critical. What is critical is, for example, figuring out that you're one of those stoic guys that even when your pain level's 10, you're not going to tell me it's a 10. That I should do. I should say, hey, in my prior calls with Alex, Alex is kind of an understater versus you've seen patients on the other end, you touch them with one little needle. They're like, ah! And you're like, okay, there's the scale of your Your perceived pain is not an absolute scale, right? So now, where do you have a risk? The people who overstate it? Not nearly as much. Maybe, I mean, there's a little bit of an opioid risk, but we have good controls for that these days in medicine.
59:15On the other hand, it's really the understated ones that are super in pain and not telling you. And if you notice that there's a pattern with Alex, you definitely want to get it. So one of the things we do at Hippocratics is we don't just try to get every last thing and every last, we're like, which ones matter and create a medical safety issue? Let's prioritize the heck out of those. Let's get those right. Let's make sure we don't mess up on those. Let's build separate models to double, triple check those. And there's something called condition specific disallowed OTCs, for example, which is the, it's a personalization.
59:49If you think about it, it says, Hey, can I have ibuprofen? Sure you can. Don't take more than 2000 milligrams. Oh, shoot, you have chronic kidney disease stage three or four. Oh, now you can't have any. Right. You got to get that right. You got to get that right every single time. And it's personalized to the fact that you have chronic kidney disease stage three or four. And so like, you can't say, oh, yeah, my model gets that right. Ninety five percent of the time. It's good. Ninety eight percent of the time is good. No, I will kill you if I tell you to take ibuprofen and you have CKD4. And so these are the different elements.
1:00:20And the last part is this is what I love. I love finding verticals where incrementalism matters. So let's say I'm doing voice AI for Taco Bell. Let's say you have a product, I have a product. Your product is 95 % as accurate, but you're half the price or one-tenth the price. And I'm 99 % accurate at getting the order right. But I'm clearly a customer. I will buy you every day of the week. But in healthcare safety, I don't buy you at all. How can I tell my boss, yeah, I saved half the money, but I knew I was going to hurt people. and then you get sued for 10 times the money you end up backwards that's right yeah listen i all i wanted was a creme brulee taco bell crunch wrap wait did they did they they haven't actually made a creme brulee they have a creme brulee crunch wrap time to go back to my stoner roots all right um a couple of other topics before we move on uh i mentioned that your company recently announced a revenue milestone well done putting you on track to be public company size whenever you'd like to be.
1:01:22Moonjong, your company, as far as I can tell, has been very close to the chest regarding its financial performance. I'm just kind of curious why and if the recent success of SpaceX's IPO is pushing founders like yourself towards maybe being a little bit more IPO friendly than they were before. We actually did recently put out an announcement. We announced that we were at a 60 million run rate in basically, we only started selling the product January of last year. So So basically 18 months, which healthcare takes a little longer. We got to do integrations. We got to roll things out. That's very impressive.
1:01:54Don't talk down that number. That's great. Well, in this world where everybody and their grandma is going to 100, I definitely am happy. But I will tell you that we're seeing tremendous adoption. I think we've now gotten 50 health systems over 2 billion. We've got five of the top seven payers. We've got seven of the top 20 pharma all in 18 months. It's been an absolute terror. But as for the IPO, I don't know. I have two words. I mean, on one hand, you're right. The market's there. On the other hand, what a pain in the neck it is to be a public company. Like, I mean, I would rather, I mean, there's like a hundred painful things I would rather do in my life than be a public company CEO.
1:02:38But I mean, it may be though, but I don't let my personal predilections drive the right thing for the business. Like if that's the right thing for the business, that's the right thing for the business. You know, we'll find a way to make that work. But I mean, I think that there is an advantage. I don't think SpaceX could have raised as much money had it not stayed private as long and invested in its product correctly. I think that you're going to see a massive focus on margins. tokens, everybody's saying, oh, you know, they've been lowering the cost of tokens continuously. Yeah, you watch once they're public if they keep lowering the token.
1:03:13That's not how it works. When everybody can see your whole entire P &L, it turns out that there's a lot more scrutiny on your margins. And I think the days of costs coming down dramatically are not, we're just not going to see them the same. And so I think when you go public, you kind of freeze a bunch of long-term investment ability or lose it. And you eventually sow the seeds of the next generation of companies that'll come get you. And so I would like to invest long-term and you can always do that better private. And today the private markets are so deep. You used to need it because you couldn't raise the$100 million we both raised privately.
1:03:57Now, I mean, there's a zillion people. Both of us probably get emails every single day from somebody wanting to put in money. Well, that's an incredibly depressing answer to my question. I was hoping that you were going to say, everyone's so excited now about going public. Elon did it twice. We're going to join the ranks. But based on what you said - I think that's the other way around that we should be really thinking about, which is how do we give ordinary people access to private markets? Because it is really not fair that the only people that have access to the highest growth companies in the world are quote-unquote accredited investors and not even really them because they don't have access to the deal flow.
1:04:35So I think the sorts of things that, you know, Robinhood is doing around this in terms of just like giving people the ability to invest in private companies more directly through these like, you know, active vehicles is really, really good. And I think it's an issue of financial fairness and freedom. Yeah. But I thought Munjal was going to say that, look, my investors are really encouraging me to just stay focused, stay private and grow. Because to me, that would allow for more venture capital or private investor, if you will, broadly take rate of the company's future success and growth and valuation.
1:05:12But instead, when he said that it's just too annoying and I just don't have to, I mean, to me, man, it just is that the common perspective amongst CEOs of AI companies, your size and age, Mujal? Is that just the kind of the de facto view? I think this is a function of, you know, I'm this is my fourth company. I've built a bunch of companies before. I've suffered through some of them. Some of them gone really great. You know, I sold a company, Google, that went great. I had a company that didn't work out so well. That didn't go so great. I had another company that, you know, we did well on. But when you've suffered, I think you know that that it's just never something for nothing.
1:05:50And so, you know, right now, I mean, it is probably what you said. I mean, we are investing. We are able to invest. We are able to do smart things. But, you know, like our burn rate's very reasonable. Like we've kept it pretty tight despite raising so much money. So I have a ton of money sitting on the balance sheet because I have seen what happened. I mean, like the wind is blowing at our backs. Right. And it's a gale force wind at the moment. I've never seen ever before. But I've also seen what happens when the wind stops blowing or when the wind blows at your face. And I just, you know, so we're running as fast as we can while the going is good.
1:06:30But I don't I think you need to plan for a day when margins are going to matter, when cash won't be so easy to raise. And. And so, I mean, I think I think the public markets are interesting. I think they have opportunity, but I think one should just think very carefully about whether you're that. The other part is, I remember I invested in a startup that became a unicorn. And I won't say which one it was. And it went public. But in going public, it was one in the advertising space, online advertising. So it was buying ad inventory and selling it on the other end as actions. And what happened was on both sides of its marketplace, everybody saw its margins.
1:07:15Oh, no. Right? Because they went public. And so then every quarter, actually, when every renegotiation came up, they actually lost margin, lost margin, lost margin, lost margin, lost margin. And so the public company process is one that degrades competitive advantage by definition. This is what Google was saying. The advantage it used to give you was access to capital, which gave you a reverse competitive, it increased your competitive. But now, you can get almost as much capital, maybe even, I mean, maybe you can't raise $70 billion in the private markets. And so he needed it, but he's also got a very capital intensive business that we don't have.
1:07:58OpenAI's last round was what, 122 earlier this year? Right, also a very capital intensive business. For sure. For sure. But I mean, just you said 70 might be the cap. I think the cap is really uncapped currently. You're right. They didn't do it. They didn't do it private. So you're right. Your point's made. Well made. So Anastasios, before we jump to the show, you wanted to talk about Trump accounts. And I was kind of curious how we'd fit that in. But you mentioned, you know, the fairness factor of this. And if companies don't want to go public and that does preclude people like, you know, me, Mr.
1:08:25Index Funds, from getting access to their upside as an American taking part in this economy. me. What about private companies donating some of their shares to the Trump accounts of the kids? You know, I absolutely think that's a great idea. It's something that I've even considered myself because I think that it's just such a universal positive to get young people in this country financially educated, interested in investing, interested in owning businesses of the future. I think there's a lot of questions about what's going to happen in terms of our socioeconomic structure in the age of AI, you know, wealth inequality, blah, blah, blah.
1:09:03There's some things I feel good about. There's some things I don't feel so good about. I don't feel that great about the way the government is using our money. But I would feel great about giving part of my money to the children of America. I know you're really glad we ended up here in the conversation, but you were at the White House. You know, a lot of people in the policy spaces that are making the sausage here. What do you think the appetite would be for something akin to a let's all do the 1 % pledge, but instead of giving it to charity, put it into Trump accounts for kids? I'd probably say two different things on it.
1:09:38Not a bad idea, actually. Giving it to the kids feels better than giving it to the government by 100x. I'm not a big fan of nationalization of assets. I think we tried that in the past. It isn't the greatest way to build any economic model. What the is Trump doing buying shares in all these American companies? Leave them alone. Sorry. Back to you. Anyways, well, at least he's buying them instead of just saying, I take over all of your oil producing. I mean, the other way was worse. Where you're buying is still, it's fine. But I think if buying it is not any different than a GIC or some of these national sovereign funds buying assets and startups, which they do.
1:10:24And buying is fine. Buying is clean. But nationalizing, And when does 5 % given to the government become, well, you know, now you have to give us 5 % more next year. And now you got to give us 10%. Now you got to give us all. I mean, like it's a slippery slope. But I would probably say if I were to take Hippocratic shares and have a foundational element, which we actually did put some shares into a foundation that is designed to benefit health systems and patients. When we first started the company on day one, by the way, which has now actually turned into quite a bit of money,$3.5 billion valuation.
1:10:57But I would probably find something that's more akin to, is there a Trump account for your healthcare costs for the people who can't afford it, who get in a pickle? And I would rather put some of our shares into some sort of structure like that, if I could. I don't know, there is one that exists, but I would try to find something that aligns better with the mission of the company and with really the goal we're at of kind of - Like an HSA. Yeah, if I could say everybody's HSA gets a little bit awesome. It's, you know, it's at least mission aligned if we're going to, you know, kind of add that dilution to the company and do that.
1:11:38You know, it's interesting. I don't necessarily see it exactly the same way. I understand why you're saying that it needs to build a mission and, of course, respect your mission. But, you know, I think that it's also very good for both sides, both for the company and for the U.S. population, for people to be educated about these businesses. that's like literally but i've written them to like understand them i have i have a portion of hippocratic ai let's say i have a you know one unit of this i'm like i'm gonna want to understand it if i'm like you know 15 years old i don't have that much else to do with my time maybe i'm doing my homework but i have like a thousand dollars and here's my stocks i'm gonna go look up these businesses i'm gonna go it's just a matter of who you get i mean anyways like i said it's not a bad idea to give it to the Trump accounts basically for future college funds.
1:12:32But if there's a mission-aligned way to give it, then I might want to look for that as well. But I think in general, this is not... Again, there's a difference between giving it and buying it. And we should just be... The nationalization of our assets has never in the history of economics shown to create better, stronger companies, better, stronger economies or more efficiency. I don't know why we've forgotten all that. I feel like there are many lessons from my upbringing on capitalism being not the perfect system, but still better that are being lost. Well, yep, I think, well, I have a lot of thoughts about that.
1:13:13Well, I mean, just kicking this can one bit further down the road. what if, if buying is better than taking, which we all agree, why not allow the Trump account universe to purchase 5 % of each startup's round in the approximate investment opportunity, and therefore buy shares that way and then appreciate on the upside, and therefore no one's getting socialized and the kids get a chance? That would be fine. But remember, these are very volatile assets. I even tell people, I'm like, do not exercise your options with your kids' college funds when you work here at Hippocratic. That is not the right answer, my friends.
1:13:53That money needs to be there when it needs to be there. You put in your money you were saving for pure fun, your vacation fund for that trip, that if you lose it, you're not going to. I mean, you have to remember, these are still very high volatility assets. They're not the place to put money in that you 100 % need. Yes, you're missing out on tremendous, crazy growth. But there's a reason there are different asset classes with different risk profiles. This is probably possibly the highest risk profile possible. But I think there's also one thing we're all missing. We're jumping on this because we think there's this runaway train and we want to catch the train.
1:14:37and everybody should catch the train. But I think that we're missing something that's the biggest idea in it. And I see this in healthcare the most, which is, and everybody says it, but they don't have good examples of it, but I do. Like we're actually doing it, which is the real power of all of this AI is not replacing work you do today and making it cheaper. That's actually not working super well. What is working super well is things you never thought to do until you had an infinite supply of clinicians at a much lower cost. And I'll give you an example. There's been a heat wave in the East Coast, right?
1:15:13I know. Yesterday, I think it was, we called 50 ,000 people and did a heat stroke assessment. And if they're having issues, educated them on where they could go or even called them in Uber to get on to a cooling center. You could have never done that without AI. You'd have to find how many clinicians at a moment's notice. And they're busy in the hospitals because the hospitals are overflowing because people are overheating. This is the power of AI. And healthcare is one of the few places that can absorb that abundance. If I give you an infinite supply of accountants, is your business going to really get better?
1:15:48Probably not. An infinite supply of lawyers? Probably not. But an infinite supply of clinicians? Yes, because the ideal staffing is one clinician to one person. Education has the same property. The ideal staffing is one teacher to one student. like there is a power here we should focus on this power's impact on society instead of just trying to divide up the chips at the moment i think it's very short-sighted i actually love your example because i think it's also a great example of how the use of ai can increase demand for human workers because by virtue of the fact that you called you know 50 000 people identified heat stroke or whatever, you're actually increasing load on the hospital system because those people are going to the hospital now.
1:16:38Well, no, no. What we do is we send them to the cooling center, which are set up all over the country. Because the idea is to not get them to the hospital because, A, it costs a ton of money. B, by the time you're having that level of an issue, you're in a bad space. And if we could have just sent you to the air-conditioned cooling center, we could have avoided paying for you. I mean, it's good for everybody. That's even better. Even better. Yeah, absolutely. The thing that I'm concerned about, because we have to close here in a second, is that the AI abundant future that we're all describing here is far enough away that people who aren't equity holders at the family level are going to be pretty mad.
1:17:17And I'm concerned that when I talk to people outside of our world, that the anti-AI sentiment is sufficiently high at the moment in many places, that it could lead to regulations or otherwise changes to the economy that could preclude us reaching that relatively more abundant future. So that's why I care. Yeah. I mean, look, like I said, it's not off the table in my mind. I think we just construct it properly, think through it, and just understand the implications of it. Just be thoughtful about it. I don't think it's such a simple answer. Is it why we just give 5 % and we're done? I mean, It's fine.
1:17:51We sell 5 % to VCs left, right, and center. It's not a big thing on our front. But I mean, I think if we sold that same to the government, fine. Money's money. But I think we should just think through the implications. But the real power is I don't think the dividend for society is the capitalization of these companies. The real dividend is the abundance of healthcare workers, at least in our vertical context, that will be there. I mean, my mother had high blood pressure and they gave her meds and she didn't take them. And that was five years ago, six years ago. And a year ago, she didn't take them, they made her dizzy and they made her mouth dry and she didn't like that.
1:18:35And then a year ago, you know what happened? We ended up in the ER for 220 blood pressure one night at the beginnings of congestive heart failure. And I was like, okay, but her doctor knew she wasn't refilling for five years because she didn't ask for a refill. Right? Like somebody knew. Now they do not staff to call out to every patient that doesn't refill and say, or they call once and my mom blew them off and didn't. Thanks to HIPAA, I didn't know she was even on the meds, which, you know, makes me, I feel so bad as a child. I'm like, I should have probably snooped in her medicine. I mean, I should have done something.
1:19:11I, you know, she would have taken care of me when I was young. And so there's an abundance here to call Mrs. Shaw and say, Mrs. Shaw, you didn't do it. Can you take your blood pressure right now? Can we get you a refill? Can we have a courier to your house? This is what we can do with abundance. That is the real dividend that AI is going to bring. And it's not far off. We are doing it today. We just did it in New York yesterday. Yeah. And I just hope that people are able to see that and that it will reflect how they approach voting and who they elect in the future, because I do think that really matters.
1:19:44And I think this is not a one party point. It's more of a multi-party point. Moonjel, Anastasios, such a pleasure to talk with you both. I'm optimistic about the future. I'm educated on benchmarks and I'm really worried about the lack of American open source AI, but we'd love to have you back. In the meantime, where can people find you on the internet? And is there a job you're looking to fill that you want to shout out to the audience we have here today? And Moonjel, start with you. We're looking to fill every job. If you're interested in our company, please come and apply. I think we are a 210 person company with 130 open recs.
1:20:14So in every last thing you can imagine. So please come because we're growing pretty fast. And then, and you can find us at hippocratica.com. Thank you. You can find us at arena.ai. And if you want to apply for a job, arena.ai slash jobs. We're looking for all types, especially if you are an incredible machine learning researcher, scientist, engineer. We would love to work with you on the future of benchmarking. Amazing. Thank you both so much. This has been This Week in AI. We'll bring both these guys back as soon as we can. In the meantime, see you next time. Cheers. Cheers.
From the publisher
This Week In Startups is made possible by:
PAYPAL
Today’s show:
China is weighing a ban on foreign access to its top AI models… the same low-cost open-weight systems on which half of the Valley’s startups are building. So we asked our two all-star panelists — Arena CEO/Co-Founder Anastasios Angelopoulos and Hippocratic AI CEO/Co-Founder Munjal Shah — what happens to US companies when their favorite cheap, powerful models vaporize overnight.
PLUS why Hippocratic runs 31 models in parallel to ensure clinical-grade safety, the latency-vs-intelligence tradeoff no one is optimizing for, and why a company’s benchmarks might become their most valuable IP.
Guests:
Anastasios Angelopoulos on X: https://x.com/ml_angelopoulos
Arena AI: https://arena.ai/
Agent Arena: https://arena.ai/leaderboard/agent
Munjal Shah on X: https://x.com/munjalshah
Hippocratic AI: https://hippocraticai.com/
Relevant Links:
Reuters: “Beijing is looking at curbing overseas access…”: https://www.reuters.com/world/beijing-is-looking-curbing-overseas-access-chinas-top-ai-models-sources-say-2026-07-07/
Alibaba’s Qwen: https://qwen.ai/
Google’s Gemma: https://ai.google.dev/gemma
Reflection AI: https://reflection.ai/
Nvidia: “How Open Models are Driving Research”: https://blogs.nvidia.com/blog/open-models-icml-2026/
DoorDash: “How we learned to trust our AI code reviewer”: https://careersatdoordash.com/blog/how-we-learned-to-trust-our-ai-code-reviewer-at-doordash
Harvey AI: https://www.harvey.ai/
ElevenLabs: https://elevenlabs.io/
Anthropic: “A global workspace in language models”: https://www.anthropic.com/research/global-workspace
CNN: “Trump rings opening bell to mark first day of trading for Trump Accounts”: https://www.cnn.com/2026/07/06/business/trump-accounts-launch
Brown University Health: https://www.brownhealth.org/
Timestamps:
0:00 Will US companies lose access to Chinese models?
1:42 Open source makes running 31 models viable
5:01 How deep is the West's open source bench?
13:20 Jagged Intelligence and why post-training matters
25:22 Benchmarks as living quality checks
32:49 Why models still struggle with drug names
42:58 Fine-tuning TV, cough, and sob detection
46:08 What is Arena actually selling?
1:08:29 Nationalization vs. buying stakes in AI companies
Subscribe to the TWiST500 newsletter: https://ticker.thisweekinstartups.com
Check out the TWIST500: https://www.twist500.com
Subscribe to This Week in Startups on Apple: https://rb.gy/v19fcp
Follow Lon:
Follow Alex:
LinkedIn: https://www.linkedin.com/in/alexwilhelm
Follow Jason:
LinkedIn: https://www.linkedin.com/in/jasoncalacanis
Check out all our partner offers: https://partners.launch.co/
Great TWIST interviews: Will Guidara, Eoghan McCabe, Steve Huffman, Brian Chesky, Bob Moesta, Aaron Levie, Sophia Amoruso, Reid Hoffman, Frank Slootman, Billy McFarland
Check out Jason’s suite of newsletters: https://substack.com/@calacanis
Follow TWiST:
Twitter: https://twitter.com/TWiStartups
YouTube: https://www.youtube.com/thisweekin
Instagram: https://www.instagram.com/thisweekinstartups
TikTok: https://www.tiktok.com/@thisweekinstartups
Substack: https://twistartups.substack.com
