Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah-Hill Smith

8 Jan 2026 · 1 h 18 min · 34 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Latent Space Podcast Episode Summary: Artificial Analysis with George Cameron and Micah-Hill Smith

Podcast Title

Latent Space: The AI Engineer Podcast Description: A podcast for AI Engineers covering the latest in AI technologies, with interviews from industry leaders.

Episode Title

Artificial Analysis: The Independent LLM Analysis House Description: George Cameron and Micah-Hill Smith discuss their journey with Artificial Analysis, an independent benchmarking platform for AI models.

---

Key Discussions

Origin Story

  • Initial Concept: George and Micah started Artificial Analysis as a side project while pursuing other AI endeavors. The project was launched publicly in January 2024 and gained traction after a significant social media mention.
  • Purpose: To provide independent evaluations of AI models, addressing the lack of objective analysis in the rapidly evolving landscape of AI.

Methodology

  • Independent Evaluations: They personally conduct evaluations to avoid biases and discrepancies in reported results from AI labs.
  • Mystery Shopper Policy: They test models incognito to ensure labs do not manipulate performance based on their awareness of the testing.

Business Model

  • Revenue Streams: They offer an insights subscription for enterprises and private custom benchmarking for AI companies, ensuring independence by not charging for public leaderboard placements.

Benchmark Development

  • Intelligence Index (V3): A comprehensive score combining several evaluation datasets to assess model performance with high confidence intervals.
  • Omissions Index: A metric to score models based on their tendency to hallucinate or provide incorrect answers.

Economic Implications

  • Cost Trends: The cost of AI models has significantly decreased; GPT-4-level intelligence is now 100-1000 times cheaper than at launch.
  • Sparsity and Efficiency: Discussions on the future of AI models indicate a trend towards larger, more sparse models, leveraging efficiency in token usage.

Emerging Benchmarks

  • Critical Point: A benchmark assessing models' performance on complex physics problems.
  • GDP Val AA: A benchmark for evaluating AI models on real-world tasks, emphasizing practical applications over theoretical performance.

Openness and Transparency

  • Openness Index: A new index assessing AI models based on their transparency regarding training data, methodology, and licensing, allowing users to evaluate the "openness" of various models.

Future Directions

  • Versioning of Intelligence Index: Plans to release a new version of the Intelligence Index that incorporates new benchmarks, ensuring continuous relevance in the fast-evolving AI landscape.

Key Takeaways

  • Independent Benchmarking: The need for unbiased assessments in AI is critical, and Artificial Analysis aims to fill that gap.
  • Continuous Improvement: The AI landscape is dynamic, with models improving rapidly, necessitating constant updates to benchmarking methodologies.
  • Market Trends: Despite decreasing costs for certain intelligence capabilities, the overall spending on AI is rising due to the complexity and demand for advanced models.

Conclusion George Cameron and Micah-Hill Smith’s work with Artificial Analysis highlights the importance of independent evaluation in the AI industry. Their dedication to transparency and accurate benchmarking positions them as essential players in navigating the complex landscape of AI models.

Additional Resources

  • Website: [Artificial Analysis](https://artificialanalysis.ai)
  • Social Media:
  • George Cameron: [X](https://x.com/georgecameron)
  • Micah-Hill Smith: [X](https://x.com/micahhsmith)

---

For detailed show notes and more episodes, visit [Latent Space](https://latent.space).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Birth of Artificial Analysis

0:45 to 2:10

The guests reflect on the origins of Artificial Analysis and its impact.

“I did give you shit for missing fireworks.”

Monetization Strategies Explained

2:10 to 5:10

Discussion on how Artificial Analysis monetizes its services while maintaining independence.

“about models and technologies for building stuff.”

Understanding AI Tools and Benchmarks

5:10 to 8:00

Exploring the challenges and nuances of AI benchmarking and evaluations.

“model for each bit was, trying to optimize every bit of it.”

Technical Challenges in AI Benchmarking

8:00 to 11:30

Deep dive into the technical considerations and methodologies for effective AI evaluations.

“at the same time you guys started, I would have used EleutherAI's eval framework harness.”

Ethics and Integrity in AI Performance

11:30 to 14:01

A discussion on the ethical considerations and integrity issues facing AI benchmarking.

“And again, that just adds a straight multiple to the cost.”

Evaluating Industry Shenanigans in AI

14:01 to 16:04

Explore industry practices that can distort performance evaluations of AI models.

“I've been in a database data industry prior, and there's a lot of shenanigans around benchmarking.”

The Impact of AI Grant on Startup Growth

16:04 to 18:14

Learn how the AI Grant program influenced the development of Artificial Analysis.

“evals, but now you're coming up with zero.”

Building the Intelligence Index

18:14 to 22:22

Understand how the Intelligence Index was developed and its relevance to AI capabilities.

“actually makes some of them pretty archetypical power users of artificial analysis.”

Evolution of AI Evaluation Metrics

22:22 to 27:04

Discover how evaluation metrics have evolved to keep pace with AI advancements.

“This would be a pretty good way to chat about a few of the new things we've launched recently.”

Introducing the Omniscience Index

27:04 to 28:00

Learn about the Omniscience Index and its metrics for assessing AI's knowledge accuracy.

“Once upon a time, we did call it Quality Index.”
Show all 34 chapters

Evaluating Model Hallucination Rates

28:00 to 29:50

Learn about the methods used to evaluate models' tendencies to provide incorrect answers instead of admitting uncertainty.

“So the metric that we use for omniscience goes from negative 100 to positive 100, because we're simply taking off a point if you give an incorrect answer to the question.”

Public Datasets and Evaluation Metrics

29:50 to 31:10

Explore the importance of public datasets in evaluating AI models and the challenges involved in ensuring their reliability.

“and we did one with Clementine of Hugging Face, who maintains the open source leaderboard.”

Correlation Between Intelligence and Hallucination

31:10 to 33:00

Discover the surprising lack of correlation between a model's intelligence and its hallucination rate.

“It leads us to a bunch of really cool things, including breakdown quite granularly by topic.”

Understanding Hallucination in AI Models

33:00 to 35:00

Delve into the nuances of model hallucinations, their implications, and the approach taken by artificial analysis.

“And there's times and a place for that, I think.”

Omniscience Metric and Its Implications

35:00 to 36:50

Learn how the omniscience metric relates to model accuracy and its connection to parameter counts.

“that actually everyone agrees is the rate, right?”

Parameter Counts and Model Performance

36:50 to 38:40

Discuss the significance of parameter counts in model performance and their implications for developers.

“We're not looking at the index and the hallucination rate stuff that we think is much more about how the models are trained.”

Exploring GDP Value Tasks

38:40 to 40:20

Get insights into the GDP value tasks and their role in assessing AI capabilities for white-collar work.

“I mean, obviously, if you're a developer or company using these things, exactly as you say, it doesn't matter.”

Evaluating AI Models with Gemini 3 Pro

40:20 to 42:00

Examine how the Gemini 3 Pro evaluates model outputs and the challenges of using LLMs as judges.

“subtasks that are the level that we run through the agent acanus.”

ELO vs. Percentages in Model Evaluation

42:00 to 44:00

Explores the reasoning behind using ELO scores over percentage metrics for evaluating AI model outputs.

“Yeah, the thing that you have to watch out for with LM Judge is self-preference, that models usually prefer their own output.”

Agentic Harness and Model Performance

44:00 to 47:10

Discusses the performance differences of AI models in various frameworks and their implications.

“And so there's no kind of ground truth necessarily to compare against, to work out percentage correct.”

Integrating AI with Tools and Data Sources

47:10 to 50:00

Covers the capabilities of AI models in integrating with tools like Google Drive and databases.

“What tools and what data connections come to mind when you say, what's interesting?”

Openness Index: A New Metric for AI Models

50:00 to 53:20

Introduces the Openness Index as a measure of transparency and collaboration in AI model development.

“I mean, we can equip it with more tools, but by default, yeah, that's it.”

Evaluating Contributions to Open Source AI

53:20 to 56:00

Discusses the complexities of defining contributions to open-source AI models and the implications for the industry.

“We bring them together to score an openness index for models so that you can, in one place, get this full picture of how open different models are.”

The Importance of Openness in AI Development

56:00 to 57:20

Learn about the significance of openness in AI model training and licensing.

“that we train a model that's not that smart, maybe even not at the frontier for a particular size category, but we chose to open up all the data, all the training code.”

Understanding the Openness Index

57:20 to 59:10

Discover how the openness index categorizes AI model licenses.

“Let's talk a little bit or at least end the pod on just the trend reports that you guys do, which is kind of a bit of the bread and butter, how you make money.”

Trends in AI Cost Efficiency

59:10 to 1:02:30

Examine the declining costs of AI intelligence and its implications.

“times cheaper than gpd4 was at launch right now i think my number is a thousand actually if you look at um the amazon nova models which are very very cheap yeah like my my conservative statement is normally like 100x.”

Hardware Efficiency and AI Performance

1:02:30 to 1:05:00

Explore the relationship between hardware efficiency and AI cost dynamics.

“One of the reasons that's important is that there's a trade-off between the throughput per GPU that you can achieve and the per user speed that you can achieve.”

Token Efficiency in AI Models

1:05:00 to 1:09:20

Learn about the importance of token efficiency in reasoning versus non-reasoning models.

“Yeah, but I remember thinking like, this must be it.”

Benchmarking Challenges and Innovations

1:10:01 to 1:11:00

Learn about the complexities and innovations in multi-turn benchmarking for AI models.

“There's a trade-off in benchmarking here where most benchmarks need to be one turn to be autonomous, to be parallelized and all that.”

Multimodal and Creative Evaluation

1:11:01 to 1:12:04

Explore the importance of multimodal benchmarking and its implications for creative fields.

“We also do speech benchmarking, image benchmarking, video benchmarking, hardware.”

Audience Engagement and Model Measurement

1:12:05 to 1:13:16

Discover how audience requests can drive the development of evaluation categories for AI models.

“We might be able to get a category going on it.”

The Future of AI Intelligence

1:13:17 to 1:14:06

Discuss the ongoing demand for AI intelligence and its implications for development.

“We're going to do a lot and do a lot to be as useful as possible to developers and companies to measure what's important on every one of those and along those lines.”

Upcoming Changes in the Intelligence Index

1:14:07 to 1:16:18

Get insights into planned updates for the intelligence index and their significance.

“I think to that, I ask people, have they ever worked with or managed someone in a work environment and wouldn't press the button that they were smarter to make them smarter or better at their job?”

Community Impact and Collaboration

1:16:19 to 1:17:44

Understand the role of community engagement in promoting projects and innovations in AI.

“Once it's v4.1, those numbers won't be compatible with v4.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:06This is kind of a full circle moment for us in a way. because the first time artificial analysis got mentioned on a podcast was you and Alessio on Land of Space. Amazing. Which was January 2024. I don't even remember doing that, but yeah, it was very influential to me. Yeah, I'm looking at AI News for Jan 17 or Jan 16, 2024. I said, this gem of a models and hosts comparison site was just launched. And then I put in a few screenshots And I said, it's an independent third party. It clearly outlines the quality versus throughput trade-off. And it breaks out by model and hosting provider. I did give you shit for missing fireworks.

0:51How do you have a model benchmarking thing without fireworks? But you had together, you had perplexity. And I think we just started chatting there. Welcome, George and Micah, to Lanespace. I've been following your progress. Congrats on an amazing year. You guys have really come together to be the presumptive new garner of AI, right? Which is something that... You can't pay us for better results. Yes, exactly. Kicking off straight into it. Let's go. Start off with a spicy take. Okay. How do I pay you? Let's get right into that. How do you make money? Well, very happy to talk about that. So it's been a big journey the last couple of years.

1:28Artificial analysis is going to be two years old in January, 2026, which is pretty soon now. We first run like the website for free, obviously, and give away a ton of data to help developers and companies navigate AI and make decisions about models, providers, technologies across the AI stack for building stuff. We're very committed to doing that and tend to keep doing that. We have along the way built a business that is working out pretty sustainably. We've got just over 20 people now. and two main customer groups. So we want to be who enterprise look to for data and insights on AI. So we want to help them with their decisions about models and technologies for building stuff.

2:12And then on the other side, we do private benchmarking for companies throughout the AI stack who build AI stuff. So no one pays to be on the website. We've been very clear about that from the very start because there's no use doing what we do unless it's independent AI benchmarking. But it turns out a bunch of our stuff can be pretty useful to companies building AI stuff. And is it like, I'm a Fortune 500, I need advisors on objective analysis, and I call you guys and you pull up a custom report for me, you come into my office and give me a workshop? What kind of engagement is that? So we have a benchmarking insight subscription which looks like standardized reports that cover key topics or key challenges enterprises face when looking to understand AI and choose between all the technologies.

3:04And so, for instance, one of the report is a model deployment report. How to think about choosing between serverless inference, managed deployment solutions, or leasing chips and running inference yourself is an example kind of decision that big enterprises face, and it's hard to reason through. Like this AI stuff is really new to everybody. And so we try and help with our reports and insight subscription companies navigate that. We also do custom private benchmarking. And so that's very different from the public benchmarking that we publicize. And there's no commercial model around that. But for private benchmarking, we'll at times create benchmarks, run benchmarks to specs that enterprises want.

3:52And we'll also do that sometimes for AI companies who have built things and we help them understand what they've built with private benchmarking, you know, through the expertise, mainly that we've developed through trying to support everybody publicly with our public benchmarks. Yeah, let's talk about TechStack behind that. But OK, I'm going to rewind all the way to when you guys started this project. You were all the way in Sydney? Yeah, well, Sydney, Australia for me. George was in SF, but he's Australian, but he moved here already. Yeah. And I remember I had that Zoom call with you. What was the impetus for starting artificial analysis in the first place?

4:29You know, you started with public benchmarks. And so let's start there and we'll go to the private stuff. Yeah. Why don't we even go back a little bit to like why we thought that it was needed? Yeah. The story kind of begins like in 2022, 2023. Like both George and I have been into AI stuff for quite a while. In 2023 specifically, I was trying to build a legal AI research assistant. So it actually worked pretty well for its era, I would say. But I was finding that the more you go into building something using LLMs, the more each bit of what you're doing ends up being a benchmarking problem. So I had this multi-stage algorithm thing, trying to figure out what the minimum viable model for each bit was, trying to optimize every bit of it.

5:18As you build that out, right, like you're trying to think about accuracy, a bunch of other metrics, and performance and cost, and mostly just no one was doing anything to independently evaluate all the models and certainly not to look at the trade-offs for speed and cost. So we basically set out just to build a thing that developers could look at to see the trade-offs between all of those things measured independently across all the models and providers. honestly, it was probably meant to be a side project when we first started doing it. We didn't get together and say, hey, we're going to stop working on all this stuff.

5:54I'm like, this is going to be our main thing. When I first called you, I think you hadn't decided on starting a company yet. That's actually true. I don't even think with Paul's like, George didn't quit his job. I didn't quit working on my legal AI thing. It was genuinely a side project. We built it because we needed it as people building in the space and thought, oh, other people might find it useful too. So we'll buy domain and link it to the Vercel deployment that we had and tweet about it. But very quickly, it started getting attention. Thank you, Swix, for I think doing an initial retweet and spotlighting it there, this project that we released.

6:31And then very quickly, though, it was useful to others, but very quickly it became more useful as the number of models released accelerated. We had Mixtrel 8 times 7b, and it was a key... That's a fun one. Yeah, like an open source model that really changed the landscape and opened up people's eyes to other serverless inference providers and thinking about speed, thinking about cost. And so it became more useful quite quickly. Yeah. What I love talking to people like you who sit across the ecosystem is, well, I have theories about what people want, but you have data. And that's obviously more relevant.

7:12But I want to stay on the origin story a little bit more. When you started out, I would say, I think the status quo at the time was every paper would come out and they would report their numbers versus competitor numbers. And that's basically it. And I remember I did the legwork. I think everyone has some version of Excel sheet or Google sheet where you just copy and paste the numbers from every paper and just post it up there. And then sometimes they don't line up. because they're independently run. And so your numbers are going to look better than, your reproductions of other people's numbers are going to look worse because you don't hold their models correctly or whatever the excuse is.

7:49I think then Stanford Helm, Percy Liang's project would also have some of these numbers. And I don't know if there's any other source that you can cite. The way that if I were to start artificial analysis at the same time you guys started, I would have used EleutherAI's eval framework harness. Yep. That was some cool stuff. At the end of the day, right, running these evals, it's like, if it's a simple Q &A eval, all you're doing is asking a list of questions and checking if the answers are right, which shouldn't be that crazy. but it turns out there are an enormous number of things that you've got control for and i mean back when we started the website like one of the reasons why we were we realized that we had to run the evals ourselves and couldn't just take um results from the labs was just that they would all prompt the models differently and when you're competing over a few points then you can um you can put the answer into the model i mean yeah that in the extreme and like you get crazy cases like back when i'm googled a gemini 1.0 ultra and needed a number that would say it was better than gpd4 and um like constructed um i think never published like chain of thought examples 32 of them in every topic in mlu to run it to get the score like there are so many things that you they never shipped ultra right um that's the this one never yeah yeah i mean i'm sure it existed but yeah so we were pretty sure that we needed to run them ourselves and just run them in the same way across all the models yeah and we were we also did certain from the start that you couldn't look at those in isolation you needed to look at them alongside the cost and performance stuff yeah okay a couple technical questions i mean so obviously i also thought about this and i didn't do it because cost did you did you not worry about costs were you funded already clearly not but you know no we well we definitely weren't at the start so like i mean we're paying for it personally at the start the numbers were nearly as bad a couple of years ago so we like certainly um like incurred some costs but we were probably in the order of like hundreds of dollars of spend across all the benchmarking that we were doing okay so nothing yeah it was like kind of fine yeah these days that's gone up an enormous amount for a bunch of reasons that we can talk about um but yeah it wasn't that bad because you can also remember that like the number of models we were dealing with was hardly any.

10:11And the complexity of the stuff that we wanted to do to evaluate them was a lot less. Like we were just asking some Q &A type questions. And then one specific thing was for a lot of evals initially, we were just like sampling an answer directly without letting the models think. We weren't even doing chain of thought stuff initially. And that was the most useful way to get some results initially. Yeah. And so for people who haven't done this work, literally parsing the responses is a whole thing. right like because sometimes the models the models can answer any way they feel fit and sometimes they actually do have the right answer but they just return the wrong format and they will get a zero for that unless you work it into your parser and that involves more work yeah and so i mean but there's an open question whether you should give it points for not following your instructions on the format so it depends what you're looking at right because you can if you're trying to see whether or not it can solve a particular type of reasoning problem and you don't want to test it on its ability to do answer formatting at the same time, then you might want to use an LLM as answer extractor approach to make sure that you get the answer out no matter how it answered.

11:18But these days it's mostly less of a problem. Like if you instruct a model and give it examples of what the answers should look like, it can get the answers in your format. And then you can do like a simple regex. Yeah. Yeah. And then there's other questions around, I guess sometimes that if you have a multiple choice question, sometimes there's a bias towards the first answer so you have to randomize the responses all these nuances you'll like once you dig into benchmarks you're like i don't know how anyone believes the numbers it's so it's so dark magic you've also got like the um uh different degrees of variance and different benchmarks right so if you if you run four question multi-choice on a modern reasoning model at the temperatures suggested by the labs for their own models the variance that you can see on a four and multi-choice eval is pretty enormous if you only do a single run of it and it has a small number of questions, especially.

12:07So one of the things that we do is run an enormous number of all of our evals when we're developing new ones and doing upgrades to our intelligence index to bring in new things so that we can dial in the right number of repeats so that we can get to the 95 % confidence intervals that we're comfortable with so that when we pull that together, we can be confident in intelligence index to at least as tight as like a plus or minus one at a 95 % confidence. Yeah. And again, that just adds a straight multiple to the cost. Oh, yeah. So there's one of many reasons that cost has gone up a lot more than linearly over the last couple of years.

12:44We report a cost to run the artificial analysis intelligence index on our website. And currently that's assuming one repeat in terms of how we report it because we want to reflect a bit about the weighting of the index. But our cost is actually a lot higher than what we report there because of the repeats. Yeah. And probably this is true, but just checking, you don't have any special deals with the labs. They don't discount it. You just pay out of pocket or out of your customer funds. Oh, there is a mix. So the issue is that sometimes they may give you a special endpoint. A hundred percent. Yeah, yeah, yeah.

13:22Exactly. Exactly. So we laser focus on everything we do, on having the best independent metrics and making sure that no one can manipulate them in any way. There are quite a lot of processes we've developed over the last couple of years to make that true for the one you bring up right here of the fact that if we're working with a lab, if they're giving us a private endpoint to evaluate a model, that it is totally possible that what's sitting behind that black box is not the same as they serve on a public endpoint we're very aware of that we have what we call a mystery shopper policy and so and we're totally transparent with all the labs we work with about this that we will register accounts not on our own domain and run both intelligence evals and performance benchmarks that's the job without them being able to identify it and no one's ever had a problem with that because like a thing that turns out to actually be quite a good factor in the industry is that they all want to believe that none of their competitors could manipulate what we're doing either.

14:23That's true. I never thought about that. I've been in a database data industry prior, and there's a lot of shenanigans around benchmarking. So I'm just kind of going through the mental laundry list. Did I miss anything else in this category of shenanigans?

14:37The biggest one that I'll bring up is more of a conceptual one actually than direct shenanigans. it's that the things that get measured become things that get targeted by labs they're trying to build right exactly so that doesn't mean anything that we should really call shenanigans like i'm not talking about training on test set but if you know that you're going to be great or not a particular thing if you're a researcher there are a whole bunch of things that you can do to try to get better at that thing that preferably are going to be helpful for a wide range of how actual users want to use the thing that you're building, but will not necessarily do that.

15:14So for instance, the models are exceptional now at answering competition maths problems. There is some relevance of that type of reasoning, that type of work to like how we might use modern coding agents and stuff, but it's clearly not one for one. So the thing that we have to be aware of is that once an eval becomes the thing that everyone's looking at, the scores can get better on it without there being a reflection of overall generalized intelligence of these models getting better. That has been true for the last couple of years. It'll be true for the next couple of years. There's no silver bullet to defeat that other than building new stuff to stay relevant and measure the capabilities that matter most to real users.

15:58Yeah. And we'll cover some of the new stuff that you guys are building as well, which is cool. You used to just run other people's evals, but now you're coming up with zero. And then I think obviously that is a necessary path once you're at the frontier, you've exhausted all the existing well-known ones. I think the next point in history that I have for you is AI Grant that you guys decided to join and move here. What was it like? I think you were in like batch two? Batch four. Batch four. Okay. I mean, it was great. Nat and Daniel are obviously great. And it's a really cool group of companies that we were in AI Grant alongside.

16:34It was really great to get Nat and Daniel on board. Obviously, they've done a whole lot of great work in the space with a lot of leading companies and were extremely aligned with the mission of what we were trying to do. We're not quite typical of a lot of the other AI startups that they've invested in, and they were very much here for the mission of what we want to do. Did they say any advice that really affected you in some way, or were one of the events very impactful? That's an interesting question. I mean, I remember fondly a bunch of the speakers who came and did fireside chats at AI Ground.

17:09Which is also like a crazy list. Yeah. Oh, yeah. Yeah, yeah, yeah. There was something about, you know, speaking to that and Daniel about the challenges of working through a startup and just working through the questions that don't have like clear answers and how to work through those kind of methodically and just like work through the hard decisions. And they've been great mentors to us as we've built artificial analysis. Another benefit for us was that other companies in the batch and other companies in AI Grime are pushing the capabilities of what AI can do at this time. And so being in contact with them, making sure that artificial analysis is useful to them has been fantastic for supporting us in working out how should we build out artificial analysis to continue to be useful to those, like, you know, building on AI.

17:59I think to some extent, I'm mixed opinion on that one because to some extent, your target audience is not people in AI grants who are obviously at the frontier. Yeah, to some extent. So a lot of what the AI grant companies are doing is taking capabilities coming out of the labs and trying to push the limits of what they can do across the entire stack for building great applications, which actually makes some of them pretty archetypical power users of artificial analysis. Some of the people with the strongest opinions about what we're doing well and what we're not doing well and what they want to see next from us.

18:36Because when you're building any kind of AI application now, chances are you're using a whole bunch of different models. You're maybe switching reasonably frequently for different models and different parts of your application to optimize what you're able to do with them at an accuracy level and to get better speed and cost characteristics. so for many of them no they're like not commercial customers of ours like we don't charge for all our data on the website but they are absolutely some of our power users so let's talk about just the the evals as well right like you start out from the general like mmlu and gpqa stuff um what's next how do you how do you sort of build up to um the overall index what was in v1 and how did you evolve it okay so first just like background like we're talking about the Artificial Analysis Intelligence Index, which is our synthesis metric that we pulled together currently from 10 different eval data sets to give what we're pretty confident is the best single number to look at for how smart the models are.

19:40Obviously, it doesn't tell the whole story. That's why we published the whole website of all the charts to dive into every part of it and look at the trade-offs, but best single number. So right now, it's gotten a bunch of Q &A type data sets that have been very important to the industry, like a couple that you just mentioned. It's also got a couple of agentic data sets. It's got our own long context reasoning data set and some of the use case focus stuff. As time goes on, the things that we're most interested in that are going to be important to the capabilities that are becoming more important for AI, what developers are caring about, are going to be first around agentic capabilities.

20:17So surprise, surprise, we're all loving our coding agents and how the model is going to perform like that and then do similar things for different types of work are really important to us. The linking to use cases, to economically valuable use cases are extremely important to us. And then we've got some of these things that the models still struggle with, like working really well over long contexts that are not going to go away as specific capabilities and use cases that we need to keep evaluating. But I guess one thing I was driving was like the V1 versus the V2 and how bad it was over time. Like how we've changed the index to where we are.

20:55And I think that reflects on the change in the industry, right? So that's a nice way to tell that story. Well, V1 would be completely saturated right now by almost every model coming out because doing things like writing the Python functions and human evil is now pretty trivial. It's easy to forget, actually, I think, how much progress has been made in the last two years. Like we obviously play the game constantly of like the today's version versus last week's version and the week before and all of the small changes and the horse race between the current frontier and who has the best like smaller than 10B model like right now this week, right?

21:32And that's very important to a lot of developers and people and especially in this particular city of San Francisco. But when you zoom out a couple of years ago, literally most of what we were doing to evaluate the models then would all be 100 % solved by even pretty small models today. and that's been one of the key things, by the way, that's driven down the cost of intelligence at every tier of intelligence. We can talk about more in a bit. So V1, V2, V3, we made things harder. We covered a wider range of use cases and we tried to get closer to things developers care about as opposed to just the Q &A type stuff that MMLU and GPQA represented.

22:12Yeah, I don't know if you have anything to add there or we could just go right into showing people the benchmark and clicking around and asking questions about it. Yeah, let's do it. This would be a pretty good way to chat about a few of the new things we've launched recently. Yeah, and I think a little bit about the direction that we want to take it. And we want to push benchmarking. Currently, the Intelligence Index and evals focus a lot on kind of raw intelligence, but we kind of want to diversify how we think about intelligence. And we can talk about it, but kind of new evals that we've kind of built and partnered on focus on topics like hallucination.

22:52And we've got a lot of topics that I think are not covered by the current eval set that should be. And so we want to bring that forth. But before we get into that. So for listeners, just as a timestamp right now, number one is Gemini 3 Pro High, then followed by Cloud Opus at 70. Just 5.1 High. You don't have 5.2 yet. and Kimi K2 thinking, wow, still hanging in there. So those are the top four. That will date this podcast quickly. Yeah, yeah. I mean, I love it. I love it. No, no, 100%. Look back this time next year and go, how cute. Yep. Totally. A quick view of that is, okay, there's a lot. I love this chart.

23:30This is such a favorite, right? Yeah. And almost every talk that George or I give at conferences and stuff, we always put this one up first to just talk about situating where we are in this moment in history. this i think is the the visual version of what i was saying before about the zooming out and remembering how much progress there's been if we go back to just over a year ago before 01 before claude sonnet 3.5 we didn't have reasoning models or coding agents as a thing and the game was very very different if we go back even a little bit before then we're in the era where when you look at this chart, OpenAI was untouchable for well over a year.

24:10And you would remember that time period well of there being very open questions about whether or not AI was going to be competitive, like full stop, whether or not OpenAI would just run away with it, whether we would have a few frontier labs and no one else would really be able to do anything other than consume their APIs. I am quite happy overall that the world that we have ended up in is one where multi-model. Absolutely. And strictly more competitive every quarter over the last few years. Yeah. This year has been insane. Yeah. You can see it. This chart with everything added is hard to read currently.

24:48There's so many dots on it, but I think it reflects a little bit, you know, what we felt like. How crazy it's been. Why 14 as the default? Is that a manual choice? Because you've got ServiceNow in there that are less traditional names. Yeah, it's models that we're kind of highlighting by default in our charts, in our intelligence index. Okay. You just have a manually curated list of stuff. Yeah, that's right. But something that I actually don't think every artificial analysis user knows is that you can customize our charts and choose what models are highlighted. Yeah, it's super important. Yeah, yeah.

25:22And so if we take off a few names, it gets a little easier to read. Yeah, yeah. A little easier to read. Yeah, but I love that you can see the all one jump. Look at that. September 2024. And the DeepSeek jump. Yeah. Which got close to OpenAI's leadership. They were so close. I think, yeah, we remember that moment. Around this time last year, actually. Yeah, yeah, yeah. Yeah, well, a couple of weeks. It was Boxing Day in New Zealand when DeepSeek v3 came out. And we'd been tracking DeepSeek and a bunch of the other global players that were less known over like the second half of 2024 and had run evals on the earlier ones and stuff i i very distinctly remember boxing day in new zealand i because i was with family for christmas and stuff running evals and getting back result by result on deep seek v3 um so this was like the the first of their v3 architecture the 611b moe um and we were very very impressed like That was the moment where we were sure that DeepSeek was no longer just one of many players, but had jumped up to be a thing.

26:33The world really noticed when they followed that up with the RL working on top of E3 and R1 succeeding a few weeks later. But the groundwork for that absolutely was laid with an extremely strong base model, completely open weights, that we had as the best open weights model on Boxing Day last year. Boxing Day is the day after Christmas for those. I mean, I'm from Singapore. A lot of us remember Boxing Day for a different reason, for the tsunami that happened. Oh, of course. Yeah, but that was a long time ago. So, yeah, so this is the rough pitch of AAQI. Or is it AAQI or AII? AII. So, okay. Good memory, though.

27:12I don't know. I got used to it. Once upon a time, we did call it Quality Index. Okay. And we would talk about quality performance and price. But we changed it to IntelliJS. There's been a few naming changes. We added hardware benchmarking to the site and serve benchmarks at a kind of system level. And so then we changed our throughput metric to we now call it output speed and then throughput makes sense at a system level. So we took that name. Take me through more charts. What should people know? Obviously the way you look at the site is probably different than how a beginner might look at it.

27:42Yeah, that's fair. There's a lot of fun stuff to dive into. Maybe so we can hit past all the like, we have lots and lots of emails and stuff um the interesting ones to talk about today that'd be great to bring up like a few of our recent things i think um that probably not many people be familiar with yet so first one of those is our omniscience index so this one is a little bit different to most of the intelligence evals that we run we built it specifically to look at the embedded knowledge in the models and to test hallucination by looking at when the model doesn't know the answer, so I'm not able to get it correct, what's its probability of saying, I don't know, or giving an incorrect answer.

28:28So the metric that we use for omniscience goes from negative 100 to positive 100, because we're simply taking off a point if you give an incorrect answer to the question. We're pretty convinced that this is an example of where it makes most sense to do that because it's strictly more helpful to say, I don't know, instead of giving a wrong answer to factual knowledge question. And one of our goals is to shift the incentive that evals create for models and the labs creating them to get higher scores. and almost every eval across all of AI up until this point, it's been graded by simple percentage correct as the main metric, the main thing that gets hyped.

29:15And so you should take a shot at everything. There's no incentive to say, I don't know. So we did that for this one here. I think there's a general field of calibration as well, like the confidence in your answer versus the rightness of the answer. Yeah, we completely agree. On that, and one reason that we didn't do that or put that into this index is that we think that the way to do that is not to ask the models how confident they are. I don't know, maybe. It might be, though. You put it at JSON field, say confidence, and maybe it spits out something. Yeah. We have done a few evals podcasts over the years, and we did one with Clementine of Hugging Face, who maintains the open source leaderboard.

Read the full transcript

29:57And this was one of her top requests, which is some kind of hallucination slash lack of confidence calibration thing. And so, hey, this is one of them. And I mean, like anything that we do, it's not a perfect metric or the whole story of everything that you think about as hallucination. But yeah, it's pretty useful and has some interesting results. One of the things that we saw in the hallucination rate is that anthropocene clawed models at the very left-hand side here with the lowest hallucination rates out of the models that we've evaluated amnesia on. That is an interesting fact. I think it probably correlates with a lot of the previously not really measured vibe stuff that people like about some of the cloud models.

30:40Is the dataset public? Or what's, is it, is there a held out set? There's a held out set for this one. So we have published a public test set, but we've only published 10 % of it. The reason is that for this one here specifically, it would be very, very easy to like have data contamination because it is just factual knowledge questions. We will update it over time to also prevent that. But we've kept most of it held out so that we can keep it reliable for a long time. It leads us to a bunch of really cool things, including breakdown quite granularly by topic. And so we've got some of that disclosed on the website publicly right now.

31:18And there's lots more coming in terms of our ability to break out very specific topics. Yeah, I would be interested. Let's dwell a little bit on this hallucination one. I noticed that Haiku hallucinates less than Sonnet, hallucinates less than Opus. And would that be the other way around in a normal capability environment? I don't know. What do you make of that? One interesting aspect is that we've found that there's not really a strong correlation between intelligence and hallucination rate. That's to say that the smarter the models are in a generalist sense isn't correlated with their ability to, when they don't know something, say that they don't know.

31:55So it's interesting that Gemini 3 Pro preview was a big leap over here, Gemini 2.5 Flash and 2.5 Pro. But, and if I add Pro quickly here. I bet Pro is really good. Actually, no. So I meant the GPT Pros. Oh, yeah. Because GPT Pros are rumored, we don't know for a fact, that it's like eight runs and then with the LM judge on top. Yeah. So we saw a big jump in, this is accuracy, so this is just percent that they get correct. and Gemini 3 Pro knew a lot more than the other models. And so big jump in accuracy, but relatively no change between the Google Gemini models between releases. And the hallucination rate.

32:38Exactly. And so it's likely due to just kind of different post-training recipe between the clawed models that's driven this. Yeah, you can partially blame us on how we define intelligence, having until now not defined hallucination as negative in the way that we think about intelligence. And so that's what we're changing. I know many smart people who are confidently incorrect. Look, that is very human. Very true. And there's times and a place for that, I think. Our view is that hallucination rate makes sense in this context where it's around knowledge, but in many cases people want the models to hallucinate, to have a go.

33:18Often that's the case in coding or when you're trying to generate newer ideas. One eval that we added to artificial analysis is critical point, and it's really hard physics problems. Is it sort of like a human eval type or something different? Or like a frontier math type? It's not dissimilar to frontier math. So these are kind of research questions that academics in the physics world would be able to answer. But models really struggle to answer. So the top score here is 9%. And when the people that created this, like Minhui and actually Ofia, who was kind of behind Sweebench. What organization is this?

34:00Oh, is this Princeton? Kind of a range of academics from different academic institutions. Really smart people. They talked about how they turn the models up in terms of the temperature. As high a temperature as they can when they're trying to explore kind of new ideas in physics as a thought partner. Just because they want the models to hallucinate. Yeah, sometimes it's a feature. Maybe get to something new. Yeah, exactly. So not right in every situation, but I think it makes sense to test hallucination in scenarios where it makes sense. So the obvious question is, this is one of many, that every lab has a system card that shows some kind of hallucination number and you've chosen to not endorse that and you've made your own.

34:41And I think that's a choice. Totally. In some sense, the rest of artificial analysis is public benchmarks that other people can independently rerun. You provide it as a service. here you have to fight the, well, who are we to do this? And your answer is that we have a lot of customers. But I guess, how do you converge the industry on one number that actually everyone agrees is the rate, right? Because you have your numbers, they have their numbers, never the two shall meet. I mean, I think for hallucinations specifically, there are a bunch of different things that you might care about reasonably and that you'd measure quite differently.

35:16Like we've called this AA amnesian's hallucination rate, not trying to declare like it's the man is last in the hallucination you could uh you could have some interesting naming conventions and all this stuff um the biggest picture answer to that and something that i actually wanted to mention just as george was explaining critical point as well is so as we go forward we are building evals internally we're partnering with academia and partnering with ai companies to build great evals we have pretty strong views on in various ways for different parts of the AI stack where there are things that are not being measured well or things that developers care about that should be measured more and better.

35:52And we intend to be doing that. We're not obsessed necessarily with that. Everything we do, we have to do entirely within our own team. Critical Point is a cool example of where we were a launch partner for it, working with academia. We've got some partnerships coming up with a couple of leading companies. Those ones, obviously, we have to be careful with on some of the independent stuff, but with the right disclosure, like we're completely comfortable with that. A lot of the labs have released great data sets in the past that we've used to great success independently. And so it's between all of those techniques, we're going to be releasing more stuff in the future.

36:25Cool. Let's cover the last couple and then we'll, I want to talk about your trends analysis stuff, you know? Totally. Before that, actually I have one like little factoid on Omniscience. If you go back up to accuracy on Omniscience, an interesting thing about this accuracy metric is that it tracks more closely than anything else that we measure the total parameter count of models. Makes a lot of sense intuitively, right? Because this is a knowledge eval. This is the pure knowledge metric. We're not looking at the index and the hallucination rate stuff that we think is much more about how the models are trained.

36:58This is just what facts did they recall? And yeah, it tracks parameter count extremely closely. Okay. What's the rumored size of GPT-3 Pro? So, and to be clear, not confirmed for any official source, just rumors, but rumors do fly around. I hear all sorts of numbers. I don't know what to trust. So if you draw the line on Amnesty and Zaki receivers total parameters, we've got all the open-weights models. You can squint and see that likely the leading frontier models right now are quite a lot bigger than the one trillion parameters that the open-weights models cap out at and the ones that we're looking at here.

37:35There's an interesting extra data point that Elon Musk revealed recently about XAI that 3 trillion parameters for Grok 3 and 4, 6 trillion for Grok 5 but that's not it yet. Take those together, have a look. You might reasonably form a view that there's a pretty good chance that Gemini 3 Pro is bigger than that that it could be in the 5 to 10 trillion parameter range. To be clear, I have absolutely no idea but just based on this chart, that's where you would land if you have a look at it. And to some extent, I actually kind of discourage people from guessing too much, because what does it really matter?

38:12As long as they can serve it as a sustainable cost, that's about it. Yeah, totally. They've also got different incentives in play compared to open weights models who are thinking to supporting others in self-deployment. For the labs who are doing inference at scale, it's, I think, less about total parameters in many cases when thinking about inference costs and more around number of active parameters. And so there's a bit of an incentive towards larger, sparser models. Agreed. Understood. Yeah. Great. I mean, obviously, if you're a developer or company using these things, exactly as you say, it doesn't matter.

38:45You should be looking at all the different ways that we measure intelligence. You should be looking at cost to run index number and the different ways of thinking about token efficiency and cost efficiency based on the list prices, because that's all that matters. It's not as good for the content creator rumor mill, where I can say, oh, GPT-4 is this small circle. Look at GPT-5 is this big circle. and that used to be a thing for a while yeah i mean but that that that is like a on us own actually very interesting one right that is it well just purely that chances are the last couple years haven't seen a dramatic scaling up in the total size of these models and so there's a lot of room to go up properly in total size of the models especially with the upcoming hardware generations.

39:28Yes. So, um, you know, taking off my shit posting phase for a minute. Uh, yes, yes. At the same time, I do feel like, you know, especially coming back from Europe's people do feel like Ilya is probably right that the paradigm is, doesn't have many more orders of magnitude to scale out more and therefore we need to start exploring at least a different path. GDP value, I think it's like only like a month or so old. Um, I was also very positive when I first came I actually talked to Tejo, who was the lead researcher on that. Oh, cool. And you have your own version. It's a fantastic data set. Yeah.

40:02Maybe I will recap for people who are still out of it. It's like 44 tasks based on some kind of GDP cutoff that's meant to represent broad white-collar work that is not just coding. Yeah. Each of the tasks have a whole bunch of detailed instructions, some input files for a lot of them. Within the 44, it's divided into 220, 2 to 5 maybe, subtasks that are the level that we run through the agent acanus. And yeah, they're really interesting. I will say that it doesn't necessarily capture all the stuff that people do at work. No avail is perfect. There's always going to be more things to look at, largely because in order to make the tasks well enough to find that you can run them, they need to only have a handful of files and very specific instructions for that task.

40:47And so I think the easiest way to think about them are that they're like quite hard take-home exam tasks that you might do in an interview process. Yeah, for listeners, it's no longer like a long prompt. It is like, well, here's a zip file with like a spreadsheet or a PowerPoint deck or a PDF and go nuts and answer this question. OpenAI released a great data set and they released a good paper which looks at performance across the different web chatbots on the data set. It's a great paper, encourage people to read it. What we've done is taken that data set and turned it into an eval that can be run on any model.

41:23So we created a reference agentic harness that can run the models on the data set. And then we developed evaluator approach to compare outputs. That's kind of AI enabled. So it uses Gemini 3 Pro Preview to compare results, which we tested pretty comprehensively to ensure that it's aligned to human preferences. One data point there is that even as an evaluator, Gemini 3 Pro interestingly doesn't do actually that well in GDP Val AA. Yeah, the thing that you have to watch out for with LM Judge is self-preference, that models usually prefer their own output. And in this case, it was not. Totally. I think the way that we're thinking about the places where it makes sense to use an LLM as judge approach now, like quite different to some of the early LLM as judge stuff a couple of years ago, because some of that, and MTV was a great project that was a good example of some of this a while ago, was about judging conversations and like a lot of style type stuff.

42:29Here, we've got the task that the grading model is doing is quite different to the task of taking the test. When you're taking the test, you've got all of the agentic tools, you're working with the code interpreter and web search, the file system to go through many, many turns to try to create the documents. Then on the other side, and we're graining it, we're running it through a pipeline to extract visual and text versions of the files and be able to provide that to Gemini. And we're providing the criteria for the task and getting it to pick which one more effectively meets the criteria of the task out of two potential outcomes.

43:01It turns out that we proved that it's just very, very good at getting that right, matched with human preference a lot of the time, because I think it's got the raw intelligence, but it's combined with the correct representation of the outputs, the fact that the outputs were created with an agentic task that is quite different to the way the grading model works, and we're comparing it against criteria, not just kind of zero shot trying to ask the model to pick which one is better. Got it. Why is this an ELO and not a percentage, like GDP value? So the outputs look like documents. And there's video outputs or audio outputs from some of the tasks.

43:40It has to make a video? Yeah, for some of the tasks. Some of the tasks. What task is that? I mean, it's in the data set. Like be a YouTuber? It's a marketing video. Oh, wow. What? Like model has to go find clips on the internet and try to put it together. The models are not that good at doing that one for now, to be clear. It's pretty hard to do that with a code interpreter. And the computer stuff doesn't work quite well enough and so on and so on. But yeah. And so there's no kind of ground truth necessarily to compare against, to work out percentage correct. It's hard to come up with correct or incorrect there.

44:13And so it's on a relative basis. And so we use an ELO approach to compare outputs from each of the models between the task. You know what you should do? You should pay a contractor human to do the same tasks and then give it an ELO. And so you have human. I think what's helpful about GDPVal, the OpenAI one, is that 50 % is meant to be normal human. Yes. And maybe domain expert is higher than that. But 50 % was the bar for like, well, if you've crossed 50, you are superhuman. Yeah. So we haven't grounded this score in that exactly. I agree that it can be helpful. but we wanted to generalize this to a very large number of models.

44:57It's one of the reasons that presenting it as ELO is quite helpful and allows us to add models, and it'll stay relevant for quite a long time. I also think it can be tricky looking at these exact tasks compared to the human performance, because the way that you would go about it as a human is quite different to how the models would go about it. Yeah. I also like that you included Llama 4 Maverick in there. Is that just one last... Well, no, no, no, no, no. It is the best model released by Meta. And so it makes it into the homepage default set still for now. Other inclusion that's quite interesting is we also ran it across the latest versions of the web chatbots.

45:38And so we have... Oh, that's right. Oh, sorry. Yeah, I completely missed that. Okay. No, not at all. So that which has a checkered pattern. So that is their harness, not yours, is what you're saying. Exactly. And what's really interesting is that if you compare, for instance, Clawed 4.5 Opus using the Clawed web chatbot, it performs worse than the model in our agentic harness. And so in every case, the model performs better in our agentic harness than its web chatbot counterpart, the harness that they created. my backwards explanation for that would be that well it's meant for consumer use cases and here you're pushing it for something the constraints are different and the amount of freedom you can give the model is different also you have a cost goal we let the models work as long as they want basically do you copy paste manually into the chatbot?

46:30that was how we got the chatbot reference we're not going to be keeping those updated at quite the same scale as on the hundreds of models I don't know, talk to browser base they'll automate it for you True. We should. I have thought about, well, we should turn these chatbot versions into an API because they are legitimately different agents in themselves. Yes. Yeah. And that's grown a huge amount of the last year, right? Like the tools that are available have actually diverged, in my opinion, a fair bit across the major chatbot apps and the amount of data sources that you can connect them to have gone up a lot, meaning that your experience and the way you're using the model is more different than ever.

47:10What tools and what data connections come to mind when you say, what's interesting? What's notable work that people have done? Oh, okay. So my favorite example on this is that until very recently, I would argue that it was basically impossible to get an LLM to draft an email for me in any useful way, because most times you're sending an email, you're not just writing something for the sake of writing it. Chances are context required is a whole bunch of historical emails. Maybe it's notes that you've made. Maybe it's meeting notes. Maybe it's pulling something from your, any of like wherever you at work store stuff.

47:45So for me, like Google Drive, OneDrive, and our super-based databases, if we need to do some analysis or some data or something, preferably, model can be plugged into all of those things and can go do some useful work based on it. The things that like I find most impressive currently that I am somewhat surprised work really well in late 2025 that I can have models use Superbase MCB to read-only, of course, run a whole bunch of SQL queries to do pretty significant data analysis and make charts and stuff, and can read my Gmail and my Notion. Okay, you actually used that. That's good. Is that a Cloud thing?

48:24To various degrees of order, both chatgpd and Cloud right now, I would say that this stuff barely works in fairness right now. because people are actually going to try this after they hear it if you get an email from Micah odds are it wasn't written by a chatbot so yeah I think it is true that I have never actually sent anyone an email drafted by a chatbot yet but you can feel it right this time next year we'll come back and see where it's going totally Superbase shout out another famous Kiwi I don't know if you have any conversations with him about anything in particular on AI building and AI infra?

49:03We have had Twitter DMs with him because we're quite big Superbase users and power users. And we probably do some things more manually than we should in Superbase. So he's just the support line because you're QE's? A little bit, yeah. Been super friendly. One extra point regarding GDPVAL AA is that on the basis of the overperformance of the models compared to the chatbots. It turns out we realized that, oh, like our reference harness that we built actually works quite well on like generalist agentic tasks. This proves it in a sense. And so the agent harness is very minimalist. I think it follows some of the ideas that are in Claude code.

49:50And all that we give it is context management capabilities, a web search, web browsing tool, code execution environment. Anything else? I mean, we can equip it with more tools, but by default, yeah, that's it. We give it for a GDP valid tool to view an image specifically because the models can just use a terminal to pull stuff in text form into context, but to pull visual stuff into context, we had to give them a custom tool. Yeah, exactly. You can explain the next, but... No, so we turned out that we created a good generalist agentic harness And so we released that on GitHub yesterday. It's called Stirrup.

50:30So if people want to check it out. And it's a great base for building a generalist agent. It is kind of cool. For more specific tasks. I'd say the best way to use it is Git clone and then have your favorite coding agent make changes to it to do whatever you want because it's not that many lines of code and the coding agents can work with it super well. Well, that's nice for the community to explore and share and hack on it. I think maybe in other similar environments, the Terminal Bench guys have done, started to harbor. And so it's a bundle of, well, we need our minimal harness, which for them is terminus.

51:08And we also need the RL environments or Docker deployment thing to run independently. So I don't know if you've looked into harbor at all. Is that a standard that people want to adopt? Yeah, we've looked at it from an EVALS perspective and we love Terminal Bench. and host benchmarks of terminal bench on artificial analysis. We've looked at it from a coding agent perspective, but could see it being a great basis for any kind of agents. I think where we're getting to is that these models have gotten smart enough, they've gotten better tools that they can perform better when just given a minimalist set of tools and let them run, let the model control the agentic workflow rather than using another framework that's a bit more built out, that tries to dictate the flow.

51:56Awesome. Let's cover the openness index and then let's go into the report stuff. So that's the last of the proprietary numbers, I guess. I don't know how you classify all these. Yeah, or let's call it the last of the three new things that we're talking about from the last few weeks. Because we do a mix of stuff where we're using open source, where we open source and what we do, and proprietary stuff that we don't always open source. Like long context reasoning data said last year, we did open source and then all of the work on performance benchmarks across the site some of them we looking to open source but some of them like we're constantly iterating on and so on and so on so there's a huge mix i would say just of like stuff that is open source not across the site so that's a lcr for people yeah yeah yeah but let's talk about open let's talk about openness index this here is call it like a new way to think about how open models are we for a long time have tracked where the models are open weights and what the licenses on them are.

52:55And that's pretty useful. That tells you what you're allowed to do with the weights of a model. But there is this whole other dimension to how open models are that is pretty important that we haven't tracked until now. And that's how much is disclosed about how it was made. So transparency about data, pre-training data and post-training data, and whether you're allowed to use that data, and transparency about methodology and training code. So basically, those are the components. We bring them together to score an openness index for models so that you can, in one place, get this full picture of how open different models are.

53:32I feel like I've seen a couple other people try to do this, but they're not maintained. I do think this does matter. I don't know what the numbers mean apart from, is there a max number? Is this out of 20? It's out of 18 currently. And so we've got an openness index page, but essentially these are points. You get points for being more open across these different categories and the maximum you can achieve is 18. So AI2 with their extremely open OMO 332B think model is the leader in a sense. With Hugging Face. Oh, with their small model. It's coming soon. I think we need to run, we need to get the intelligence benchmarks right to get it on the site.

54:13You can't have it open in the next and not include Hugging Face. We love hugging face. We'll have that up very soon. You know, RefineWeb and all that stuff. It's amazing. Is it called FineWeb? FineWeb. Yeah, totally. One of the reasons this is cool, right, is that if you're trying to understand the holistic picture of the models and what you can do with all the stuff the company's contributing, this gives you that picture. And so we are going to keep it up to date alongside all the models that we do intelligence index on the site. And it's just an extra view to understand. Can you scroll down to the trade-offs chart?

54:46Yeah, that one. This really matters, right? Obviously, because you can be super open, but dumb. I mean, the slum obviously goes the wrong way here, right? A lot of people would like to see labs hill climb and target the openness index. This is the access to hill climb. Unfortunately, it might be fundamentally true that the slum will always go this direction, because once you open something up, then everyone else can get to the level of what you opened up. Well, so let me tweak your point system, right? Like you have these like numbers on the point system and it go up to 18, you know, but like just because I have a little bit of open data doesn't mean I'm necessarily that much better in someone who put a lot of effort into their open ways that it's smarter.

55:28So I might just mess with the point system to make sure that like I'm accurately representing the contribution to the openness. It is hard to wait for the materiality of the contribution to open source. Like we tried to make it so that it is quite well-defined and no one can disagree about like which category things should be. And so we're not saying like this was a big contribution or a small contribution in terms of impact on the industry or anything. It's just like how much of your data did you release? I would say that it is still valid to say that we train a model that's not that smart, maybe even not at the frontier for a particular size category, but we chose to open up all the data, all the training code.

56:12That is a very useful exercise for the industry and we want to recognize that even if the smartest model in the category. Yeah, and also a special shout out to NVIDIA Nemotron which doesn't get enough credit for the amount of stuff that they do. And honestly, it's a sales enablement for NVIDIA as well. Like the fact that they can do this is a side project. Totally, but I mean, it is true that NVIDIA have actually put an enormous amount of effort over the last year, especially into the Nemotron models. Yeah, and so many people actually use it for synthetic data and stuff. It's a pretty interesting secret of the industry that NVIDIA holds up all these guys.

56:45I mean, it's in their interest for there to be more AI. So obviously, I think you want to push openness as having an index. Every index that you push has encodes some kind of opinion or value. Yes. I think one of the openness questions for this year was people messing with the license. and so llama had this like if you have 700 million daily active users you're not allowed to use our model or you have to talk to us something like that so basically like what are your customers telling you about the kind of licensing worries that they have right because obviously most people will never hit 700 million users we have like a detailed breakdown of that in the openness index and that was actually one of the initial questions like took us down the route of wanting to do this um because yeah the simplest thing that like our opinion is is that there's a lot of advantage to having like an official osi license like mit or apache 2 because then the box is just checked you don't even need to read it because it's just apache 2 and you can do it ever you want and it's fine there are often very good reasons that companies don't want to release language models with those completely open licenses the index tells you so if you get the top category that's one of those licenses you're totally good and then we've got um some lower categories for when attribution is required and then when commercial use is not allowed.

58:03Yeah, they're there. So that's the openness index. Thank you for doing all those works. Let's talk a little bit or at least end the pod on just the trend reports that you guys do, which is kind of a bit of the bread and butter, how you make money. I highly encourage everyone to see George's talk at World's Fair, which gives a little bit of a preview. And you were very excited about talking about the smiling curve, or I don't know what you call it. Yeah, yeah, yeah. Yeah, let's talk about that one let's explain it for people and and i might i might actually put put it up um because i don't have it yeah you've got to copy the slide better be that'd be excellent it's important for people to have in their head because yeah people only get the marketing message from the labs that oh we're cutting costs all the time yep yep but it's true it's not the whole picture so okay a couple of like the big trends that we track at artificial analysis over time and that like we're always showing charts of on the trends page in these reports and stuff one that the cost of intelligence has been falling dramatically over the last couple of years the best way to think about that is that the cost for each terror of intelligence has been dropping the like one fact on that is that you can get intelligence at the level of gpd4 for over 100 times cheaper than gpd4 was at launch right now i think my number is a thousand actually if you look at um the amazon nova models which are very very cheap yeah like my my conservative statement is normally like 100x.

59:24But in fairness, this slide, like we were actually saying before the podcast, right, is like maybe six months old now. And it's conceptually still correct, but like could actually probably do a tweak on the exact numbers because like the market's moving so quickly. If you're feeling to kick it off, I mean, we'll have this chart. I told people to watch the World's Fair talk, but let's introduce what context makes you make something like this. There are two trends that seem to not make sense together, both of which we talk a lot about artificial analysis and are very important to developers building stuff in AI.

59:56The first is that the cost of intelligence for each level of intelligence has been dropping dramatically over the last couple of years. We track the cost run artificial analysis intelligence index for each bucket of intelligence index scores and each bucket you just see the line go down really really quickly and actually go down more quickly for each new level of intelligence that's been achieved over the last couple of years. So the rate of that cost decline has actually been going up. So we've got that being true. And yet it is clearly possible to spend quite a lot more on AI inference now than it was a couple of years ago.

1:00:33NVIDIA stock go up. It's going really up. I just heard from a friend's startup that just went to the shift zero. They're spending$5 ,000 per employee on coding agents spend alone. that's uh that's an impressive number we need to get our numbers up we're uh we're not hitting well i was like it's so high down i'm like are you doing something wrong because there are some efficiency questions along the way but like you can make ai inference useful to that level in a bunch of ways that i can imagine right yeah um i i don't think that's that nuts um but basically the reason we made this slide to answer the question right is to to show that the crazy thing is that it is actually true.

1:01:15We've had this 100x to 1000x decline in the cost of GPT-4 level intelligence on the left-hand side. And yet on the right-hand side, because the multipliers are so big for the fact that even though small models can do GPT-4 level now, we still want to use big models and probably bigger than ever models to do frontier level intelligence. We've got reasoning models using tokens, and then we're throwing them in these agentic workflows where they're consuming enormous numbers of input tokens and making enormous numbers of output tokens, working for a really long time, those two things taken together get you back to, we can spend enormously more today than we could a couple of years ago.

1:01:49Yep. I think that's right. There's a number of drivers at play and we kind of outline kind of six key ones here. But as complex, it's changing quickly. All of these have changed very dramatically in the last 12 months. Let's pick on hardware efficiency since you also track hardware stuff. And I think the general assertion or the message is that the efficiency from next-gen NVIDIA chips is actually not 4X. You have what, 3X or 4X? You have 3X in here. And it's like 2X maybe, or it's more of like a power story rather than like a sheer sort of compute tokens efficiency story. But yeah, what's going on in hardware?

1:02:30where okay so the the the answer unfortunately uh is it depends and it just depends massively on like so many things across a bunch of different types of workloads and ways to think about it so one of the simplest ways to think about this is to take single relevant model to think about serving it at speeds that are realistic for what you actually might want to hit and can afford to hit and then think about the throughput per gpu that you can achieve serving the model at those speeds. One of the reasons that's important is that there's a trade-off between the throughput per GPU that you can achieve and the per user speed that you can achieve.

1:03:07And it costs more to serve stuff fast to users. When you run all of that, for especially big sparse models, you can get a lot better than 2 or 3x gain going from hopper to Blackwell generation to video. I shouldn't be too controversial to say, but I'm pretty confident that Blackwell has delivered pretty enormous gains and that the next couple of years of NVIDIA's roadmap are going to continue to deliver quite enormous gains and that those will actually come through as lower total cost per token to the companies that are running models on them and will allow bigger models, will allow way more tokens to be made for lower cost and that that's going to continue.

1:03:49These things also stack on all of the software and model improvements being made. So basically my prediction across both sides of that smile chart that we're going to see the left-hand side continue to be true and probably for another order of magnitude and the right-hand side continue to be true for another order of magnitude. And that's going to enable a whole lot of things. Okay, well, I'll push on. Let's go back to the small chart. I'll push back on sparsity, right? We've gone a long way on sparsity. DeepSeek was a major pusher of fine-grained experts, let's call it. Yep. Right? I have a mental number of sparsity in terms of, let's say, active params versus total params.

1:04:29And that number went from 25%, let's say, down to like 15, right? You obviously can't really go below, I don't know, five. So there's a lower limit to sparsity is what I'm saying. I don't know that that's that obvious, actually. There must be a limit somewhere, right? Yeah, exactly. But we've got numbers in the wild that are quite a lot lower than that right now. So the GBD OSS models, like the big ones at about 5 % active. Kimi K2, is it like 3 % active? Oh, okay. I think, pretty sure. I've looked at those numbers. I calculated them. I don't remember. Yeah, but I remember thinking like, this must be it.

1:05:10Your 5 % is exactly like around the ballpark for the open weights models of what's released today. I think one interesting that gives me kind of pause when thinking that it won't go, the sparsity won't go higher or the number of percentage of active parameters lower is that we in our benchmark see a lot of performance correlated more with total parameters than active. And not that correlated with how sparse the models are. Our accuracy benchmark is part of AA Amniscience. It's very correlated with total. It's not correlated with active parameters, which I think is very interesting. And so I think, yeah, there could be quite a bit to go here.

1:05:54Awesome. Well, we don't have that much time, but I did want to leave some room to cover reasoning and non-reasoning models and token efficiency. Let's do that one. So at a super high level, people have to classify this binary thing of reasoning versus non-reasoning. People who are insider have some discomfort with that because basically you just have the think tag or no think tag. How have you guys decided to approach this? And also, how is that laid out over the course of the year where we have things like GPT-5, which is a model router? Let's say GPT-5 and ChatGPT, the consumer experience, is a model router.

1:06:27When you're hitting the API, you can pick the different versions and you can pick reasoning strength of the different versions. But that goes to why this is now such a complex thing. So earlier this year, and probably when you and George last spoke for the AI Engineers World's Fair, we had this great slide that was super easy where we would show that the average reasoning model is using 10 times the number of tokens per query in our intelligence index as the average non-reasoning model and there was this moment where that was a pretty clear distinction and extremely useful to look at it just like that definitely no longer the case not least because you can think about reasoning strength for a bunch of these different models but particularly because different models have wildly different token efficiency now, more than an order of magnitude in difference, that means that the way that you probably need to think about cost for any application is to use something like our cost to run intelligence index metric as the starting point for what it's going to look like for these different models, these different reasoning strengths, and this continuous spectrum from non-reasoning to reasoning.

1:07:30That's basically like where we're at. So we will still show reasoning and non-reasoning and define reasoning as when there is that separated chain of thought that you're getting at a different parameter in an api normally but it doesn't necessarily anymore mean that that model is actually going to have longer end-to-end latency that is going to use more tokens than something that is branded a non-reasoning model for the same task that's true i think 5.1 was it and then 5.1 codex had these these chart which is super nice of this like let's say bottom 10 percentile query being faster but top 10 percentile being longer and that's a kind of the efficiency chart you want to see right yeah that so that is a an extra thing let's say let's say that we've got that's a really important extra thing though right that you've got not just the average number of tokens being used by the model which we cover really well right now but the behavior that you want in the model is it to use more tokens when it needs more tokens and not to use more tokens when it doesn't need more tokens.

1:08:30So that's what OpenAI, we're basically claiming that 5.1 codex is better at. We don't actually publish anything on this right now, but have tracked it a bunch internally in our internal analytics on evals across all the models that we run, where we look at the difficulty, the questions, and the correlation between token usage and difficulty. And net-net, surprise, surprise, models have got better at doing that over the course of this year. I think going into next year, that's going to be really important, especially as you multiplied by the number of steps in an agentic workflow that a model has to take to get to an answer we are going to care a lot about token efficiency and number of turns efficiency for getting to what we want which would you rather have token efficiency or number of turns efficiency um or like which is more important to work on it's like it depends on the application and both are going to be really important uh yeah because your total total cost retail top edge airline Yeah, interestingly, in Tau2Bench Telecom, it's cheaper to run, you know, on a per token basis, more expensive models like a GBD5 compared to some smaller open source models because some of the GBD5, for instance, got to the answer faster.

1:09:41And so it was able to resolve the customer's query faster in fewer turns. And maybe it used more tokens per turn, but it suddenly cost more per token. So you would always rather use GPT-5 in that scenario. And so I think that's where we're getting to. I think number of turns is going to be a metric that we're going to be talking about a lot more. And I think it'll be something that people want to really start to think about a lot more. There's a trade-off in benchmarking here where most benchmarks need to be one turn to be autonomous, to be parallelized and all that. But most, a lot of real-life use cases need to be multi-turn, especially like quick multi-turns.

1:10:17So you can align. Yeah. Yeah. I mean, I would say that historically benchmarks have been single turn, but I wouldn't say they need to be at all into the future, right? Like we have a couple of agentic benchmarks in the index right now and GDPVal that we were talking about. We led the models to up to a hundred turns and our stirrup agentic harness to do that eval. And we're going to build similar stuff like that in the future. It definitely is hard and you've got whole kinds of infrastructure problems to run that and exactly as you say, parallelize it because we need to run that on hundreds of models and we want to do that really fast when new models come out and where labs want us to run it on their models.

1:10:52But you can do it. We've put it in the work to build that stuff and it's going to be great. Okay, so we've covered, I mean, there's a lot more to cover. I haven't even touched on multimodal, which is huge. We also do speech benchmarking, image benchmarking, video benchmarking, hardware. I like the way that you've done it because it's very smart, which is video takes a long time. So you pre-generate, right? So then people just pick their preferences and you can see the overall arena results. And you also avoid any sensitivity issues around unsafe content that is being generated. Yeah. And you can see it as a good thing or bad thing depending on what your view is, but it means that we have a quite active creative direction approach to trying to understand what creative professionals and users want to do with those image and video models.

1:11:42And so that we can be directing the arenas in our categories toward gathering votes on what people care about. One call out actually to listeners, like if you are using our arenas, is that you can submit requests to us for things that we should cover. I didn't know that. Yeah. Understudied categories, areas that you think the models are bad at and the labs don't focus on enough. Like if you want something solved, one of the levers that you have is send us a couple of prompts on it. We might be able to get a category going on it. And this thing that we talking about earlier, right? That once things get measured, they can get targeted.

1:12:16You can make that work for you. For me as a content creator, infographics, very needed. I took the latest DeepSeek paper and, uh, I, you know, they had some descriptions of their search agents and their coding agents and I put it in and I created an infographic. And, um, I just think like this, an industrial use case that doesn't require a lot of, I guess, design taste, but just requires some, like you need to conform to some preset references, which is something that, uh, that is increasingly important, especially in the Nano Banana series. But yeah, I think OpenAI is releasing Image 2 soon, which is going to have that.

1:12:51So I think it's all of a kind where people need to incentivize workhorse use cases and not just art. I don't know. Totally, yeah. What are we going to be talking about next year? What's emerging that you're seeing and maybe not in the discussion? The first answer that I'll give to that is the boring answer is that on most of our charts, the lines go in a particular direction and our overall prediction is the lines are going to keep going that direction. We're going to do a lot and do a lot to be as useful as possible to developers and companies to measure what's important on every one of those and along those lines.

1:13:24But I think we're going to talk about similar stuff. It's just that we're going to have continued on the trajectory for another year and things are going to feel pretty different because of that happening. I know this is the boring answer to that question. No, no, I mean, I'm a fan of things that, truths that don't change because you can build and plan for that. And I think in media in general, in the podcast business, newsletter business, Twitter business, people are addicted to change. Like, oh, everything's breaking, everything's... No, like there's some truths that are just constants that you can plan on and build.

1:13:56And yeah. I think one of the truths is that the demand for AI intelligence and smarter AI intelligence is going to be insatiable. Some people disagree that, okay, once we reach certain thresholds, then you don't need more intelligence. I think to that, I ask people, have they ever worked with or managed someone in a work environment and wouldn't press the button that they were smarter to make them smarter or better at their job? Or would they never press that for themselves? And I'm not sure that that's the case. But I think for artificial analysis, we'll keep benchmarking raw intelligence, but we also want to think about it and explore models more deeply across other axes as well.

1:14:37I think hallucinations, the start of that, but we're getting into wanting to support people and understanding, okay, the behavior, the personalities of the models to help people make more nuanced decisions. You're going to have a personality bench. Maybe. That is a direction that Chad G, OpenAI is leaning into a lot. So if you manage to solve that, you should definitely talk to Fiji and Roon. Oh, okay. Yeah. So what is going to be included in, let's say, like a V3 of the intelligence index? because obviously you're going to saturate in March. Why don't we break it now? How soon is the podcast going to come out?

1:15:12Wherever you want. Okay, so we're at v3 right now. So the version that's going inside is v3. v4 is what we're going to call the next major update. Surprise, surprise. We're going to be adding several of the things that we've actually talked about today that we've launched over the last few weeks. So that's not going to be wildly shocking, but some of the things that are most exciting is that adding GDPVAL is going to give us this general agentic performance in a really strong way in intelligence index. And in critical point, the physics eval George was talking about similar to Frontier Math, that gives us completely new view with a brand new data set of very, very hard research problems.

1:15:49We are going to be using omniscience and we are going to be using hallucination rate. The exact ways that all those are going to come together. The waiting is going to be hard because the numbers are different. Yeah, we're going to make sure that we don't do anything to cause odd distortions and stuff that could be misleading. But every time you version it, you have a one-time reset of them. Exactly. Yep. That's exactly how we think about it. We will make sure that within each version number that there's no drift in any of the scores so that people can rely on them and reference them. You just have to watch out for that version number.

1:16:19Once it's v4.1, those numbers won't be compatible with v4. Of course. There's a little bit of debate over the accuracy of TileBench. I don't know if you're clued in to what's going on. Apparently, a very high number of TileBench tests are impossible. potentially for the earlier versions TAL2 Bench Telecom we're pretty convinced is pretty good if anything the only issue there is that models have got very good at doing it and so like anything TAL3 yeah on we go yeah on we go okay well thank you so much for providing such a great service to the industry I'm glad to at least know you guys before you got famous and now you are famous oh look our pleasure and we really appreciate your support along the way like I wasn't kidding at the start right that it was a quite material moment for us like when artificial analysis was covered on Latent Space.

1:17:08Some random guy in San Francisco mentions you. I was a fan of Latent Space for like a year before you mentioned us. So I'd been listening. I don't think I was like familiar with like you personally yet at that point, but like I listened to your voice probably for many, many hours. And so once like you mentioned, then like got to get to know you and like meet you for the first time nearly a couple of years ago. It was really cool, honestly. So yeah, it's great to be here. And thanks for being such a great member of the community and kind of spotlighting projects which don't have attention and bringing them to your audience.

1:17:43Yeah. Well, actually, so it wasn't me, right? Someone in the Discord dropped it in our Discord. And I rely on our community and it kind of feeds itself, right? Nice. So someone brought it to my attention. I don't know who. We should probably go back and check. But once I saw it, I was like, this looks good. This is something I always wanted. I wanted to build it. I was too shy or dumb or lazy to build it. And you guys did. And that's a whole thing. So thank you for doing it. I built some really cool other stuff, like this pod. Yeah, yeah. Totally. So thank you. That's it. Great. Thanks.

From the publisher

don’t miss George’s AIE talk: https://www.youtube.com/watch?v=sRpqPgKeXNk

—-

From launching a side project in a Sydney basement to becoming the independent gold standard for AI benchmarking—trusted by developers, enterprises, and every major lab to navigate the exploding landscape of models, providers, and capabilities—George Cameron and Micah-Hill Smith have spent two years building Artificial Analysis into the platform that answers the questions no one else will: Which model is actually best for your use case? What are the real speed-cost trade-offs? And how open is "open" really?

We discuss:

The origin story: built as a side project in 2023 while Micah was building a legal AI assistant, launched publicly in January 2024, and went viral after Swyx's retweet

Why they run evals themselves: labs prompt models differently, cherry-pick chain-of-thought examples (Google Gemini 1.0 Ultra used 32-shot prompts to beat GPT-4 on MMLU), and self-report inflated numbers

The mystery shopper policy: they register accounts not on their own domain and run intelligence + performance benchmarks incognito to prevent labs from serving different models on private endpoints

How they make money: enterprise benchmarking insights subscription (standardized reports on model deployment, serverless vs. managed vs. leasing chips) and private custom benchmarking for AI companies (no one pays to be on the public leaderboard)

The Intelligence Index (V3): synthesizes 10 eval datasets (MMLU, GPQA, agentic benchmarks, long-context reasoning) into a single score, with 95% confidence intervals via repeated runs

Omissions Index (hallucination rate): scores models from -100 to +100 (penalizing incorrect answers, rewarding \"I don't know\"), and Claude models lead with the lowest hallucination rates despite not always being the smartest

GDP Val AA: their version of OpenAI's GDP-bench (44 white-collar tasks with spreadsheets, PDFs, PowerPoints), run through their Stirrup agent harness (up to 100 turns, code execution, web search, file system), graded by Gemini 3 Pro as an LLM judge (tested extensively, no self-preference bias)

The Openness Index: scores models 0-18 on transparency of pre-training data, post-training data, methodology, training code, and licensing (AI2 OLMo 2 leads, followed by Nous Hermes and NVIDIA Nemotron)

The smiling curve of AI costs: GPT-4-level intelligence is 100-1000x cheaper than at launch (thanks to smaller models like Amazon Nova), but frontier reasoning models in agentic workflows cost more than ever (sparsity, long context, multi-turn agents)

Why sparsity might go way lower than 5%: GPT-4.5 is ~5% active, Gemini models might be ~3%, and Omissions Index accuracy correlates with total parameters (not active), suggesting massive sparse models are the future

Token efficiency vs. turn efficiency: GPT-5 costs more per token but solves Tau-bench in fewer turns (cheaper overall), and models are getting better at using more tokens only when needed (5.1 Codex has tighter token distributions)

V4 of the Intelligence Index coming soon: adding GDP Val AA, Critical Point, hallucination rate, and dropping some saturated benchmarks (human-eval-style coding is now trivial for small models)

—

Artificial Analysis

Website: https://artificialanalysis.ai (https://artificialanalysis.ai (\"https://artificialanalysis.ai\"))

George Cameron on X: https://x.com/georgecameron (https://x.com/georgecameron (\"https://x.com/georgecameron\"))

Micah-Hill Smith on X: https://x.com/micahhsmith (https://x.com/micahhsmith (\"https://x.com/micahhsmith\"))

Chapters

00:00:00 Introduction: Full Circle Moment and Artificial Analysis Origins
00:01:19 Business Model: Independence and Revenue Streams
00:04:33 Origin Story: From Legal AI to Benchmarking Need
00:16:22 AI Grant and Moving to San Francisco
00:19:21 Intelligence Index Evolution: From V1 to V3
00:11:47 Benchmarking Challenges: Variance, Contamination, and Methodology
00:13:52 Mystery Shopper Policy and Maintaining Independence
00:28:01 New Benchmarks: Omissions Index for Hallucination Detection
00:33:36 Critical Point: Hard Physics Problems and Research-Level Reasoning
00:23:01 GDP Val AA: Agentic Benchmark for Real Work Tasks
00:50:19 Stirrup Agent Harness: Open Source Agentic Framework
00:52:43 Openness Index: Measuring Model Transparency Beyond Licenses
00:58:25 The Smiling Curve: Cost Falling While Spend Rising
01:02:32 Hardware Efficiency: Blackwell Gains and Sparsity Limits
01:06:23 Reasoning Models and Token Efficiency: The Spectrum Emerges
01:11:00 Multimodal Benchmarking: Image, Video, and Speech Arenas
01:15:05 Looking Ahead: Intelligence Index V4 and Future Directions
01:16:50 Closing: The Insatiable Demand for Intelligence

More from Latent Space: The AI Engineer Podcast

All 247 episodes
Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah-Hill SmithLatent Space: The AI Engineer Podcast · 1 h 18 min
Listen in VO