Do AI Tokenomics Matter More Than Model Benchmarks? with Chris Potts - #776

9 Sep 2026 · 59 min · 28 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode argues that AI progress and value should be measured by token economics (“tokenomics”) and system-level efficiency, not just model benchmarks. As reasoning/agent use increases token consumption, providers are shifting pricing toward true costs, creating “tokenflation” where tokens buy less value.

Guest

Chris Potts is a Stanford professor (linguistics background; PhD work on swearing, semantics/pragmatics, and corpora) and co-founder of Big Spin. He focuses on interpretability, data and architecture efficiency, and how users interact with AI systems.

Key claims

(1) Evaluate “what tokens are actually buying us” using an economist-style “consumer price index” over a basket of engineering outcomes (e.g., PRs, bug fixes, code survival, knowledge discovery) relative to tokens spent. (2) Token efficiency gains flatten due to inference-time scaling; big improvements require system changes, not just more tokens. (3) User expertise and “AI fluency” (iterating, pushing back, collaborating) causally drives success; novices delegate and accept wrong outputs.

Notable examples

Copilot billing sticker shock (e.g., $500/month rising to $11,000); CPI experiment using SWE-bench-like data and “code survival >4 days”; Opus 4.6 CloudCode sessions where token “purchasing power” declined from Feb to mid-April; adaptive thinking/product changes causing ~5x token usage.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Chris Potts' Background and Journey

1:31 to 3:00

Discover Chris Potts' background in linguistics and its connection to NLP.

“I'm Sam Charrington, and this is the TwiML AI podcast.”

Swearing and Linguistic Identity

3:00 to 6:09

Explore the significance of swearing in communication and its cultural implications.

“And I feel like if I had to, I could trace the lineage of every one of my current projects back to my fascination with why we care when someone drops an F-bomb.”

Transition of Linguistics in NLP

6:09 to 8:15

Understand the shift from linguistics to statistical models in NLP.

“These questions are on my mind all the time, yeah, because I operate at the intersection of all these different fields.”

Challenges in NLP Research Today

8:15 to 11:00

Examine the impact of large language models on NLP research and the uncertainty faced by researchers.

“For NLP, I think it's a more uncertain prospect because pre the arrival of like pre-trained models, which for me would be like the ELMO model back in 2017, 2018.”

Future Directions in AI Research

11:00 to 14:00

Discuss innovative approaches in AI research focusing on architecture and data.

“And that felt like a very productive choice.”

The Bitter Lesson in AI Scaling

14:00 to 15:00

Explore the lessons learned about scaling AI models and efficiency.

“This is the bitter lesson-pilled thing, right?”

Understanding Model Mechanisms

15:00 to 17:00

Understand the deep intuitions and analysis driving AI model advancements.

“I think it also really calls out the relationship between data, MEG and TURP, and efficiency, like core themes that you've been focused on and how they, you know, interrelate and support one another.”

The Evolution of Scientific Contributions

17:00 to 18:50

Learn about rethinking scientific contributions beyond traditional papers.

“And what a meta strand I could offer you, because we were talking about being strategic with research.”

Modularity and Prompt Optimization

18:50 to 21:00

Discover the importance of modular design and prompt optimization in AI.

“Yeah, I was thinking not too long ago the degree to which model strength as a correlate to model size, I suppose, and capability has kind of overcome the need for an explicit framework like DSPY, DSPY.”

Variability in Model Responses

21:00 to 23:08

Examine the inconsistencies in AI model responses and their implications.

“They will be blown away by the amount of variation that still exists.”
Show all 28 chapters

Exploring Tokenomics in AI

23:08 to 25:50

Delve into the implications of AI tokenomics and usage costs.

“we were writing a grant and we needed a title and you want to be strategic with these titles.”

AI Cost Shock: A Realization

25:50 to 26:20

Uncover the surprising costs associated with AI services like Copilot.

“And if they keep up the way they are with Copilot's new billing, it will be$11 ,000 in the next month.”

Valuing AI Contributions

26:20 to 28:00

Discuss the complexities in measuring the value of AI contributions in coding.

“We start to wish for those easier stories.”

Understanding Token Economics and Code Metrics

28:00 to 30:28

Explore the relationship between token usage and coding metrics to assess AI efficiency.

“But is that a good thing or a bad thing?”

The Decline of Token Purchasing Power

30:28 to 33:15

Examine recent trends in token usage and its impact on the productivity of AI outputs.

“So we could make it longer, and maybe the adjustment would be different.”

Scaling Laws and Their Implications

33:15 to 35:36

Discuss the implications of scaling laws on token efficiency and AI model performance.

“And this is independent of the approach to inference time scaling you're taking, whether it's multiple parallel inferences or some kind of Oracle or any number of other schemes?”

Task Variability and User Interaction

35:36 to 37:59

Investigate how different types of tasks influence token use and user expectations.

“one kind of pushback that comes up for me is in the real economy, eggs is different than milk, is different than beef, et cetera, et cetera.”

Future Projections in Token Economics

37:59 to 40:36

Speculate on future trends in token economics and their impact on AI cost structures.

“Sometimes you're partnering with the AI.”

The Emergence of Tokenflation

40:36 to 42:00

Analyze the concept of tokenflation and its significance in the current economic landscape.

“I'm wondering, does this research also attempt to project forward?”

Understanding Token Economics and Tokenflation

42:00 to 43:30

Explore the complexities of token economics and the notion of tokenflation.

“predict what token spend is going to be like when you have all these exogenous events in addition to changes that we don't even know about and questions about where the value actually lies.”

Evaluating Model Improvements and Systems Thinking

43:30 to 45:20

Delve into the implications of model improvements and the importance of systems thinking.

“How would you articulate what you're seeing, the models are getting quote unquote better.”

The Role of Expertise in AI Utilization

45:20 to 47:20

Discuss how expertise influences the value derived from AI technologies.

“And so that shows you that even for a fixed model, we can get very different outcomes for these things because they really are sophisticated engineered systems at this point.”

Fluency and Causal Factors in AI Success

47:20 to 50:20

Examine the connection between user fluency and success in utilizing AI.

“because I think probably most things are getting designed for those experts implicitly or explicitly.”

Challenges and Opportunities in AI Accuracy

50:20 to 52:50

Address the challenges of ensuring accuracy in AI responses and the implications of expertise.

“to be more successful with AI what do you think are the key lessons of this fluency work?”

Future Directions in AI Research and Innovation

52:50 to 56:01

Consider the future of AI research, focusing on architecture innovation and data safety.

“But as soon as we leave that and go even into something like the legal realm where the requirements are strict, but they're not codified in code and they have ambiguity about them, this whole picture falls apart.”

Connecting Data and Model Capabilities

56:01 to 56:49

Explore the relationship between data integrity and model safety.

“but it also checks a box for me on connecting interpretability to safety.”

Innovative Research in AI Architecture

56:50 to 57:30

Discuss emerging underappreciated research in AI model architecture.

“And so, again, it's just a data-oriented question that's very alive for me in the current moment.”

Rethinking Gradient-Based Learning

57:31 to 58:26

Consider alternatives to gradient-based learning in AI research.

“But also Julie's perspective is that this is speculative, but I think there's something to this, that it's a kind of inference time scaling because you do more compute at test time because you have more tokens.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00We've started to enter a phase of AI that's not just about making models smarter. It's also about making them economically sustainable. As reasoning models consume more tokens, context windows continue to grow, and agents become embedded in more products and workflows, the economics of these systems are becoming impossible to ignore. That's given rise to a new conversation around tokenomics, how we think about the costs, incentives, and trade-offs shaping the next generation of AI. One person who's been thinking deeply about this is Stanford professor and Big Spin co-founder Chris Potts. His recent work argues that measuring AI progress requires looking beyond model benchmarks to ask a different question.

0:42What are our tokens actually buying us? Here's Chris explaining how he thinks about tokenomics. Another interesting moment to be in as we're all being made aware of the true costs of all this AI usage. The analogy here is like it used to cost me$20 to take a ride share to the airport, Uber or Lyft, and now it costs$90. But it's more like$20 to like$500 or something, right? I think what's happening is that the big providers are testing the waters on charging us the true costs, plus whatever profit they need to make as they all try to gear up for IPOs and so forth. And in turn, that is very quickly leading people to ask questions like, what is the return on investment for all these tokens that we have purchased?

1:25And it's a very tricky area to be in because what does it mean to think about value in this context? I'm Sam Charrington, and this is the TwiML AI podcast. For over a decade, I've been exploring the ideas and innovation shaping the future of AI through conversations like this one that help you understand what's real, what's next, and what matters. Let's jump in.

1:55I want to say thanks for coming on. I've been looking forward to this conversation. And I think where I'd love to start us off is to really dig into your background and how it kind of got you, you know, to where you are now. Yeah, my background is in linguistics. Linguistics proper, not even natural language processing. I did my PhD on, among many other things, swears. What swears are like, why we swear, what information they encode, what kind of taboos exist around them and so forth. And that was actually the trigger that got me into NLP because I wanted a lot of data of people swearing. I wanted to know what the context was like, what their intentions were.

2:40So I turned to corpora. And from there, you start using NLP toolkits to add structure to those corpora. And then after a few years, maybe you're writing your own tools for doing that work. And then when you look back after 18 years or whatever it's been, you're just an AI person or an NLP person. But that is the true story. And I feel like if I had to, I could trace the lineage of every one of my current projects back to my fascination with why we care when someone drops an F-bomb.

3:11So are you an F-bomb dropper or did you come at it from the perspective of trying to understand these others? I think very infrequently in my life, for my linguistics class, semantics and pragmatics, which is about linguistic meaning, on the final day, we always do a class on swearing. And I review the history and we kind of tie all the course themes together. My handouts for that are full of swears. But I only swear once in the lecture. I present the result that people remember things better if the utterance contains a swear because it has a kind of emotional resonance, very primitive reaction.

3:51And so in that moment, I pick some fact from the course, some trivial thing, and I restate it with a swear. And then I say, all of you will remember this for eternity. But other than that, I'm very shy about it in the class. That's funny. I'm sure there is loads of research on this, but I grew up in New York City and as a New Yorker, I think that swearing is just kind of part of my natural language and way of communicating. And I married a Midwestern girl and she doesn't tolerate it at all. She doesn't do it. She doesn't tolerate it. She won't tolerate it for me. and it made for we've been married for 30 years so I adapt quickly apparently but uh it you know for a long time it took a lot of restraint to like change that way of communicating particularly when I'm communicating about something that I'm excited about or emotional about or you know want to convey the importance of.

4:54It's a really interesting topic. And well, we're not going to turn the podcast into a podcast about swearing, but I imagine there's enough research there that we could if we wanted to. It's a fascinating area, yeah, because it gets right to the heart of the culture that we've constructed and how it relates to our usage and everything else about us. Yeah, it's fascinating that we have swears. When the old swears lose their power, We invent new ones. We pretend like nobody should use them. But as you say, people use them all the time. And it feels like an important part of being a language user that we've got them available to us.

5:30Yes. Endless string of questions. I'd love to hear your take on kind of a linguist in the age of modern AI, you know, transformers, statistical models. You know, this is a NLP used to be kind of coming from a linguistic perspective. and now the entire field is shifted to a statistical perspective. And I'd love to hear your reflections on being on the other side of that transition, as well as maybe more importantly, ways that you think that kind of the traditional foundational linguistics is still important to the way we think about AI today. These questions are on my mind all the time, yeah, because I operate at the intersection of all these different fields.

6:15And I will say it's useful to distinguish in this context linguistics. And people in my department at Stanford study language and social identity, historical linguistics, the structure of language. And they're just doing scientific investigation of language as a human phenomenon. And they are not technologists and they're not trying to inform technology. So their project is interestingly impacted by technological developments. For NLP people who are, of course, participating directly in the engineering project, they're affected in a very different way by the rise of Gen AI and the kind of homogeneous nature of the solutions that people now adopt in that space.

6:53So for the linguists, I feel like this is the most exciting moment that anyone could have dreamed of. I feel incredibly privileged to be alive in this moment where humans encounter for the very first time non-human creatures that use our language very fluently. I think it's weirding us all out, but from the point of view of understanding the human capacity for language, what a gift, because you can ask about the mechanisms, which are different from humans, but obviously sufficient for achieving a certain kind of behavioral performance. we can think about them as investigative tools i mean we train them on the internet they're basically incredibly powerful distributional learners and we can learn a lot from them about the true structure of language by just looking at the kinds of things that they learn and it really gets at the heart of core questions in linguistics about how much of language learning is innate and the nature of our capacity and whether it's statistical or symbolic all those things come flooding in in a completely fresh way.

7:57And so whatever your reaction to language models is, it should be a significant one, right? This should be causing you to rethink key questions. And that's all you could hope for as a scientist, that you have new angles, new perspectives, new questions reopened. That's been incredible. For NLP, I think it's a more uncertain prospect because pre the arrival of like pre-trained models, which for me would be like the ELMO model back in 2017, 2018. Before that, there was still a lot of statistical work, of course, and we were in the deep learning era. But you could still, for example, do a PhD that was entirely about some specific phenomenon and maybe some very specific tweak to a model.

8:43So you could say, I'm going to work on summarization, and I've got a new idea about how to do that well. using deep learning models. And that could be your PhD. And what we started to see 2018, 2019, 2020, especially with the arrival of GPT-3, that that was a very uncertain prospect because you might wake up one morning to find that you had been completely scooped, that with essentially no effort, one of these large pre-training runs had done better than you at the thing that you'd worked so hard on. And that caused an interesting, probably overall productive, but interesting and challenging crisis for people, especially students who are trying to figure out what to do next with their PhD research.

9:21But I think all of us felt a kind of real uncertainty in that moment. Yeah, I remember the anxiety of that time. And I always felt it was kind of expressed as, you know, is research in NLP fundamentally like scale limited or do you need a certain degree of scale that only a handful of organizations have to do foundational research? And is everyone else going to be relegated to like poking the beast and seeing what it does? And I'm curious, do you feel like that was an anxiety that's passed or is it still very present? Has it panned out quite like that? How do you, you know, how has it been resolved for you?

10:07Also fascinating, not resolved. It's something I discuss a lot with my collaborators and with my students. We're all trying to figure this out in this moment. I will say one concrete thing we did was orient a lot of our research toward interpretability, just the project of understanding how these models end up being so good at such hard tasks. And the reason we did that is it's relatively inexpensive, and it's also an area where clearly you would be explicitly hoping that models would get better because then there would be more to explain. Versus if you were doing that summarization project, you might quietly be hoping that there wasn't going to be so much progress so that you could make the progress.

10:47Let's hope the next model isn't good at summarization. I want to be the star of that show. That's, as I said, very uncertain. But if you're doing Mechintrip, you're like, let's get the new model released because now we're going to have even more structure to find, even more to explain. And that felt like a very productive choice. I don't want to leave out the fact that it's also cheaper to do this research, and that is significant. And then I would say that right now, a lot of us are in a moment of thinking we should do stuff that is weird and creative and out of the mainstream. We should be thinking about trying to achieve the next big thing because competing with these massively resourced, incredibly creative and talented teams is just not a winning game.

11:27So let's play a different game and hope that that's, as they say, where the puck is going, not where it is. And what are some examples of that kind of thinking? We've been thinking a lot about architectures because I have a lot of complaints about current architectures. And I would say the other main theme right now for us in my group is thinking about data. You know, data have strange and wondrous properties. I think we don't understand how data affect models. And that has all sorts of implications for security and safety and also the nature of the learning that these models do. It really, data is fundamental.

12:00It's all data-driven learning. and so telling the full causal story from data to final model state feels like it will just be significant for lots of questions. But I wouldn't want to leave out the architecture one because I feel like the architecture everyone has arrived at, these stacked transformers that we make very deep and very large, are tremendously inefficient. You would hope they were using all that depth and all that representational power to learn modular, recursive functions for things and all sorts of exciting stuff. It is not what we find, and that seems like a real opportunity to just level up and do better.

12:40And maybe we could get massively more capable models with half the depth and a quarter of the representational width. And that would be transformative for the economics of AI in addition to leading to all sorts of exciting things for capabilities. It's funny and maybe a bit validating for me to hear you say that because whenever I articulate a thought in that direction, particularly with folks that are coming from the frontier labs or essentially the frontier labs, I get back this kind of feeling that, yeah, you're just not bitter lesson-pilled enough. Like structures, that's old school thinking.

13:23You're just trying to train some features, just collect a lot of data, throw it at the model, and that's all you need. Okay, but here's my response to them. Let's say rewind to 2017. We've got the transformer. It's got absolute positional encodings, and it's got a particular structure for its MLP layer, which is pretty narrow and pretty dense, and a certain structure to its activations and its layer norms. That's 2017. The bitter lesson-pilled thing to do would be to scale that up. But just consider, for example, how much it would cost to use the n-squared attention and the absolute positional encodings but have a context window of 1 million.

14:06This is the bitter lesson-pilled thing, right? Just keep scaling. But it would be absurd. It would cost trillions of dollars to produce models that we all interact with right now. What did people do instead? They thought hard about locality and they thought about how positional encoding should be favoring local relationships. ships. They completely rethought the MLP so that it's now wide and sparse. Everyone did careful work on the activation functions to make sure there weren't weird outliers so that they could quantize in a good way. And so forth and so on. All of this analysis work built on intuitions about data and learning led to the model that we have now, which is like a ship of Theseus compared to the 2017 transformer.

14:44The only thing that survives is attention and the feed-forward layer. And I claim for you that none of that stuff is bitter lesson-pilled. That was all analysis work that was meant to save based on priors and the data and priors about how they knew learning would happen. So I go back at them. You're not bitter lesson-pilled enough, apparently. Although this is a reductio, I think. Oh, I love this. That's such a great response. I think it also really calls out the relationship between data, MEG and TURP, and efficiency, like core themes that you've been focused on and how they, you know, interrelate and support one another.

15:24Yeah, absolutely. And this relates to one of my hot takes. You know, it's very fashionable, especially among INTERP researchers. But I think in general for people to say, we don't understand how these models work, it is also very mysterious to us. But the truth is that people in the field have very deep intuitions about how these models work. And that is the causal factor in us making so much progress because they could think analytically, what would the structure of positional encodings and attention be so that I could do this at million context scale? You can only achieve that kind of thing based on deep analysis and insight, not by just guessing.

15:59And so when people say, oh, we don't understand, I say, I think you understand much better than you're letting on. I think you understand at least as well as my car mechanic understands how my car works. There are mysteries, but you can take a lot of action and be very effective in proving things. Why do you think they say that? Why do you think they say that they don't understand the models? There's got to be some payback there. It's probably a paradox of expertise, right? So the more you do know, the more you feel like there are also mysteries. And it's It's hard to step back from that and be objective and say, yeah, well, we did make a phenomenal amount of progress.

16:33And that can't be just because of happenstance. That was because we know a lot. But all you see as an expert is all the things that are still to be explained. Partly also, it's just a narrative in the field. And it does stretch back to days when I think we had very little understanding of how these models worked. And possibly because a lot of them weren't that good, there was very little to explain. And so that's just been slow to catch up with how much progress we have made in understanding the kind of intuitive human level mechanisms that these models are operating with. I also wanted to ask you about DSPY.

17:07I forgot about this as we were talking earlier, but you were involved in DSPY, which, well, I'll let you talk about it, but I'm curious how it connects into your research and like, you know, some of these pillars that we've talked about. Oh, there's lots of wonderful strands. And what a meta strand I could offer you, because we were talking about being strategic with research. This does stem from Omar Khatab, my student. He's the visionary behind DSPi and still its lead. And he just had the intuition early on that we should rethink what it means to make a scientific contribution. Previously, we thought in terms of papers as the beginning and the end of all of this kind of thing that you would contribute.

17:51We should instead, he said, think about projects and about empowering people. And so for him, the paper is one part of a broader contribution that might actually be centered on an open source or open weights release that would allow people to do big things. And that's where you find impact. And that's the nature of a contribution going forward. And DSPi is a kind of embodiment of that, although he made a similar investment with the Colbert retrieval model. And then people built on what he did. And then you really saw it take off where open source contributions made it easier and easier to use that technology, leading to more and more impact.

18:29And of course, DSPy is another wonderful example because in investing in this community and in the open source resource itself, he built a huge following. There are lots of startups, mine included, where the core tech stack for the LLMs is built on DSPy. And that has made life so much easier. And then, of course, it was a platform for him and for us to really think in an innovative way about prompt optimization and agentic workflows and all of those things. Yeah, I was thinking not too long ago the degree to which model strength as a correlate to model size, I suppose, and capability has kind of overcome the need for an explicit framework like DSPY, DSPY.

19:19Yeah, there are kind of two levels to that. The one would be just the engineering side where DSPy is great four years ago because it's kind of hard to construct the code around one of these systems in a way that's modular and reproducible and so forth. Because pecking out something where you've got a prompt string in the middle of your code with some slots in it, it's very error prone and it leads to bad system designs. And DSPy solved that. And you could think that the need for that is diminishing somewhat because now we all specify these systems in English and have the coding agents do them.

19:51And even before that, there were, you know, another hundred frameworks that solved that particular part of the puzzle. Oh, yeah, there's always competition. And I think at that level of just thinking about programming interfaces and APIs, they can all learn from each other. And so like, you know, DSPy learned a lot from PyTorch in terms of layer-wise design and the kind of modularity that introduced. And then, of course, you would hope that everyone kind of slurps up all these interesting innovations and it leads to everyone being better. There's lots of evidence of that at the level of interfaces.

20:21I would maintain for you that even if we have agents actually writing the code for these systems, it's great for us and for them if they write it in something that actually expresses these systems as modular components so that we can audit them, so that they can change them. It just feels like good engineering practices for any agent to think in a modular way. And that's what DSPy encodes. The other side is like the prompt optimization side and a belief people have that the need to be careful with your prompts is diminishing over time. I understand that narrative, but people should also, for example, just run like a simple annotation study where they use a few different models or the same model a few times on slightly different data.

21:01They will be blown away by the amount of variation that still exists. to be charitable. Let's say that these LLMs disagree about fundamental facts about how to label certain texts or what kind of response to give. We all kind of slip past this because we feel like, hey, they're smart and they're good and they're getting better. But if you quantify it, it's pretty disturbing. And the next step from that is to think about having all those agents optimize a prompt so that their behavior is at least consistent. And then you're right back at that DSPy vision. It's interesting that you say that because I don't feel like that necessarily aligns with my recent experience.

21:40And in particular, one thing that I've noticed that's been surprising is how well aligned, I guess, maybe that's not the right word, but how similar the responses I get to a query across different models. So for example, these are often kind of what I would call like a casual prompt, a casual query, something that I might type into Google and it will now generate an LLM response for me and it's kind of AI mode. And I'll take the same thing and put it into ChatGPT and maybe Claude. And it surprises me that the responses are often very, very similar, like, you know, very similar structure, very similar facts, very similar citations.

22:36And, you know, stepping back, like there are lots of ways that they could answer or approach these different questions, but it seems like, you know, the models or the training or the system prompts or something is all kind of converged on, something that makes the models express themselves, you know, very similarly, which, you know, seems to be at odds with, you know, the last thing you said about the need to optimize prompts or the impact of the individual prompt. I'm open-minded. But for example, like we just did it, we did, we did a thing recently, we were writing a grant and we needed a title and you want to be strategic with these titles.

23:12So we come up with a whole bunch of them ourselves and then we all disagree on what would be the best. So let's find out what the agents think. So ask a few anthropic models and a few GPT models, which of these five titles, which is the best? So you get a different answer from all of them, along with a detailed rationale about why obviously, of course, the choice that the model has made in that moment is the best one. This is great because then we can think about which one of these arguments is most persuasive. But if you were hoping for consistency at a subjective labeling task, which this is one, you can see right there that you're going to have a real problem unless you give very specific criteria and then you're kind of also constructing a prompt for them and you might want to manage them differently.

23:55There is a real, I don't have evidence for this yet, but we have an intuition at Big Spin in the research we've done that you get a kind of paradox that the more requirements you add, actually the more variation you'll see because the different models will key into different sub parts of the requirements. And since they do it very concertedly, you can actually get systematically biased behavior from something that you thought was a very good specification. And that again calls for this idea that what you need to do is figure out what the labels ought to look like and then have some automatic optimization process get the model there.

24:30And that's what things like JEPA and MeetBro were for. A topic that I really wanted to, a topic that I would really like to dig into with you based on our previous conversation was the idea of tokenomics. It's something that people are talking about a lot recently. I think, you know, folks that use cloud code, for example, have like a very visceral experience with anthropic changing the terms around usage. But it's happening under the covers with all of these large providers. And so I think way more now than, you know, six months ago, like we're all a little antsy with the relationship we have with these, you know, big model providers and the value that we get.

25:19You recently wrote an article about this. Talk a little bit about AI, how it ties into kind of your broader research, but also some of the things that you found when you started to dig into this area. Yeah, another interesting moment to be in as we're all being made aware of the true costs of all this AI usage. I saw a tweet from Ed Zitron, just a screenshot from someone who was noticing that Copilot was telling them that their bill last month was$500. And if they keep up the way they are with Copilot's new billing, it will be$11 ,000 in the next month. Wow, wow. Which is real sticker shock. And, you know, the analogy here is like it used to cost me$20 to take a ride share to the airport, Uber or Lyft, and now it costs$90.

26:12I use that analogy as well. But it's more like$20 to like$500 or something, right? Right. If only the slope will be as shallow as Uber, right? That's right. We start to wish for those easier stories. Yes. And so what will happen? I mean, I think what's happening is that the big providers are testing the waters on charging us the true costs, plus whatever profit they need to make as they all try to gear up for IPOs and so forth. And in turn, that is very quickly leading people to ask questions like, what is the return on investment for all these tokens that we have purchased? And it's a very tricky area to be in because what does it mean to think about value in this context?

26:54Even if we focus in on people who are doing just coding with coding agents, can we agree on what it means to add value? Maybe we have a few measures in mind like making a pull request or committed lines of code that last in the repo for a while or documentation touched or skill files created. But we might also worry that that's not capturing the value from any kinds of sessions we have, which are more open-ended and about discovery. So that's the first question is just solving this value issue, right? Let's just agree on what it would mean to add value for a coding agent. I like this line of inquiry because to me, it's the response to this thing that drives me crazy, which is, oh, you know, big tech company CEO.

Read the full transcript

27:47This year, 95 % of our code will be generated by, you know, AI. It's like, yeah, A, what does that really mean at that level? Like, what are the details beneath there? But is that a good thing or a bad thing? That's another dimension, right? Which is, is that code a liability or an asset? Right, right, right. And this idea of like, you articulate it as kind of code longevity in the code base. That's an interesting way to think about it. There's probably a lot of interesting ways to think about it that very few are thinking about right now. And all of these fall victim to the standard thing that once you make it a metric, it's no longer useful to you.

28:27Like if we said, oh, it's completion of projects, right? Well, then everyone would just have many projects that they completed, but they could all be liabilities and add very little value. So, but one framework we could offer that we did in the research you alluded to is, let's think about this like economists might. So we might have like a consumer price index. And the first step will be, what's the basket of goods that we're going to consider? You know, in that standard land, it would be like the price of eggs and the cost of rent and other kinds of tangible goods. what are engineering goods that we might track?

28:59But eggs might be a summary or a pull request, a bug fix or something like that. Or we could think broadly because we both use these coding agents. Requirement discovery, right? Knowledge accumulation. These are things that we don't currently track, of course, even as engineers, but might be behind our intuition that these coding agents are making us productive, even if it's not reflected in the PR accounts or whatever, right? I mean, in a sophisticated approach, you might say, I don't want more PRs because this is just a certain kind of busy work that doesn't relate to the actual goals I have.

29:34What are the actual goals? It's completing valuable projects and so forth. If I could do it with fewer PRs, but I had, you know, really robust code, I'd be possibly happy with that. So we got to figure out what the basket of goods is, but then we could start to track it relative to token usage. And that would be the consumer price index. So for any time period, we could just say I've got my tokens spent and I've got my goods produced. Tokens divided by goods produced is a pretty rough measure of the purchasing power of the tokens in those time periods. Then you would do the standard consumer price index thing of making what they call a hedonic adjustment.

30:10So you could just say maybe quality is improving over time. So you'd pick some measure for that and make an adjustment to the line. And when we did that study, we did code survival. So the number of lines of code that survives more than four days in the repository. We made an adjustment upward because that rate is going up. That's surprisingly short. Four days? Four-day survival? So again, all this is around measurement, and I'm happy to just be starting this discourse because we can see it's important to the economics of AI, and it seems like the work isn't being done at a high enough rate for us to get a clear picture.

30:44So we could make it longer, and maybe the adjustment would be different. I think currently for the data we have, which is this SWE chat benchmark, which was released by researchers at Stanford, it's about 6 ,000 real coding sessions, all the metadata, everything you'd want. What we see with Opus 4.6 usage in the time period we have, which is February to mid-April of this year, a decline in the purchasing power of tokens. That CPI is going down. And again, I just want to open the question, is it because we have the wrong basket of goods or is it because we're actually getting less value from these tokens?

31:21The one thing I can say that's kind of definitely a causal factor here is that in February of this year, most of the tokens went to producing code, which relates to the outcomes we just talked about. by mid-April, it was quite split between code generation thinking and also explanation to the user. And so that split now is going to have an effect on the things we're measuring, and that might be cause for reflection. There's value in those explanations that's not reflected in PRs, but might be reflected in something like knowledge discovery. And I see that coming up within the same time frame, it's become very common to now talk about the token efficiency of a new model that's been released with the implication being tokens of, you know, internal use tokens, thinking tokens versus, you know, per token of output, I guess is maybe a way to think about it.

32:18Another fascinating dimension. And this actually relates all the way back to the theme of efficiency for these architectures. So here's a claim I'll make for you. based on my read of the literature on inference time scaling, what's sometimes called test time scaling, which is just having the models generate lots of tokens at the moment that you ask them a question. So those scaling trends, everything we're seeing now is completely in line with those predictions, which is you get pretty good gains for a while with the more tokens you spend on a log scale. So this is jumping up quite a lot, but you do see it reflected in performance improvements, but it flattens out over time.

32:56And it's not like this curve skyrockets. It's sobering. You got to spend a lot of tokens for small gains in performance. We all knew this. We all knew this and we're just seeing it now play out. And when people talk about token efficiency and worry about this, I think what they're seeing is just the real lesson of what we already projected from inference time scaling. And this is independent of the approach to inference time scaling you're taking, whether it's multiple parallel inferences or some kind of Oracle or any number of other schemes? It's just fundamental to inference time scaling? That's a great question, right?

33:38I think we know that it's independent of some of those things, like the parallel work versus having it do lots of long chains. But some of the other factors you mentioned, I think we just don't know. And that's why I said it relates back to the question of efficiency for these architectures, if we made a fundamental change to how the models work, maybe these trade-offs would be very different. I mean, after all, so all of this stuff is a kind of patch job on the fact that there's no recursion in the depth. It's a fixed depth. And so the only recursion we can get, the unopen-ended notion of computation is by generation.

34:12But if we had models that could be recursive, maybe fewer tokens for larger gains, I think we don't know. So, yeah, I mean, in the end, we're going to spend the cost on compute or tokens. So this might not affect our bills in the end, but it is a fascinating question. What are the true scaling laws and what's possible in this space? And you're right to push back. We talk about these things like they were like platonic ideals of laws. Scaling law invokes that, right? But even for the scaling laws for pre-training, you know, there's lots to discover there. And many of the stories of progress are actually like transcending the scaling law.

34:49And we see like better improvements than those laws predicted because everyone worked so hard behind the scenes to do very innovative things, which maybe relates to our bitter lesson discussion. Any particular example come to mind of that? Data usage and the nature of the data really matters and overtraining the models really matters, which is kind of pushing up against the standard scaling law presentation. And now I'm just going to speculate. I should check on this, but things like mixture of experts might have really flipped the script on what it means to count parameters and in turn how these laws relate.

35:18And then I think maybe even also stuff like the context window and so forth. This is another thing to check, but I just speculate that we've seen larger gains from pre-training than you would have predicted by those early scaling laws papers, suggesting that there is some innovative thing that was happening on top of pure scaling? Thinking about the concept of a market basket, one kind of pushback that comes up for me is in the real economy, eggs is different than milk, is different than beef, et cetera, et cetera. And they're all influenced by different factors. uh you know production for example whereas what you've done with the this kind of cpi basket with tokens is kind of like more like analogies like here's the typical bundle of work and you know what it requires from a consumptive perspective but the tokens aren't fundamentally different like they're the same tokens it's just like how much it takes to do this versus how much it takes to do that versus how much it takes to do that.

36:30Tell me what I'm missing there and what does kind of characterizing these products give you in your analysis? Yeah, fascinating to think about. One thing I could insert there is the tokens are different at the level of being used for code generation or skill file writing or explanation or thinking, right? Those are different kinds of tokens that probably do feel tangibly different to us. So is that an element in your thinking? I think I was thinking from our conversation that you had 10 different, almost like tasks, like 10 different types of tasks from the domain of code generation, which, you know, if they were all kind of largely code generation, you know, that is the part that had some dissonance for me.

37:17But if you're talking about, like if your basket is like creative writing versus, you know, a few code generation things that are kind of different versus, you know, summarization versus editorial commenting feedback, those, you know, may be more fundamental. Yes. So I think this is very significant. when we have done some research on this as well at the level of what kinds of session types exist and in turn, what kinds of users are there. So you might notice of your own behavior, I guess this is reflected in your comment, that sometimes you want a quick check-in on a question. Sometimes you want a quick PR to get fired off.

37:57Sometimes you want to be in a mode of deep collaboration. Sometimes you're partnering with the AI. Sometimes you're delegating the work and so forth and so on. And the outcome measures that we choose should be sensitive to this. We shouldn't penalize the agent if your chat interaction with it about some scientific question didn't lead to a PR. It was never on the table in the first place. Whereas if you're trying to get some work delegated that's actually a coding task and all it does is chat with you, that would feel quite unproductive. So we need to bring that in and that would be a higher level discovery process of what people are trying to do and so forth.

38:33In thinking about the notion of value, is this something that you're anticipating? Like, it strikes me that that's an entire, you know, research thread that, you know, one could go into. I don't know if that's a linguistics or a linguist or an economist or a computer scientist, you know, probably interdisciplinary, like most interesting questions. but is that something that you're working on or was it something that you put out there for someone to take up and run with? I am not sure. I can tell you the lineage of this idea is that we founded this startup Big Spin because we would like to see more people benefit from AI.

39:19Whether you love it or hate it, it's here and I would like the benefits to be more evenly distributed And I can tell that that will mean bringing on board many more people than currently benefit from AI. Right now, I would say that it's mostly experts deriving real value. And a lot of the world is currently even trying to figure out what this is all about as a tool or an entity in their lives. So we would like to have more access and more productivity. And that implies making the user experiences much better. Figuring out what interactional patterns lead to success for people, meeting them where they are in a kind of adaptive way, the whole list of things that you might worry about if you were a product manager who had some deployed AI product.

40:04And I think by that route and from that perspective, we just ended up worrying about our own token usage increasing and wondering whether there's real value there. And it just happened to collide, actually, just like three weeks ago, with this emerging narrative on the back, I think, of all these rumors about IPOs, about what the return on investment was. And then all these CEOs came out and said, oh, our spend was enormous and we want to scale back and we're walking back our claims from a few months ago. And that is just a fascinating thing to witness in the narrative here. I'm wondering, does this research also attempt to project forward?

40:41In theory, you could create a model for Anthropics cost and spend based on publicly available data and some presumptions and give us a sense for how close we are to paying full freight for our tokens versus if we're only paying 10 % for our tokens, you could then project what that cost might look like. over time as we're paying more and more of the full cost? Yeah, I don't have, again, fascinating questions. I don't have resolving answers. I am glad I am not tasked in some organization with projecting spend on all of this stuff because I think it would be basically impossible. For the time period that I was describing for our little CPI experiment, Anthropic changed the default reasoning on the model at least two times.

41:34So we see like they launched it with default reasoning high. we have a mysterious sudden rise in the token usage which we cannot explain and there's a new baseline they lowered it to medium as the default they patched a bunch of bugs that were related to context management and then turned it back up to high and all of these things have an effect on the total token output as you can imagine they also changed the default context window which meant people could swallow up much more stuff at any given moment so imagine trying to predict what token spend is going to be like when you have all these exogenous events in addition to changes that we don't even know about and questions about where the value actually lies.

42:15Very difficult. And then, you know, the true cost of a token, the estimates vary wildly for every dollar we spend. It could be as low as two and as high as 20. And I think this is just because it's hard to factor in things like R &D and future build out and depreciation and all of that stuff. I think at the current moment, we just don't know, but there couldn't be a more significant question for the global economy, basically, than where the value is and who's going to pay and how much. Yeah. In your article, you coined the term tokenflation to describe at least the recent behavior of token economics.

42:54I imagine you see that continuing. Seems to be continuing. Yeah. Yeah. That's certainly the picture that we get from the CPI, a picture of tokenflation. Yes, your token is not buying you what it once did according to everything we can think to measure here. And even adjusting for models getting better, right? That's critical there. Because if it was just a story of models thinking more and being more robust, and we were all getting exponentially better outcomes from this, then the spend would look completely rational. But that's not the picture that we see. And so we have to do some hard thinking about what's going to happen and how to improve the situation.

43:28Let's dig into that a little bit more. How would you articulate what you're seeing, the models are getting quote unquote better. There's a set of open questions about are the reported ways that models are better actually reflective of some intrinsic betterness? And that question is often about benchmarking and learning the benchmarks, overfitting, that kind of thing. And then there's the kind of question of chattiness and the volume of thought that it requires a given model generation to produce an answer. What are other factors that you see? Yeah, we could pick that apart as well. And this relates to a line I've had consistently, which is that we should think in terms of systems, not in terms of models.

44:26So in the data that we've got, Opus and Sonnet 4.5 versus 4.6, those two generation changes, those are real model changes, I assume. I think they did something very substantive at the level of the weights. And everybody immediately saw that that led to like a 5x increase in token usage. And this was related to the introduction of adaptive thinking. Now, fix that. That's the level shift that we already took, and maybe we're seeing improvements there that are worthwhile. It gets hard to say, but let's assume there was a level up in improvement. Then for the period that we did our CPI experiment for, that's a fixed model, Opus 4.6.

45:07So all the code improvements that we saw in the data relate to the product. This has to relate to things like them turning the knobs on the adaptive thinking, changing things about the system prompt, changing things at the level of the product. and that's where the improvements were. And so that shows you that even for a fixed model, we can get very different outcomes for these things because they really are sophisticated engineered systems at this point. Yeah. And so was the product in this case specifically CloudCode or? Oh yeah. And so we don't, there's tons of stuff there. Yes, I believe we know that these are all CloudCode sessions that we kept in our data.

45:43SweetChat is broader than that and involves a couple of other coding agents. But I think I can say that all our data our Claude Code sessions using Opus 4.6. Have you seen any evidence that changes via API usage experience, similarly dramatic variation in performance? Oh, fascinating. To kind of control for a lot of that product level stuff, all the prompts that are hidden from us, all of those affordances. I don't know, but that's a nice thing to think about because it gives us more things that we can control for and more things that are knowable. So kind of in parallel to the model evolution, there's also evolution of the user.

46:28You've alluded to this a little bit about kind of your concept is that most AI users now are experts. Talk a little bit about the role of expertise. I think this is also kind of echoing back to our conversation about DSPy and like prompt optimization. You know, you've done some research into how folks are using these models and the role of, you know, AI fluency. Tell us about that research. Oh, yeah. First, I should say, so the distribution of users across expertise levels. So I guess the nuanced picture I'd offer is that the people deriving a lot of value from AI in the current moment tend to be experts.

47:09It must be the case that most users of AI are beginners, just because the numbers are so large and expertise can't be that widely distributed yet. And that's a very interesting thing because I think probably most things are getting designed for those experts implicitly or explicitly. But for the whole economic picture to work out, many more people need to derive value from this via one avenue or another. And so that does shine a light on this expertise thing as a real factor. And the headline result there actually builds on something that Anthropic did. They have this AI fluency index and their core observation in that work is that experts display an augmentative style.

47:50They iterate with the AI, they push back, they complain, they change their requirements. It's really a collaborative mode, whereas novices, low-fluency users, delegate. So they trust in the AI, they let it do its thing, they accept the responses uncritically. and our contribution is to show that this is a causal factor in success with these products right now. Experts can do harder things more reliably as a result of all that friction they introduce, all that pushback. Whereas novice users, they accept, but they end up accepting the wrong thing and they're not able to level up from the basic tasks that they think to start with.

48:34That's obviously significant. And it feels so tantalizing because pushing back is a natural human behavior. I feel like we could encourage everyone in the world to do this. We probably need to get them out of the mode of thinking, it's a super intelligence. You should just trust it. That has been the narrative for a while. What we're seeing in the current moment and possibly for the foreseeable future is that you got to complain, collaborate, introduce yourself, push back, all that stuff that I think we do. We take that for granted, right? Yeah. Yeah. And so from a methodology perspective, how did you approach exploring this?

49:10We built on the work that Anthropic did, which they set up a nice framework with some independent research who were doing this kind of usability stuff. And we just have an annotation protocol. We can talk in detail if you want about this, but at BigSpin, we have lots of these best practices around having language models essentially collaborate on annotation projects to kind of triangulate on the truth and factor out their individual biases. So we do that stuff and we apply all these fluency markers. And then separately, we do a thing of estimating task complexity and looking for signs of visible and invisible failures.

49:47And so it's the connection between the fluency markers and the task complexity success metrics that was our contribution there. And that's where you can see high fluency users are the ones doing harder tasks. Paradoxically, there's more signs of failure for them because they complain, they push back, they're trying harder things. but as part of all that friction they're successful with harder things as well. And if you were to try to apply this insight from the perspective of someone in an organization that's trying to help or guide their organization to be more successful with AI what do you think are the key lessons of this fluency work?

50:28If it's an org that's just starting out and wants people to figure out how this could be part of the organization's mission it would just be that pushback message. And you could do an experiment where you interact with it about something where you're a world expert. We're all an expert in something. Engage in a discourse with one of the best models about something you're an expert in and see how often you feel you have to push back. And this could be a kind of lesson. You say, aha, for other spheres where I don't know the answer, it might be just as errorful. That could be a good visceral thing.

50:56If the org is very far along, I think the main thing to do right now is to have a team of these LLMs interacting to improve things. For example, at Big Spin, I didn't set this up. Our founding engineer is very future forward on agents, and he's incredible at this. And when we do PRs now, the first round of review is the agents all interacting, collaborating, disagreeing. They do the first round of comments. They do the first round of code updates. Only after they've resolved things do we look at a PR. So the final human stage should be very high value. And the agents did all that work. But when you have one agent do it, they often just reinforce themselves and you don't get good outcomes.

51:37It's that team of rivals thing that is transformative. You know, I think it's interesting because, you know, on the one hand, like, of course, that makes sense. But on the other hand, there's something, you know, it also implies that you shouldn't be using these things in areas where you don't have enough expertise to evaluate the answer. Yet, that's where you most need the assistance, the support. um so and again and this is a little bit worrisome about the overall narrative around ai the place where we can get around this is with software development because let's say that i'm trying to accomplish something in a language that i don't know how to code in i can have the agent do work for me because probably in the end i can run the program and look at the results and that's what mattered to me is that i run the results and i see and if i don't see what i wanted I can complain and we can iterate.

52:36That verification step that doesn't imply I have comprehensive knowledge, it just implies that I know what I want to see in the end, is so critical. And I think this is a causal factor in models being so good at coding because it's like the ultimate verifiable domain for them. But as soon as we leave that and go even into something like the legal realm where the requirements are strict, but they're not codified in code and they have ambiguity about them, this whole picture falls apart. And you then are back at what you just said, which is this awful kind of paradox is like, yeah, use AI, but in the end, unless you're expert enough to evaluate every single one of its responses, you might be in real trouble.

53:16I don't know how to get out of this because the verification step is like we go to trial, but this is very consequential. That's expensive. Yeah, that's funny. I mean, it does make me think a little bit about, you know, some of the types of errors that we're trying to avoid are factuality. And, you know, there is a temptation to say, well, let's just throw more tokens at it. Like I'll have a critic model that evaluates everything that the, you know, know, is generated by the primary model. But then you go back to my observation that these models tend to correlate in their responses as well. Yeah, it's super interesting.

54:01That's a good point. Yeah, for my picture, we want real diversity of perspectives. This is just like, you know, red teaming for humans. This is most successful when you have a really diverse team of people who think creatively and differently. And if every one of the members of that team is thinking in a homogeneous way, they miss all of the crucial things. Same exact issue. If all the code review agents are biased in the same way, they will miss exactly the same class of bugs and then we're all sunk. Yeah, I don't know how you'd encourage this diversity in the ecosystem. We're probably, as you say, converging towards some kind of one model.

54:33But I think for my picture, we need diversity. Yeah, we got to keep those open weights models going or something because they're the weird players in the space. For sure, for sure. So we've talked about efficiency, interpretability, tokenomics, fluency. You're involved in a lot of different research directions. Excellent. Excellent. What's next for you? Where do you see either where do you see this all going kind of externally, but also like where is your research going? Yeah, this is great. And as I said before, we're trying to think in weird and creative ways about what the future could hold.

55:13And I encourage my students to do this and they're smart. So they say, all right, Chris, I'll think along those lines, but what's your answer to this question? So I do have an answer. And it's really shooting for the moon here, which would be, what about the architectural innovation that would upend the whole story around the stack transformer and the way we need to do data center build out to even get incremental gains in performance that could be upended. And it would come from some very innovative thing around maybe recursive use of the building blocks that we've got. So architectures, we should think.

55:45And when people say, oh, no, we don't need more architectures. The transformer is good enough. That's where we should push back as academics doing something more clever and more scrappy that could change the world. And the other one is thinking in the interp space much more about data. And that's just because I want to tell the true story of how we go from data to model capabilities. but it also checks a box for me on connecting interpretability to safety. It has been hard for me to connect those two things. We have found some ways to do it, but it's not a slam dunk as a narrative, even though it's the dominant narrative.

56:20But I will say that when we get into things like data poisoning from innocuous examples, this is probably a growing societal concern. There is evidence that with very few examples planted in a pre-training data set, you can have a significant influence on the outlook and preferences and quirks of the final model. So can we detect those examples? What's the nature of those attacks? How well hidden could they be? What's the smallest number of examples? And why does it happen? These are all going to be very pressing questions. And so, again, it's just a data-oriented question that's very alive for me in the current moment.

56:57On the architecture front, is there research that you're seeing or doing that is as yet under the radar that you think is promising and or underappreciated? I think you had my student Julie Kalini on, and she is an advocate for byte level models, essentially tokenizer free models. I think that's a big part of the future. It's a critical thing if you want to have truly multilingual models that are also equitable in terms of how many tokens they charge us for, getting back to that earlier theme. But also Julie's perspective is that this is speculative, but I think there's something to this, that it's a kind of inference time scaling because you do more compute at test time because you have more tokens.

57:44and therefore more opportunities to build on interesting things. So that could be a big part of the future. And the other one would be recursive architectures, as I said. But if you want to go all the way out, you could think, why do we always assume we're going to do gradient-based learning? There are lots of alternatives to that, and nobody is exploring them because everyone takes it as a truism. We're all in our very narrow row here without even really realizing it. Who knows what's outside in this garden? And it's very risky as a research bet because only one in a thousand of these ideas will pay off.

58:20But what's the point of being an academic researcher if you're not going to take that kind of risk? That's what we're positioned to do. Well, Chris, thanks so much for jumping on and sharing a bit about what you're working on. It's very cool stuff. Thank you. What a wonderful conversation. It gave me lots of new things to think about. Awesome. Awesome. Thanks so much.

58:42Thank you.

From the publisher

As reasoning models consume more tokens and AI systems become more expensive to run, understanding what those tokens actually buy is becoming increasingly important. In this episode, Stanford professor and Big Spin co-founder Chris Potts joins us to discuss AI tokenomics and his research into “tokenflation”—the possibility that token usage is growing faster than the measurable value those tokens produce.

We explore how to measure the return on AI spending, why benchmarks alone provide an incomplete picture of model progress, and what inference-time scaling means for the economics of increasingly capable models. Chris also explains why expert AI users tend to get better results by challenging and iterating with models, how AI fluency affects outcomes, and why more efficient architectures could change the underlying economics. We also discuss DSPy, interpretability, the limits of today’s transformer architectures, and where Chris sees opportunities for more fundamental innovation in AI.

🗒️ Full show notes: ⁠⁠https://twimlai.com/go/776.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
Do AI Tokenomics Matter More Than Model Benchmarks? with Chris Potts - #776The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 59 min
Listen in VO