Investing on the Front Lines of the AI Arms Race | Nathan Benaich

10 Nov 2025 · 54 min · 19 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Nathan Benaich (Airstream Capital; State of AI Report) discusses recent AI breakthroughs, shifting scaling narratives (from more compute to “inference-time scaling”), reasoning models (chain-of-thought), and the AI arms race’s commercial/geopolitical implications—especially China’s push for open-weight models—plus where value accrues across the AI stack.

Guest background

Early-stage VC focused on AI-first companies; studied biology/bioinformatics (cancer research, microRNAs) before VC (since 2013). Creator of the open-access, 300+ page annual State of AI Report (8th edition; widely reviewed/validated by major research and companies).

Key claims

Scaling still matters, but “thinking more” (reinforcement learning at inference time) is increasingly important. DeepSeek’s low headline training cost is misleading because it omits broader R&D, data, infrastructure, and compute. User prompting/context strongly affects output quality; longer chats can degrade performance.

Notable examples

DeepSeek (R1, V3; reinforcement learning); OpenAI O1 Preview; “think step-by-step” prompting; voice cloning services (e.g., ElevenLabs); transformer attention; export controls, energy constraints, regulatory red tape.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The State of AI Report Overview

0:45 to 2:34

Nathan shares the origins and significance of the State of AI Report.

“You can also do that on our subscriber page.”

AI as a Force Multiplier

2:34 to 5:48

Discussion on AI's role as a force multiplier for technological progress.

“please enjoy this deeply informative and timely conversation with my guest, Nathan Banesh.”

Understanding AI Terminology

5:48 to 9:00

Clarification of key AI terms, including AI, AGI, and superintelligence.

“This is the most recent report that came out a few weeks ago.”

Transformers and AI Development

9:00 to 12:40

Deep dive into how transformer architecture has changed AI modeling.

“But people are at this point, if they interface with this subject area remotely, they will have heard terms like transformer architectures, neural nets, large language models.”

Current AI Breakthroughs and Expectations

12:40 to 14:00

Nathan discusses the latest AI breakthroughs and the community's perception.

“So we're going to get to this later because I have some deeper philosophical questions, one of which is, how do these models determine what is interesting in a data set?”

Understanding Exponential Change in AI

14:00 to 16:44

Explore the challenges of perceiving exponential growth in AI advancements.

“None of these tasks, if you would have asked me or others in the community 10 years ago, would this be possible?”

The DeepSeek Moment and Its Implications

16:44 to 20:00

Discuss the significance of the DeepSeek moment and its impact on AI scaling narratives.

“For me, the deep seek moment, which was run in January this year, was really an example of how there's information asymmetry in the market.”

Recent Breakthroughs in AI Technology

20:00 to 22:39

Evaluate major AI breakthroughs and their implications for future development.

“Just to clarify for people that may need it, when we're referring to DeepSeek, so DeepSeek itself is a quantitative hedge fund, correct?”

The Transformative Power of AI

22:39 to 25:41

Analyze why AI is viewed as the most transformative technology compared to the internet.

“And then it's a bit of corporate priorities of there's one or two edge cases worth solving given other focus areas.”

Perception and Reality in AI Systems

25:41 to 28:00

Delve into the challenges of AI's perception of reality and its implications.

“Because digital information is a representation of the analog world.”
Show all 19 chapters

Navigating AI's High Dimensional Space

28:00 to 30:18

Learn how AI processes information and the importance of effective prompting.

“So I think a way to think about it is that the AI has a crazy amount of information And that information is like stored in an extremely like complicated high dimensional space.”

The Challenges of Context in Conversations

30:18 to 32:55

Understand how context affects AI responses during conversations.

“So in the meantime, it's important to give it as much context as you can.”

Innovations in AI Reasoning

32:55 to 36:43

Explore recent advancements in AI reasoning and performance improvements.

“So let's talk a little bit about some of these innovations that you mentioned earlier, like inference time scaling and chain of thought.”

Experiences with AI System Changes

36:43 to 40:36

Discuss personal experiences with changes in AI performance over time.

“And so would that explain why, for example, I've actually seen ChatGPT get worse at having a conversation with me about my interviews if I give it my transcript, for example?”

Understanding AI Compute and Transparency

40:36 to 42:00

Learn about compute allocation in AI and the importance of transparency.

“in the ChatGPT app, you might get a slightly different experience.”

Understanding Inference Time and Compute

42:00 to 44:41

Explore the concepts of inference time compute and its implications for AI systems.

“and this thing's going to be flying in the air and I'm going to end up in Tokyo in no time.”

The Evolution of AI Scaling Narratives

44:41 to 48:26

Discuss the transition from scaling through data to focusing on inference time for better AI solutions.

“these systems by throwing more compute and data at them.”

Expertise and Progress in AI Development

48:26 to 50:26

Analyze how the influx of talent from various fields is driving AI advancements and innovations.

“Whether it'll be across every single task is another question.”

Geopolitics and the Future of Open Models

50:26 to 52:44

Examine the implications of open AI models in the context of US-China competition and investment opportunities.

“So Nathan, I'm going to move us to the second hour.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Demetri Kofinas:What's up, everybody? My name is Demetri Kofinas, and you're listening to Hidden Forces, a podcast that inspires investors, entrepreneurs, and everyday citizens to challenge consensus narratives and learn how to think critically about the systems of power shaping our world. My guest in this episode of Hidden Forces is Nathan Banesh, founder and general partner of Airstreet Capital and the creator of the annual State of AI Report, an open access compendium that tracks advances across AI research, industry, policy, and geopolitics. Nathan and I spend the first hour of our conversation today exploring some of the most important AI breakthroughs of the year.

0:41Demetri Kofinas:We unpack the deep seek moment, dig into some of the advancements made by the latest reasoning models, and why there appears to be a regression in capabilities across certain domains in artificial intelligence at the same time as we are seeing marked improvements in reasoning heavy use cases like coding and scientific research. The second hour turns to a conversation about the commercial implications and geopolitical dynamics of the AI arms race, including China's strategy to become the leader in open-weight models and tooling, the industries, sectors, and professions most ripe for disruption, where the investment opportunities are, whether we're in a bubble comparable to the 1990s internet boom, and how export controls, energy constraints, and regulatory red tape could play an outsized role in shaping the trajectory of the current arms race.

1:32Demetri Kofinas:Lastly, we look at where along the AI stack most of the value is likely to accrue, from the underlying picks and shovels, through the foundation models, to the apps that ride on top of them, and what all this means for labor markets education, and the cadence of scientific discovery. If you want access to all of this conversation, go to hiddenforces.io slash subscribe and join our premium feed, which you can listen to on your mobile device using your favorite podcast app, just like you're listening to this episode right now. If you want to join in on the conversation and become a member of the Hidden Forces Genius Community, which includes Q &A calls with guests, discounted access to third-party research and analysis and in-person events like our intimate dinners and weekend retreats.

2:20Demetri Kofinas:You can also do that on our subscriber page. If you still have questions, feel free to send an email to info at hiddenforces.io and I or someone from our team will get right back to you. And with that, please enjoy this deeply informative and timely conversation with my guest, Nathan Banesh.

2:45Demetri Kofinas:Nathan Benesh, welcome to Hidden Forces.

2:49Nathan Benaich:Thanks for having me.

2:50Demetri Kofinas:It's my pleasure to have you on, Nathan. I'm actually very excited to have you on the podcast. One of our genius members, Thanos Papadopoulos, introduced us. He's a VC based out of Athens who we also need to get on the podcast at some point. So before we get into the substance of your latest annual report, which is, I should mention, over 300 pages long, so there's a lot to get to, tell me a little bit about you. Who are you and what is the origin story of this report? How long have you guys been publishing it? And who else has been involved with researching and writing it?

3:21Nathan Benaich:Yeah. So I'm an early stage venture capital investor. I founded a fund called Airstream Capital, which invests early stage in AI first companies across lots of different sectors. And I've been working interested in AI for pretty much my whole career. I went to undergrad in the US studying biology and bioinformatics, got to work on cancer research and microRNAs, and then did grad school in the UK, similar topics. It was really through that experience that I got to see just how much value you can derive if you have computers that can understand complex data sets. And I started working in VC in 2013 and have been doing that ever since.

3:58Nathan Benaich:And AI has been the sort of theme that I've been most interested in since that time. And when it comes to the state of AI report, I've just generally been a big believer of like writing and sharing ideas. I think there's so many fascinating people out there that you just never get the chance to meet unless you have sort of nuggets of shared interest that, you know, they can consume and find on the internet or other places. And, you know, working in a small firm, that was a big way that I could accelerate my learning. And so in 2016, 17 or so, I was already a fan of the Mary Meeker internet trends report, which was like a fantastic kind of compendium of progress and just internet more generally, and felt that a product thematically similar to that in AI would be really valuable because there was more and more interest in the early waves of deep learning in 15, 16, 17.

4:50Nathan Benaich:But there wasn't really like a canonical document available online for free that was like not marketing, but actually deep in the research and also deep in policy and industry and all the other facets that you might have to kind of know about if you were working in AI, but might not have the time to look at on a day-to-day basis. And so this Data to the AI report is like my effort to create this open access document that sort of captures the zeitgeist over a 12-month period and dives into all these different facets of the technology so that people working in each one of those facets can appreciate what everybody else is doing.

5:23Nathan Benaich:And yeah, it's become kind of like the most trusted, I think, source in AI and got lots of interesting contributors to the report from major research major companies, a lot of reviewers from the same organizations to make sure that the opinions are valid and correct. And yeah, every year I work with a couple of researchers to help produce it, and it's been eight years.

5:47Demetri Kofinas:Exactly. So this is the eighth report. This is the most recent report that came out a few weeks ago. The 2025 report is your eighth report. And yes, it's a very successful report, very widely read. So for people that don't know, in the report, you write that AI acts as a force multiplier for technological progress in our increasingly digital data-driven world. And then you go on to say that this is because everything around us from culture to consumer products is ultimately a product of intelligence. I'd like to take as much time as you think we need to expound on this argument. And let's begin, I think, with terminology.

6:22Demetri Kofinas:So the term AI itself, forget AGI, which I do want to understand the distinction between AGI and superintelligence, because superintelligence now is being increasingly used as a term in place of artificial general intelligence. But even the term AI is very fuzzy. It's very fuzzy in terms of what we should expect for its capabilities. It's fuzzy in terms of what it is that we mean and how we assess whether something is artificially intelligent or not artificially intelligent. So how would you answer that question? First of all, on a terminology level, what is AI?

6:53Nathan Benaich:I think in simplest form, AI is some kind of machine program, so not a biological system that can exhibit capabilities that one would normally associate with an intelligent biological system. And so probably the most obvious application or instantiation of that is an AI system can look at prior data and learn statistics and make predictions about the future using that prior data. And we've been basically on a journey of bringing in more and more sophisticated data, more and more sophisticated predictions to amalgamate an overall machine system that can kind of climb up the human capabilities intelligence curve using those very same principles over the years.

7:38Demetri Kofinas:And what about the distinction between AGI and super intelligence? And why am I hearing people use the term super intelligence increasingly now in place of AGI?

7:48Nathan Benaich:So AGI would be different to what was previously called artificial narrow intelligence, ANI, which people don't seem to use very much anymore, in that general system can do many tasks and a narrow system can do one task. And so that might be a narrow system can predict sentiment on Amazon reviews and a general system can do that, but then it can also generate pictures of cats or something. And then the superintelligence piece, which has really ramped up in the last probably six months or so, is driven by researchers recognizing that in their pursuit of AGI, there are already certain sort of characteristics or tasks that models can already kind of outperform humans on.

8:30Nathan Benaich:And you could argue that as a user of ChagyBT, the system can do a lot of things that we're not capable of doing just because we don't have that expansive knowledge about every single subject in the universe. And so in some sense, within one's own capabilities, the systems we have today are super intelligent. But so that's really the arc. It's one task to many tasks, generally as good as human beings, and super intelligence is better than human beings.

8:58Demetri Kofinas:So just to wrap up this discussion about terminology, I'm sure we'll have more terms thrown out. But people are at this point, if they interface with this subject area remotely, they will have heard terms like transformer architectures, neural nets, large language models. Without necessarily getting into an especially nerdy conversation about the core architecture of these systems, what is it important for people to understand about how they work?

9:25Nathan Benaich:I think the most important thing is that in pre-transformer AI, the community was very focused on trying to hand engineer or sort of bring our prior knowledge of how a specific problem might be solved and encode that in the network. So for example, what was very popular in computer vision pre-transformers was called convolutional neural networks. And those networks basically exploit the fact that in an image, a pixel that you might pick is very similar to a pixel really adjacent to it, just based on the statistics of how things look in the world. And so on that basis, you could design a neural network that would exploit that local similarity and basically have patches.

10:11Nathan Benaich:so it would operate on little squares across the entire image and process the image that way.

10:18Demetri Kofinas:So let me just understand something correctly here. What's operating here is the same principle that allows us to compress information. We find redundancies in data and we eliminate them knowing that we can recreate them later because we understand what the overall statistical structure looks like. That's the prior knowledge you're referring to.

10:33Nathan Benaich:Pretty much, yeah. So there was some pretty cool work back then showing how every layer of this multi-layered system, hence deep learning, because it has lots of layers, so it's deep, you would see these like features, these kind of visual representations of what the neural network was learning. And it was going from like, oh, it's learning edges to several edges put together look like a nose and then like a nose with other edges of a face, and then it goes into a human being. And so that's all exploiting like the local similarity of pixels, because it's just like how images look. So that's what nerds would call a prior because it's just an established way that data looks like.

11:12Nathan Benaich:And so the big change is instead of going for these priors, which is encoding what we already know about a system, or instead learning everything from scratch with a more general purpose way of treating input data, and that's the transformer, which basically processes input data in the form of a sequence. And so in that image example, instead of treating the input sequence as patches, you would treat each pixel as an input token in this case in a sequence. So you could imagine like starting from the top left of the image, and then you have one pixel and another pixel, and you kind of go to the right, and then you skip the line, you go back to the left, you're sort of like reading a text.

11:52Nathan Benaich:But instead of reading words, you're reading pixels. So you feed that sequence of pixels into a transformer. And then the core thing that a transformer does is it basically figures out where in history, where in the input sequence is important to focus on. It's what in NerdSpeak is called attention. Where should I attend in order to best predict what's going to come out in the future? And so just in the same way as human language has patterns, which can be learned at a large scale on the internet, images also have patterns. So it turns out we'll kind of learn similar statistics that we would otherwise be exploiting naturally because we treated images as squares.

12:34Nathan Benaich:But we can learn that just through the statistics embedded in sequences. And this sequence architecture is applicable to a ton of different tasks as opposed to just one. Okay.

12:43Demetri Kofinas:So we're going to get to this later because I have some deeper philosophical questions, one of which is, how do these models determine what is interesting in a data set? But before we even get there, one last question about terminology. You mentioned tokens. Is it fair to compare tokens in some sense as an analogy to bytes of information? Are they kind of like the smallest subdivision of information in this context?

13:06Nathan Benaich:Pretty much, yeah. All right.

13:08Demetri Kofinas:So let's get into the report. And as I said, we'll get into, I have other categories. Again, we have a limited amount of time, but there's a geopolitical section I have here in my rundown. There are, of course, sections about economic applications and stuff like this, but let's get into the report itself. So, you've been doing this, as we said, for eight years. What feels qualitatively new today versus the, let's say, transformative breakthroughs of the 2018, 2020 period or the chat GPT moment of 2022?

13:38Nathan Benaich:For me, the biggest point in all this is magic. just the fact that we have a software system that can produce lifelike images, produce increasing lifelike video, clone your voice, generate voice, translate anything, consume a ton of PDF pages and then reason about what's in it. None of these tasks, if you would have asked me or others in the community 10 years ago, would this be possible? that say like, yeah, but in a long, long time. And if you showed them what we have today, go back 10 years, they'd say it's magic. So I think it's really one of those examples from that wait, but why post about how humans perceive the exponential.

14:21Nathan Benaich:And when you're on this like kind of curve that's growing, I don't know, 5 % compounded every year, and then there's an exponential that's gonna happen. It's right in front of you and you can't perceive it sort of on the precipice. That's I think what it feels like in any sort of month increment. but when you zoom out, you really realize this is magic.

14:38Demetri Kofinas:So is that what explains the mismatch between your expectations and reality, largely the inability to think non-linearly about change? Or was it actually also something fundamental that you guys missed as a community about how these systems would scale?

14:57Nathan Benaich:Yeah, it's a good question because there was early evidence that scaling could work many years ago. Probably the most obvious one was back in 2005, 2006 or so. And then again, like a couple of years thereafter, this was when speech recognition was starting to be applied to GPUs, high performance computing at a lab run by Baidu. And they were showing that you could take the same model and give it more GPUs. In this case, more meant going from two to 32 chips. and the model would learn the same thing, but learn it way faster. And the second thing that they showed was what the industry calls scaling laws, which effectively means a model has less mistakes, produces less mistakes if it has more parameters, like more knobs to tweak, and is fed more data and more compute.

15:50Nathan Benaich:So this was to some degree known. It wasn't broadly appreciated. And also, there wasn't really much of a capital source that was willing to throw enormous sums of money to take those ideas to their next logical conclusion. But it's important to know that the people who are professing that scale works today, the people who run OpenAI, who will demine, who run Anthropic, same people who were saying scaling was working back then. So now they have the cash to go prove their thesis.

16:19Demetri Kofinas:Now you're opening the door to another conversation about whether or not the current narrative around scale, which is that we just need to throw more and more compute and acquire more and more data, and eventually we'll get super intelligence. There's a real question about whether in fact that actually is going to prove to be correct. And would you say that the deep seek moment, which happened less than a year ago, was the first moment where that narrative began to crack in the industry?

16:45Nathan Benaich:For me, the deep seek moment, which was run in January this year, was really an example of how there's information asymmetry in the market. Because for AI research people who work in the space all the time, they've been tracking DeepSeek for a while. In particular, like end of last year, there was a paper, DeepSeek V3, which was sort of like their first like large language model, base model. And what they showed in January was basically that, but with reinforcement learning on top. And the thing that like Wall Street got wrong was that it read in the paper that the final model that was produced, DeepSeek R1, cost DeepSeek$5 million to train.

17:27Nathan Benaich:But why that's a bad take is that it's the equivalent of saying in a Formula One race weekend, the$4 million or$5 million that DeepSeek R1 cost was basically the cost of the fastest qualifying lap. And it ignores everything else that gets thrown into getting there.

17:48Demetri Kofinas:Including data and training it was able to do as a result of all the money that was spent by some of these closed models.

17:53Nathan Benaich:Absolutely. And I mean, the R &D cost to get to the final data mixture, all the annotated data, the software infrastructure to even run this thing, and then the GPUs to go do the training, and then the knobs that you're eventually going to decide are the final write knobs settings for your perfect model, all of that's excluded from the$5 million.

18:18Demetri Kofinas:So was that the most significant breakthrough of the past 12 months? And if not, what was? And if so, what are some additional major changes in the last 12 months that we should be focused on? Yeah.

18:30Nathan Benaich:I think DeepSeq was a meta trend. What could be branched off of it was the rise of China and open source, which we can talk about. Pretty exceptional rise in developer adoption today. The second thing is that there's still a lot of juice to squeeze from this transformer architecture we discussed. DeepSeq brought in a bunch of innovations that make it cheaper to run and cheaper to train, which now other labs are using. And then the third one is that reasoning has been brought to the table. And this started end of last year with O1 Preview from OpenAI. And reasoning is effectively a way of saying the model outputs its step-by-step thinking process in order to generate a response.

Read the full transcript

19:13Nathan Benaich:Whereas previously it would just take your question and then spit out a response. Wouldn't like explicitly think or plan different routes. So this is like thematically a little bit similar to AlphaGo and, you know, D-Bind's work solving chess and other games where, you know, the model is trying to roll out different futures of, oh, if I say this or if I take that path, like, am I going to get an answer that's better or worse than the other? And doing that like at a large number and then it's getting like rewarded for which path ultimately gave the best response. So it can now like use similar paths in the future.

19:49Nathan Benaich:So this general idea of models thinking before they answer was a really big trend in the last year. It feels old already today because every model has this, but that was a big breakthrough end of last year.

20:01Demetri Kofinas:Just to clarify for people that may need it, when we're referring to DeepSeek, so DeepSeek itself is a quantitative hedge fund, correct?

20:08Nathan Benaich:The quant fund I think is called High Flyer Capital, and then it spun out at a separate AI company by the name of DeepSeek, yeah. And it's fully funded by the HFT firm. Great.

20:18Demetri Kofinas:So they released a large language model, as we said, less than a year ago that appeared to achieve near parity with leading US closed models at a fraction of the cost, which we've already established was a bit misleading, but still it was impressive. So that doesn't, I mean, I think we probably could both agree that as impressive as that was, it still wasn't a moment comparable to the release of ChatGPT 3.5 that happened in November of 2022, right?

20:42Nathan Benaich:Yeah.

20:42Demetri Kofinas:Do you foresee anything like that on the immediate horizon?

20:47Nathan Benaich:I mean, it's so hard to say. People can't say this.

20:50Demetri Kofinas:Why are you laughing? Yeah, I'm laughing because it was just like such an epic event.

20:56Nathan Benaich:I still remember where I was when I read Sam Waltman's tweet of like, hey, we're recently releasing this research preview. Let's see what people think. And then just to think that within, I don't know, two years or something, almost a billion people would be using this every day.

21:09Demetri Kofinas:Well, that's the thing too. And that's important to emphasize because I don't know how many people actually understand that this is by far the most successful product launch in history and the most rapidly adopted technology in the history of humankind.

21:22Nathan Benaich:Yeah. And so when you ask me, what's the next one? I'm like, I don't know. Yeah.

21:28Demetri Kofinas:No, no. I get it. I get it. Well, this is the media. So we have to do that. Yeah. Yeah.

21:31Nathan Benaich:But I'll tell you which ones I think are pretty amazing. and maybe you're not on that scale. But I'm continually amazed at how we've pretty much solved voice synthesis and voice cloning. I mean, it's truly astonishing how now with services like 11 Labs and others, you could input a whole amount of text, use your voice clone, generate audio. And for people who don't know you extremely well, maybe outside of your immediate family, it's hard to tell the difference.

22:00Demetri Kofinas:I have a question about that. So this is one of those questions that I think many people will find relatable. If we've been able to do that, why is my iPhone still so bad at doing speech to text? And why are automatic transcription services still so bad at transcribing my podcast?

22:17Nathan Benaich:I think it depends. I would say OpenAI's transcription service is pretty good. It might not be embedded in the tools that you use, but I'd say they're very good. 11 is very good. And Apple is definitely better than what it was. But I agree, they've been really dropping the ball over pretty run-of-the-mill use cases. So I'd say it's the case of you have to wait. It will get better. And then it's a bit of corporate priorities of there's one or two edge cases worth solving given other focus areas.

22:50Demetri Kofinas:Well, I think we'll have a chance to maybe dig into that a bit more later when we talk about the brittleness or robustness of these systems and the wide variability in the quality of responses to prompts and why that is. Let's go back to my very first question where I quoted the report and you said that it was going to be a force multiplier for technological progress. I saw Bill Gates recently talk about how... I don't remember exactly the language he used and maybe you know this interview and you can tell us, but it was essentially... I'm paraphrasing here that AI is by far the most transformative, the most remarkable technology he has ever witnessed.

23:27Demetri Kofinas:So he's basically saying that the internet was nothing compared to artificial intelligence. If you agree with that general sentiment, I'm curious to understand why you feel this way and why you're so confident that it'll become the most transformative general purpose technology in the history of humankind.

23:46Nathan Benaich:I think I agree with him. The internet has amalgamated as much information as humans have generated in our entire history. And yes, there's still a lot of information that's not online and in company databases and museums and other repositories that will eventually make their way into digital form. But if we look at just the internet, that being the exhaust of all of our knowledge and opinions and global events, etc. And AI is a way of basically learning and to some degree memorizing. Even, frankly, a system that can memorize all the information, retrieve it is pretty damn useful. But so having access to all that information in one place is A, pretty magical.

24:30Nathan Benaich:There's no way that any human could ever parse all information on the internet and figure out what's relevant to them in different cases. And then I think the second part where I think AI is quite special is, if we look at anyone's profession, even the people who are executing at their peak ability, that's still probably a local maximum compared to where we could get to. Because again, they are the product of who they've studied, who they've been trained by, what they read. And now with AI, we have the opportunity to explore probably global maxima, depending on how much compute and innovation we want to throw at the problem.

25:04Nathan Benaich:But you can synthesize like all the best experts into a single system that would never be possible otherwise. And so I think for that reason, it's like a very, very strong enabler to any profession that has to like think deeply about a task or, I mean, sometimes not even think particularly deeply. But if the world is digital, AI is the accelerant on anything that's digital and is increasingly kind of going into the offline world as well, which we can talk about. Yeah, no.

25:36Demetri Kofinas:So that actually raises a philosophical question, which has to do with how accurately these models perceive reality. Because digital information is a representation of the analog world. Yeah. Now, the argument that some will make is that eventually they will perceive it perfectly because we're going to have sensors in the real world that are actually capturing that data. So robots will be walking around and we'll eventually have all the data we need to replicate anything. And the fidelity will be for all intents and purposes, perfect. How would you respond to someone that would say, we can't capture reality perfectly and therefore we cannot rely on these systems to automate critical parts of our infrastructure or to play the role of a global answering machine in the future?

26:32Nathan Benaich:Yeah, for sure. There's an element of sort of like perceiver bias almost.

26:38Demetri Kofinas:When you say there's a perceiver bias, you mean from on the side of people that are actually working in the space or epistemological gap on the part of the systems.

26:47Nathan Benaich:Yeah. As in like, you know, we're all kind of, we're living in the same world. Like theoretically, we're all living in the same world, but we have different vignettes that we experience it through, you know, based on socioeconomics, like our, where exactly we live in the world and our knowledge. And I think in some ways, like AIs are not too dissimilar. I think the difference with the AI though, is they're pretty good at like rule following, or at least you could tell them like, you are a person of X description. How would you answer this question? And because the internet has a lot of that information, the system is able to endow that persona and answer on that basis.

27:25Demetri Kofinas:That's fascinating because I've found the same thing to be true when I... Because I use ChatGPT, I use Claude a little bit. Actually, as imperfect as ChatGPT is for the type of work that I do, I still have found it to be the most useful. And maybe that's just because I've just gotten used to using it. But this is so interesting because I found this to be true. And I've just intuitively done that, which is to say that I've tried to create a scene and I put it in that scene in order to get it closer to the answers that I want. Why is it that that works so well when it comes to prompting these systems?

27:59Nathan Benaich:Yeah. So I think a way to think about it is that the AI has a crazy amount of information And that information is like stored in an extremely like complicated high dimensional space. And like us as users, we're basically poking the system with like a stick to try to nudge it in a part of the high dimensional space that answers our question. And so I think the more support, like scaffolding, guidance, documents you can upload, descriptions of how it should behave, the more you give it support to figure out where in this high dimensional space it should live to answer your question. And so, yeah, it's giving it a bit of guardrails and walking sticks.

28:43Demetri Kofinas:Is it fair to... So I love that when you think of this in terms of three-dimensional terms, is another way of describing that, calling it the answer space, so that there's a huge answer space. And the more context you give these systems, the more you can narrow that answer space. And so when you pull on that lever on the slot machine to get an answer, you're more likely to get an answer that has less variability and is closer to what you're looking for.

29:11Nathan Benaich:Exactly. It's like the more you give it as information for a question, the better navigation system you're giving it to eventually get an answer that's coherent, correct, and useful.

29:21Demetri Kofinas:So for someone listening to this, is it fair to say that a lot of the variability and the poor responses that they're getting are just as much, if not more reflective of their failure as a prompter than it is actually what these systems are capable of doing?

29:40Nathan Benaich:I'd argue in large part, yes. now we can say is that really the user's fault or is it that the product is not sort of designed in a way to have affordances for the fact that not every user knows that they should do this and we can also ask like should we really force a user to know how to do that or should the thing just figure it out i think the field in general is going towards the direction of you shouldn't really expect people to know and give all these like prompt complexities and context and stuff, it should be able to figure it out, but we're still nudging our way there. So in the meantime, it's important to give it as much context as you can.

30:22Demetri Kofinas:But it does mean that people that are aware of that have an advantage. So that's a good thing for people to be aware of.

30:28Nathan Benaich:Correct. Another way that this is experienced is when you start a chat, whether it's in Claude or ChatGPT, et cetera, you have basically a memory that the model has, which is not full. And then you're giving it context, you're giving it examples, files, you have your conversation. And over time, that memory fills up. And the more full it is, the harder of a job it has at figuring out what part of your historical chat or what part of the information you've inputted and what part of answer space it should really focus on. So it tends to get a bit distracted, maybe performs poorer the more your chat extends.

31:05Nathan Benaich:And so it's generally a good practice to anytime you deviate on topics significantly, just start a new chat and you'll see like the responses are much better.

31:15Demetri Kofinas:So you're basically building up errors over time and those errors are compounding and there's more noise in the channel as a result of the fact that you're... I think so.

31:25Nathan Benaich:It's less like errors. It's more like you're at school and then you've taken 10 different classes and 10 different subjects in a given day. And then at the end of the day, you're asked to perform an exam on one of those subjects. I mean, your brain's probably fried. Whereas if you can do one class and take an exam versus the next day, take another class with another exam, you can focus a lot more because your brain's not fried.

31:52Demetri Kofinas:But is that the right analogy? I want to just poke at that a little bit. Is that the right analogy though? Because its brain isn't being fried. It actually says something about the way that it understands the world and reality and its propensity to become confused when you give it more information, which for people is the opposite is true. If you give people more information across various disciplines, they'll make connections and understand something more deeply. So that's a really interesting distinction between how humans learn and how these models to perceive reality.

32:21Nathan Benaich:Yeah, I would say this is also a bit of like a here and now. It's very likely that in a year or two years, the community has figured out different improvements to bring about the system design that kind of overcome these kinds of problems. And there are some services like Gemini that have a really, really long memory, what's called context window. And so it can get a bit less distracted. And then some models that have much shorter ones. So because everything is evolving very fast, I think this is true in the here and now and might not be in the future. All right.

32:55Demetri Kofinas:So let's talk a little bit about some of these innovations that you mentioned earlier, like inference time scaling and chain of thought. So the report opens with a discussion about how OpenAI's O1 preview model, which you mentioned earlier in this conversation, demonstrated inference time scaling with reinforcement learning using chain of thought as a scratch pad. and how that's led to marked improvements in reasoning, heavy domains like software coding and scientific research. Help me understand what exactly the innovation was here and also what is chain of thought and inference time scaling?

33:28Nathan Benaich:Yeah. So actually two years ago or so, there was evidence in the community in the form papers that were showing if you just told Chad GBT to think step-by-step or to think carefully before it was doing any reasoning, it would perform better.

33:43Demetri Kofinas:Is that because you're prompting it to use reason in its thinking? You're narrowing the context window in a sense.

33:51Nathan Benaich:Well, I think you're almost like decomposing the overall task into smaller tasks and systems are just better at solving those kinds of sequence of smaller tasks because again, like sort of like the jump is less far, sort of like making a little hop and then a little hop, a little hop, and then you solve the problem. And as a result of that, a system can debug a bit better because they can figure out where it might have made a mistake in its thinking process. So like once the community figured out, like you just tell ChatGPT to think step by step, it does better. I think that amongst other things prompted developers to figure out like, how can we actually train a model to do that?

34:29Nathan Benaich:As opposed to, like you said, sort of invoke that capability that would be like maybe late present in a latent way in the model based on its training from internet data. And so what happened was, you know, model developers would request data from experts in a variety of different domains, history, physics, math, and then give them tasks and explicitly ask them to annotate and just write out their thinking process, like the key steps that they might take to logically conclude a specific answer. And, you know, there might be different ways to develop a reasoning path to answer the very same question, and then use that as training data for a model, and then basically encourage the model when it spells out its reasoning, which can be, you know, in the English language, sometimes there's some weird nuances where sometimes a model will like start reasoning in Chinese or something and then switch back to English.

35:25Nathan Benaich:There's different ways you can incentivize the model to produce like legible or like shorter or longer reasoning traces. And sometimes this improves performance or degrades performance. And so then we started creating these like reasoning models because they were trained with complex reasoning traces from experts. And the first generation of these systems were demonstrated to work super well on math, physics, what was overall called like verifiable domains, which is to say like, this is a mathematical proof and we know the answer and I can check that the answer is correct or not versus like poetry where there's no correct answer.

36:05Nathan Benaich:It's very like subjective and that's unverifiable. So 01 preview did like super well on math and verifiable domains. And then they had the whole like deep seek work, which we described, which was very similar, except that model in order to give it reasoning capabilities was trained to encourage reasoning explanations, which were in cogent English. And then there's like other variations of this over the ensuing months. And overall, we've just seen these kinds of systems perform better on math benchmarks, chemistry benchmarks, physics, real world knowledge tasks and legal or banking, et cetera.

36:43Demetri Kofinas:This is really fascinating because I'm thinking right now about my own experience using ChatGPT and moving especially from 4.5 to the most recent model, ChatGPT5, because I saw an actual decline in efficacy or in its ability to be useful in my particular use cases of either speaking to it about my upcoming conversations or asking it questions and interrogating its answers. So it seems that these systems do very well, as you said, in formal settings where you have definitions that you're confident about and you can structure the question answer or the prompts in a sort of step-by-step format. But if you have an open-ended question where you yourself don't fully understand what you're asking, or if what you're asking about cannot be easily converted from your experience into language and from language into the language that the computer understands or that the AI understands, something gets lost.

37:48Demetri Kofinas:And so would that explain why, for example, I've actually seen ChatGPT get worse at having a conversation with me about my interviews if I give it my transcript, for example?

38:01Nathan Benaich:Yeah, a few things there. I don't think your experience is anomalous. I also see that ChatGPT in particular loves bullet points, loves to be succinct.

38:12Demetri Kofinas:Why is that? I love that. Thank you. You're 100 % right. I mean, it's also bad at formulating coherent sentences. And the other thing that it does, Nathan, you know when you speak to somebody who isn't especially intelligent, but feels insecure about themselves, they're being interviewed maybe on the street, they're just an everyday person, they're being interviewed about a topic and they want to appear intelligent, they will talk in a certain way. They'll use words that are unnecessary because they sound smart, they don't economize their language, they over-explain. That's what it does.

38:44Nathan Benaich:That's what

38:44Demetri Kofinas:it feels like I'm reading when I'm reading its responses.

38:47Nathan Benaich:Yeah. It's like pseudo-intellectualism.

38:48Demetri Kofinas:And it's very superficial. When you begin to poke into it, you can see there's a superficial level of understanding. So go ahead. I just wanted to throw that in there as you respond.

38:58Nathan Benaich:Yeah. I see in my own experience that 4.0 was much better at writing than 5. Like 5, for the reasons we said are too succinct. It's like bullet point notes, like shorthand. And that's especially true for 5.0 Pro. Yeah. Yeah. So the thing to, I think, appreciate with building these systems is We talk about a model, but these systems are just a gargantuan amount of work. To get a sense, if you go and see the Gemini 2, Gemini 2.5 paper from Google DeepMind, it has 1 ,000 authors or something, or maybe more. And to eventually get the fully baked behavior that we're seeing in our applications, there's probably hundreds of different signals, or maybe even more.

39:42Nathan Benaich:like little tweaks of behaviors across every single domain that any human being who has access to this tool might ask. And like little rewards or ways that, you know, the team that's developing the model can grade whether the system does well for their specific case. And then you're throwing like all this data into the soup, all these evaluations into a soup, and like hoping that every new model generation doesn't regress on some capabilities that people care about. And while, you know, improving on other capabilities people care about. I think it's just generally magical that like the thing works as well as it does, given like all these competing directions in which the model is being pulled.

40:23Nathan Benaich:And at the end of the day, like there is not really like a free lunch with, you know, being good at everything while not like regressing in one direction or another. And so this is another reason why I think, you know, every time you hit like update in the ChatGPT app, you might get a slightly different experience. And, you know, we might I'd say right now in October 2025, it's not good at writing like we used to, but then in a few months, a new update comes and then we love it again. So these are dynamical living systems and it's a game of trade-offs. And it's also important to point out that it's only been three years since ChatGPT launched

41:00Demetri Kofinas:its... Not even three years since it launched that model that we talked about as being sort of that breakthrough moment where it hit the public zeitgeist. Yeah. And that New York Times reporter was having a conversation with Sydney. Remember that? Yeah, there's a Microsoft spot, right? Yeah. It's amazing how amenable human beings are to normalizing conditions that would have been totally abnormal just a few years ago.

41:26Nathan Benaich:For sure. It's just like this human thing of we sort of ignore the 99 % that works and we focus on the 1 % that's really crappy. Right.

41:36Demetri Kofinas:And then adjust our expectations accordingly so that we become frustrated and annoyed that the system is unable to perform at the level that we expect. It reminds me of this Louis CK joke where he described being on an airplane on his phone and the wifi wouldn't work. And he was like, this stupid piece of crap. And then he had a realization, this thing, I'm getting in this aluminum tube and there's got giant tires that no one knows how they got the air into those tires, and this thing's going to be flying in the air and I'm going to end up in Tokyo in no time. And I'm complaining about the Wi-Fi.

42:07Demetri Kofinas:How does chain of thought and inference time compute compare to what we're seeing in terms of visible thinking, which is to say these systems are actually showing us their reasoning? Is this essentially them just opening up the hood and letting us see what's happening so that we improve the transparency of these systems?

42:24Nathan Benaich:Yeah, that's the general idea. So with inference time compute, what we're saying is we're not taking compute and then allocating it to pre-training where what's happening with pre-training is that the weights, like the parameters, knobs, if you will, on the neural network are being updated, changed. It's a bit like some very high dimensional pinball machine. And you want to get the ball on the goal and you're adjusting the pins to make sure the ball gets in the hole. That's kind of what you're doing with updating weights. You're instead allocating a lot of compute towards the model thinking when you ask it a question.

43:00Nathan Benaich:And when you do that called inference, which is like getting a trained model with fixed weights to answer a question, you're not updating the weights. You're just getting it to expend more time to explore different paths in the like search tree of potential responses. And then, you know, model providers, I think originally weren't too keen to expose these reasoning traces, as we call them, because you could effectively copy those reasoning traces and then use them to train your own system. So really what you see is not the raw reasoning trace, because that would be insanely long and sometimes very complicated.

43:39Nathan Benaich:You're seeing a sort of pre-processed version of it that makes it sometimes more legible for a normal person using the app. And so the benefit from doing this chain of thought and inference time compute appears to be the more inference time or the more compute you apply to thinking time, the better performance you get for reasons that we discussed previously. And then the other benefit you get is monitorability, which is like, why is a model behaving a certain way? Why did it bug out? Why did it display behavior that we don't want? And also, can we determine if a model is being hacked or tampered with?

44:16Nathan Benaich:Not necessarily because it's saying in this chain of thought, I think I'm being tampered with, but because there are certain patterns of tampering which can appear statistically within the reasoning trace.

44:26Demetri Kofinas:So let's take a moment to discuss scaling here because it's something we touched on earlier and I want to make sure we cover it before we move to the second hour. As I mentioned, there was a sort of broad narrative in the last few years that the path to artificial general intelligence or super intelligence was through physically scaling these systems by throwing more compute and data at them. And eventually that magic would produce super intelligence. It now seems that we're moving or beginning to move from that kind of scaling compute narrative to one of inference time scaling, where the systems spend more time thinking through the problem in order to provide better solutions, which I think is also, if not the innovation of DeepSeek, it was something that DeepSeek relied on to make the gains that it did.

45:17Demetri Kofinas:right? So walk me through where you think we are in terms of how we're going to scale these systems and also like what is the general consensus in the industry? Is there no longer a consensus and we're kind of somewhere in the middle? And how is that informing the CapEx spending of these companies, of these foundation models like OpenAI or Google, et cetera? Yeah.

45:38Nathan Benaich:So you're right. That's the narrative of scaling, which is more compute, more data, bigger models over time yields better performance. There was a time late last year, early this year, where a moniker was bandied around on Twitter and in the press on the lines of deep learning is hitting a wall. And what people meant by that was, oh, look, OpenAI has delayed their model release. Oh, look, this latest model release from other vendor is not as impressive as what the prior model release was. Therefore, there's less juice to squeeze. Therefore, capabilities are tapering off. I think the challenge with that is, one, related to your point around humans just get used to things very fast.

46:21Nathan Benaich:And so our bar for expectations is now higher because we've witnessed so many magical capabilities. The second thing is we're generally as a field moving away from static benchmarks, which were created for academic purposes. for example historically categorizing images in one of thousand different categories and then a model that did really well on that was deemed to be good at computer vision to now saying well a model that actually does useful work that for example mckinsey or goldman sachs are buying and spending a lot of money on and saving human time that's the eval benchmark that matters and i'd say the other part is yes there are like new axes now to do scaling like it used to be on pre-training originally might have a bigger and bigger corpus of unlabeled data like the internet and maybe internal corpuses to do training, then throw in image, video, audio, all encoded in the same way as text, sequence by sequence, and then building gigantic GPU clusters.

47:28Nathan Benaich:And then on top of that, throwing reinforcement learning when the domain in which you're working is verifiable, which generally I think teaches logic, which can be extrapolated usefully in other domains, which are not verifiable. And so, yeah, there's been a different number of axes and different ingredients that you throw in the soup, which has improved capabilities. And I think the consensus in the community is there's definitely more to squeeze. This is even excluding modalities like synthetic data, where a model is now very smart and can generate very smart answers. that it can use itself to self-improve, basically.

48:07Nathan Benaich:And then there's like a whole axis of like kind of new architectures which are being worked on concurrently. But in general, the bets like, yeah, more scale along what we discussed will bring about better capabilities. And, you know, if you follow people from the major labs, there's general belief that this recipe will yield some form of superintelligence. Whether it'll be across every single task is another question.

48:30Demetri Kofinas:I'm curious what you think. First of all, do you think that we've reached a plateau? Do you think we've reached a plateau in terms of scaling the systems based on throwing more compute and data at them? Do you think that we can continue to find sort of quote, software hacks that help scale the system or that help continue to scale these systems? Do we need a new architecture? I mean, where do you fall on this in terms of what is going to push the frontier here? i'm i'm positive on progress because it has only been you know two and a bit years since chat gpt

49:05Nathan Benaich:and it's important to know that basically everybody who worked in ai up until chat gpt were ai people ai research people and so it's quite magical that those individuals managed to produce an artifact that's this useful and then after that ai became the like new new thing and maybe the only new new thing, which has then subsumed all the energy of every other industry into it. And so you've seen people like the individual who invented WebRTC, which allows the browser to do so many things, has moved into OpenAI. People who've done crazy optimizations to run crazy large cloud services and basically build the internet and bidding exchanges and gambling exchanges, all these systems that need to work very quickly and at a low cost have moved into AI.

49:55Nathan Benaich:You have all the chip people who are working on non-AI things moving to AI. So I think as you get all this expertise moving into this field and bringing basically their capabilities that they've learned in different disciplines that are generally around robustness, cost efficiency, speed, AI is going to be a beneficiary of this. And when you apply that to the three axes we discussed around scaling, I think it's a fair bet that the systems we have today are probably the worst that we'll see.

50:26Demetri Kofinas:So Nathan, I'm going to move us to the second hour. I want to discuss this move towards open models and what that means both in geopolitical terms, because China's really leading the field in that area. I'm also curious to dig into why exactly they decided to do that, what the strategy is there. There are obviously huge implications commercially for investors. One, if we move away from the current scaling paradigm. But also, if in fact the push eventually leads to more and more of these foundation models going from closed to open, what that means in terms of where the profit opportunities are, are they further up the stack?

51:05Demetri Kofinas:And is that why we're also seeing companies like OpenAI begin to commercialize new products to sit on top of OpenAI's foundation model, because that's ultimately where the profit is going to accrue and be able to be captured. I'm also curious to hear your thoughts about how this is going to accelerate innovation in science. What are we seeing there? Because I think there's a lot of hope for breakthroughs. And then I mentioned China. There's the broader question of the geopolitics because up until now, companies like OpenAI have very cynically tried to exploit the US-China competition to sort of build regulatory moats around their companies and protect themselves and their profit streams.

51:50Demetri Kofinas:I'm curious how this plays out geopolitically in your view, and if the US should in fact be concerned about the progress that China is making. And then what does that mean for the debate about AI safety? When I started digging into or learning about this field, it really started with Nick Bostrom's book, Superintelligence. I read it in 2015, but I think it was published in 2014. And at that time, that was the most interesting part of this discussion. And that has really just faded from view. And it seems like it's faded from view largely because people have decided that it's just too hard, that the game theory doesn't support it.

52:28Demetri Kofinas:So I'm curious to ask you about all of those things in the second hour, Nathan. For anyone who is new to the program, Hidden Forces is listener supportive. We don't accept advertisers or commercial sponsors. The entire show is funded from top to bottom by listeners like you. If you want access to the second hour of today's conversation with Nathan, head over to hiddenforces.io slash subscribe and sign up to one of our three content tiers. All subscribers gain access to our premium feed, which you can use to listen to the rest of today's conversation on your mobile device using your favorite podcast app, just like you're listening to this episode right now.

53:05Demetri Kofinas:Nathan, stick around. We're going to move the second hour of our conversation onto the premium feed. If you want to listen in on the rest of today's conversation, head over to hiddenforces.io slash subscribe and join our premium feed. If you want to join in on the conversation and become a member of the Hidden Forces Genius community, you can also do that through our subscriber page. Today's episode was produced by me and edited by Castilianos Nicolaou. For more episodes, you can check out our website at hiddenforces.io. You can follow me on Twitter at Kofinas, and you can email me at info at hiddenforces.io.

53:46Demetri Kofinas:As always, thanks for listening. We'll see you next time.

From the publisher

In Episode 448 of Hidden Forces, Demetri Kofinas speaks with Nathan Benaich, founder and general partner of Air Street Capital and the creator of the annual State of AI Report, an open-access compendium that tracks advances across AI research, industry, policy, and geopolitics.

Nathan Benaich and Demetri spend the first hour of their conversation exploring some of the most important AI breakthroughs of the year. They unpack the DeepSeek moment, dig into some of the advancements made by the latest reasoning models, and discuss why there appears to be a regression in capabilities across certain domains in artificial intelligence at the same time as we are seeing marked improvements in reasoning-heavy use cases like coding and scientific research.

The second hour turns to a conversation about the commercial implications and geopolitical dynamics of the AI arms race, including China's strategy to become the leader in open-weight models and tooling. They look at what industries, sectors, and professions may be most ripe for disruption, where the investment opportunities are, whether we're in a bubble comparable to the 1990s Internet boom, and how export controls, energy constraints, and regulatory red-tape could play an outsized role in shaping the trajectory of the current arms race.

Lastly, Kofinas and Benaich examine where along the AI stack most of the value is likely to accrue—from the underlying picks and shovels, through the foundation models, to the apps that ride on top of them—and what all this means for labor markets, education, and the cadence of scientific discovery.

Subscribe to our premium content—including our premium feed, episode transcripts, and Intelligence Reports—by visiting HiddenForces.io/subscribe.

If you'd like to join the conversation and become a member of the Hidden Forces Genius community—with benefits like Q&A calls with guests, exclusive research and analysis, in-person events, and dinners—you can also sign up on our subscriber page at HiddenForces.io/subscribe.

If you enjoyed today's episode of Hidden Forces, please support the show by:

Producer & Host: Demetri Kofinas
Editor & Engineer: Stylianos Nicolaou

Subscribe and support the podcast at https://hiddenforces.io.
Join the conversation on Facebook, Instagram, and Twitter at @hiddenforcespod
Follow Demetri on Twitter at @Kofinas

Episode Recorded on 10/29/2025

More from Hidden Forces

All 76 episodes
Investing on the Front Lines of the AI Arms RaceHidden Forces · 54 min
Listen in VO