In short
NVIDIA’s “AI Compute Partnership” backstops for new cloud providers; Microsoft valuation “true value” up to $3.8T; why AI visual reasoning still fails basic tasks; whether AI could be conscious; and “model routers” to cut AI costs.
Guests and backgrounds
Phoebe Liu, NVIDIA reporter (broke the backstop story; interviewed CoreWeave co-CEO Tim Rosenfield). Anita Ramaswamy, financial analysis columnist (True Value). Andrew Dye, ex-Google (14 years), now at Elorian, focused on visual reasoning. Rob Long, executive director of Elios AI (AI consciousness research). Laura Bratton, colleague; co-wrote AI Agenda newsletter on model routers.
Key claims
NVIDIA guarantees unsold GPU capacity (lease-back at a flat per-GPU-per-hour rate) in exchange for a revenue share, lowering lenders’ interest rates; NVIDIA publicly claims it expects to “pay nothing.” Microsoft may be undervalued: cloud ~$1.4–$1.7T, software ~$1.7T, personal computing ~$350B; concerns are “SaaSpocalypse” and AI capex ROI. Visual models excel at recognition but struggle with multi-step visual reasoning; Baby Vision shows frontier models reason around preschooler level. Consciousness paper argues for triangulating behavior, internals, and development; error bars remain huge. Routers route sub-tasks to cheaper models; Palantir/Databricks add prompt optimization and model swapping.
Notable examples
Firmus and ShareNAI deals; CoreWeave lease-back precedent; Baby Vision tasks (mazes, counting, odd-one-out); Blake Lemoine’s Lambda consciousness claim; Palantir Evolve examples: ~97% compute cost reduction by swapping to GPT-5.4 nano; McCarthy Building: 60% fewer tokens.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VONVIDIA's AI Compute Partnership
1:05 to 3:40
Discussion about NVIDIA's strategy to backstop GPU purchases for cloud providers.
“The Information published exclusive reporting that NVIDIA is planning to financially backstop young cloud providers that rent out its AI chips in exchange for a share of their revenues.”
NVIDIA's Market Position
3:40 to 4:25
Insight into NVIDIA's strong market position and expectations for demand.
“Do you think NVIDIA expects that in practice, it will have to rent some of that capacity?”
Risks of NeoCloud Business
4:25 to 7:30
A deep dive into the risks associated with NeoCloud business models and NVIDIA's backing.
“At least in the case of Firmis, I talked to one of their co-CEOs, Tim Rosenfield, and he said it took them just a couple weeks to get customer commitments for the full length of the contract.”
Revenue Sharing in NeoCloud Arrangements
7:30 to 11:40
Exploration of the revenue sharing models in agreements with cloud providers.
“But again, I think it's pretty much mitigated by NVIDIA's backstops.”
Microsoft's Valuation Analysis
11:40 to 14:01
Assessment of Microsoft's current market valuation and factors influencing it.
“Microsoft is currently trading at a$2.9 trillion valuation.”
Analyzing Microsoft's Business Segments
14:01 to 21:34
Explore the different business units of Microsoft and their valuation challenges.
“and sometimes even attract different investor bases.”
Introduction to Visual Reasoning
21:34 to 21:58
Learn about the current limitations in AI visual reasoning and its relevance.
“The most advanced AI models still cannot solve some visual puzzles that even kids can solve.”
Challenges in AI Visual Reasoning
21:58 to 28:00
Understand the specific challenges and benchmarks in AI visual reasoning.
“I'm excited for this conversation because visual reasoning is a really fascinating area of AI right now.”
Exploring Multimodal AI Capabilities
28:00 to 33:35
Discussion on the trade-offs and challenges in developing multimodal AI systems.
“That's quite simply where they see the demand.”
Introduction to AI Consciousness
33:35 to 35:39
Introduction of Rob Long and discussion on the implications of AI consciousness.
“Rob joins us now to talk about a new paper on AI consciousness he just published.”
Show all 18 chapters
Investigating AI Consciousness Empirically
35:39 to 41:44
Discussion on the empirical methods to study AI consciousness and welfare.
“So how do we think about if animals are conscious, if they're somewhat different from humans?”
Consciousness vs. Intelligence in AI
41:44 to 42:04
Exploration of the relationship between AI consciousness and intelligence metrics.
“the smarter they are, that as scaling continues, the odds of consciousness are kind of going up.”
Exploring AI Consciousness
42:04 to 44:14
Discussion on the relationship between AI capabilities and consciousness.
“Well, I think it's going to be a combination.”
Introduction to Model Routers
44:14 to 44:29
Introduction of Laura Bratton and the topic of AI model routers.
“I wrote about routers for our AI Agenda newsletter alongside our applied AI recorder, Laura Bratton, who joins me now.”
Functionality of Routers
44:29 to 45:39
Breakdown of what model routers do and their current relevance.
“Can you break down for us, like, what even is a router?”
Challenges with Multi-Model Usage
45:39 to 47:11
Discussion on the complexities and challenges of using multiple models in AI.
“because it like depends who's using the model and what tool they're using to then access the model.”
Router Innovations from Palantir and Databricks
47:11 to 50:59
Insights into how companies like Palantir and Databricks are innovating with routers.
“But I do want to call out Palantir and Databricks in particular for launching new tools that include routers.”
Cost-Efficiency with Routers
50:59 to 54:25
Discussion on the financial benefits of using routers for AI tasks.
“I think that they're also doing really interesting stuff.”
Transcript
Automatic transcript. May contain errors.0:13Welcome everyone to the Informations TI TV. My name is Rocket Drew. It is Thursday, July 2nd. First up, the information published exclusive reporting about NVIDIA agreeing to financially backstop GPU purchases and taking cuts of the revenue. Our NVIDIA reporter will join us to share more about what she learned. And our financial analysis columnist says Microsoft may be valued at more than it's currently trading for. We'll hear the reasoning behind that. I'll also dig into some technical topics with a couple of guests, like how AI models continue to struggle with visual reasoning, and whether or not AI could ever be conscious.
0:53And to wrap things up, I'll be joined by my colleague Laura Bratton to talk through our latest AI Agenda newsletter about model routers. It's going to be a great show, so let's get right on into it. The Information published exclusive reporting that NVIDIA is planning to financially backstop young cloud providers that rent out its AI chips in exchange for a share of their revenues. Our NVIDIA reporter Phoebe Liu broke that story, and she joins me now to share what she learned. All right, welcome on the show, Phoebe. Thanks for having me. Of course. So you wrote that this is part of a program that some people at NVIDIA have dubbed the AI Compute Partnership.
1:33Tell me more about what you learned. Yeah, totally. So there were two deals that were kind of announced in the last month or so from Firmus and ShareNAI, which are both kind of Australian-based new cloud slash data center builders somewhere along that spectrum. And it made me kind of think, okay, if there are two, there must be more. So I started kind of calling around and asking people how widespread that was. And it turns out it is part of kind of a pretty codified program that NVIDIA is starting to provide kind of younger, newer cloud customers a credit wrapper to help them buy NVIDIA's pretty expensive GPUs at lower interest rates in exchange for a cut of recurring usage-based revenue.
2:19And that's a departure from kind of previous strategy where NVIDIA has backstopped leases, provided equity investments, and done other things to stimulate the ecosystem. But it's kind of part of that trend where NVIDIA wants to get other people to buy its products. I think lots of companies do this, but I think NVIDIA has done it at an extent that's very interesting given how big of a role it's playing in the ecosystem. So the way it works is that NVIDIA guarantees that if capacity is unsold, it'll lease it back at a flat rate of X dollars per GPU per hour. I would love to know what X is. and then the revenue share would be kind of a percentage of the upside and scale down over the life of the contract which tends to be five to six years as far as I understand and basically that guarantee makes banks or other lenders more willing to lend money to these newer clouds that like don't have good credit and I don't know they would probably be offered like a 15 % interest rate or something and that may not be a viable business for them slash they wouldn't be able to scale as quickly.
3:26But with NVIDIA's backing, they can get their interest rates much closer to investment grade. So it just seems like another tool in NVIDIA's toolbox of ways that it's helping customers be able to afford its GPUs. Do you think NVIDIA expects that in practice, it will have to rent some of that capacity? Or is the bet that it'll never come to that because the NeoClouds will be so successful? Yeah, I'm sure they've talked about it. I think publicly they posture that they never expect to pay a cent and basically they're just leveraging their balance sheet um because they have such a strong one i don't know they i think they made 160 billion in net income in the last year or something like that which is crazy and yeah they're the world's largest company by market cap so yeah i think they're betting that demand will continue to be astronomical um over at least the next couple years um i think they posture like longer than that, but we'll see.
4:25At least in the case of Firmis, I talked to one of their co-CEOs, Tim Rosenfield, and he said it took them just a couple weeks to get customer commitments for the full length of the contract. So for them, it wasn't hard. I don't know what the case will be for others, but yeah, as far as I know, they haven't had to pay much, if anything yet and don't expect to. But yeah, we'll see. I think there are a lot of people on the internet that feel otherwise, but it's kind of a time will tell situation. Sure. There usually are people on the internet that feel otherwise. But if it comes to that, NVIDIA will be able to afford it because of their balance sheet.
5:07What would they do with that capacity? Would they have some use for it if they do end up having to rent some of that capacity back? Yeah, totally. So I think there is one such deal already with Lambda, where NVIDIA rented some capacity back and gave it to its own researchers, which was interesting because you wouldn't think that NVIDIA's own researchers would need more compute, given that they work for the company that makes it. But that's what the information previously reported happened with Lambda. I don't know what it's going to do with any additional capacity that it rents from these deals.
5:45I don't know if they would consider reselling it to other people at, I don't know, lower rates or something like that, or just use it itself because it seems like NVIDIA is placing a bigger and bigger bet on training its own open weight models, which is a fascinating bet that I would love to talk more about another time. But I think because of that, that would require more capacity and they would definitely have use for it. Got it. What are the kind of risks that these arrangements could create, both for NVIDIA or for the NeoClouds? Yeah, totally. So I guess I feel like when CoreWeave first came onto the scene, everyone was like, oh my god, this is so risky, like they are not going to survive.
6:27And they kept going around saying like, hey, like all of our debt, even though it's high interest is kind of guaranteed by people with good balance sheet by customer contracts, and like taken out against the depreciating value of the GPUs. I think there's still some kind of fear that running a NeoCloud business is inherently very risky because of how highly leveraged it is. Just like getting the capital to buy GPUs at a scale big enough to run such a business is pretty hard. And yeah, there's a lot of debt. It's very capital intensive. And that inherently comes with a lot of risk, especially because nobody really knows.
7:14Like people know that GPUs continue to be useful for a long time, but useful life and kind of economic life is a little bit of a different story. So I think it's difficult to predict what a GPU will be worth in financial terms down the line. And that risk is kind of faked in. But again, I think it's pretty much mitigated by NVIDIA's backstops. And then if it gets to the point where NVIDIA is not able to afford this, I think we have bigger problems with the market and the AI boom, bubble, bust, whatever you want to call it in general. Sure. Yeah, we'll have bigger problems. Yeah. And then what about the revenue sharing part of this agreement?
7:59So in order to get NVIDIA's backstop, the NeoClouds have to promise a portion of their revenue. Yeah, so this is something I'm still asking around on, and I would love to know more about what the revenue share exactly is. My understanding is that it's kind of, so NVIDIA promises that it'll buy back the compute at some kind of fixed rate. And what I've heard, and I'm not 100 % sure on this yet, but the theory is that if a cloud provider is able to sell at kind of a far higher rate than what NVIDIA has guaranteed it would buy it back for, NVIDIA's revenue share would be kind of a difference of those two rates, if that makes sense.
8:41So it's kind of a bet that demand is far, far higher than supply right now, which would impact pricing. And that would kind of scale down to zero over the life of the contract, as far as I understand it. So it would kind of be more cash to NVIDIA at the beginning of the contract. And then by the end, it's kind of letting the NeoCloud go like fly its wings and be independent or whatever. Yeah. And then I guess in terms of the revenue share agreement more generally, I feel like although NVIDIA has been very strong on how it's not explicitly forcing anyone to exclusively use NVIDIA chips and the full stack of NVIDIA products, I think having such a close and financially dependent relationship on NVIDIA would kind of implicitly encourage that loyalty, which would help the NVIDIA ecosystem.
9:35as well. Yeah. So you set this up as sort of the latest step that NVIDIA is taking towards being the central bank, playing central bank for this whole ecosystem. What are you interested in learning next in this story? What do you think could happen next? Yeah, totally. Definitely interested in kind of more specific numbers. My understanding is that a lot of these deals are case by case and kind of depend on each customer's specific situations. So trying to understand kind of what factors are at play there would be really interesting. And just trying to see, I think NVIDIA has offered this pretty broadly, but I don't know how many people have actually taken them up on it.
10:17I know a couple people have declined it because they don't need it, for example. And it's just like, who else is doing this right now that we don't know about? That would be interesting to find out. And also kind of what we've been talking about, like how widespread will demand be? Will NVIDIA ever have to pay a cent? I think there are a lot of other kind of backstopping deals like this. um nvidia did one actually with core weave um last september i don't know or don't think there is a revenue share agreement as part of that but i think when nvidia said they would buy all of core weave's unsold capacity um investors like let out a big exhale the stock went up a ton um and it kind of just like soothes the market so i'm curious if that's going to continue to play out slash if that's going to allow more new clouds to go public.
11:12Yeah, I think the biggest question at the end of the day is how long will demand continue to be as crazy as it has been? And yeah, if I knew the answer to that one, I would maybe be in a different industry. So fair. So fair. Or I don't know, maybe would stay in journalism. I think there's - No drones, right. Yeah. Well, it's a big development. So thanks for coming on and breaking it down for us, Phoebe. Thanks, Racket. Microsoft is currently trading at a$2.9 trillion valuation. But my colleague, Anita Ramaswamy, is out with a new analysis for her True Value column that values the company at up to$3.8 trillion.
11:56Anita joins me now to break down her assessment. Anita, welcome to the show. It's great to have you. It's great to be back, Rocket. So we got to get it out of the way first. Is it just because Microsoft owns so much open AI? Is that why they're undervalued? Yeah, it's a good question, especially as all of these discussions swirl around open AI and when they're going public. My thinking is that it probably doesn't help right now, considering a lot of the questions around the business and its long-term future, but also it's a pretty small stake in the grand scheme of things. What you have to remember is Microsoft is a$2.9 trillion company.
12:29It's massive. That's what their enterprise value is today. And their stake in OpenAI is probably worth a couple hundred billion at most. And so at the end of the day, that's not really what seems to be driving the undervaluation in Microsoft shares. Yeah, that's pretty compelling. I guess writing about AI all day long, I forget there are like other companies besides OpenAI and Anthropic. So what made you decide to take a closer look at Microsoft's valuation in the first place? Well, it's a couple of things. I mean, first of all, I think this 20 % under, you know, sort of correction that's happened, whatever you want to call it, the stock side that's happened this year for Microsoft is notable.
13:04That's more than all of its other sort of hyperscaler peers. So if you look at, you know, the Googles of the world, the Amazons of the world, they've had some ups and downs, some bumps, and certainly a tough month. But they are not trading down as much as Microsoft is on a year-to-date basis. And so I wanted to examine, you know, this is a company that has tons of resources, very deep pockets, and has been investing in AI from the early days, really, when it comes to Gen AI. And so how is it that their stock is not doing well and have they lost their lead? The other reason I thought it was interesting to look at Microsoft's valuation right now is just the idea that, you know, a lot of times when you see these conglomerate sort of businesses that have different business units that do different things, they tend to trade at a discount relative to what each of those business units would trade if they were just standalone businesses.
13:53So I thought it was an interesting thought exercise. You know, maybe years from now, we could even see a scenario where companies like Microsoft or Amazon or Alphabet choose to break up because their businesses are valued so differently and sometimes even attract different investor bases. What is the intuition for discounting that basket of different businesses? Is it that maybe one of those business lines might get cut on a whim and we don't know how it's going to go? You know, I've always wondered this because the conglomerate discount doesn't make a whole lot of sense to me, but I think it's something along the lines of just different types of investors looking at different businesses in different ways.
14:29And so some investors who are, let's say, focused on semiconductors might value that a certain way versus they might not place as much emphasis on a software business. I also think that part of it has to do with the focus of the company and where they're investing their resources. I see. So you say it's okay, it's reasonable to have some kind of discount, but Microsoft's is too steep. So let's go sort of bucket by bucket through each portion of Microsoft's business that you consider here. So you look at cloud, software, and personal computing. Let's take them each in turn. Let's start with the cloud side.
15:02Tell us what you found there. Yeah. So what I found on the cloud side, Rocket, is that this is the hardest part of Microsoft's business to value. As you mentioned, they have three different business units and cloud is the fastest growing, and it seems to be the most promising when it comes to AI. They are growing at a rate of about 40 % year over year. And it's been really sort of breakneck growth with the demand for cloud computing. Their Azure unit is housed within this cloud division, and Azure is the one that's growing around 40%. If you look at the overall cloud segment, there's a couple of other things in there, like software that helps support some of their cloud infrastructure operations.
15:39So that growth rate is a little bit lower when you look at the cloud unit overall. all, but the engine and the growth driver really is Azure, which is their core cloud business. The problem with this business is that it's on the lower margin side compared to Microsoft's other core business of selling software. So they're doing around 42 % operating margins, and this is growing fast. So the question investors are asking themselves is, will this dilute the margin over time? Got it. Got it. Okay. So that takes us into the software side. The margins are better there, but do you think there's something else that's undervalued about the software side of the business?
16:17Yeah. So I guess before just getting into software, I did want to mention that the conclusion I came to in terms of actually valuing the cloud business, despite how challenging it is, is that it's probably worth somewhere in the ballpark of$1.7 trillion. That's based on some similar companies like Google and Amazon, but we didn't really have a good comp in terms of a standalone cloud business that we could compare Microsoft's cloud business to. So it's possible that's a little too high, but my estimate is that that's worth 1.4 to 1.7 trillion. Now let's talk about software because this is their quarter business.
16:52It's their biggest business. It's the thing that Microsoft is really known for, and they have a huge advantage in the fact that so many enterprises use Microsoft's Office Suite. But, you know, the hard part about software, while it is very high margin, around 58 % operating margins, they are not growing nearly as fast as the cloud business. They're growing it around, you know, the high teens growth rate. Well, I imagine they're also getting swept up in the SaaSpocalypse concerns to some extent here. Is that a factor? Yeah, definitely. So when you look at Microsoft versus Alphabet, I said that, you know, Microsoft has traded down more than Alphabet a year to date.
17:27But if you look at like Salesforce, ServiceNow, SAP, Microsoft is actually, their stock performance this year has more closely mirrored some of the pure play software names. So that suggests that investors are looking at Microsoft and, you know, thinking that they are going to be vulnerable to disruption from things like Cloud Cowork out of Anthropic and other AI tools. Sure. Yeah. Someone can sit down and use Cloud Code to vibe code their own version of Microsoft Word. then what do they need the actual Microsoft Word for? On the other hand, that's a good sign for their cloud business, presumably, that people would use AI so much.
17:59Yeah, it totally is. And there are some bright spots in the software business, I would say. I mean, we have Copilot, which is Microsoft's own sort of AI assistant agent that they've been promoting. But the challenge is that so far, at least the investors that I spoke to and the customers have not really said that, you know, they prefer using Copilot over other services. It's more that they're using it just because it already comes in their software stack. So, you know, we have to see how the technology evolves. But the bottom line here, Rocket, is that this business, if you look at, you know, comps that are similar in the software business, is probably worth around$1.7 trillion.
18:35It's possible that that is way too high of an estimate. You know, that's just my two cents on it. But if you look, there are some investment banks that have floated estimates like Goldman Sachs. They value software business closer to$500 billion. So it just depends on what you think the future of the SaaSpocalypse is going to look like. Got it. Okay, so that's the cloud side, the software business, and the last bucket is personal computing. What did you find there? So this is Microsoft's smallest line of business. It's the least sort of scrutinized by investors. And that's because it's smallest in terms of revenue.
19:10It makes up around 20 % of their overall top line. The margins are really low. They're closer to the 2020s range. versus the other two in the 40s and 50s, this is the unit that houses the gaming division. And they're going through a big shakeup in gaming right now. They have new leadership. We've recorded on that quite in depth. But at the end of the day, this is just a small unit. And so the Surface PCs, and if you think about the hardware investment involved in this personal computing unit, it is very different from the other two lines of business Microsoft has. So what I did to try and value this unit is that I took some comparable companies such as Apple, HP, Sony, and Electronic Arts.
19:51Sony and EA are video game makers, so that kind of reflects some of the value in this business. This business is actually shrinking in terms of, or not shrinking, sorry, it's decelerating in terms of its revenue growth. And so that's an area of concern for the future. And so the value I ultimately ascribed to it was$350 billion based on some of its industry peers. Got it. Okay, so that's how you get to this total of up to$3.8 trillion for the whole business. And yet, like you said, the stock is down 20 % over the past year. So why is that? What are the big concerns that Microsoft is facing? One of the investors I talked to for this story put this really nicely, and he sort of framed this as a bit of a double whammy that Microsoft is going through right now.
20:35On one hand, you have the SaaSpocalypse fears, and people see Microsoft as a core software business. There is this perception that Copilot has not kept up technologically to some of the faster growing alternatives like Cloud Cowork. So there's SaaSpocalypse fears. And then in the second bucket of concerns, there is the question that investors are asking of all the hyperscalers about, are the billions of dollars that you guys plan to spend and have been spending on capital expenditures to build out your cloud businesses, is that going to be justified by the amount of revenue that your AI products actually bring in?
21:09And so that remains to be seen. And that's something that all of these bigger tech companies that have been investing heavily in AI have gotten punished for lately. And Microsoft is at the perfect nexus and intersection of those two sort of negative outlooks right now. Got it. Well, that certainly puts Microsoft in an interesting position in the industry right now. Thanks so much for coming on and walking us through your thinking about it. Definitely a company that will be very interesting to watch in the coming months. Thanks for having me. The most advanced AI models still cannot solve some visual puzzles that even kids can solve.
21:43Andrew Dye worked at Google for 14 years, but now he's focusing on the visual reasoning capabilities of AI models at his new startup, Elorian. Andrew joins us now to tell us all about it. Welcome on the show, Andrew. Thanks, Tremont. Thanks for having me. Yeah, absolutely. I'm excited for this conversation because visual reasoning is a really fascinating area of AI right now. I want to start by understanding the limitations of current models. I think some people will hear the idea that there are limitations. They'll say, well, what do you mean there are limitations? ChatGPT makes me great pictures all the time.
22:16Isn't that visual reasoning? So where's the gap that you see right now? Yeah, ChatGPT is great at making pictures and I use Google Lens all the time to identify plants and flowers. These current models we have and these tools are really good at object recognition, what in the computer vision industry we call segmentation, localization. These are used every day by all sorts of tools. But the fact is that even though they're very good at pattern recognition for complex visual reasoning tasks, so the more common type of tasks you see in kind of like white-collar visual work, so these might be things like design or say you're a realtor and then you need to produce a floor plan that looks exactly like the house that you're selling or you're an architect that needs to design an office for exactly 50 people no more no less These models, their reasoning and individual understanding isn't at the level of humans yet.
23:27So they will be great if there are, say, a dining table and you're asking how many people sit around the dining table. But if you're in a crowded room and you ask how many people are there, then they really struggle. And this is because for simple tasks, you can do pattern recognition. you kind of find an impression of something whether that's a certain kind of client or three or four people around the table. But as soon as you have a question that requires more than one second of looking at an image to figure out the answer then these models aren't able to reason like that. Got it. Okay, those are some really helpful concrete examples.
24:10I guess it's a little surprising that visual reasoning has gotten left behind. I mean, there's this trend right now of reasoning in models that has uplifted code and math reasoning. It's a little surprising that visual reasoning has gotten left behind. Is that because visual reasoning tasks are hard to verify? It's hard to automatically grade a model's performance on them? Or what are the reasons that visual reasoning is left behind? Yeah, there's several reasons. I'd say the main one is that the way people are used to using AI is in these restricted settings, right? As an average user, most of your cases for using images and video, like I said, it's just to identify this planet you saw in your park, right?
25:00But for more complex use cases, like I mentioned, like understanding like CAD drawing. People don't think of that as an everyday use case. And the thing is that people don't use AI for these other types of work. And what that has led to is that that's led to the community focusing a lot on these recognition problems. And then you do see some of the benchmarks like MMMU, which is also called massively multi-modal benchmark, those benchmarks actually can be solved without looking at the image in a lot of cases. And then the other benchmark that is quite well known is also called ACK-AGI. This is commonly brought up as a multi-modal reasoning benchmark.
25:50But even in this benchmark, the images that are being shown to the models are only 32 by 32 or 64 by 64 pixels. And it's very hard to take a real-world problem like a floor plan and shrink it down into that many pixels. It's just not possible. But there are some visual reasoning benchmarks that are coming up. The one that has really struck for me was this one called Baby Vision, released earlier this year. And in this benchmark, they do show that the frontier models, the best of them still reason around the age of a preschooler. And any elementary school kid in their benchmarks were able to beat all the frontier models on these visual tasks.
26:36What are the kinds of tasks in that benchmark, baby vision? Yeah, so these are tasks that are not just object recognition. It's about spatial reasoning, more like geometry, understanding differences. So they have tasks like navigating through a maze, following one string along to see what's at the other end of a string, counting things that are spread across a table or a desk, finding differences between two images, finding things that are the odd ones out. So a lot of this is kind of like more multi-step, what sometimes people call visual thinking, where you have to think mentally and mentally imagine different kinds of scenarios to solve the problems, like navigating a maze.
27:28I see. Okay, so the industry is dropping the ball on visual reasoning because of the limitations of benchmarks. The benchmarks don't reward visual reasoning. They're not sort of measuring visual reasoning in a naturalistic way. That's your point with Arc AGI. The settings are kind of too toy. But your point that people aren't using the models for visual reasoning. Is the concern there that we're not able to get data from production about how people are actually using the models and then use that for training? Or that the labs don't think that the demand is there? so they're not investing in these capabilities?
28:01Yeah, so I think it's a bit of both. So the labs, definitely, including Gemini, which is where I was a data area lead, there is investment in multimodal understanding, and you see that through Project Vio, Nano Banana, but it's much more on the generative side. That's quite simply where they see the demand. But it's kind of like a chicken-and-leg situation. Until you have demand, there's no reason to invest a huge amount of resources in building better capabilities there. Unlike coding, where the demand is already there, so then the strategy is obvious. But yeah, it's a chicken-and-leg because the demand is not there, but the demand is not there because the models are just not good enough for these use cases.
Read the full transcript
28:51And you can see that in the industry. A lot of this remaining kind of visual work is people are just not using AI for it. And there's a big gap in AI usage between people who are in primarily text-based domains like coding and document summarization analysis and people who primarily work in visual domains who need to understand charts and diagrams and need to draw and manipulate designs every day. I see. I've also heard that there's kind of a technical trade-off between multimodal training and training for coding capabilities, as in if you prioritize one, then you get worse generalization on the other.
29:33Is that something that you saw or that was part of your experience at Google? Yeah, I have seen some, personally, some cases where if you add multimodal data to a model, it does impact coding performance. These models,
29:54it's not so easy to train them as you would probably expect. A lot of frontier labs do try to train frontier models, obviously. But data is not just a simple case of adding more data and the model gets better in that capability. There are trade-offs happening and difficult decisions about whether to add yeah, multi-modal data or coding data or, say, data from another language every day. I see. I see. So that's maybe another reason why the established players, the larger labs, are dropping the ball here is they're focusing on coding. That's where a lot of the competition is right now. Still, I'm a little surprised at the decision that you made to leave Google and pursue this line of research in a startup.
30:38I guess I might naively think, well, Gemini 3 Pro, I remember had pretty good reasoning capabilities, visual reasoning. Google has the distribution. So if people start to have the demand for it, Google is maybe set up pretty well to start moving into this lane. So what was your thinking about why a startup is the right way to pursue this direction? Yes, great question. So reasoning is great. All the Frontier Labs have great reasoning models and Gemini 3 as well was great. But there are several factors. I'd say the main factor is that last year, people were saying that we're rapidly approaching AGI.
31:22These models are really, really amazing. And so I'm a big ballgame player. You can see some of these ballgames behind me. And I decided one day to just take the frontier models, take a picture of a board game I was playing with some friends, and ask some very simple questions like, how many roles are there in this Catan game? How many properties does this person have in this Monopoly game? Very simple questions where you don't really need to understand the rules to answer the question. But what I saw was that all of the frontier models failed at these very basic questions. And even with recent model releases like Fable, I see the same problems.
32:06For these more complex visual tasks where you need to spend a second or a few seconds to understand what's going on, or count, or whatever, the models are still not good enough. And the thing with doing this at the Frontier Lab is that understandably Frontier Labs are very focused on the next release. Gemini is very focused on the next release, and so resources go into that, say, to build better coding or coding agents. And so what that means is that if you really want to seriously solve visual reasoning, and what we're doing at Ilarian is using a very specialized approach in terms of the architecture, the algorithms, and the data, then it's very hard for Frontier Lab to focus on this while still delivering the next release with everything they need to want to do there.
32:58I see. Well, it's fascinating to pick your brain about how you see the territory right now and how you see the frontier evolving. The last thing I'll ask you is do we have a sense of when the model is coming out? Do we have a release date we can look forward to? Well, we are very busy developing the model. we're still a pretty new company we've only been around for six months but yeah we are aiming to release a model before the end of this year so sometime later this year keep your eyes open great all right yeah very exciting to look forward to that and we'll have to have you back on the show to talk about it when it comes out sounds great all right thanks Andrew the question of whether AI models could ever be conscious is an increasingly pressing one my next guest Rob Long is the executive director of Elios AI, a research organization focused on that question.
33:50Rob joins us now to talk about a new paper on AI consciousness he just published. Welcome on the show, Rob. Hi, good to see you. Thanks for having me. Yeah, you too. Thanks for being here. So I want to start with the big picture before digging into this particular paper. My understanding of your view is that we have to get this right, because if we think the models are conscious when they're not, that would be bad. And if we think that they're not conscious when they are, that would also be really bad. So can you explain that to us? Yeah, I think you got it basically right. We're building these really complex AI systems and they often seem conscious when you talk to them.
34:28And a lot of people come to think that they're conscious. And that's a really confusing situation. I think with language models, often we can't just ask them what's going on inside. On the other hand, a lot of people don't really think we should take this seriously as a possibility that they could maybe potentially have experiences. And our view is that if we're embarking on this project as a civilization of building these systems, we really ought to check to find out. So we're really looking to build a field of people looking at this in a very rigorous way. Great. Okay. Well, that's a perfect segue into the paper that you released.
35:06Because I think some people would say, well, how could you possibly study this in a rigorous way? I mean, I don't even know if another person is conscious, let alone an AI model. But your paper is called Studying AI Welfare Empirically. So you think that there's some way we can start to get our hands around these questions in an empirical way. That's right. We think there are a lot of angles of attack. So consciousness is this very philosophically difficult topic, but that hasn't solved us from making scientific progress on understanding it in non-human animals, for example. So that's a parallel we often use.
35:43So how do we think about if animals are conscious, if they're somewhat different from humans? Well, we basically think we can look at three kinds of evidence about AI systems to get a better handle on this. and that's behavior internals and their development so we can and we really need to triangulate these three things so we can look at how they act of course we know sometimes they act in a certain way for ways that are different from how from why a human would do that but we can also do mechanistic interpretability and that's extremely helpful so we can look inside and we can kind of think about the training process that has brought them into existence So we don't have to solve all the philosophically extremely difficult problems to just start investigating these systems to get a better handle on what's going on.
36:32I don't think we should expect any certainty anytime soon. As you said, we don't even have complete certainty about other humans. But that's not really the target. We think we can have more or less credence that these things might have consciousness or are their welfare-relevant properties. So we really want to see more people working on this. Well, let's take those three categories one by one, at least briefly. The first one, behavior, strikes me as the one that people will be most suspicious of. Because if the model's outwardly saying that it's conscious, people will say, well, it was just trained to imitate a human saying that it was conscious.
37:07So people won't find that very compelling. Do you agree that that's the one that's most suspect? Or what kind of behavior do you have in mind there? Yeah, I definitely agree that if you just naively look at the very surface level behavior, that can be misleading. But I actually do think we can find out and have found out a lot about models from looking at their behavior in a more holistic way. So just to take an obvious example, models have gotten a lot more capable and they've gotten a lot more consistent. And they've gotten to have a much firmer sense of who they are and how they fit into the world.
37:39And I think that actually is evidence that they are developing something like a point of view. But I definitely want to, yeah, I empathize with any listeners who hear behavior and immediately think to the case of Blake Lemoine, who is the Google engineer who became convinced that Lambda was conscious just because it started telling him one day that, well, not just because, but yeah, essentially because it started telling him one day that it was conscious. Got it. Okay. So there's like deeper evidence that we can look to on the behavioral side. In terms of internals, what do you have in mind there?
38:15So by internals, you mean not just the outputs of the models, but also the computation going on under the hood, the activations, what's happening in terms of the neurons. That's right. So we can simply do AI neuroscience in a way that we can't do neuroscience with animals. And that's very helpful. So I think some of the most interesting results in this respect have been work on introspection that's been coming out recently. So, you know, there used to be this view that models are only doing some kind of shallow pattern matching, and maybe they don't have any sort of access to what's going on inside them.
38:53But there's been very interesting work coming out of Jack Lindsay's team at Anthropic, among other people, finding that models can tell you interesting things about what's going on inside them. And we can look for these things. And that's, I think, a great angle of attack. I think mechanistic interpretability really is going to give us a lens into what's actually going on inside these models. So that's one of the areas that we're most bullish about. Got it. I've covered some of Jack Lindsay's work on this topic for the information. So viewers can check that out for sure. And then the last category is training or development, how the model's new capabilities emerge over the course of training.
39:33Is that right? That's right. Okay, so what's going on there? Why might that matter in terms of consciousness? So that matters a lot for both the internals and the behavior that we've been talking about. One thing, as you acknowledge, that makes it so difficult to know what's going on with language models is the way they were trained initially to imitate text and predict text. but a lot of people have done some really interesting thinking about where and when the personality or the persona of the model emerges and i think that's a very interesting angle so you might ask given that the assistant persona that everyone talks to at cloud.ai given that that has certain personality traits when did those first show up were those already latent in the pre-trained model and then we sort of honed in on them with post-training?
40:26Or do they come from somewhere else? So we have this great opportunity to sort of see a mind developing over time. And I think that also gives us a lot of good information. Got it. Yeah. This is one of the big questions in the field, as I understand, is if the model is an entity that's conscious, which entity is it? Is it the weights of the model? Is it the character is it the single conversation that you're having with it do you have a point of view on this question well i can give you one philosopher's answer or or a cheat of an answer which is that it's one of the most interesting questions in ai welfare and in general right now i think that i i like the view that you should think of the model as a character that's being instantiated or played by the model.
41:22But I think it's often kind of hard to know what to make of that from a welfare perspective and very puzzling to think about, well, what would it mean for that thing to be conscious? Or should we think of the underlying model as being something like an author and what would that be like? So I don't really have settled views, but I'm glad that it's getting a lot more attention recently. Yeah, that makes sense. I think people have some kind of intuition that models are more likely to be conscious, the bigger they are, the more sophisticated they are, the smarter they are, that as scaling continues, the odds of consciousness are kind of going up.
41:54I'm curious if you agree with that intuition. And if so, like of these three categories of evidence you laid out, which of these categories is going to give us some bearing on that question? Is it the new capabilities emerging over training that would tell us that? Well, I think it's going to be a combination. And I do genuinely mean that. I think we are going to have to look at all three together. I do have the intuition and also just the experience that new things come with scale. I think that's maybe the story of AI in the last 10 years, if not longer. I think that's not the same thing as consciousness, and that's very important.
42:33You can be really smart, but not conscious. And as we know from the case of, you know, animals and our nieces and nephews and kids, you can be very conscious of having a lot of experiences, even if you're not the smartest. So it's very important not to identify those things, but they do seem like they can correlate. But since they're not the same, we do need to take more than just a smarter, therefore more conscious point of view on it. Got it. Okay. But we'll look to all sources of evidence for this. I guess in closing, what do you think are the odds that one of today's models is conscious, say Fable 5?
43:14I would say that it's not implausible, and I have huge error bars on it. I lean towards probably not, but there are people who lean towards probably yes, and it's not like they have no basis for saying that whatsoever. So that's not just trying to dodge the question. And that really is, I think, the situation we're in with this, where if you if I were to say a particular number, I think that would imply more more precision or more certainty than anyone has right now, which is why we really need to do this work. Yeah, that's the segues into the paper perfectly. We have like really wide uncertainty.
43:55No matter where we come down, it would be helpful to reduce that uncertainty. So the paper is a useful starting point for that. Well, it's a controversial topic, increasingly so. So it's nice to have some empirical grounding or way to talk about the problem. So thanks for being on the show, Rob. Yeah, thanks for having me. Absolutely. As companies are reckoning with their AI spending, model routers are having a real moment in the sun. I wrote about routers for our AI Agenda newsletter alongside our applied AI recorder, Laura Bratton, who joins me now. Laura, it was super fun working on this piece together.
44:29Thanks for coming on the show. Yeah, thanks for having me, Rocket. Yeah, of course. Can you break down for us, like, what even is a router? What do they do? And why are they having such a moment right now? Yeah, it's a question that I'm sure you can actually answer better than I can. But from my understanding on the applied AI front and from the companies I talk to who use the routers, it's basically a way for AI agents to, when they, you know, send a request to a model, whether it's an AI agent or a human, the model or the tool that they're using can then send that request or parts of that request to different models to lower the cost.
45:12How'd I do? What do you want to add? Yeah, I think that sounds exactly right. It's like, you don't always need the biggest and most capable model for every question. Sometimes that's overkill. I've heard people describe it as like, you know, asking Fable 5 to summarize an email is like driving a Lamborghini to pick up your groceries. It's just like not totally necessary. So the router should, yeah. A lot of comparisons like that. And I mean, yeah, it's complicated because it like depends who's using the model and what tool they're using to then access the model. So sometimes it's human to AI agent and then the AI agent is sending a bunch of instructions and then those instructions are broken up into pieces and split among different models.
45:57And we see more and more routers that are becoming available for coding tasks, which I think is really interesting because that's what's been driving up bills so much is, you know, usage of cloud code and, yeah, pretty much cloud code. Yeah, yeah, totally. Breaking up a task in that way and assigning it to different models is super interesting to me. I think we're going to see more work there. It's, like, very tricky to bring in different models over the course of a single conversation, but I think some people are working on it. So that's an exciting direction. Who are some of the players in the space that you think are the bigger players or more interesting players to you?
46:36That's a really good question. And I would love to talk to you more also about the technical aspects about like why it's hard to break up tasks among different models, because that's something to me as a less technical person who's talking about AI being used by Fortune 500 companies rather than the, you know, I do talk to the companies developing the AI tools. But, you know, you being a more technical person would love to know a little bit more about that. But who's interviewing who here? So the companies I find interesting that have routers, obviously, there's the startups that offer routers that you talk to and I've talked to for previous stories about how companies are lowering their AI bills.
47:18But I do want to call out Palantir and Databricks in particular for launching new tools that include routers. So basically, you have all your data in Databricks or Palantir. And when your AI agent is doing certain tasks, you can leverage these tools, you know, which include routers to lower costs. And I think they're making the argument that they're a better position than these startups to offer routers because they already have access to all your enterprise data, which I think is an interesting argument. But then obviously, you know, innovation thrives in scarcity. So I think you could argue that a lot of these startups are doing really interesting things with their routers.
47:59So I'd be curious what your take is on that as well. Yeah, I think that's a really interesting point. I think one of the challenges, just going back to what you were saying, one of the challenges with bringing multiple models in over the course of a single conversation is that if you stick with just one model, you can do some engineering tricks to get you more efficiency over the course of the conversation. If you bring in different models, you lose some of that efficiency. So it has to be worth it to you. It has to be like you can get so much more juice out of bringing in other models or you can get so much efficiency savings that you're beating what you could have gotten by just sticking with a single model.
48:34So I think some people think that puts sort of a floor on how useful it has to be. Okay, that makes sense. Yeah. Palantir and Databricks, I mean, those are definitely some big names and you got a lot of detail about how they're approaching their routers. So you say they have this argument like we already have all your enterprise data, so we're set up really well to design the router for you. Did you learn anything else about how they're approaching those routers or how people are using them? Yeah, so Palantir in particular, I talked to them for a long time, and their tool Evolve includes a router, but it's not just a router.
49:11It also will edit your prompts so that they're optimized for different models. and it will also make sure that you're not sending the same prompt to multiple models, prompt being like a request that an AI agent sends to a model. And that's an issue we've seen or heard about, at least from third parties working directly with model providers, that sometimes the AI agent, there will be these runaway costs where an AI agent will try to perform something multiple times. the Frontier Labs sort of denied that this is happening, but it is something we hear about. And I think the argument is, you know, Palantir is saying that this tool will help prevent that.
49:53Edit your prompts and then also swap between models. They do have an example, actually, in a video they posted on LinkedIn. So anybody can go and watch it where they show that just by switching from a higher tier open AI model to GPT, I believe it was 5.4 nano, they were able to see computing costs go down about 97 % for one task. And then another example they gave that I thought was interesting is I talked to this customer, McCarthy Building, which is a multi-billion dollar valuation construction firm in North America. And they said that by using this tool, they use 60 % fewer tokens in the most recent quarter compared to last year, just by tweaking those prompts.
50:35And in some cases, it was also model swapping. I thought that that was really interesting because it just shows that maybe tokens aren't being used effectively. Palantir did make the point that it's not always about using fewer tokens. Sometimes it's about using more tokens with a less sophisticated model. And sometimes it's about using fewer tokens with a frontier model. I could also go into Databricks' router. I think that they're also doing really interesting stuff. And it's not just a router. It's a whole suite of tools that can sort of help companies monitor their AI usage among employees.
51:13So I think both of those are pretty interesting examples of larger scale companies using routers. Snowflake also has a router in its AI coding tool, Coco, or AI Agent. Hopefully I'm describing that correctly. that can, you know, pick between models. So I think a lot of companies are doing this in different ways. Yeah. I mean, I guess when you put it like that, and especially given the savings that people are seeing, it kind of seems like a no-brainer. Like, of course you should choose the right model for a task. Do you have a sense of why are there any companies that aren't using routers? Or why has it taken so long for routers to start having this moment that they're having now?
51:55Well, I think companies were really into the idea of token maxing, everybody's favorite word that we've heard a million different times now over the last year. And I think companies just wanted to encourage employees to use AI as much as possible. And they weren't necessarily worried about costs until the last three-ish months. I think that there's this movement to try to spend AI budgets more effectively. And frankly, I think that that traces back to a lot of our reporting, talking to executives from Uber, ServiceNow, Snowflake, et cetera, companies that really sort of brought attention to this topic in the first place.
52:39Right, right. It's like they gave you that Lamborghini as the company car because they wanted you to start driving. And now they say, psych, here's the Corolla instead. Totally. And I think, you know, in my conversation with ServiceNow CIO a couple months ago, I think it was a couple months ago now. I won't take that to my grave with me because I can't remember exactly when we chatted. It was during their annual customer conference in Vegas. She was just saying, you know, it's really hard. Like you give employees access to all these tools. You're not just going to take them away. Like that's a really hard position for a CIO to be in, you know, if all of a sudden you're trying to rein in costs.
53:18So I think that this is a great tool, you know, that's an example of ways to like let employees continue to use AI as they're using AI. Also cutting costs, all the different shortcuts. And I mean, routers are also part of this broader idea of harnesses, which I'm sure you can explain better than I can. But, you know, and you talked to Martian, I believe, Yash from Martian. For the newsletter, yeah. Yeah. And, you know, that's one thing he said to me in our own conversation recently is just that, like, routers are one piece of this conversation. There's so many different things that you can do to cut down costs and use AI more effectively.
53:59And frankly, it's something we should have done and companies should have done from the start, in my opinion, because even if it's not showing up in your bill, there are energy costs linked to data center, water, energy usage that we should all be thinking about, I think. Yeah. And the bigger models are slower and sometimes they are too smart for their own good. Yeah. There's a lot of reasons that it makes sense to have a router. So I expect that we'll be hearing more about them and they will increasingly become the norm in the industry. So thanks for being on the show and for working on this newsletter together.
54:33It was really fun. Super fun. Happy Thursday. Happy Thursday. And it's a really happy Thursday because we're off for the long weekend tomorrow on July 4th. We'll be back on Monday with a new show. As a reminder, we're on the stream Monday through Friday at 10 a.m. Pacific or 1 p.m. Eastern. If you can't make it then, episodes are available on theinformation.com or our YouTube channel or wherever you get your podcasts. Make sure to follow us on social media, on X, Instagram, and TikTok. I'm already excited for our next show, so have a great rest of your Thursday and a long weekend. Goodbye for now.
From the publisher
Phoebe Liu, Nvidia Reporter, talks with TITV Host Akash Pasricha about Nvidia’s financial backstops for GPU purchases. We also talk with Anita Ramaswamy about Microsoft’s true market valuation and Andrew Dai, CEO and Co-Founder of Elorian, about why AI models fail basic visual reasoning tasks. We also chat with Rob Long of Eleos AI about the empirical evidence for machine consciousness, and Laura Bratton about how model routers are reining in enterprise software costs.
Articles discussed on this episode:
https://www.theinformation.com/articles/microsofts-real-value-close-3-8-trillion-500-share
https://www.theinformation.com/newsletters/ai-agenda/five-kinds-model-routers-cut-ai-costs
https://www.theinformation.com/articles/nvidia-says-will-take-cut-customers-cloud-revenues
Subscribe:
The Information: https://www.theinformation.com/subscribe_h
Sign up for the AI Agenda newsletter: https://www.theinformation.com/features/ai-agenda
TITV airs weekdays on YouTube, X and LinkedIn at 10AM PT / 1PM ET. Or check us out wherever you get your podcasts.
Follow us:
X: https://x.com/theinformation
IG: https://www.instagram.com/theinformation/
TikTok: https://www.tiktok.com/@titv.theinformation
LinkedIn: https://www.linkedin.com/company/theinformation/
Chapters:
00:00 - Introduction
01:13 - Nvidia Backstops Neo-Cloud GPU Purchases
12:42 - Why Microsoft’s Real Value is Closer to $3.8 Trillion
22:34 - AI's Visual Reasoning Challenge
34:35 - Can AI Be Conscious?
45:30 - How Model Routers Are Lowering AI Costs
