Inside Oura’s IPO Postpone, Anthropic’s AI Discount Cuts, Anthropic Launches Claude Sonnet 5.5

29 Sep 2026 · 44 min · 18 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode covers: Oura postponing its IPO due to uncertain IPO-market demand; implications for Anthropic amid AI financing uncertainty; AMD’s $8.2B acquisition of World Labs to push “physical AI”; Anthropic cutting usage-based discounts (about 15% off list) once customers hit caps, opening space for OpenAI; Claude Sonnet 5.5 release; and model benchmarking and AI travel booking.

Guests and backgrounds

Martin Pierce, Information co-executive editor (wrote an IPO analysis column). Sriram Viswanathan, founding managing partner at Celesta Capital (chip/AI investing). Laura Bratton, author of The Information’s AI Agenda applied AI newsletter. Ryan Krishnan, co-founder/CEO of Vals AI (AI benchmarking). Steve Hafner, CEO/co-founder of Lola (AI travel assistant; formerly OpenTable and Kayak CEO).

Key claims/examples

IPO delays are market-wide (rising rates, weak demand), not Oura-specific; Anthropic may delay to manage fundraising confidence. AMD buys World Labs for physical-AI architecture and talent, including Fei-Fei Li. Anthropic’s discount-cap change may drive switches (CodeRabbit; Replit moved to OpenAI citing costs). Vals AI says Claude Sonnet 5.5 ranks #2 on its private benchmark and is frontier-capable but token-hungry; it predicts Anthropic could match human researchers on RSI by Aug 2027. Lola grounds LLM recommendations with real-time inventory APIs; it launches with Booking Holdings inventory and partners (e.g., SeatGeek, GetYourGuide, ResX) and uses human concierge via 10 Group.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Oura's IPO Postponement and Market Conditions

1:10 to 3:59

Discussion on the postponement of Oura's IPO and the broader market implications.

“It's going to be a fun show, so let's get right on into it.”

Anthropic's Funding Challenges and AI Market Impact

3:59 to 6:34

Exploration of Anthropic's situation in the AI funding landscape and its effects.

“There's a lot of uncertainty in the entire AI financing market that will affect their ability to grow.”

AMD's Acquisition of World Labs Explained

6:34 to 8:15

Analysis of AMD's acquisition of World Labs and its implications for AI.

“That is Martin Pierce, our co-executive editor here at The Information.”

Physical AI and Industry Shifts

8:15 to 14:00

In-depth look at the evolution of physical AI and AMD's strategic positioning.

“Can you just remind us of her body of work and, you know, how much of a catch this is for AMD to bring her on board?”

Data Center Infrastructure and Software Opportunities

14:00 to 15:42

Discusses challenges and opportunities in the data center infrastructure and AI space.

“we see incredible number of companies that are solving real data center infrastructure problem in the traditional data center space.”

Anthropic's Pricing Strategy and Market Competition

15:42 to 16:40

Analysis of Anthropic's pricing changes and its impact on competitors like OpenAI.

“Well, Shriram, I want to thank you for coming on.”

Customer Reactions to Pricing Changes

16:40 to 21:40

Explores customer responses to Anthropic's pricing adjustments and the competitive landscape.

“And, you know, more than 1 ,000 firms spent over a million each.”

Introducing Val's AI and Model Benchmarking

21:40 to 22:48

Introduction to Val's AI and their approach to benchmarking AI models.

“It's been another busy week of model news.”

Model Performance and Competitive Analysis

22:48 to 27:24

Discussion on the performance of AI models, including Sonnet 5.5 and Astra, and their market implications.

“So even though it's supposed to be more of a mid-tier model, it's actually competing at the frontier.”

Evaluating Safety and Future of AI Models

27:24 to 28:04

Examination of safety concerns surrounding AI models and implications for future developments.

“what the eval is that they're testing it against, like what the rubric is essentially that they're ranking the model on safety-wise.”
Show all 18 chapters

Anthropic's Research on Recursive Self-Improvement

28:04 to 29:11

Learn about the research predicting when Anthropic's AI models might match human capabilities.

“I want to ask you about something you wrote here.”

Concerns About AI Development and Risks

29:11 to 30:34

Discuss the implications and potential dangers of AI reaching human-level capabilities.

“There's a similar extrapolation done with OpenAI models, but I think the timeline for them to achieve RSI is a little bit longer.”

Auditing AI Models and Ensuring Independence

30:34 to 32:49

Explore the complexities of auditing AI models and the importance of independence.

“all this talk about pacing the frontier, like, I mean, what I'm, the read I'm getting from you is This is like, we're talking about months here.”

Challenges in AI Model Auditing

32:49 to 34:33

Understand the historical failures in auditing and the need for robust structures.

“that people are so worried about with these auditing organizations, the question around, well, what is true independence?”

Overview of Lola AI Assistant

34:56 to 37:10

Discover what Lola is and how it integrates various services for users.

“We use natural language on the interface to help our members discover and book everything that's worth doing.”

Challenges Faced by Lola and LLMs

37:10 to 39:26

Learn about the challenges Lola faced and the importance of accurate data.

“weekend, but we also make it very easy to book it.”

Building Relationships with Travel Companies

39:26 to 42:04

Discuss how AI agents like Lola maintain relationships with travel companies.

“out$408, but thankfully the hotel staff were nice about it.”

Building Relationships for Bookings

42:04 to 43:33

Learn about the importance of relationships and customer support in booking services.

“to get those restaurant relationships, to get those hotel relationships, to provide great customer support.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:13Welcome everyone to the Informations TI TV. My name is Akash Pasricha. It is Tuesday, September 29th. We are tracking President Trump's meeting with top tech execs in Washington today. AI safety is likely to be on the agenda. We'll have more coverage for you on our website. Today on the show, Aura is postponing its IPO, citing market conditions. We'll talk about the conversations that could be happening among prospective investors behind the scenes. We're also breaking down AMD's World Labs acquisition for$8.2 billion. We also have exclusive reporting for you around Anthropik's more aggressive stance on pricing.

0:54and we're going to end the show with a conversation with two separate founders. The co-founder of Vals AI, an AI benchmarking company, and the co-founder of Kayak, who is now starting another company. Lola is an AI travel booking assistant. It's going to be a fun show, so let's get right on into it. Aura announced it is postponing its IPO, which was set for Wednesday. The company says there was strong demand but cited uncertainty in the IPO market. I want to bring on the information's co-executive editor, Martin Pierce, who published a column on the questions facing the company's IPO over the weekend.

1:31Martin, welcome back to the show. It's great to have you here. Akash, how are you? I'm good. Look, you wrote one column and they called it off. You broke it. I would love to take credit for this, but unfortunately, I don't think I can. All you did is reflect the questions that people might have about this IPO. So walk us through. What's... Akash. Akash. I don't think, I honestly think that this hasn't got anything to do with Aura as much as we would like to think it has. This is at least the third or fourth company to have delayed its IPO in the last week. All of them have cited the uncertain market demand.

2:14The other companies are mostly outside of tech, but they include this metals company called Amira, I think it's pronounced. There's a Holtec Nuclear Energy. Now you've got Aura. We don't really know about SB Energy, which is the soft bank backed energy company that was trying to develop or wants to develop data centers. They had filed to go public on September the 1st, but their paperwork is not advanced. So, you know, I think there's, as I said, There's probably at least three or four companies in this position, and I really don't think it's the companies. I think it is the overall market. Holtec, when it announced it was delaying its IPO, it laid out a whole series of issues, and they relate to the war in the Middle East, most obviously rising interest rates.

3:14I mean, you've got right now, you've got the long bond yield is hitting levels that hasn't hit since 2002. And as that increases, the value of stocks goes down. So this is really not an ideal time for companies to be going public. So, I mean, I guess we don't know then how long it'll be before these companies. No, but really the issue is what does this mean for Anthropik? Right. Well, that was my next question. Yes, that's the most important issue right now. Well, they have to delay because really they're obviously trying to raise a huge amount of money. There's a lot of uncertainty in the entire AI financing market that will affect their ability to grow.

4:10So you really couldn't blame them at this point, particularly as they're spending half of their time warning that AI, the product that they make, could kill all of us. You couldn't blame them if they said, maybe this is not the right time to go public. Maybe you should just wait a little while. So I think that's something that we have to wait and see. Yeah. So, I mean, this all makes sense because I have to tell you, you and I, when Aura's S1 came out, you and I both looked at it. And I mean, I want to ask - Now you're coming back to Aura. You know - Coming back to Aura. That's what the news is.

4:47Don't obsess about Aura. I mean, people obsess about Aura a lot. I don't think we should obsess that it's an Aura-specific thing. I really think this is a - What is most interesting is that this is a market-wide issue, which has significance for all sorts of companies. And that's the thing that we should really focus on. Right. Okay. Let me just add you then about Anthropic. I mean, if Anthropic does delay, I mean, let's just talk a little bit about the repercussions here on the rest of the AI story, given the size of the capital is trying to raise and the effect on the AI story overall. I mean, let's say it pushes, I don't know, six months.

5:31What impact does that have then on the rest of the AI story? That's a good question. I mean, it's coming at the same time that the cost of raising money for all sorts of AI-related things is increasing. If Anthropic has to delay, and look, as far as I know, they've got a lot of cash in the bank, so they're probably going to be fine. But I mean, look, I don't know. Maybe they have to put off some projects. I'm not sure. I think the bigger question is not so much on them, but on investor confidence in the AI world. It's really hard to say, but I don't think it's a positive sign right now. Right, right.

6:17And I mean, it's more time for investors to be nervous. There's more time for these essays to stoke anxiety about the future of the sector. So, okay, well, fine. We won't talk about Aura. But thank you for coming on, Martin. That is Martin Pierce, our co-executive editor here at The Information. AMD is buying World Labs, the AI model company focused on the physical world for$8.2 billion. It is the latest in a streak of M &A for the chip sector. I want to bring on Sriram Viswanathan, founding managing partner at Celesta Capital for his thoughts on it all. Shriram, welcome back to the show. It's great to have you here.

7:00Thank you for having me. Okay, so let's talk about this deal. Why do you think AMD is making such a big purchase here? Well, you know, physical AI is the next big frontier after, you know, what has happened in the generative AI space with text and predictive, you know, models and all of that. The physical AI space requires a grounds-up architecture. I mean, this is really because, you know what you and i are used to in sitting in front of a computer or on your mobile phone and interacting with a plot or whatever is very different than what you have to do when you're interacting with a drone or a robot or or any kind of physical infrastructure it needs to be you know very low latency it needs to be you know high compute it needs to actually support sort of the physical aspects of the world that we live in, AMD does not have any of that capability today.

7:56So to me, I think this is a natural for them to sort of evolve into, to try and create a pole position and potentially try to match, if not get ahead of FNVIDIA. Right. So let's come back to the physical world thing in a minute here. I mean, Fei-Fei Li is an esteemed figure in the AI research community. Can you just remind us of her body of work and, you know, how much of a catch this is for AMD to bring her on board? Well, you know, I mean, look, I think there are some pivotal moments in the whole evolution of AI itself, as we're all aware. And, you know, people talk about Jeff Hinton and Ayanna Kuhn and, you know, Yasuo Benjio and folks like that as being pioneers in this area.

8:42And in my view, you know, Fei-Fei Li is equally, if not, you know, one of the leading pioneers that started the whole thing. you know, remember this happened as a result of her work in ImageNet, which led to, you know, AlexNet, which led to sort of the back propagation work that got prominent. I mean, of course, you know, Jeff Hinton worked on it for a long time, much before that. But to me, I think the whole idea of image search and image sort of categorization and sort of extracting intelligence out of it was really pioneered by Fefe Li. And to me, I think, you know, she was the pivotal transition point for the traditional text-based intelligence getting implemented in various things that we use to image-based and multimodal applications.

9:36And I think, you know, she's immense in terms of her prowess and her research work is phenomenal. So, I mean, how much of this do you think is an acqui-hire for AMD and how much of this is actually them wanting the World Labs technology? Well, you know, it's hard for any outsider to really, you know, understand that. And that's something that Lisa and Faye Faye Lee probably, you know, talked about and did a handshake on. I have to believe that the work that she's got out of the world labs is pretty substantial. And I don't know what their IP position is and all of that. But there has to be an element of sort of attracting superstars.

10:21Right now, we're in a place where sort of think about what happened with DeepMind. And to me, I think this is akin to what Google did with DeepMind is how I see it for the physical world. Wow. So, I mean, this, look, if what you suggest is true, this is a deep mind moment for AMD. I mean, this could be a huge inflection point then for them. And I'm thinking about NVIDIA making the Hugging Face Hack with all of the work that they're doing, trying to get into stimulating the model environment, you know, this being sort of a robotics play here. So, I mean, this could really be, you know, an inflection point for the company then.

11:08Yeah, I mean, look at it. I think the key thing here is that there are lots and lots of people that are actually working in this area. I mean, we have companies that we have invested in. You probably have heard of them. You know, Bellora is one of our companies that's actually very actively looking at building the infrastructure for physical AI and robots and drones and, you know, medical robotic sort of environments and such. So you're going to see a lot of people go into this space. I think what's interesting here is that getting to see the workload early for AMD is really important because that helps them define the architectural direction of what they have to build.

11:48And it's not dissimilar to what, you know, NVIDIA has done with Hugging Face. And it also gives them an opportunity to sort of close the gap from a software standpoint with respect to NVIDIA. As you know, CUDA really is the moat for NVIDIA. This potentially can create something similar for AMD from a physical AI standpoint. And last but not least, as you mentioned, there's no price you can put on really hiring this incredible talent. And Faithfully comes in with that background, and it's going to be huge. So let's talk about the M &A environment, broadly speaking. I mean, we know the deals that NVIDIA has done.

12:33AMD, I mean, they've done a number of acquisitions before this as well. This is just the latest for them. M &A is heating up. This is very much the story right now. So where do you expect to see more transactions here from the chip companies? I mean, you know, NVIDIA, we've talked on the show about energy being one area they might look to make acquisitions in. Where else are you thinking? Yeah, well, look, I think, you know, as we have seen, the constraint in the entire chain of AI workload, you know, continues to move. I mean, it is initially GPU, and then there's power, and then there's networking, and then there's, you know, chip-to-chip interconnect.

13:17and there's all kinds of constraints that exist. And you would naturally expect some of these large players to look for point acquisition that actually moves them up the stack. I mean, it's not very complicated because if you look at the model guys, they look for greater control and their control of their destiny, if you will, from a chip architecture. So they have their own sort of development plans and all of that. Now, if you look at the chip guys, they're obviously going up the stack, if you will, in terms of getting, you know, creating into the model world. And that's why you can explain hugging face and the rest of it.

13:58I think from where I sit, you know, we see incredible number of companies that are solving real data center infrastructure problem in the traditional data center space. And then comparably, there are interesting opportunities for the physical AI space. Like for instance, you can think about, the memory architecture is very inefficient. The power and thermal issues are inefficient. Safety and security is a big, big opportunity for gap-filling type acquisitions to occur. So those are some of the areas. I guess what I'm wondering is, Broadcom, Intel, Qualcomm even, I mean, are these companies going to also move up the stack with more software acquisitions, do you think?

14:45I mean, it's hard to tell, but I think I can only say that the space is moving so fast that the modes are getting redefined, right? I mean, CUDA as a mode for NVIDIA, you know, was great. And it's hard for me to see that being sustained advantage for NVIDIA in the long term. So they have to do something, you know, to sort of continue to establish their pole position. everybody else is you know it's like as they say you know you rob the bank because that's where the money is and everybody's going to hit on this space and they've seen the margin profile and everybody wants a piece of it so they're going to sort of you know complement their offering with software in a very big way they're going to you know they're going to worry about power they're going to worry about networking worry about you know interconnects these are all areas where you know we are very active in as you know and and you know not just us the entire industry is focusing on it.

15:40I would expect some moves over there. Great. Well, Shriram, I want to thank you for coming on. That is Shriram Biswanathan from Celesta Capital here on TIATV. Anthropic is taking a more aggressive tone to its pricing, which has made an opening for OpenAI to take advantage of this moment and win over customers. My colleague Kevin McLaughlin wrote that story with an assist from Laura Bratton. I want to bring on Laura to hear more about what they found. Laura, welcome back to the show. It's great to have you here. Hey, Akash. Okay, so the headline here is that Anthropic is cutting off discounts when customers hit their cap.

16:16So that seems to be pretty straightforward here on usage limits. I mean, how big were these discounts in the first place? The discounts were around 15 % off of listed model prices, is what my colleague Kevin reported. And that's pretty substantial if you think about the fact that more than 100 firms spent over$10 million each with Anthropik in the 12 months through June. And, you know, more than 1 ,000 firms spent over a million each. And why are they removing the discounts now? What do we know about that? Well, I think it's important to consider the fact that Anthropik's IPO is coming up. They have these substantial compute commitments and, you know, maybe under some financial pressure as they think about what their S1 filing is going to look like.

17:08But, you know, and how they're going to pitch in themselves to customers on their roadshow. So I think that, you know, on the flip side, we could think of this as Anthropics pricing power and the fact that, you know, they don't necessarily have to keep offering these discounts to customers after a certain point in order to keep them, or at least they might believe that right now. And we'll see if that holds true. Are customers switching to OpenAI or how are they reacting? Well, we do have evidence of at least one customer, CodeRabbit, switching to OpenAI. and we've seen other companies sort of publicly say, OpenAI is cheaper and more flexible.

17:48And so we're gonna switch. For example, the startup Replit shifted to OpenAI publicly in kind of a substantial way this summer, citing model costs. And that's interesting because earlier in the summer, they'd done a deeper integration with Anthropic and Claude. So I think that we are seeing some switching beginning to happen. I think this is kind of interesting, Laura, because, you know, as you mentioned, the IPO is coming out, coming up. Meanwhile, there are rivals coming out with competitive products. I mean, we already know Meta is making a big play here. And so I kind of see this as a bit of a hard game for Anthropic because once you change the pricing once, you can't keep changing it really.

18:40I mean, you can offer discounts and stuff like that. But to a certain extent, if you keep changing it, then I don't know. How can customers have any predictability with it, right? They might just switch just for that reason. Well, I think it's not necessarily the prices themselves changing, but the terms around discounting. And I think that that's interesting because, you know, we see companies, basically every company that is not Anthropic really aggressively trying to discount. If you think about any big enterprise software firm, you know, I reported with Aaron and Kathy a week ago that, you know, software firms are adding new discounts to try and lure customers from Anthropic and OpenAI.

19:23We see OpenAI becoming more aggressive with discounting and offering cheaper prices. I think, if anything, it's just a sign that we are in a competitive landscape. And as customers turn to open source models and become more price sensitive, any company that is not anthropic and didn't have that early lead in enterprises with their AI products is going to try and use this as a tactic to entrench themselves within business teams and engineering teams and organizations. Because when we think about it, like AI adoption is still pretty early. There are a lot of blue chip companies that are very early in their AI adoption journey, you could say.

20:04Right. I guess, I mean, what I was trying to get at, though, is so with these discounts, it seems like it's a very fluid situation, as you said. I mean, the price is, they set a price, right? Right, right, right. They come and they go. It's all kind of made up is the point that I'm making. And I just wonder from a buyer's perspective, right, if I'm dealing with a company that I have unpredictability around, when they're going to discount, if there are more discounts coming. I mean, yes, I like the product, but I wonder if it's better to just stick with one, you know, so that buyers sort of have a better relationship with the economics of it all.

20:48yeah i definitely think as you said it leaves room for more players like meta um like you know xai um that maybe haven't gotten as big a share of the pie as open ai and anthropic um and i think it'll be interesting to see how this all plays out over the next year right and you know the other way i'm thinking about too is they could come in and um again price predictability, like they have more data as well. However much Anthropic offers discounting or not, I mean, they kind of know, well, maybe this is their cost basis. This is what the price was. This is what the discount was. So it all makes for a very interesting landscape.

21:31I want to thank you for coming on, Laura. That is Laura Bratton, author of our AI agenda, applied AI newsletter, I should say. There's so many newsletters. here at The Information. It's been another busy week of model news. Claude released Sonnet 5.5, and meanwhile, OpenAI said it is not releasing the latest iteration of its Astra model because of safety concerns. All of this makes the business of benchmarking all the more important, and that is exactly the bet that Andreessen-backed startup Val's AI is making. I want to bring on co-founder Ryan Krishnan for a conversation. Ryan, welcome to the show.

22:09It was great to have you here. Hey, thanks so much for having me. So I want to ask you about the business that you're building. But I mean, look, the news this week, we got to do the model review here. So we've got 5.5. It's out. I take it you've been messing around with it a little bit? Yeah, that's right. We had a chance to run it through our valuations. We did a bunch of quantitative and both qualitative valuations. So we've had a lot of interesting discoveries the last few days. And what were those discoveries? Yeah, well, one was literally a novel discovery, but in terms of the actual numbers behind it, we found that the model is incredibly capable.

22:47It's actually scoring number two on our benchmark. So even though it's supposed to be more of a mid-tier model, it's actually competing at the frontier. Who's at number one? Number one is Opus 5.5. Oh, wow. Okay. So this is pretty impressive. I mean, this is not meant to be the the best of the best here, right? It was supposed to be a cheaper alternative. Yeah, that's right. In fact, overall, it's actually outperforming Fable 5.1 in a bunch of places. 5.1 also sees refusals, so it may also count for a lower refusal rate that we see for Sonic 5.5. But yeah, overall, it's an incredibly capable model.

23:27It's also, on a per-token basis, much cheaper than those others. We find that in practice, it is quite token-hungry, so in expectation, we find it actually performs double the cost of its predecessor, Sonnet 5. But needless to say, it's a very capable model. And sorry, this is the bench. I mean, you're comparing it to the OpenAI models too, right? Like how does it compare to Astra? That's right. I believe Astra on this benchmark is place number four. So it's not breaking the top three. Okay. And I mean, we'll get into it. But so this benchmark, I mean, how do you, what's your competitive advantage with this particular benchmark that you put together?

24:07Yeah, so I think the key point is that we're responsible for building all the benchmarks ourselves. And so we've assembled a bunch of benchmarks that have coverage across the economy, looking at places like finance, legal, healthcare, coding. And this forms a more accurate picture into the expected performance of the model for a particular user in these domains. And I think it's also important to note that the way we structure the business, it's in such a way that we're not going to leave that test set or even sell any of the training data to labs which may allow them to artificially score higher on this benchmark.

24:40So this is really a high signal way for people to get a sense of the intelligence of these models. So your approach here is you're creating industry specific benchmarks? Is that how you're hoping to compete here? Yeah, the two main components are we're trying to measure the real world capabilities and risks of the models. So as much as possible, it's closer to can you perform this financial research task as opposed to a quiz style question. So that's the first, let's try and imitate real capability and risk. And the second is keep the test set truly private and held out. And this may seem like a simple detail, but the way that a lot of evaluation has been done historically is it's akin to having students grade their own tests or self-proctor their exams at home.

25:24And that's just the way that academia has worked in releasing open source benchmarks that labs are self-proctoring. But by comparison, we keep our test sets completely private. We make sure we run the models with zero data retention policies. And so that makes sure that our test set results are uncontaminated. They have a higher signal on the true model performance. Now, you mentioned the risks of these models. We saw the news in the last 24 hours that OpenAI is delaying or scrapping 6.1, say GPT 6.1 Astra. I think I have that name, right? Forgive me, they're not the best name products. What was your reaction to that?

26:06Yeah, I mean, I think it's a natural conclusion, to be honest. I mean, I think there were some concerning incidences going back to hugging face, but also since then. And so I think, although there were some previous calls to slow down training runs, especially on a large scale, those seem to have been ineffective. And so I think this is causing OpenAid to take another step back to understand what are all the risk points of risk and how do they actually mitigate them. I think they also have laid it out in a couple of blog posts very clearly that the places where they're going to spend more time and attention, you know, looking at doing more offline alignment evaluations, doing complete audit of RL environments, trying to do more work on monitorability, containment, and even just taking responsibility and accountability when mistakes go wrong, when mistakes happen.

26:55They've also laid out a way for auditors to be involved in the testing process as well. But I think the key question in my mind is, what is the threshold that they want to see across the board, which will allow them to continue with training and continue to push the frontier? And it seems like without that well-defined, we end up in another situation where although we understand the points of failure and the risk, a lab will continue to try and stay competitive, keep up with the frontier, and then just go ahead and scale up their RL training runs. So when you say threshold, you mean the question in your mind basically is what the eval is that they're testing it against, like what the rubric is essentially that they're ranking the model on safety-wise.

27:37That is what's unknown to us right now. Yeah, I think there's some sort of an if X, then Y. If we meet an X condition, then we are comfortable scaling up to this Y set of experiments, or it necessitates an additional set of evaluations to be run. And right now, I think a lot of the threshold questions are still open in my mind of what they actually want to see to give them enough confidence to decide to continue to scale up. Right. I want to ask you about something you wrote here. You wrote, our study on RSI shows that at their current pace, anthropics models could match human researchers by August 2027.

28:17Tell me a little bit about the research you did that led you to make that claim. Yeah. So this is looking at a benchmark on recursive self-improvement, or RSI, and that measures a model's ability to make the successor version of itself. So that's almost akin to cloud making the next version of itself without any anthropic researchers. And the way we constructed that is by collecting five proxy tasks, which measure the work that machine learning engineers do within these labs. So that covers things like pre-training, post-training, harness level engineering. And what we're doing is comparing the model's ability to perform these tasks with respect to human baselines.

28:57And so at the current pace, looking at the last year of model releases, we're projecting forward that if this pace continues, we'll see models start to meet human level researchers by August 2027. And you only found that with Anthropic, not with OpenAI? RSI? There's a similar extrapolation done with OpenAI models, but I think the timeline for them to achieve RSI is a little bit longer. I think maybe by the end of 2027, the soonest one we see is Enthropic getting there first. Okay. And so you're making this prediction, which is kind of, I mean, it's kind of nice that people have actually put a timeline on this because otherwise people just sort of say, if we get there, when we get there, you know, it'll be dangerous.

Read the full transcript

29:42So what's your view then on what happens post-2027? I mean, you're saying that it'll be as good as human researchers. Is it risky? Is it dangerous? How are you thinking about this? Yeah, I mean, I think it's important to have a timeline because what's being called for is a pace to the frontier and investment in our ability to understand the models before they're actually released. And so this is really the time that we have. Once we get to full RSI, it's going to be incredibly unpredictable how model intelligence will emerge. We will have to have completely automated approaches in order to keep up with the pace of development.

30:21And so I'm somewhat concerned that we may charge ahead towards this timeline, towards full RSI, without fully understanding the true risks of these models. So this is not a long time. I mean, you know, all this talk about pacing the frontier, like, I mean, what I'm, the read I'm getting from you is This is like, we're talking about months here. What's the point of even having the conversation, I guess? I mean, honestly, I'm pretty optimistic. I think even 10 or 11 months in the AI world is an eternity. That gives us plenty of time. I mean, if we're talking about the end of the world, I mean, you know, like I'd love.

30:56I guess to be clear, it's not, I wouldn't say it characterizes the end of the world. I'm just saying that model development will happen in a very unpredictable way. If we don't understand fully what new models are capable of or the new risks that they're opening up. You know, I think it also is possible that there may be, is a, I think people are afraid that RSI is akin to a fast takeoff. And that means that, you know, models eclipse all humans in all domains in a very short period of time. But in reality, I think there will be other resource constraints, which mean that doesn't actually happen.

31:27So, for instance, we have limitations on chips or energy. And although models may also make progress in those domains, there are more physical constraints that we might have, which would limit model progress in those areas. So are you also involved in the business of, you know, auditing these model companies? And I mean, we've heard about Meter, we've heard about the debate about whether or not it should be the government, should be an independent organization. Are you also involved in that operation as well? We are, yeah. I mean, we already do a lot of pre-release testing with release candidate checkpoints.

32:02Now labs have called for more embedded evaluators to work with them in earlier points in the training process. And so a lot of the skillset that we've developed to do our evaluation extend to those places as well. So if you're involved in that, then it sounds like you probably would want the government to not have a role in this? I mean, I think it's going to take a bunch of diverse approaches, and I would love to see folks in the government more involved in this. We already work pretty closely with Casey, and I think Casey's doing a lot of great work. AC in the UK has also released some really comprehensive reports.

32:38And so I think it's going to take a variety of approaches. I would be foolish to believe that we alone or any one evaluation alone is going to solve everything at the wide surface area of model risk. And so let me ask you about these conflicts of interest that people are so worried about with these auditing organizations, the question around, well, what is true independence? Can we actually achieve it? How do you strike that balance with your own company? And have you gotten any questions from people saying, well, what are your ties to these labs? Who's paying you? Who's funding you? How do you deal with all that?

33:16Yeah, I think there's a lot of failures in history to look to and try and identify what we can do differently to try and prevent for those. I think, for one, making sure that the same group responsible for auditing is not the same group responsible for selling the solution is a really key part of how we're structured. When you look at Enron, that was a case in which the same company was responsible for auditing as well as consulting and providing the solution. And so they had a failure to audit and actually maintaining integrity of that. And so we take that to the machine learning world where although we build these benchmarks, we would never sell training data to the labs in order to enable them to score better or artificially get a higher score on a benchmark that's not realized by an end user.

34:03I think there's also going to be a more complex set of accreditation that probably has to happen, especially if IVOs go into effect. And so that's a mechanism whereby the federal government would actually need to give a license to do the evaluation and make sure that's done in a truly independent way. So I think in Limit, that's actually the best solution. But until then, we're being very honest and disclosing the ways that VALS makes money and contracts we're refusing in order to maintain our independence. Great. Well, Ryan, I want to thank you for coming on. That is Ryan Krishnan, the founder and CEO of Vals AI here on TI TV.

34:42Lola, an AI assistant backed by Booking Holdings, is launching today. For more on that, I want to bring on Steve Hafner. He is the CEO and the co-founder of Lola. He was also formerly the CEO of OpenTable and Kayak. Steve, welcome to the show. It's great to have you here. Great to be here, Akash. Thank you. uh i have to say tell you that uh i have had few i'm trying to think if there are ceos or executives i've had whose product i use more frequently than than open table um but it's you know for someone who lives in new york well it's it's a way of life i mean you know it's like it's like the game that everyone plays in their spare hours um okay let's talk about lola though So what is it?

35:27Tell us about his assistant. Yeah, Lola is an AI assistant. We use natural language on the interface to help our members discover and book everything that's worth doing. So, yes, we have open tables, so you can do dining and restaurants on Lola. But you can also do sporting events, concerts, hotels, travel, basically anything that's worth doing. And you mentioned that we're owned by Booking Holdings. That's the world's leading company of a whole portfolio of brands across travel and related services. So we've got all of their inventory, but Lola's much broader than that. We've actually partnered with a lot of the world's leading experience and lifestyle brands.

36:07You know, folks like SeatGeek, GetYourGuide, ResX, and others to have a really awesome experience for both discovering and booking anything that's worth doing. What about Airbnb? Airbnb is not on the platform, but anytime Brian Chesky wants to put that inventory on us, we'll gladly take it. Right, right, right. And tell me, so it's basically an assistant that can, I mean, it's a chatbot, effectively a chat platform that hooks up to all these different inventories and you can sort of build a customized booking. Is that the idea? Yeah. So basically it combines two technologies, right? The best of an LLM, which can do like chat TVT and Anthropic, who can do a pretty good job recommending things to do to you, but they usually fall down when it comes time to actually turning that recommendation into a reservation.

37:01And what we do is we take all the live rates and availability from our partners that we access in real time via API and combine the two. So not only do we help you decide what you want to do tonight or this weekend, but we also make it very easy to book it. So why do you think that this is more effective than what Muse advertises to do? Because Muse and Instinct, I mean, booking restaurants, hotels, these are like the most popular use cases, I want to say. Absolutely. So we're not going to let Muse and Instinct have all the fun. So if you think about both of those platforms, they're horizontal agents who try to do everything they can for you.

37:45And as a result, they do many things okay, but not anything very well. Lola is a vertical agent. So we're just focused on doing experiences really, really well. And that focus allows us to do it much better. And the partnerships that we've struck across the industry provide us that real-time availability and rates so that we can conduct a booking. So if you want to try something, try booking an open table restaurant on Muse or Instinct for tonight here in Miami Beach where I am. It won't do a very good job because, you know, they don't have that direct API. Try the same prompt on Lola and you'll be delighted, I think, with the result set.

38:27Are you guys using open-weight models in the background, frontier models? What are you using? We're model agnostic. We're using a lot of different models in the background. And what we try to do is we optimize for the latest release for the frontier models on the edge. And then for the stuff that we think is low booking intent, we'll use old models that are a lot cheaper. But as we look at our architecture, when new models come out, right now we're using Gemini a lot, to tell you the truth. We'll easily swap it out to get better result sets. And that's on the front end on the LLM side. have you had any uh issues with and and the platform just is being launched today so you know there there will be a lot more customer data to work with but i mean agents i don't want to say going rogue but we had a story i mean one of our editors uh they used uh muse to do a booking and um i i think it double booked him he gave him two rooms instead of one so he uh he could have been out$408, but thankfully the hotel staff were nice about it.

39:36Tell me about some of the challenges that you've had to overcome during the testing of this. Sure. Again, this goes back to approach. So we've been building for five months, so we've had plenty of users go through the platform and we've caught a lot of stuff. But there's two big problems with LLMs who don't have partnerships. The first is accuracy of the data, right? So they can hallucinate. They can create a hotel with a pool where the hotel doesn't have a pool. The second is permissions, right, and security. If you are relying on a browser interface to go search for a table on OpenTable, you don't have permission or direct API into OpenTable's inventory, you're going to make mistakes.

40:20The risk of mistakes goes up. So on both of those topics, we were very keen on grounding the data. So we get a hotel inventory from Booking.com, for example, among others, who has great accuracy on what a hotel amenities actually are. And we'd have direct connectivity with OpenTable, etc. So when we present a result or a recommendation to a user, we know it's real. We know the hotel is this. We know the restaurant availability is there. And when we actually make the reservation, we confirm it directly with the provider. So we've eliminated both of those big gaps that the LLMs can have. Yeah. So have we given up on any of these hotels, travel companies?

41:06I mean, you know, the relationship that any of these brands have with customers, I mean, certainly with booking and kayak, the relationship sort of became a little bit more diluted because you have an aggregator model. Now, you don't even see the options. It's just the agent just does it for you. I mean, how do you think about how any of these companies will develop a relationship with the end customer at all? Yeah, I mean, there's a lot of people poking around at that. I follow my CEOs, Glenn Fogel, who's the CEO of Booking Holdings. I share the same kind of podcast. This isn't an either-or situation.

41:45We've had fragmentation in the industry for many, many years. People were talking about Google disintermediating all of us a long time ago, the online travel agencies. It never happened because the reality is the services that the OTAs and others provide require lots of feet on the ground to get that content, to get those restaurant relationships, to get those hotel relationships, to provide great customer support. You know, I don't think you're ever going to be in a situation where you pick up the phone and you have an issue with a booking and you made it via Muse and you try to call Muse. You can't call Muse.

42:19You know, you can test instinct, but instinct's not going to help you on a booking. So I think you have to have all those relationships. You have to have the last mile. Will we be able to call Lola? You can chat with us right now. we can route you and service your reservation. The other thing we have is we have a partnership with a company called the 10 Group, where we have actual human concierge access for our members. So yes, there is a human in the loop with Lola, which is one of the many things that differentiates us from some of the other agentic AIs out there. Great. And before I let you go, I mean, I guess I have to do everyone a public service and ask the question, what is the secret to getting a good booking on OpenTable?

43:04Because it's not easy, I have to tell you. Well, the first thing you need to do is make sure that - Time, you need time to do it. You got to monitor the system. You also got to show up for your reservations because if you don't, you get penalized. But if you sign up for Lola and you're - Maybe the agent can show up. Maybe the agent can, that's one thing the agent cannot do is show up for the - That's right. But be a Lola VIP and you'll be an OpenTable Gold member. so you'll get access to the tables that others will. Great. All right. Well, Steve, I want to thank you for coming on. That is Steve Hafner, the CEO and co-founder of Lola here on TITV.

43:40That does it for today's show. Reminder, we are on this stream Monday through Friday at 10 a.m. Pacific, 1 p.m. Eastern. If you can't make it that, episodes are available on theinformation.com, on our YouTube channel, or wherever you get your podcasts. Make sure to follow us on social media, on X, on Instagram, on TikTok, and on LinkedIn. in. I am already excited for our next show tomorrow. Have a great rest of your Tuesday. Bye-bye for now.

From the publisher

The Information’s Co-Executive Editor Martin Peers talks with TITV Host Akash Pasricha about delayed tech IPOs and market uncertainty. We also talk with Sriram Viswanathan about AMD’s $8.2B acquisition of World Labs, Laura Bratton about Anthropic cutting AI discounts, and we get into AI model benchmarks and safety with Vals AI Co-founder Rayan Krishnan. Lastly, the Lola CEO, Steve Hafner, discusses his AI travel booking agents.


Articles discussed in today’s show:


Subscribe: 


Sign up for the AI Agenda newsletter: https://www.theinformation.com/features/ai-agenda


TITV airs weekdays on YouTube, X and LinkedIn at 10AM PT / 1PM ET. Or check us out wherever you get your podcasts.


Follow us:

X: https://x.com/theinformation

IG: https://www.instagram.com/theinformation/

TikTok: https://www.tiktok.com/@titv.theinformation

LinkedIn: https://www.linkedin.com/company/theinformation/


Chapters:

00:00 - Introduction

01:13 - Delayed Tech IPOs & Anthropic Market Uncertainty

08:00 - AMD Buys Fei-Fei Li’s World Labs for $8.2B

17:08 - Anthropic Cuts AI Discounts & Usage Caps

22:44 - Claude Sonnet 5.5, OpenAI Safety Snags & AI Benchmarking

35:41 - Lola AI Assistant & The Future of AI Agent Travel Booking


More from The Information's TITV

All 304 episodes
Inside Oura’s IPO Postpone, Anthropic’s AI Discount Cuts, Anthropic Launches Claude Sonnet 5.5The Information's TITV · 44 min
Listen in VO