Anthropic’s Invasion of Slack, OpenAI Cuts Inference Costs in Half, Amazon’s Higher Anthropic Costs

30 Jun 2026 · 41 min · 16 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Four-part AI/infra news roundup: (1) Helix Infrastructure Partners’ plan to build/scale AI data centers via acquisitions, led by former AWS CEO Adam Solipsky; (2) Anthropic renegotiating pricing with Amazon by switching from compute-hours to tokens, raising Amazon’s costs; (3) Anthropic’s Claude Tag in Slack confusing Salesforce employees about cannibalization vs Slackbot; (4) OpenAI developing inference optimizations cutting model run costs by more than 50%.

Guests (backgrounds)

Dakin Campbell (AI and finance reporter, The Information); Catherine Perloff (Amazon reporter, The Information); Laura Bratton (author, Applied AI newsletter); Stephanie Palazzolo (author, AI Agenda newsletter).

Key claims & examples

Helix raised $10B with NVIDIA, Vistra Energy, Kuwait Investment Authority, plus KKR; will pursue “build-to-suit” deals and leverage Vistra power bottlenecks and NVIDIA efficiency. Anthropic’s token pricing affects Alexa, AWS products (e.g., CodeWhisperer/Cura-like coding, Quick/Workplace assistant). Claude Tag is “at Claude” shared in Slack; Salesforce employees worry it competes with Slackbot. OpenAI’s exact optimization is undisclosed, but engineers reported >50% inference cost reductions; likely boosts gross margins rather than passing savings to customers.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Adam Solipsky's New Role at Helix

0:20 to 1:30

Discussion of Adam Solipsky's new position and the company's goals.

“aiming to compete in the AI infrastructure race.”

Insights from Dakin Campbell

1:30 to 3:42

Dakin Campbell shares insights about Helix's strategy and partnerships.

“What is Helix Infrastructure Partners and what did you find?”

Interview Highlights with Adam Solipsky

3:42 to 6:10

Key questions and responses from Solipsky regarding the company's vision.

“Now, Adam Solipsky is the man at the center of this company.”

Private Equity and Data Centers

6:10 to 10:00

The role of private equity in the data center industry and future trends.

“NVIDIA has a somewhat new product out, from what I can tell, helping to build data centers more efficiently so that the power in a data center most efficiently powers their chips.”

Anthropic's Pricing Changes with Amazon

10:00 to 11:55

Discussion on how Anthropic is changing its pricing model with Amazon.

“More money's coming in, but I don't want to overlook the Adam Szilipski portion of this.”

Impact of Pricing Changes on Amazon

11:55 to 14:00

Exploring the implications of Anthropic's pricing shift for Amazon's AI products.

“Anthropic is taking a tougher stance on Amazon in terms of pricing, despite the fact that Amazon is a major investor in the company.”

Amazon's Relationship with Anthropic

14:00 to 19:20

Discussing Amazon's pricing and leverage over Anthropic in cloud services.

“And, you know, another thing to kind of point out is the previous unit, Compute Hours, Amazon AWS is obviously a massive infrastructure company.”

Introducing Laura Bratton

19:20 to 19:46

Welcoming Laura Bratton to discuss Anthropic's new AI tool for Slack.

“That is Catherine Perloff, our Amazon reporter here at The Information.”

Exploring Claude Tag

19:46 to 23:24

Delving into the functions and implications of the Claude Tag integration in Slack.

“So Claude Tag is basically a shared teammate that enterprise users can access in Slack.”

Salesforce's AI Agent Dilemma

23:24 to 26:04

Examining Salesforce's promotion of Claude Tag amidst its own AI offerings.

“And I should say, I mean, look, maybe I haven't given Slackbot the time that it deserves.”
Show all 16 chapters

Privacy and Data Concerns

26:04 to 28:01

Addressing privacy issues regarding data access and usage by Anthropic.

“But then once it's generally available, these customers are going to burn through flex credits that they have or credits that they have with these accounts to use CloudTag.”

AI Management Dashboard Competition

28:01 to 30:03

Explore the evolving landscape of AI management tools among enterprise software firms.

“And every enterprise, I should say every enterprise software firm right now is trying to compete to become this sort of AI management, AI agent management dashboard.”

Deep Dive into OpenAI's Cost Optimization Strategies

30:12 to 34:06

Detailed discussion on the strategies OpenAI employs to cut inference costs.

“I want to bring on Stephanie Palazzolo, author of our AI Agenda newsletter, who reported details on that this morning.”

Implications of Cost Reductions for OpenAI and Customers

34:06 to 36:18

Analyzing how OpenAI's cost-saving measures could affect pricing and profit margins.

“So I would say that with this OpenAI example, we really don't know exactly what the optimization is.”

Anthropic's Pricing Strategy in Context

36:18 to 38:30

Discussion on Anthropic's pricing model compared to OpenAI's cost strategies.

“I don't think the customers are going to see any of this.”

Impact of Cost Optimizations on Cloud and Chip Providers

38:30 to 40:59

Exploring how OpenAI's efficiencies affect relationships with cloud and chip companies.

“But I would say so far, yes, I think it's a fair point that like they haven't maybe done the best job for some developer.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:13Welcome, everyone, to The Information's TI TV. My name is Akash Pasricha. It is Tuesday, June 30th. First up today, The Information published exclusive reporting that former AWS CEO Adam Salipsky has a new job taking the helm of a data center company aiming to compete in the AI infrastructure race. Our AI and finance reporter spoke to Sielipski as part of the story. We will hear from him shortly. Next up, Anthropic has renegotiated one of its deals with Amazon, potentially increasing what the tech giant pays. Our Amazon reporter joins us with the scoop. Speaking of Anthropic, the company released an AI tool that integrates with Slack, but it is creating some confusion for Salesforce employees.

0:56We will hear more from the reporter who broke that news. And finally, OpenAI appears to have found a way to reduce the cost of inference by a significant margin. We'll talk about whether or not customers will see those cost savings. It's going to be a fun show, so let's get right on into it. Adam Solipsky is a name that you might recognize. The former CEO of Amazon Web Services has a new gig at Helix Infrastructure Partners, a data center company that the information's AI and finance reporter, Dakin Campbell, reported exclusive details about in a piece published this morning. I want to bring on Dakin to walk us through what he found.

1:32Dakin, welcome back to the show. It's great to have you here. Yeah, thanks so much. What is Helix Infrastructure Partners and what did you find? So this is a new data center, a new entrant into the data center market. They don't have any operating assets yet, but they have plans to acquire some very shortly. You know, it will be a company that builds data centers and then leases them to hyperscalers. Solipsky will be the CEO of the company. You know, they announced a couple of weeks ago that they'd raised$10 billion with strategic partnerships with NVIDIA and Vistra Energy. And so this is sort of we're learning about what this is going to be a little bit in real time.

2:22And so it'll be interesting to see sort of how they compete and where in the ecosystem they slot in. And so the idea here is what? They're going to buy up data centers? Is that the game plan? They're going to buy up a data center operator, an already existing operator, that might be small, that might need more money, that might have private equity or sovereign wealth backers who haven't been able to put the money in or may want their money out. So, you know, there was some question when this was announced a couple weeks ago whether they would build from scratch. And the big news that we've learned through talking to them is that they are not going to build from scratch.

3:05They are going to be on the hunt for at least one significant deal that they can then use as a foundation to build from and scale from. Got it. So they're basically, it's sort of, it's not a holding company, but it's a management company that is going to buy management companies that already manage data centers, essentially. Yes, that's exactly right. And then they believe that they can take that company that they buy and grow it much quicker, much bigger, because of their unique set of what they believe are competitive advantages. Now, Adam Solipsky is the man at the center of this company. You had the chance to speak with him.

3:48I wondered, what were the big questions that you went into that interview with? Yes. You know, I was thinking about this. I probably shouldn't be totally honest about this, but, you know, Adam Solipsky, the name doesn't mean quite as much to me as it does maybe to our viewers or certainly to other colleagues in the newsroom. I'm the finance guy. So, you know, my questions going into it were very much, you know, the announcement was pretty vague. It wasn't clear. So I wanted, you know,$10 billion is a lot of money. So I wanted to know sort of what they were planning to do with it. I wanted to know a little bit about what, you know, Solipsky being at the helm really meant for them and what they thought he could bring and what he thought he could bring.

4:40So questions much more around, you know, what is this company going to be doing? And, you know, are they going to be successful at it? Okay. And what did he say? You know, they're really focused on the strategic partnerships that they've got with NVIDIA and Vistra. Both companies, plus the Kuwait Investment Authority, each put a billion dollars in or more. into this$10 billion. And so, you know, Vistra is an energy company. There are no, to my knowledge, data center operators out there that have such a close connection to an energy provider. We know power is a big bottleneck in the industry. And so the idea, I think, is because they've got Vistra on a phone call and can talk to, you know, the CEO or the board, if they've got a data center that they need a power plant or something built for, they'll call Vistra.

5:44Or if they are looking to site a data center, they might choose to site it near some power plants or some battery facilities that Vistra already owns. So the idea there is power is a big bottleneck. If you've got one of the biggest energy companies in the country on speed dial, then you should be able to break up that bottleneck a little bit. The other strategic partner is NVIDIA. We can't overlook them. NVIDIA has a somewhat new product out, from what I can tell, helping to build data centers more efficiently so that the power in a data center most efficiently powers their chips. And so Helix will be working with NVIDIA to build super efficient data centers from the ground up, which they believe will allow them to build them more quickly and make a much more compelling package for the hyperscalers who are going to be their clients.

6:51Did he give you a sense for any of the companies that they could look to acquire? Do we have any idea there? uh he didn't um you know i anticipate that as the secret sauce of the whole company yes exactly i mean the place they're going to play is uh is what's known in the industry as build to suit so uh you know they're going to talk to hyperscalers hyperscalers will tell them exactly what they want in terms of cooling in terms of power to the facility in terms of you backup generators, and they'll build them then for the hyperscalers. And I think they'll be doing a little bit of powered shell, which is basically you build the four walls of the data center and get power to it and then give an empty warehouse basically over to a hyperscaler.

7:41So that is a business model that a lot of companies are in. Private equity to this degree has really weighed in higher, like, you know, more established companies, this looks like they'll be, they'll buy somebody in the, in the mid tier, um, and then, you know, look to grow that into an industry leader. Right. Talk a little bit more about private equity and the extent to which it has moved into data centers. Is this something that you're hearing? And by the way, we haven't even acknowledged the KKR component of this venture here. Um, so KKR is, is an investor in the new venture? Yes, they are.

8:22And they are basically the owner of Helix. And they'll have the$10 billion that was contributed by all four partners to draw down when they see a company that they want to acquire. So, yeah. Is KKR the only company here looking at data center? We've talked about other financing deals. I think Apollo's is in the mix. So is this a framework for a deal that is likely to be more common? It's a good question. I mean, to some degree, KKR is late in this. Blackstone bought QTS in 2021. Earlier this year, they acquired 49 % of a company called Rowan that, you know, looks to be what Helix is going to become.

9:19But yeah, I mean, there's lots of money. There's still money being raised. So, you know, maybe they're middle of the pack. I think we will see other entrants come in after them. There's a ton of money that's ready to be deployed and more money currently being raised. You know, what Helix, what they believe their competitive advantage to be is these strategic partnerships. So I'm not sure you'll see NVIDIA or Vistra lend their support to a lot more private equity-backed players. So in that regard, this might be unique or somewhat rare. But there's more money coming in. That's for sure. More money's coming in, but I don't want to overlook the Adam Szilipski portion of this.

10:10And I know it's not a name that you might have followed as closely as a finance reporter, but the CEO of AWS, the one-time CEO, that's not a small character to play in the space. So I assume there is some confidence there from KKR and saying, well, maybe we were waiting for the right person to come along or the right venture. Did you get a sense from Solipsky at all? I mean, on a personal level, did he tell you at all why he had chosen to commit the next stage of his career to this infrastructure category after leading AWS for so long? Yeah. I mean, he started at KKR in September of last year as a senior advisor.

10:50And from what I understand, the idea for Helix came together pretty quickly. I mean, I think he sees data centers, access to power, and basically the capacity needed to power AI as one of the biggest, if not the biggest challenge in the business world right now. So I think that from that extent was exciting to him. And he wants to be in the mix at a high level and this gives him the opportunity. And it's very clear he's going to be calling on CEOs, division leaders at the other hyperscalers, pitching Helix. So that's a Rolodex that he's got in spades and plans to use. Right. Well, I anticipate it's going to be a name that we're going to hear even more about.

11:45Dakin, I want to thank you for coming on. That is Dakin Campbell, our reporter covering all the money behind the AI boom here on TITV. Thanks so much. Okay. Anthropic is taking a tougher stance on Amazon in terms of pricing, despite the fact that Amazon is a major investor in the company. My colleague Catherine Perloff reported inside details of that partnership this week, and I want to bring her on to share more about what she found out. Catherine, welcome back to the show. It's great to have you here. Hi, Akash. What did you find out about Anthropic's pricing plan with Amazon, Catherine? So, you know, Anthropic has been changing how they charge a lot of their customers, and a lot of those customers are now paying more for their models.

12:31Amazon is no exception, but obviously Amazon is not just any customer. So that's why we were interested in this story. Basically, they, a couple months ago, changed the deal so that instead of Amazon paying Anthropik in a unit called compute hours. So basically, you know, the sort of hours of compute taken up when Amazon is using Anthropik's models, they shift to tokens. And this is going to make it more expensive for Amazon to use Anthropik's models, which is sort of a big deal because anthropics models power a lot of amazon's big ai products that includes alexa for shopping which is their chat bot um on uh their retail site some of the aws software they're selling like cura a coding product and quick um a workplace assistant so uh yeah they want to you know anthropic isn't the only model underneath the hood but it is a big one so now other costs are going to go up.

13:38Yeah. And look, I mean, from an anthropic standpoint, it certainly makes sense. We've talked about how many more tokens are being used right now with agents. And so the idea that you can consume more tokens in a per hour basis, it would certainly be beneficial for them to charge based on the consumption model there. Do we have any reporting at all on why this changed and what flipped the switch cheer? You know, I'm not exactly sure what was in everyone's heads, but my understanding is that Amazon employees for a while were sort of aware they were getting a pretty good deal and they were worried that the cost could increase.

14:21And, you know, another thing to kind of point out is the previous unit, Compute Hours, Amazon AWS is obviously a massive infrastructure company. that's something easier for AWS to manage the costs and, you know, optimize their compute. So Anthropic could, so when they have to pay Anthropic, you know, that could be cheaper, they could optimize it. But with tokens, that's the unit that Anthropic controls. There's sort of less leeway Amazon has to sort of, you know, do their best to optimize costs down. I think the other thing to really point out here, though, is Anthropic is becoming like a really big company, You know, it's sort of becoming this dominant enterprise tech company and they're going to go public soon.

15:07And despite the fact that like Amazon was one of their biggest and earliest backers, they have a bit more leverage than maybe they used to. Right. And I mean, you pointed out in the story that Amazon also has a relationship with OpenAI here through its latest funding. I can't remember if it was$40,$50 billion investment, I think, that Amazon made. It said up to$50 billion. Up to$50 billion, right. I think they haven't done it all yet, but a lot of money on the table. Right. And I certainly recall when I think Amazon said we will invest up to$4 billion in Anthropic, and then they came up with another one and stuff like that.

15:47So, okay, fine. So Amazon is cozying up to both OpenAI and Anthropic. Maybe that has repercussions here. We're not sure. What did Amazon tell you about their costs going up? Did they acknowledge that their costs could go up through this new pricing model? So Amazon's statement to us was that, you know, pricing model may have changed, but they don't feel like their costs aren't going up. My sources, you know, said it would make it more expensive. So I'm going to have to agree to disagree. But, you know, that's OK. You know, and they probably have some fair points there. And Anthropics POV that they shared with us was like, you know, while the pricing model may have, you know, may be changing, maybe they're charging all customers different, you know, with more usage based pricing and overall AI models are getting cheaper.

16:48So even if like the unit economics is changing, you know, AI, the actual cost of the models is going down. So, and I think also big picture, Amazon is conscious of costs. Like they, I've spoken to some of their executives and for example, in their product Quick, they kind of help the customers of that product use a cheaper model when the task allows for it. So they're kind of under the hood routing customers to a cheaper model to help those customers of their AI software keep costs down. So they are definitely – They're not just sitting idly by basically and taking costs. They're finding ways to – Right, right.

17:31Yeah, yeah. I wondered what other bargaining chips do you think Amazon has here knowing that this relationship is a long one and I'm sure these deals will continue to get renegotiated. I mean, are there other levers you think Amazon could pull on here given that Anthropic certainly needs Amazon for cloud services? How do you see these negotiations playing out in the long run? Yeah, it's interesting. I mean, so Amazon has internally thought about like with these rising costs, could we use OpenAI's models? Could we use our own Nova models more to power some of our products? You know, remains to be seen.

18:09So obviously, there's other providers out there. I mean, Anthropic still really relies on AWS for infrastructure. And they also really rely on AWS to help them reach new customers still. They sell Anthropic. A lot of the model sales happen on AWS. And that was a big way for Anthropic to access a lot more business customers. And through that business relationship where Anthropic sells their models in AWS, they pay like 50 % of the gross profits. We reported this earlier this year back to Amazon. So, you know, basically that means after they pay their cloud computing bill, they still have to fork over 50 % of the profits back to Amazon.

18:51So, you know, Amazon had the infrastructure. They also have some chips that Anthropic is using. So definitely not like Anthropic doesn't need Amazon. But, you know, now Anthropic can say we've got some business customers of our own. You know, we have also compute is really expensive. You know, we have to make our economics work and everyone kind of has to play ball. Great. Well, Catherine, I want to thank you for coming on. That is Catherine Perloff, our Amazon reporter here at The Information. Anthropic released an AI tool that integrates with Slack recently. But when Slack's parent company Salesforce started promoting the Anthropic tool.

19:34That confused some of the company's own employees over risks of cannibalization. To explain that more, I want to bring on Laura Bratton, author of our Applied AI newsletter. Laura, welcome to the show. It's great to have you back. Hey, Akash. What's up? What's up? You tell me. I'll tell you. Okay. Tell me about the new tool. What is Claude Tag? Yeah. So Claude Tag is basically a shared teammate that enterprise users can access in Slack. So you tag Claude, literally, at Claude in your Slack channel that you want Claude to answer a question in or do some sort of administrative or coding task in and all your teammates can see it.

20:16Over time, Claude tag builds context and memory from your Slack channel and it can follow up with you and sort of initiate actions on your behalf. Have you used it? Have you at Claude'd yet on Slack? I have not ad plotted yet, but I guess we should ask our editors if we should be testing these tools in our Slack channels. True. Very true. Well, because one of the things that you mentioned, we'll talk about in a second, but there are privacy implications to all this. But one of the tools that I have used is Slackbot, okay? And I have tried to, you know, basic functions, reminders, okay? And they do reminders pretty well, but I've tried to create some automations.

20:57It works like okay at best. It's not really the greatest tool, I will say, in my experience. But what are... I just want to ask, have you been using this since January, since they did this sort of re-imagining of Slack, hard-biothropics models and stuff? That's when you're trying? So I tried to create yes is the answer. And the reason I know this is because now the reminder that I've created, I see in the notification it says thinking. And so I assume there's some models working in the background. But it's not that great. I don't know. Am I alone here? I mean, I haven't tried Slackbot personally, but I've talked to a lot of people who have used Slackbot.

21:43And they said the ways that they're using Slackbot are primarily to interact with their CRM data in Salesforce. So that's customer relationship management software. So teams that I talk to will use Slackbot to pull data from their CRM about the deals they have in the pipeline or their customers to summarize the sort of demographics of their customers and use the reports that they can generate from that data in meetings. So a lot of administrative tasks. And Salesforce certainly wants Slackbot to be kind of like a personal assistant that can hook up to any business software applications that you use because they have MCP servers that can connect with any business software that a client uses.

22:33and then they want Slackbot to help you do whatever administrative task you want to do in those software systems. But I think it remains to be seen if customers will choose Slackbot for that. And now you have all these other AI agents that they're loading into Slack that customers might choose instead. So it's not just CloudTag. There's also AI agents from Perplexity, Linear, Cursor. This small startup, Victor, has an agent that essentially does all the same things that CloudTag does. But I think Claude Tag particularly made a splash because Salesforce also pushed it a little bit more heavily than it did other AI agents that have been launched in Slack and had a whole post on X.

23:10And Salesforce executives posted about it on LinkedIn. And I think that's what created some of the confusion is, you know, why are we heavily pushing this AI agent that kind of competes with our own proprietary AI agents? Right. Right. And I should say, I mean, look, maybe I haven't given Slackbot the time that it deserves. Okay. Maybe I don't mean to, you know, talk down on the Slackbot. We should talk to our business teams because they're the ones actually interacting. Right. And I don't have a CRM that I'm pulling for. So I anticipate I'm not exactly the ICP here that they are looking for. But I mean.

23:48But it still should work well for you in a way that, you know. Yeah. It should be easy, right? easier. So now you touched on this now, but the crux of this issue is that you have Claude tag coming along, right? And that's the at Claude function. And you have Salesforce staffers that are saying, well, why are we Salesforce promoting this Claude integration when our Slack bot conceivably could do some of what the Claude assistant could do? And I just want to confirm one thing. So when I use Slackbot, is that powered by Anthropics models under the hood? Is that the idea? Yes. Okay. Okay. So then the question that I have for you is then what did you find is the rationale here for why Salesforce would be promoting Claw?

24:37Is there a business, is there a deal under the hood here that they're getting paid for this? So Salesforce nor Anthropics would tell me any of the financial terms of their agreement. I did find out from some other providers or people close to the other providers of AI agents that have agents in Slack that they're not paying Salesforce. That was Victor and Perplexity. I found out from sources familiar with those companies that they're not paying Salesforce, which sort of puts a dent in the argument I heard from Salesforce, which is that they want Slack to be kind of like an Apple App Store for all AI agents.

25:16And I think this move just shows that they're prioritizing the stickiness of Slack, their messaging app, over the success of their individual agent Slackbot. So even if Slackbot isn't necessarily the AI agent of choice for every enterprise customer, they might be more willing to come to Slack if they can use any AI agent that they want to in the messaging platform. But then - Because they don't have to pay to access the Claude assistant or the Perplexity assistant, right? This is not an added cost for them. Well, it is in their accounts with Perplexity or Claude. So Claude tag is available in beta to enterprise and team customers through Anthropic.

25:58And there's some credits that those customers can use while it's still in beta to test Claude tag and see if they like it. But then once it's generally available, these customers are going to burn through flex credits that they have or credits that they have with these accounts to use CloudTag. And you can only imagine if CloudTag is this quote unquote multiplayer agent that can interact with lots of different team members and operate in the background and re-up tasks. I assume that it's going to burn through tokens pretty quickly, which I already did hear from a user of Claude inside of Slack who particularly tells his teammates that they can only use Claude in Slack for certain things and have to use Slack bot for other things because Claude burns through tokens so quickly in Slack.

26:44Right. Now let's go to the privacy concerns that we talked about. Is there reason to be concerned if Anthropic has access to all this Slack data? So what I learned by looking through Anthropoc terms of service is that it does not permit the data from customers with whom it has commercial agreements. So that is, you know, it's enterprise and team accounts, for example. It can't use that data to train its models. So there's no reason to think that it's going to use this data to train its models. I do think, though, this gives Anthropic itself a look under the hood at how teams are interacting on Slack in a way that felt like Salesforce's competitive advantage.

27:30Now Anthropic gets to see how all that's taking place in Slack, and it could inform its product decisions going forward. Right. Let me ask you this. So you mentioned the Perplexity app. It sounds like Slack is playing the long game here, trying to grow the stickiness, as you said. How do you see this issue playing out in the long run? Do you think eventually, do you think these integrations, I mean, they already cost money to a certain extent for the customers? I don't know. I'm sort of of two minds here. I don't know where this story goes. I think nobody knows. And every enterprise, I should say every enterprise software firm right now is trying to compete to become this sort of AI management, AI agent management dashboard.

28:22So that's Salesforce, ServiceNow, Microsoft, all the hyperscalers, they want to be the place where you come and you manage all the various AI agents you have from whichever provider you're using. And that's really the advantage that these software companies are looking for. And I think this just shows Salesforce and Microsoft are trying to extend that to their messaging apps. These AI agents are also available in Microsoft Teams. Sources told my colleague Aaron Holmes that Claw Tag is going to be coming to Microsoft Teams. And then Perplexity and Linear and some other AI agents are also available in Microsoft Teams.

29:01So I think this is just another evolution in the battle among enterprise software firms to control the place where customers come to use their AI agents. Right. And I'm almost certain that someone is going to hear this and think, this guy can't even send a reminder for himself, really? Like, that's like the easiest thing Slapbot can do. So fine, I will try harder. Okay, it was not as simple as a reminder. There were some automations involved. I think a better test is going to be trying these AI agents in Teams versus Slack and seeing if one has a better experience and if there's going to be any differentiation between these messaging platforms and sort of the battle between Microsoft and Salesforce to become this AI agent app store.

29:48And if eventually they're going to enact some sort of tolls for these AI agents to be hosted on their platforms. Right, right. Right. And yeah, I mean, this then goes to the corporate data wars question is that will they end up charging each other for all this? So, Laura, I want to thank you for coming on. That is Laura Bratton, author of our Applied AI newsletter here at The Information. Open AI appears to have found a way to reduce the cost of inference by a significant margin. I want to bring on Stephanie Palazzolo, author of our AI Agenda newsletter, who reported details on that this morning.

30:22Stephanie, welcome back to the show. It's great to have you here. Thanks. Great to be here. How is OpenAI lowering its inference costs? So as we reported this morning, short answer is we don't know exactly how, but we do know that earlier this month, OpenAI engineers that work on inference or this process of running OpenAI's models told colleagues that they had developed certain optimizations that when applied to existing models would reduce the cost of running those models by more than 50%. So that's basically cutting those costs more than half. And so, you know, obviously this is a very big deal because running these models is super expensive and it's, you know, as more and more people use ChatGPT, use these models, those costs are just going to go up.

Read the full transcript

31:12So anything OpenAI can do to basically optimize the process of inference and lower those costs is going to be very helpful for them. Right. And we, of course, had the CEO of DigitalOcean on this week, and he was talking about the ways in which they are helping their customers lower their compute costs, even without any underlying changes to the models. And we talked about the Brian Armstrong post, which I know is one that you highlighted in your agenda earlier this week. Is this the same as Brian Armstrong saying that we were able to half our AI spending without reducing the token use, or is this a slightly different equation?

31:59So I think there are some similarities, but there are also some important differences here. I think you can think about this kind of maybe as like two layers to the stack, right? Like on the lower layer, you have the people that are actually developing these models and are, you know, OpenAI like has access to these huge clusters of chips that the models are being run on. They can do all these very low level optimizations that are like in the weeds of exactly how those chips are programmed, exactly how those models run on top of those chips. So these are like very kind of low level optimizations that takes very specific kind of technical talent to be able to do.

32:39So that's kind of one layer. I think on the layer above that are things that they are doing, but also a lot of, you know, companies that are using open source models, for instance, or models in general can do. So for instance, things like, you know, something as simple as not every task needs a Fable 5 level model or like a, you know, like a 5.0, GPT 5.6 level model. Sometimes you can go with a cheaper model, with an open source model, with an older model. And so those are some of the examples that, for instance, Brian Armstrong talked about in his post where he was saying, hey, we've come up with these different ways to lower the cost of running models.

33:19There are also things that like both OpenAI and also maybe like less technical developers can do things like caching or like batching requests together. So maybe you have like many queries that you're running through the chips at the same time versus one after another. That helps to cut down on costs. So there's really a whole range of like optimizations that can be done from like the most technical in the weeds. Like you need to be like literally physically like with the chips type of thing to much higher level things that any developer can really do to lower their costs. But just to confirm, this second group of technical optimizations that you're talking about, that is the type of optimization that we think OpenAI is doing on their end here?

34:04Or am I still missing the point? So I would say that with this OpenAI example, we really don't know exactly what the optimization is. Okay. Because you talked about this term quantization in your newsletter. And I wondered if maybe that connected with all this. Yeah. So quantization is kind of more in this category of like, it's definitely something that OpenAI, Anthropic, and all these other labs are doing, but we don't know if that's exactly what's going on here. I think that along with other things like KV caching was brought up in the newsletter as like examples of, you know, optimizations that have been talked about publicly that could be going on.

34:45I think the issue here is that even though we know, like, even though the engineers were talking to other employees about, hey, we did these optimizations that cut down on cost by more than a half, the engineers at OpenAI, Anthropic, and all these labs hold the information of, like, what exactly were those optimizations, like, super close to their chest. I mean, in some ways, this is like a very important secret sauce for them that they don't even want to tell other OpenAI employees about because if these things leak, it can very quickly be picked up by other labs, which can also then use that to lower their costs, right?

35:19And as much as the AI race is about capabilities, there's also big cost components. So all these labs are racing to lower the pricing of their models as quickly as possible to compete with each other. So this is something that like they're holding very close to their chest and they're not even telling other OpenAI employees what exactly the optimizations were. Right. Well, so let's zoom out a little bit. So, I mean, how they get the optimization is one half of the story. The other half of the story is what they do with the optimization. If they're able to lower their costs, you spell out two different options in a newsletter.

35:52One is they can pass that on to their customers and lower their costs. The other option is OpenAI is still burning a ton of cash, and it'd be nice to expand that gross profit margin a little bit. I'm just asking you, based on what you know, which of those two do you think is more likely here for OpenAI? So it's hard to say. Because I think they're going to expand their gross margin, okay? I don't think the customers are going to see any of this. That's my prediction. Yeah. I think, okay, so here's my take. Like, my take is that they will not, like, OpenAI is already seen as the company that has, like, better pricing and higher usage limits and is, like, constantly, like, refreshing usage limits.

36:36Like, I think they already have that rep. So they're like, okay, like, we have that rep locked down. If I were them, I would use those savings to help my gross profit margins because of, you know, what you just mentioned. Like, they are burning a lot of cash. They're maybe closer to going public. Like, you know, as we said in the piece, like, they have projected some pretty high gross margin targets for the end of this year. And, like, to reach that, they're going to really need to improve their gross margins to be able to reach that average by the end of the year. I think if Anthropic cuts costs drastically, I think there's a good chance that OpenAI will also cut costs to kind of match or exceed those savings, just to make sure that they continue to have this reputation of being the more cost-effective model provider.

37:20But I think if Anthropic doesn't, then I think they are probably going to try to move those savings into gross profit. It doesn't really seem like Anthropic is lowering their costs, though. We've talked about the switch to consumption-based pricing, and we just had Catherine on as well. She was talking about how they have sort of flipped from an hourly-based pricing model to a token-based pricing model for their relationship with Amazon. on. So my thinking here, and you tell me what you think, but Anthropic doesn't seem to be lowering their costs. OpenAI, I mean, their projected revenue profile, it's actually lower than Anthropic now.

38:05So maybe this is an opportunity. If they can lower prices, maybe they can eat back some of that revenue from Anthropic. Yeah, I think that's a fair point. I do think Like this year so far, Anthropic has gotten a lot of flack for either, you know, raising prices or not lowering prices as much as what developers want. I do think, though, that like they have pretty recently struck some really big compute deals with companies like SpaceX. So I think there is a possibility that as they bring more compute online and kind of, yeah, basically get more access to these AI chips, There's a chance that they will have the opportunity to kind of lower costs or raise kind of usage limits for customers.

38:50But I would say so far, yes, I think it's a fair point that like they haven't maybe done the best job for some developer. Let me ask you one more question before you go. So we're talking about how OpenAI is going to find ways to optimize compute. That means they're not going to pay as much money to the cloud services companies, I imagine, or even to the chip companies because they're finding ways to do more with less. So this is bad news then for the cloud and the chip companies, right? So I wouldn't say that it's necessarily that simple. I mean, this has been like a very long running debate, right?

39:33It's like, this is kind of like the deep seek moment, right? Where we saw, okay, actually people can train models with like way less compute than what we thought. Does that mean that, you know, NVIDIA and these cloud providers are totally screwed? The answer is, it seems to be no so far because, you know, companies like OpenAI, they've already committed to these multi-year contracts. They're already building out these giant clusters of chips. So I think for them, they're kind of like, you know, rather than pulling back on our cloud spending why not take that extra compute that we have now opened up and just try to get more usage get more uh revenue from those chips i think for them it's more about like rather than thinking of it as like um let's cut back on our spending they're gonna be like okay let's try to ramp up sales and usage of our models as much as possible because now we are able to kind of do more with less.

40:29Right. Right. Well, Stephanie, I want to thank you for coming on. That is Stephanie Palazzolo, author of our AI Agenda newsletter here at The Information. That does it for today's show. Reminder, we are on this stream Monday through Friday at 10 a.m. Pacific, 1 p.m. Eastern. If you can't make it then, episodes are available on theinformation.com, on our YouTube channel, or wherever you get your podcasts. Make sure to follow us on social media on X, on Instagram, on TikTok, and on LinkedIn. I'm already excited for our next show tomorrow. Have a great rest of your Tuesday. Bye-bye for now.

From the publisher

Dakin Campbell, AI Finance Reporter, talks with TITV Host Akash Pasricha about the former AWS chief taking the helm of Helix Infrastructure Partners, a $10 billion data center company aiming to compete in the AI infrastructure race. We also talk with The Information's Catherine Perloff about Anthropic shifting its pricing model with Amazon to tokens, which will increase costs for the tech giant. We then chat with Laura Bratton about how Anthropic’s new Claude integration for Slack is sparking cannibalization concerns among Salesforce employees. Finally, we get into OpenAI's secret internal optimizations that cut model inference costs by more than half with our AI reporter Stephanie Palazzolo.


Articles discussed on this episode: 

https://www.theinformation.com/articles/amazon-pay-anthropic-technology-new-deal

https://www.theinformation.com/articles/new-kkr-venture-hunts-deals-clear-data-center-logjam

https://www.theinformation.com/newsletters/ai-agenda/openai-discovers-new-way-cut-inference-costs-half

https://www.theinformation.com/articles/salesforce-employees-worry-anthropics-invasion-slack


Subscribe: 


Sign up for the AI Agenda newsletter: https://www.theinformation.com/features/ai-agenda


TITV airs weekdays on YouTube, X and LinkedIn at 10AM PT / 1PM ET. Or check us out wherever you get your podcasts.


Follow us:

X: https://x.com/theinformation

IG: https://www.instagram.com/theinformation/

TikTok: https://www.tiktok.com/@titv.theinformation

LinkedIn: https://www.linkedin.com/company/theinformation/


Chapters:

00:00 - Introduction

01:13 - Adam Selipsky Leads $10B AI Infrastructure Play

12:56 - Anthropic Hits Amazon With Token Pricing Shift

20:46 - Salesforce Staff Clashing Over Claude in Slack

31:11 - OpenAI Sneaks Out 50% Inference Cost Drop


More from The Information's TITV

All 304 episodes
Anthropic’s Invasion of Slack, OpenAI Cuts Inference Costs in Half, Amazon’s Higher Anthropic CostsThe Information's TITV · 41 min
Listen in VO