The AI Challenges Businesses Are Actually Focused On Right Now

18 Sep 2026 · 30 min · 21 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

How enterprises are responding to the recent AI safety debate, and what practical AI challenges they’re prioritizing now (governance, agent security/identity, evals, legacy integration, architecture changes, and compute/data sovereignty). It also covers Anthropic’s proposed AI-safety metrics (R&D automation, agent oversight, and safety compute), plus market/regulatory ideas like compute concentration regulation and possible antitrust carve-outs for safety coordination.

Guests

No named guests; the episode is a solo host news-and-analysis format.

Key claims

Safety discourse is mostly “XRisk headlines” for media, while businesses focus on operational governance and cybersecurity; AI adoption continues but shifts toward monitoring agents, improving security, and considering owned/open-weight models.

Notable examples

Anthropic’s R&D Automation Index (Claude leading 26% of R&D; 30,000 agents; ~0.002% escalation); Bridgewater CIO Greg Jensen on compute concentration; Microsoft’s 15,000-word code of conduct; Hugging Face and Mistral hacks; OpenAI Astra for Law and Cooley’s GoPublic; Latham & Watkins buying NVIDIA servers; KPMG/UT Austin “AI amplifiers” study; Ramp AI index showing spend shifts.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Current AI Safety Discourse

0:39 to 1:40

Exploration of the ongoing AI safety discourse and public perception.

“And to learn more about sponsoring the show, send us a note at sponsors at AIDailyBrief.ai.”

Anthropic's Three-Axis Measurement

1:40 to 2:47

Anthropic's proposal for measuring AI development pace through three axes.

“The threat of societal destruction makes for a good headline, but the AI researchers issuing the warnings have had a difficult time describing exactly what they've seen inside the labs.”

Understanding AI-led R&D

2:47 to 3:43

Details on Anthropic's R&D Automation Index and its implications.

“Anthropic put together a measure they called the R &D Automation Index.”

AI Oversight and Safety Measures

3:43 to 5:10

Examination of Anthropic's measures for agent oversight and safety resources.

“In March, prior to Mythos, only 1 % of Anthropic's R &D was led by AI.”

Market Implications of AI Safety Costs

5:10 to 6:41

Discussion on how increased spending on AI safety may affect market dynamics.

“This is because safety research consists of individual researchers designing experiments, which is time-consuming even though running the experiments is not particularly compute-intensive.”

Global AI Landscape and Regulation

6:41 to 7:45

Insights on global AI developments and regulatory considerations.

“Will markets reward companies that spend more on safety, or will they punish them for cutting into their own margins.”

Concentration of Power in AI

7:45 to 10:10

Analysis of concerns regarding the concentration of computational power in AI labs.

“Our successors are the AI systems we are creating ourselves.”

Antitrust Issues in AI Coordination

10:10 to 11:58

Exploration of potential antitrust carve-outs for AI safety coordination.

“Either way, let's get going on sorting that out legally.”

Dismissing AI Extinction Fears

11:58 to 12:43

Discussion on contrasting views regarding AI extinction risks and practical challenges.

“Speaking at the same event, she said, This is a classic game theory problem.”

Cybersecurity Concerns with AI Models

12:43 to 13:34

Recent cybersecurity breaches affecting AI models and implications.

“In his view, the dangers posed by AI are, quote, practical engineering problems, and the industry needs to continue focusing on, quote, many of the wonderful things it can do.”
Show all 21 chapters

Gemini 3.8 and AI Model Innovations

13:34 to 14:03

Introduction to Gemini 3.8 and its advanced conversational capabilities.

“Later in the day, Mistral said that they found no evidence of unauthorized access after a thorough investigation, suggesting perhaps that this is the same material from the May break-in coming up for sale again.”

AI-Driven Voice Interfaces in SaaS

14:03 to 15:32

Learn about the potential of voice interfaces transforming SaaS into invisible operations.

“Aside from the smooth conversation style, the model also supports near-real-time visual inputs, automatic detection for 97 languages, and handoff for tool calls to allow it to complete tasks in the background.”

Impact of AI on Workplace Performance

15:32 to 16:21

Discover how early career professionals amplify AI outputs to enhance their performance.

“main episode where we talk about what, after this couple weeks of crazy AI safety discussions, big enterprises are thinking about their priorities with AI next.”

AI Safety Debate and Business Implications

17:06 to 17:32

Examine how the AI safety debate influences enterprise strategies and spending.

“Using it well, on the other hand, is a whole different story.”

AI Safety Debate and Business Implications

17:39 to 18:51

Examine how the AI safety debate influences enterprise strategies and spending.

“Forget local agents and chat workflows waiting on your laptop to be prompted.”

Enterprise AI Governance and Risk Concerns

18:51 to 21:02

Delve into the governance and risk challenges that enterprises face with AI adoption.

“Certainly the safety debate has found its way to the business leader conversation.”

Analysis of AI Spending Trends

21:02 to 23:09

Analyze recent trends in enterprise spending on AI technologies and services.

“business sovereignty points that they'd been trying to make for the past several months.”

Emerging AI Security and Identity Challenges

23:09 to 26:21

Explore the security challenges enterprises face with AI agents and identity management.

“which are already cheaper over the frontier.”

Adapting to AI Innovations in Business

26:21 to 28:01

Learn how businesses are adjusting their strategies in response to rapid AI innovations.

“Now I think this sets really interesting context for watching some of the competitive dynamics right now around how different types of actors are trying to appeal to different types of businesses.”

Adapting AI Strategies in Business

28:01 to 28:41

Learn how businesses are leveraging open-weight models for competitive advantage.

“While the labs debate how quickly intelligence should advance, software companies should be racing to commoditize the intelligence we already have.”

Navigating Regulatory Challenges in AI

28:41 to 29:39

Understand the impact of AI safety discourse on enterprise strategies.

“The pace and challenges of adoption within the enterprise were always unique and distinct to them, and the challenges of big institutional inertia that comes with them.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00It has been a heck of a last couple of weeks when it comes to the AI discussion in society. And yet in all of that, one group that's left trying to figure out if anything has actually changed for them or if they are just on the same path that they were before is the businesses and enterprises that have been trying to figure out how to maximize AI for their own value for at this point a number of years. Today we're discussing both how enterprises are thinking about, if at all, this new era of AI safety and also digging a little bit deeper to find out what the real concerns that businesses have and the real AI challenges they're facing right now.

0:33As always, we're in a moment where new challenges are creating new opportunities as well. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI.

1:03and Apple Podcasts. And to learn more about sponsoring the show, send us a note at sponsors at AIDailyBrief.ai. We are now two weeks into the AI safety discourse absolutely dominating the conversation. For those who think that awareness of these issues had been sorely lacking, it has been a very good period. Although now seeing polling numbers that suggest that something like 17 % of Americans are completely convinced that AI is going to end humanity, there is certainly some reasonable concern that we might have over-calibrated. Holding that discussion aside for a moment, one thing that I think almost everyone is looking for is increasing specificity.

1:40That is, specificity of policy, but also specificity of monitoring, and it's to that that Anthropic speaks with their new proposed set of measurements to help the public understand just how quickly advanced AI is developing. Once you move past the scariest headlines, this month's safety debate has been largely about recursive self-improvement, and the idea that AI development is A, moving too fast, but B, really, about to move much more quickly. The threat of societal destruction makes for a good headline, but the AI researchers issuing the warnings have had a difficult time describing exactly what they've seen inside the labs.

2:14To better inform the public, Anthropic has proposed a three-axis measurement to understand the current pace of AI development. The first axis is AI's ability to build the next version of itself. The second is Anthropic's ability to oversee and intervene in the actions of agents. And the third is the scale of resources that go into AI model development. Now, Anthropic explicitly notes that these measurements only look at how models are built. In other words, they are measuring only the inputs to model development, but they argue that these inputs are correlated to the growth in model capabilities.

2:46On the first measure, the use of AI models in future model training, Anthropic put together a measure they called the R &D Automation Index. It seeks to measure every aspect of model training, rate how automated each task currently is, and aggregate those rankings to express them as a single number. By Anthropic's measure, Claude now quote-unquote leads 26 % of their R &D work and quote-unquote collaborates on more than 90%. Anthropic's definition of AI-led R &D was tasks where a human provides a high-level goal and the AI completes it end-to-end with human oversight. And here you see the challenge even as they are trying to be precise.

3:23They are talking about AI's leading R &D, but that definition of leading still includes the human providing a high-level goal rather than the AI determining the goal, which is what I think some might assume when they hear AI-led R &D. For their part, Anthropic was careful to note that, quote, Claude is not operating fully autonomously for any measured subset of AI R &D work. Still, they noted that research automation has massively increased since Mythos finished its training run and has continued to ramp in recent months. In March, prior to Mythos, only 1 % of Anthropic's R &D was led by AI. Mythos boosted this to 12 % in May, and it has doubled since then.

3:57Regarding oversight of agents, Anthropic measured coverage, review latency, and escalation rates within their agent oversight systems. They said that around 30 ,000 agents are currently doing research and engineering work at any particular time. Anthropic claims to have 100 % coverage of agentic actions for this work, instant AI review of any flagged actions, and around a 0.002 % escalation rate for real-time review, meaning around 1 in 47 ,000 actions are blocked by the monitoring system. Anthropic also maintains an after-the-fact review system with 100 % coverage. This system flags around 100 ,000 agentic transcripts per week that are then parsed for false positives with around 50 escalated to human review.

4:39In other words, around 1 or 2 transcripts per thousand cause any material concern. Finally, on resource allocation, Anthropic expressed their measure in terms of a percentage of overall compute dedicated to safety systems. Around 6 % of the compute allocated to AI-assisted R &D went towards safety, and around 12 % of the compute allocated to AI-led R &D went towards safety. Anthropic noted that this measure doesn't include classifiers that constantly monitor agentic activity and is a, quote, imperfect proxy for how much a company focuses on safety. This is because safety research consists of individual researchers designing experiments, which is time-consuming even though running the experiments is not particularly compute-intensive.

5:19Overall, Anthropic's goal was not to provide perfect measures. Instead, it was to define and propose a set of metrics that any AI lab could report as part of a regulatory system. They conclude, As the world considers pacing the frontier, we should do everything possible to minimize the gap between what frontier labs know and what the public knows. This means better measuring the development of AI, reporting on it publicly, and giving society an opportunity to decide how to use this information. We hope to model that transparency by releasing these measurements, and will continue to do so. Now, the response to this announcement, I think, shows the good, bad, and difficult of this particular moment.

5:54On the one hand, it is the rare move where people on all different sides of the AI safety debate largely think that this existing is better than it not existing. AI safety researcher Jeffrey Ledish writes, Great to see this from Anthropic. A few weeks ago, I wrote that the company could be a lot more transparent, and I appreciate how they've stepped up, both with this and the recent incident investigation report and misuse report. At the same time, there are plenty of people who also wanted to see more. Prime Intellect's Eli Bakaush writes, Interesting that OpenAI gives us the breakdown of usage per R &D task, and Anthropic gives us automation levels.

6:25Would be great to combine both and get the evolution of automation per R &D task. There are also some who argue that the self-reporting, while a fine start, isn't enough. Arun Rao writes, we need standard measures across all labs to measure progress to RSI and publicly report on it weekly to monthly. Now, one discussion that we'll see a lot more of, especially as markets digest all of this, is that one way, perhaps a cynical way, perhaps a realistic way, to look at an increase in safety monitoring is to view it as really expensive overhead for R &D, meaning that these measures of percentage of spend on safety is basically a measurement of that overhead.

6:59Will markets reward companies that spend more on safety, or will they punish them for cutting into their own margins. This shows another complication for having these companies operate in a public market environment. For what it's worth, this week also reminded that recursive self-improvement is not just the province of the U.S. labs. ZAI released a blog post on Thursday called Towards Recursive Self-Improvement, How GLM Built Its Own Inference Infrastructure. The first sentence reads, as we develop GLM, the model sometimes exhibits capabilities that surprise us and even unsettle us. The most recent moment that shook us, GLM is increasingly helping build AI itself.

7:33We watched the model complete an infrastructure task that would have previously taken a team of experienced infrastructure engineers weeks. When we realized that this work would directly change how the next generation of models is trained, we became even more convinced. Our successors are the AI systems we are creating ourselves. It's certainly beyond the scope of the headlines, but it is a reminder that any slowdown discourse that does not include China is basically no discourse at all. Meanwhile, if I am correct that we have now entered this new negotiation period for the next generation of AI, where it is no longer just the AI lab that determine our pacing.

8:06Lots of people are now stepping in with proposals for what the new overall system might look like. Bridgewater CIO Greg Jensen made headlines this week by suggesting that concentration of power is the major AI risk to be regulated. Bridgewater is one of the world's most successful hedge funds and has used AI extensively over recent years. In an interview with The Information, Jensen discussed his views on where AI risk actually resides and how he would approach regulation. He said, there's a huge problem with open source models because you can train them. We do this at Bridgewater. It's extremely effective reinforcement learning training on powerful open source models.

8:38It's extremely powerful, a great technology, and at the same time, clearly a very dangerous one. You don't know what people are doing with them. There's no way to track that. So it's dangerous from that perspective. We haven't even begun the conversation on how to deal with that issue. And yet, in Jensen's view, that risk pales in comparison to the risk posed by the frontier labs themselves. He continued, we have to deal with the concentration of power. In two years, OpenAI and Anthropic are going to control 35 % to 50 % of the world's compute. That's a crazy outcome for a society to allow on something as powerful as compute.

9:11Would we let one entity control that much of some other form of energy or commodity? Now, one of the common observations over the past month has been that the Hugging Face incident and other security breaches like it require immense amounts of resources to power agent swarms. It's unclear that they could be replicated by threat actors outside of the frontier labs, at least with current models, because of limits in their access to compute. Jensen argued that part of the solution needs to be clear guidance on liability for actions taken by AI so the companies and individuals understand the risks they're taking.

9:41When asked how he would deal with the concentration of power, Jensen responded, we should take anybody that has more than X percent, let's say 5 percent, of the world or U.S.'s compute resources. say, okay, those are systemically important institutions. We're going to put them into some sort of thing like we do with systemically important banks and say there's going to be regulation. I also think it's fine to say we're going to have caps on how much you can own. There's a cap on how many commodity futures you can own. There's a cap on much less important things. We may be able to do that with current law.

10:08You may need new laws. Either way, let's get going on sorting that out legally. There should be a bipartisan recognition that you shouldn't want monopolistic control on what I think most people will agree is one of the most important resources in the world. Now, I could spend the next week's worth of episodes discussing the implications and challenges of trying to apply G-SIB-style regulation to existing compute, but what I like about this discourse is that not only does it leave behind overly simplistic binaries, it starts to ask questions about where the actual locus of power is. Is the problem in the models, or is the problem in the power to run the models?

10:39Those have very different intervention points if we're trying to stop bad outcomes, and that's exactly the sort of conversation we need to be having. Now, one additional dynamic of the safety question that has been playing out is a legal question of whether the labs can actually coordinate on safety or whether that implicates antitrust issues. The Trump administration is reportedly considering an antitrust carve-out to allow frontier labs to collude on safety. Now, the theory that this deals with is basically that an AI slowdown could mean a coordinated reduction in research spending and therefore an increase in profitability at the expense of the consumer.

11:12In the abstract, and in very loose analogy, it would not be dissimilar to if Apple and Samsung agreed to stop working on new phones. Associate U.S. Attorney General Stanley Woodward said the administration is considering an update to interagency antitrust guidance to provide a carve-out for AI safety coordination. The guidelines already allow for coordination on cybersecurity risks, so a change would extend that guidance to AI safety. Interestingly, Woodward said that although multiple AI executives have publicly called for an antitrust carve-out, no one has contacted his office on the matter as of yet.

11:44Even without guidance, though, Woodward has indicated that the DOJ has no objections to a coordinated slowdown, commenting, It doesn't occur to me that coordinating on cybersecurity or security is anti-competitive. Europe's antitrust chief, Teresa Ribera, agreed. Speaking at the same event, she said, This is a classic game theory problem. When the risks are shared, cooperation is in everyone's interest. At the same time all of this is happening, there are certainly still many out there who are trying to tone down the tenor of the conversation in general. Legendary AI researcher and Coursera co-founder Andrew Ng, in an interview with Bloomberg TV, dismissed extinction risk as science fiction and warned that AI doomers have pushed this narrative numerous times over the past decade in a bid to gain publicity for their cause and help shape regulation.

12:26He said, I was quite dismayed over the past two weeks. This wave of PR has kicked up again for probably similar purposes. Andrew acknowledged that there are genuine risks associated with AI, largely around cybersecurity, but said, this recent fear about AI leading the human extinction and so on is much more science fiction than science. It's very damaging. In his view, the dangers posed by AI are, quote, practical engineering problems, and the industry needs to continue focusing on, quote, many of the wonderful things it can do. And speaking of real here and now cybersecurity issues, Mistral has been hacked for the second time, with their intellectual property now available for purchase on the dark web.

13:03In May, around 5GB of internal source code and around 450 private repos were exfiltrated and offered for sale, and now it's happened again, with hackers offering full source code, internal development files, web app code, and additional proprietary information. An ex-user called Benny said that they contacted the seller and confirmed the file dump included model weights, post-training pipelines, and dataset construction. The hacker was asking$25 ,000, but deleted their post shortly afterwards, likely either because they found an exclusive buyer or got sloppy with their OPSEC and needed to disappear.

13:34Later in the day, Mistral said that they found no evidence of unauthorized access after a thorough investigation, suggesting perhaps that this is the same material from the May break-in coming up for sale again. Lastly today, one model release that went a little under the radar this week was Gemini 3.8 Live Extended Thinking. The model is another live speech model, meaning it can process continuous conversations rather than using a turn-based structure. It topped the Artificial Analysis speech-to-speech index, beating out GPT Live 1 Astra and Grok Voice Thinkfast 2.0. Aside from the smooth conversation style, the model also supports near-real-time visual inputs, automatic detection for 97 languages, and handoff for tool calls to allow it to complete tasks in the background.

14:14Tim Messerschmidt, a developer relations lead at Google, published a cool tech demo showing the model powering a Ricci mini-robot. Tim demonstrated the model's ability to keep up a seamless conversation that we use between English and German. Speculating about the implications, Greg Eisenberg wrote, Are invisible interfaces coming? Google just announced Gemini 3.8 Live. It can talk through a task with you and then keep working after the conversation ends. I think 90 % plus of vertical SaaS will need a voice front door. By that, I mean the way you use the software becomes talking to it. And the typing, clicking, and form filling happens on the other side without you.

14:47So a contractor standing on a job site just says what went wrong out loud, and by the time he's back in the truck, the quote is sent, inventory is checked, the CRM is updated, the customer got a text, and anything risky is flagged for him. Kind of the dream, right? The same thing works for nurses, dispatchers, recruiters, brokers, insurance agents, etc. The person talks and the agent finishes the admin. Lots of opportunities here to build voice-first businesses. I think this is how vertical software becomes invisible. Nobody logs in, nobody fills out a form, and nobody learns your interface. You just talk and the work gets done behind you.

15:17This is a glimpse of where SaaS is going. Not fully there yet, but it's coming. In this weekend's long read slash big think episode, I get a little bit more into this particular shift, as well as a bunch of other shifts in how we use AI. But for now, that's going to do it for the headlines. Let's move over on into the main episode where we talk about what, after this couple weeks of crazy AI safety discussions, big enterprises are thinking about their priorities with AI next.

15:45A new study from KPMG in the University of Texas at Austin found that when people work with AI, similar skills don't guarantee similar outcomes. Researchers studied more than 500 early career professionals and found that the best performers consistently amplified the value of AI by guiding, evaluating, and refining its outputs. These top performers, called AI amplifiers, weren't defined by what they knew alone, but by how they worked with AI. Learn more about what separates AI amplifiers from everyone else at kpmg.com slash us slash AI amplifiers. Blitzy's understanding of massive codebases unlocks autonomous security fixes, modernization, and new features.

16:26So what happens when there's no legacy code at all? Greenfield is supposed to be the easy part. Clean slate, no technical debt. But even Greenfield moves at human speed one sprint at a time. Blitzy changes the unit of work from the developer to the project, autonomously planning, building, testing, and validating entire applications from scratch. Hundreds of thousands of lines of production-ready code. One Blitzy customer stood up a brand new application, 534 ,000 lines of code, compressing a 65-week roadmap into two weeks. Another shipped an entire application with no front-end engineer. Legacy or Greenfield, the answer is the same.

16:58Software at the speed of compute. Build what's next at Blitzy.com. That's B-L-I-T-Z-Y dot com. At this point, it's no longer a question of whether companies are actively using AI. Using it well, on the other hand, is a whole different story. Robots and Pencils, though, is a company that I can point to that is actually built for this time. They're an applied AI engineering firm working directly with clients on problems that matter to the business, not experiments that live in a slide deck. Every engagement starts by working backwards from the outcome a client actually needs. If you're trying to tell real AI engineering apart from noise in this space, that's the difference maker.

17:32Head to robotsandpencils.com. This episode of the AI Daily Brief is brought to you by HyperAgent, where you run fleets of agents your team can manage together. Forget local agents and chat workflows waiting on your laptop to be prompted. Hyperagent deploys always-on agents in the cloud, doing real work across the tools your team already uses. Marketing agents turn competitor moves into landing pages. Sales agents enrich leads, draft emails, and updates the CRM. Ops agent chases the paperwork and tracks the budget. Every agent has access to shared context and follows your rules about scope and approvals.

18:03It's time you add agents that feel like teammates. Hire yours at Hyperagent. Get$100 in credits at hyperagent.com slash AI Daily Brief.

18:16Welcome back to the AI Daily Brief. The background context for today's episode is, of course, the AI safety debate, which has completely broken containment over the last couple of weeks. And yet, even as we have seen the political resonance of this issue absolutely explode and polling showing very dynamic, fast-shifting attitudes around it, enterprises and businesses of all shapes and sizes are still stuck out here, figuring out whether it has any real implications for them or whether they just need to keep on keeping on. Today we're looking at some specific answers to that question, as well as a few of the broader AI challenges that enterprises and businesses are actually focused on right now.

18:55Certainly the safety debate has found its way to the business leader conversation. When the Wall Street Journal asked for a show of hands at the WSJ Technology Council Summit on Monday, only a few people said that they were worried that AI might kill us all. At the same time, around half of attendees said that they were in favor of slowing down frontier AI research and prioritizing better guardrails. This sort of comports with my thesis that I shared I think in last Thursday's big AI safety episode, that while I believe that the markets would view any sort of slowdown as initially a big risk for AI, I actually think that there is a very compelling economic counter-argument that a slower pace of new development might actually increase the amount that companies were spending on AI.

19:36The thesis was basically that the incredible speed and pace of change actually in some ways creates a disincentive for companies to try to do comprehensive transformation on the logic that they're going to spend all this time and energy on transforming into something that isn't even relevant anymore by the time the transformation is complete. Indeed, moving back to the Technology Council Summit, on the panels, the big takeaway was that an AI slowdown actually doesn't really have that many implications for the way enterprises are using AI right now. Vishal Talwar, the president of FedEx DataWorks said, It's in our hands and it's up to us to apply AI for good, and I think it can do a lot of good for society.

20:11The discussion basically still centered around the need for prudent AI governance around normal business risks rather than the existential risks that are dominating the media discourse. Up north at the Canada Investment Summit, BlackRock CEO Larry Fink was far more worried about the data center backlash than XRisk. He warned that construction delays could make AI the, quote, domain of large firms. Fink added, the faster we can build out more capacity, the more we can democratize and make it available for everyone. In the startup world, meanwhile, founders are starting to think about the governance and monitoring tools that businesses will need as agents get more powerful.

20:44One venture investor told the information that they're beginning to focus on startups building tech that improves model security and infrastructure. In other words, private markets are responding to all of this concern by funding startups that can address specifics around this concern, which to me is part of exactly what you want to see. Microsoft also tried to bring the AI safety discussion back to some of the broader AI business sovereignty points that they'd been trying to make for the past several months. This week, the company dropped a very extensive 15 ,000-word code of conduct document that has, as they put it, a single overriding objective, that humans must retain meaningful control over AI so that it can help people live healthier, happier, and more productive lives.

21:23A lot of the discourse around this document was its outright rejection of AI consciousness and a prohibition of designing AI to even imitate consciousness. But in his post about this on X, Microsoft CEO Satya Nadella also brought up the implications for businesses themselves, writing, For firms, it's imperative that they retain full control over their unique and tacit knowledge. Every organization should be able to build its own continuous learning loop and hill climbing machine without becoming dependent on any one model provider and have the ability to embed its own knowledge into models and weights they control.

21:54Basically, part of this is power concentration in the firms, and a way to deal with power concentration in the firms is to not surrender power to the firms in the form of your unique and proprietary data. Now, meanwhile, outside of the AI safety conversation, there had started to be some debate about AI market signals in the form of enterprise spend. On September 9th, Ramp released its latest AI index that found that AI spend declined among the top 1 % of businesses spending on AI. In August, wrote Ramp lead economist Eric Karazian, the top 1 % of businesses spent$7.2 ,000 per employee per month, which was down 10 % from a July peak of 8 ,000.

Read the full transcript

22:30Now, we've talked a lot about the limits of Ramp data. Ramp has a very, very highly concentrated tech-forward early adopter type of audience. It is also pushing products that are specifically about cost efficiency. But of course, when you take all those caveats, it still provides an interesting and important signal. Now, when it comes to their analysis in this particular area, I personally think that they are underestimating summer seasonality as a driving force, but their take is that this is about the most sophisticated AI users getting more adept at complex model architectures that don't rely on the most expensive models at all times.

23:01Error wrote, My take is that it's not because of Chinese open models, it's model wars. Price cuts plus a growing share of spend is shifting to standard and light models which are already cheaper over the frontier. Now, speaking of RAMP not necessarily having exactly the right analysis all the time even though their data is really valuable, While they had previously caused a bunch of frantic headlines when they showed that businesses weren't adopting Fable, arguing all sorts of reasoning for that besides the obvious one, which was data retention policies, but to their credit in more recent reporting, they found that Fable 5.1, which got rid of the data retention requirements, had started to make up 22.5 % of enterprise spend and was rising very quickly.

23:37But what about overall? On September 10th, Box's Aaron Levy wrote a long post on X about the issues that he was hearing about from executives across industries, including banking, media, information services, and insurance, when thinking about AI and agents in the enterprise. The big trends that Aaron heard about were cyber, model battles, agent security and identity, process re-engineering, architecture adjustment, evals, and the hurdle of legacy systems. On cyber and security, he wrote, Everyone is nervous about the growing rate of vulnerabilities coming at them from AI and the implications of the OpenAI Hugging Face incident.

24:10The conversation is not as existential as it is in Silicon Valley, but still highly concerned and pragmatic about what to do about it operationally in their environments. Lots of new discoveries due to AI and still hard to keep up with all the changes they have to execute now. Adding in what he wrote about agent security, he continues, Somewhat tied to Hugging Face, there's much more awareness to the new challenges around agent security and identity management in a world where agents are trying to get into every system they can. In a perfect world, enterprises could set up identities for all their agents and control what they're doing, but of course sometimes the agent needs to act exactly as the user as well.

24:41This is certainly something that we've seen a lot in our conversations in the podcast context, as well as super intelligent. And interestingly, a lot of the security concerns aren't just about malicious actors. It is to some extent rooted in just the general power of these systems. Now that it's not just engineers who have access to agents, many companies are finding that the agents are powerful enough that they escape the containment of the non-engineers that are using them, even if those non-engineers aren't trying to do anything problematic. This is also showing up in the numbers. Once again, looking at ramp data, lead economist Eric Karazian writes, one area companies are increasing their spend is AI security software.

25:15In the wake of the Hugging Face hack, three of our trending software vendors make software specifically designed to monitor agents in production. Now, he did point out that these specific vendors might not have done anything to stop the Hugging Face hack, but it's clearly a category of focus for business buyers as well as startup builders. Another interesting area that Aaron talks about is the nascent exploration of open source and alternative model architectures. He wrote, Most companies are deploying multiple frontier models within their enterprise, too hard to standardize on anything, and seeing different preferences across their teams and use cases.

25:46But the dollars are still concentrated on just a few vendors. Open weight's still in infancy at scale in most of these organizations, often due to lack of domestic frontier open source options. Plenty of appetite for more options here, but so far, few places to go. Now paired with that, I think, is Aaron's observation that there is a ruthless adjusting of architectures. Most companies, he wrote, had examples of changing systems out multiple times just in the past year or two with different vendors. I probably haven't heard, we tried X and it didn't work, so I've gone with Y, more than in today's environment.

26:15The lesson here is that because innovation is happening so fast, no one hangs around until a vendor gets something right. They just move on to the next one. Now I think this sets really interesting context for watching some of the competitive dynamics right now around how different types of actors are trying to appeal to different types of businesses. Labs like OpenAI and Anthropic are clearly trying to keep everything consolidated in their own environment. Part of that, especially for OpenAI, has been a real focus on cheaper models and pushing the price of their models down as far as they possibly can.

26:44That's something we've seen a lot over the past couple of weeks. But they also are, as they have been all year, focused on vertical solutions, such as the newly launched Astra for Law from OpenAI. Alongside the Astra for Law launch, OpenAI and law firm Cooley also co-launched a product that they called GoPublic that's designed to draft S1 filings, which are the documents that companies must submit to the SEC before they can go public. And yet, as if to give us a perfect comparison of the different types of options that different companies are taking, another law firm, Latham & Watkins, was recently reported to be buying NVIDIA servers to set up their own in-house systems, specifically as an alternative to models from OpenAI and Anthropic.

27:20Latham, which is the U.S.'s second largest law firm, said, quote, have information that is so sensitive, client information that we really want to protect, we don't want to put it on any cloud vendor. Seeming to make the point that one of the ways that enterprises can stay out of the fray of the AI safety discourse is to own their own models, Mistral CEO Arthur Mensch posted, don't pace building and owning your own AI models and systems as an enterprise, and there will be no doomsday for you. Foundation Capital's Jaya Gupta also thinks that this pacing the frontier moment could be a good one for the software incumbents.

27:51She wrote, if you're the CEO of any software company and you're not offering open-weight models as a skew right now, you're asleep. Pace the Frontier may be the greatest invitation software incumbents have ever gotten. While the labs debate how quickly intelligence should advance, software companies should be racing to commoditize the intelligence we already have. Pharma and banks are already picking up open-weight models partly for margins, partly because a revocable lab API is a dependency they increasingly don't want. AI natives and tech companies that care about cost of goods sold are doing the same.

28:18Most software companies that tried had failed attempts because the open-weight models sucked. But now open-weight models are good enough. I believe that every major software company should become a model factory for its own vertical. Own the evals, post-train open-weights on the workload it uniquely sees, serve those models to its customers, and use production feedback to continuously improve them. And so on the one hand, to some extent, the response from businesses to the AI safety discussion seems to be business as usual. The pace and challenges of adoption within the enterprise were always unique and distinct to them, and the challenges of big institutional inertia that comes with them.

28:51But on the other, there are some ways that the conversation is reinforcing trend lines that we're already starting. Companies will need to spend more time and resources on cyber and security issues. That was always coming, but the point has been made even more crisply now. And while already there were some compelling reasons to explore and consider open weights or more owned model alternatives to just getting in bed with the big vendors, there are now even more reasons to be willing to walk down that path, not least of which is the increasing likelihood of regulatory disruption. I think that if I had to summarize my advice in a single thought, it's that while overall, the changing AI safety discourse doesn't really impact the short term for enterprises, it certainly reinforces the fact that the companies that are willing to try the hardest things, like actually investing in their own owned architectures, have even more potential to differentiate from their peers and competitors than they did before.

29:41I will, of course, continue to watch these trends as they evolve, but for now, that's going to do it for today's AI Daily Brief. Appreciate you listening or watching as always. Until next time, peace.

From the publisher

While AI safety dominates the headlines, businesses are focused on agent security, shifting model choices, and control over their own data. NLW explores how the slowdown debate could accelerate the case for companies to build and own their AI systems. In the headlines: Anthropic proposes new transparency metrics, Washington considers an antitrust carve-out for AI safety coordination, and Google advances conversational AI with Gemini Live.

Multiplayer AI Sprint - ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://multiplayerai.ai/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Brought to you by:

KPMG – Research from KPMG and the University of Texas at Austin shows the highest-impact AI users treat AI like a reasoning partner — and those skills can be taught at scale. Learn more at ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://kpmg.com/us/Sophisticated⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Harbor - Invest in the AI ecosystem. ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.harborcapital.com/aidaily⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Hyperagent - Hire a team of always-on agents. New users get $100 in free credits. ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠hyperagent.com/aidailybrief⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Rackspace Technology- One accountable partner to build, operate and run your full enterprise AI stack ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.rackspace.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Section - Section turns AI investment into workforce transformation and ROI - ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.sectionai.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Blitzy - Want to accelerate enterprise software development velocity by 5x? ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://blitzy.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Robots & Pencils - Cloud-native AI solutions that power results ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://robotsandpencils.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

The AI Daily Brief helps you understand the most important news and discussions in AI.

Newsletter: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://aidailybrief.beehiiv.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Interested in sponsoring the show? sponsors@aidailybrief.ai


More from The AI Daily Brief: Artificial Intelligence News and Analysis

All 1,099 episodes
The AI Challenges Businesses Are Actually Focused On Right NowThe AI Daily Brief: Artificial Intelligence News and Analysis · 30 min
Listen in VO