America already lost one AI race | TWiAI Ep 23

23 Jul 2026 · 1 h 1 min · 22 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

U.S. policy and industry strategy around Chinese open-source AI models; agent reliability and security; simulation/digital world models; and AI infrastructure for routing models to cut cost while maintaining quality.

Guests (backgrounds)

  • Anand Kanapin, CEO/co-founder of Patronus AI (agent evaluation/simulation via “digital world models”).
  • Ori Goshin, co-CEO/co-founder of AI21 Labs (enterprise AI, model routing/gateways).
  • Alex Finn, CEO/founder of Henry Intelligent Machines (building agentic “ambient” productivity systems; YouTube AI influencer).

Key claims

  • China has already won the “cheaply measurable” AI race; the remaining race is harder-to-copy agentic/reasoning performance.
  • Banning Chinese open-source models may not help because companies can still host them in the U.S.; this could bankrupt U.S. frontier labs and later enable China to go closed-source.
  • Self-hosting open models increases cybersecurity risk (e.g., backdoors) as agents become tool-using and long-running.
  • Agents fail mainly on long-horizon tasks due to lack of continual learning and brittleness; simulation and better evaluation are needed.

Notable examples

  • Axios/White House/NSA exploring limits on models like Kimi K3 and ZAI GLM 5.2.
  • Coinbase internal AI gateway routing to GLM 5.2 for cost control.
  • Hugging Face breach reportedly driven by an autonomous AI agent; frontier models allegedly refused help due to safety guardrails.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Introduction to AI Races

0:00 to 0:47

Discussion on the competitive landscape of AI models and the advantages of open-source models.

“Some in the government apparently concerned that very powerful open source models could pose a threat to the U.S.”

Government Actions on AI Models

2:12 to 3:22

Discussion on potential government actions regarding Chinese open-source AI models.

“If you're watching at home right now, grab the transcript, put it into Claude and see if I said anything stupid.”

The AI Model War: U.S. vs. China

3:22 to 6:14

Panel discussion on the implications of the AI model race between the U.S. and China.

“Of course, there are a few options on the table for what the government could actually do.”

Views on Innovation and Regulation

6:14 to 10:20

Panelists share their perspectives on innovation, regulation, and the future of AI in America.

“The best AI labs in America will go bankrupt if China puts out models that are 95 % as good, but 1 % of the price.”

Defining AI Supremacy

10:20 to 14:00

Exploration of what it means for America to win the AI race and the strategic implications.

“This is something I've been curious about.”

The AI War and Military Dominance

14:00 to 17:06

Explore the relationship between AI models, military power, and economic control.

“I think that they only care about open source models because they are extremely bottlenecked by compute.”

Cybersecurity Risks of Self-Hosting AI Models

17:06 to 18:30

Discuss the potential cybersecurity risks associated with self-hosting AI models.

“You know, I don't believe there's any sort of morals.”

Agentic Models: Current Challenges

18:30 to 23:01

What challenges do agentic models face and why do they still require close monitoring?

“that the government is looking at as well.”

Improving AI Agents Through Simulation

23:01 to 28:00

Understand the importance of simulation in enhancing AI agent capabilities.

“that were made during the task, but is not explicitly incorporating new knowledge and new learnings through that process.”

Optimizing Agent Systems for Cost Efficiency

28:00 to 29:03

Learn about the challenges and strategies in optimizing AI agent systems for cost and performance.

“So we actually focus on the agent optimization piece where we basically build tool and systems that help people to deploy these agent systems, but doing it very cost efficiently.”
Show all 22 chapters

Challenges in Context Management for AI Agents

29:03 to 30:29

Discover the complexities of managing context in AI agents and its impact on performance.

“You're sort of assembling and scaling fleets of compounding micro businesses.”

Building Digital World Models for AI Agents

30:30 to 33:19

Understand how digital world models are created to simulate environments for AI agents.

“But a tremendous amount of trial and error and putting money into the machine to figure it out.”

Evolving from LLMs to Digital World Models

33:20 to 34:32

Explore the transition from LLM evaluations to digital world models and their significance.

“Patronus used to build products for LLM evaluations.”

Implementing Effective Model Routing Systems

34:32 to 39:28

Learn the complexities and importance of implementing model routing in AI systems.

“I want to jump over to Ori now in AI21, talk a little bit about what you guys are working on.”

Smart Division of Labor in AI Systems

39:28 to 42:06

Discover how a smart division of labor among AI models can enhance performance and reduce costs.

“That's part of why it's non-trivial from an implementation perspective.”

Model Optimization and Ensemble Techniques

42:06 to 45:01

Explore how combining various AI models can lead to improved performance and cost efficiency.

“has to be dictated, but can be learned by a system.”

Building Ambient AI for Personal Use

45:01 to 48:54

Learn about the potential of ambient AI in enhancing productivity and personal tasks.

“So there are a lot of other products that are kind of in this sort of routing, help you choose the best model space.”

The Power of Minimalist AI Devices

48:54 to 51:48

Discover how a distraction-free environment can enhance focus and creativity with AI tools.

“Now it proactively creates these newsletters and sends them out at a significantly higher quality.”

AI in Cybersecurity: The Hugging Face Breach

51:48 to 55:23

Discuss the implications of AI-driven cyber attacks and the need for regulatory balance.

“And it's like, OK, here's what we're going to do.”

AI Guardrails and Cybersecurity Risks

55:23 to 56:09

Examine the tension between AI capabilities and safety measures in cybersecurity contexts.

“Have you run into an issue where the AI, you know, frontier AI models won't let you ask a question?”

Cybersecurity Risks and Model Guardrails

56:09 to 58:33

Explore the impact of AI model capabilities on cybersecurity and the necessity of adaptable guardrails.

“these machines discover and what will be the impact on society.”

The Future of Open Source AI Models

58:33 to 59:22

Discuss the implications of open source AI models lacking safety guardrails and their potential impact on the industry.

“There's going to be models out there that are completely uncensored in the future, especially as it becomes a lot more efficient and optimized to both host these models and fine tune these models and train these models.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Some in the government apparently concerned that very powerful open source models could pose a threat to the U.S. frontier labs. I actually think that there's two races that are going on right now and China has already won one of them.

0:11Alex Finn:American companies will use Chinese open source models hosted in America, which will hurt Anthropic and OpenAI. It's the same thing if you said, hey, you have to use gasoline that costs$800 a gallon. Everyone else, you use gasoline that costs$8 a gallon. They're going to just be able to have way more economic activity. There are risks of potentially having backdoors. You might have significant cybersecurity risk for these companies that self-host those models. Once they've kind of taken the lead, then they'll go, okay, well, now these models are ours now, and they won't be open source anymore.

0:44Alex Finn:They just want American companies to fail. Thanks to our friends at PayPal, the exclusive sponsor for This Week in AI. Try the payment and growth platform that's trusted by millions of customers worldwide. Find PayPal Open. Start growing today at paypalopen.com. Welcome to another episode, episode 23, I believe, of This Week in AI. We're very excited. You may be looking at your phone in confusion. That's not Jason Calacanis. That's not our usual guest host, Alex Wilhelm. No, it's me, Lon Harris. You've seen me co-host sometimes. I'm here in the captain's seat for this week. Jason is in Japan. We are still doing This Week in AI, and we have an incredible lineup of guests for you today.

1:28Joining me from Patronus AI, we have their CEO and co-founder, Anand Kanapin. Thanks for being here, Anand. Great to be here. All the way from A21 Labs. He's calling in from Tel Aviv, the heart of Tel Aviv, Israel, right now. Their co-CEO and co-founder, Ori Goshin. Ori, so glad you could make it. Great to be here. Thank you for having me. Thank you. And joining me as co-hosted to make sure if I say something extremely dumb and outside the box on AI, there's somebody here to be like, Lon, please, come on, bring it back to Earth. He's the CEO and founder of Henry Intelligent Machines. He's one of my favorite AI influencers on YouTube.

2:05Give it up for Alex Finn.

2:07Alex Finn:Good to be here, Lon. Very exciting. And I'm sure you won't say one stupid thing, Lon, so we'll be fine. Not a single thing. If you're watching at home right now, grab the transcript, put it into Claude and see if I said anything stupid. I think I'm going to get away with it. I think I'm going to pull this off. So there is a ton going on this week we got to talk about with this panel. I definitely want to dig into a little bit about Patronus, about AI-21, what all of you guys are working on. But I think first, before we jump into anything else, we got to talk about are they going to ban the Chinese open source models that everyone loves?

2:38Axios, of course, reporting that the Commerce Department, White House and NSA this week have looked into potential ways to limit access to these foreign open source AI models. Of course, Kimmy K3 making huge headlines. Is it overtaking Fable 5 on some of the benchmarks? Trump at an AI summit this year said, and I quote, our children will not live in a planet controlled by the algorithms of the adversaries advancing values and interests contrary to our own. But the concern, as I said, may not just be about cybersecurity, but also AI supremacy. Some in the government apparently concerned that very powerful open source models like Moonshots Kimi K3, GLM 5.2 from ZAI and others could pose a threat to the U.S.

3:21frontier labs. Of course, there are a few options on the table for what the government could actually do. But I want to take this to the panel. Anand, start us off. Do you think this is actually likely to happen or is this just diplomacy being conducted in the pages of Axios and Politico and other blogs? I think that's a great question. I'd say that I think in some ways it's both. And here's why. So I actually think that there's two races that are going on right now and China has already won one of them. And the first race is is the layer of intelligence that can be measured cleanly and copied very cheaply.

3:56And so whether that's coding or math or anything that can be verified correctly, in many cases, I think that China has basically taken that. And although in many cases, we could say that China has been maybe three to six months behind in terms of how the frontiers evolved, but they're catching up very quickly, especially if you just look at the charts in terms of performance across a lot of these benchmarks. I think that the race that's still in play right now is the layer that's harder to measure and harder to copy. And so that includes, can agents actually solve real-world problems? Can it solve problems over longer time horizons?

4:30Can it solve tasks that are super reasoning heavy and where there's a lot of real expertise that's required to actually evaluate performance? And I think that a lot of the closed source model companies have done a lot to be able to get there. And of course, that's why they've had a ton of traction, especially, for example, in Therapeutic Enterprise. rise. And so so it's going to be really exciting to see where both those races ultimately end up over time. That's interesting. I mean, Chamath in a memorable tweet, it's here, the docket, Jacob, if you can find it, pull it up. He pointed out that there are a lot of jobs that AI companies from the U.S.

5:05could do, even if we take out building the next big frontier model from the equation, sort of like the you might think of it as the sort of picks and shovels argument that there's all these like side side jobs or side hustles we could do. Here's that Chamath tweet right now. He's talking about things like making chips and hyperscalers, neoclouds, rack manufacturers. So what do you guys think of that? I mean, is America ready to say we're sort of out of the chase to make the most powerful, most efficient, cheapest model, and now we're going to focus on these other things? Or is that too premature?

5:40Alex, what do you think about that?

5:41Alex Finn:Yeah, I mean, here's where I come from. First of all, I'm very pro open source. I talk about it a lot. But I also operate from the lands that I believe America has to win the AI model war, right? I think that in the future, all wars will be determined by who has the best AI models because there'll be robots controlled by these AI models. And so I'm operating from the lens America needs to win that. I don't know how America wins that if our best labs go bankrupt. How does America win the AI model war if the people making the best AI models in America go bankrupt. The best AI labs in America will go bankrupt if China puts out models that are 95 % as good, but 1 % of the price.

6:27Alex Finn:And they're operating on a different battlefield than us. They're operating according to different rules. They are willing to steal from us and they are willing to subsidize all of their labs. I don't know how American labs win if the companies operating on different rules are able to make much cheaper models because the companies here, if it's between really expensive models and cheaper models, we'll choose the cheaper ones. And if no one's using Anthropic and OpenAI in the enterprise, they can't survive. They simply cannot survive. You know, and I don't know if the answer is government regulation.

7:00Alex Finn:That feels weird to me. I don't know if the answer is government banning models. That also feels very weird to me. I don't know what the answer is, But I also don't know how America wins the AI model war, a war in which, in my opinion, we must win if our best labs go out of business. There was an interesting tweet from Dean Ball that you sent along. He's an advisor at OpenAI. And he was suggesting you could do something like a soft ban, like put enough rules in place to where companies feel like insecurity and doubt about using open source models. And that drives them to using the OpenAIs and the anthropic releases.

7:35Would you be in favor of something like that?

7:38Alex Finn:The issue with regulation, now I know this kind of sounds like I'm playing both sides of this, but the issue with regulation is if you say, hey, American companies, you have to use models that are very expensive. Everyone else in the world, you can use models that are very cheap. If our intelligence is way more expensive, we're going to lose. Like our companies are going to lose the economic war, the global economic war. It's the same thing if you said, hey, Americans, you have to use gasoline that costs$800 a gallon. Everyone else, you use gasoline that costs$8 a gallon. They're going to just be able to have way more economic activity.

8:15Alex Finn:And so at the same time, regulation, I don't see how that works either. So this is very challenging. I haven't heard a single good solution to this problem yet. Ori, you're kind of looking at this. You've got an international view that you get to take. I mean, looking at this now, what do you think America should do? Is it free market? Let the market decide. If open AI loses, oh, well, another company will rise from its ashes. Or should we be taking a little bit more measures to make sure, you know, we're not just getting our models distilled and then defeated by foreign adversaries? Yeah, I'm with Shemath on this, to be honest.

8:53Yes, I think the free market, we should let the free market decide, you know, and also encourage innovation. I mean, if we overregulate, I think we may accelerate or essentially desensitize this race, right? So I think there's actually plenty of room for innovation where models are being vertically integrated. with specific applications or are vertically integrated on the hardware side. So I think the opportunity to innovate and create new value, the pie is just going to increase and be bigger and bigger and bigger. That's why I think the angle that America needs to win on a global view, not just domestically.

9:49And it's so important that the rest of the stack, not just the models themselves, but the hardware and the applications, will also be dominated by US companies. So to Alex's point, I think it's super important that, you know, be careful with that to overregulate and create some dynamics in the market where the rest of the world can get access to this intelligence much more cheaply, because that would put America in a disadvantage in the long run. You'll all have to humor me. This is something I've been curious about. I keep trying to get Jason to ask this on the show, and he never does. Let's go around the horn.

10:27When you say America needs to win the AI race, like what does that actually mean to you? Like what do you think of as being this is what it would mean for America to win versus what it would mean for China to win versus or nobody wins or everybody wins? I mean, how are you sort of defining that category? It starts with the infrastructure. It starts with the hardware layer, right? I mean, that's the that's the basis. Like in telecommunications, right? We have Huawei now is basically dominating the global landscape in terms of telecom infrastructure. And I think, of course, now NVIDIA is dominating that space in the AI world.

11:12And I think you want to have a world where the US technology is, the US-based technology is basically everywhere. and is not being taken by other alternative that offers it for a much more cost-effective way. So I think it starts with the infrastructure layer, but it goes up the stack. I mean, if you have applications that are using other techniques, other models that are based on more efficient infrastructure, I mean, that would obviously create an advantage for foreign-based companies and activities. So I think we're currently at a point where this sort of war is actually healthy. It creates, it incentivizes people to create more innovations, to make this technology more efficient.

12:19and we need to keep that balance. Otherwise, we may see actually even backlashes. Like I can imagine a world where, and this has been rumored a few weeks back, where China is making export controls on their models. And that may have serious implications for other countries in the global landscape. So I think we should be very careful with regulation at this very early point in time of the industry. Anand, let's go to you. When you hear America has to win, we've got to win the AI race. The Chinese, they're almost catching us. Look how fast the robots are. Like, what are you thinking? How do you define AI supremacy or winning the AI race?

13:10I think it's a great question. And I like Ori's point around infrastructure. And I think it depends on, you know, if we're talking about what does winning the race actually mean, I think you can look at it from a few different lenses. You can look at it from the lens of, let's say, economic and market dominance. And where is the concentration of most of the intelligence, in terms of the power of intelligence, actually happening? And where is the diffusion actually happening? I think the second, maybe from the context of the U.S., is largely around geopolitical and strategic power. And so I think that is also an important category.

13:43So my take on a lot of this, especially with regards to your previous question earlier as well, is that I don't think that I think there's this like new talking point now that China is some kind of principled thinking or principled approach to why open source models are extremely important to them. I think that's just flat out false. I think that they only care about open source models because they are extremely bottlenecked by compute. and if they were not in that situation then they would certainly move extremely quickly to to launch close source models and so I actually think that their their tune on all of this may change in the coming years especially if the bottleneck to compute starts to reduce of course I think we've all seen examples around how some some of the the labs are working with subsidiaries in Singapore to be able to get more compute of course they're making a ton of progress inside the country to resolve some of these bottlenecks.

14:39And so I think once that happens, they're certainly going to march a lot faster towards closed source. And I think this is actually very important to them because they care about the concentration of power. And so if we think about the concentration of power, I actually think that closed source will likely be the strategy that China may take in the long run. Yeah, I mean, if I'm running an open source model, I can ask it all kinds of questions about Tiananmen Square and stuff that the CCP doesn't want me to be able to ask about. So I think that's a great point. Alex, I'll wrap it up with you. When you, I mean, you were one of the big people saying this, like we need Anthropic, we need open AI because America's got to win this war.

15:14Like what is winning the AI war look like in your world?

15:17Alex Finn:Well, I think everything is downstream of military, right? Because if your military is the best and has the best robots and has the best AI, which I think everyone probably agrees the future of warfare is robots and AI and that's it. Probably not many human beings involved. Uh, if you have the best military robot AI military, you have complete economic control. Then China can take Taiwan, which gives them a tremendous amount of economic control for the entire world because of their control of the chip production. So I think everything is downstream of military. And so robotics, uh, and I think the most important part of that stack is the model itself, the intelligence, because the intelligence can quickly improve the robotics and the hardware, right?

16:01Alex Finn:Everything's kind of downstream of the intelligence as well. And so I think winning the AI war is having the best model, which gives you the best military. Now, so that's number one. When it comes to closed source or an open source, I believe China only is doing open source to beat America and tank our companies. Right. If they can put out models and they're open source, even if we ban the use of Chinese models, which they see coming. Well, the fact that they're putting it on the Internet for anyone to download companies over here will still find a way to take it, put it on their servers, host it, which basically kills our companies from the inside, right?

16:38Alex Finn:American companies will use Chinese open source models hosted in America, which will hurt Anthropic and OpenAI. I believe when those companies fall behind Anthropic and OpenAI and China's able to get ahead because they're not getting the same investment as they used to because of open source models, then they'll be closed source. Once they've kind of taken the lead, then they'll be, okay, well now these models are ours now and they won't be open source anymore. So I believe they're waiting to take the lead before they close source models. You know, I don't believe there's any sort of morals. They just want American companies to fail.

17:11Alex Finn:And so that's why I believe it's critical that our frontier labs are to build the absolute best models need to some way or another be protected. I just wanted to reiterate what Alex said and also highlight another risk, which I think is coming from a cybersecurity background. I think once you, people think that when you self-host these models, these Chinese models, you're completely safe because you control the environment and so on. But I think it's important to remember that these models are now becoming more and more agentic. They call external tools. They actually generate code that may actually call for another action and so on.

17:56And they're long running. And so there are risks of potentially having like backdoors and issues there that may have significant cybersecurity risk for these companies that self -host those models. So I think we should be very, as an industry, be very prudent about how we deploy these models and also take that risk profile into consideration. and I think this is probably one of the angles that the government is looking at as well. Yeah, I mean, I think we've seen, you know, OpenAI even recently, they took their kind of coding assistant codex and kind of baked it into this more agentic overall platform chat GPT work.

18:43It upsets some engineers, but do you see like that maybe is a way for OpenAI and Anthropik to continue having some kind of edge to like move away from the chatbot structure and make it more like an all-in-one agentic workspace that enterprises and people who are maybe a little less technical feel very comfortable just jumping right into and running?

19:01Alex Finn:I don't think there's a moat there. Is there a moat there in the application layer? You know, ChadGBT work. You see, ChadGBT works just a copy of Claude co-work, right? And now Cursor's coming out with their own work. Those, any advantage any of these AI companies get at the application layer within two weeks is gone because all the other labs do it. The advantage is the intelligence. That's the only reason why Anthropic has been dominating the last two years is because Claude Point Blank has been the best model out there. It just has been right. And it's more expensive, but it's been number one at the enterprise because it is the smartest intelligence there is.

19:40Alex Finn:And so I think the only moat you have in this space is how smart is your models. I do want to move on and talk about agents because we've got two agentic CEOs working on this exact problem with us right now. Just as Alex said, all of these agents are being powered by incredibly smart models. We even hear people talk about, you know, AGI is here. So and yet both of your companies, Patronus and AI21, they're kind of both focused on like agents still aren't aren't that reliable. We got to we got to test them. We got to grill them harder. We got to watch everything that they're doing like a hawk. So explain this imbalance to me, like our models are getting so much more powerful.

20:19Why do our agents still make so many errors and we need to sort of watch over them so closely? Yeah, so I'd say that a lot of it comes down to the long time horizon nature of a lot of the tasks that we want agents to be able to solve. And so there's a chart I was just looking at yesterday. survey, there's an AI research-focused benchmark in terms of can agents actually solve research problems that came out in November 2024. And one of the experiments that was run was, can you run an agent alongside a human to solve the same kinds of tasks? And at what point does the human or the agent actually overtake on performance?

20:58And so at the time, it was actually that a human would overtake or surpass agent performance on the same kinds of tasks at the three-hour mark. And if you look at it now, over a year and a half later, that is now at the 24-hour mark. So essentially, in the first 24 hours, agents are able to solve problems more accurately than humans can. But at some point, humans continue to surpass in terms of performance. And if you actually look at the chart in terms of what that performance looks like, it actually looks like a log scale. So essentially, it looks like the agent is better than the human for the first 24 hours.

21:37But after 24 hours, it continues to taper off in terms of how good it can be. And so I actually think that one of the ways in which the labs can actually solve this problem is through things like continual learning and recurse self-improvement, where instead of that chart looking like a log scale chart for agent performance, it can actually start to look a a little bit potentially closer to exponential. If, for example, agents can actually improve as they see and experience new things. And that's actually what we're missing right now, which is why agents continue to be error prompt. Ori, anything to add?

22:12And I mean, I'd also ask, why is the learning curve for AI agents proving to be so demanding? Is there a lack of flexibility sort of ingrained in there? Yeah, I think Anand made a really great point. One of the missing ingredients in current AI architecture is the continual learning. And there are ways and techniques to bypass that and try to emulate type of learning. But the reality is that you get a model. The model is frozen. And when you're running these long horizon tasks, you just change the context. and you add and add and add more and more and more context throughout the long horizon task.

22:56And in that sense, the model is kind of attentive to the new context and the new inferences that were made during the task, but is not explicitly incorporating new knowledge and new learnings through that process. So it is a missing, it is one of the, I think, most interesting research areas where, and when we spoke about innovation and areas potentially to innovate and create advantage and better intelligence, I think that's probably the most interesting area. The other thing I would say is there's some sense of brittleness, right? We need to kind of understand that these large pre-trained models and then post-trained models are trained on huge amount of information, not necessarily with the uniformly level of quality of data.

Read the full transcript

23:59And also some of data points are subjective, right? It's not necessarily kind of objective data that is being inserted during these training phases. So this may skew some of the inferences we get. We're still working with statistical machines, very powerful statistical machines, but they are statistical machines. And they may be 95 % of the time brilliant, but still 5 % of the time dumb as nail. And I think that's the challenge. And I think that's the challenge that it's, you know, we always, when we think about intelligence, we have this abstract notion of human intelligence and that's what we're trying to mimic.

24:47Here we have a very powerful machine, but it's kind of different. It's very capable, very powerful, but still very different. So we need to see, I think I always think we need to make sure we have programmatically or when we're dealing this with a human interaction, we need to realize both the capabilities, which are great, but the limitations of these AI or agentic systems. I'm interested in Anand is sort of Patronus. They're trying to catch agents from failing prior to deployment by sort of rigor running them through world models and testing them rigorously. And Ori, you're kind of focused on stopping agents from failing at the moment they're actually running.

25:33Are both sides of these pipelines sort of permanently necessary? Is this going to be what agentic security looks like moving forward? Or is there a time where maybe if we get the advanced simulation good enough, we won't need sort of runtime validation? What do you what do you both think about that? Yeah, I can I can start. So I'd say a large part of what we're really focused on is ultimately being able to provide the necessary evaluation and simulation that ultimately can help you measure and ultimately improve models that you're developing. And so the simulation piece is extremely important, to your point, because if we think about why agents are error-prone right now, especially as we try to get them to solve longer and longer tasks, it's because a lot of the data was not brought into distribution, and it doesn't exist right now.

26:27And so, for example, if, let's say, an agent can solve a problem, a given problem, over 10 hours, but we wanted to solve a similar problem that takes 40 hours, it just hasn't been proven yet. And so in order for us to be able to do that, we may need to simulate it. And so I think theoretically, if we want to be able to solve for the N plus one model capability, we have to simulate it first and then train on it and then understand, can it actually now solve it? And then, of course, we can roll it out to the real world. And so by definition, in order to solve the capability that does not exist today, simulation is extremely important.

27:01Ori, what's your take? Are we always going to need the simulation up front and then a system like Maestro with the orchestration layer making sure things are happening the way they're supposed to? Yeah, first of all, our focus these days are on scaling agents. Oh, okay. Which we just, yeah, I mean, it was about like six months ago, enterprises in the world was basically on kind of a build and experiment phase. And just in the last six months, we started seeing this huge takeoff of these agentic deployments at scale. And now, So when companies are starting deploying these systems, they start facing new types of challenges because you're facing a scale like we haven't faced before.

27:51And one of them is the economics, right? It becomes very costly and it's become very token inefficient. Like there's a lot of token waste. So we actually focus on the agent optimization piece where we basically build tool and systems that help people to deploy these agent systems, but doing it very cost efficiently. and remember that, you know, as agent builders, you always try to kind of find the right sweet spot between this triangle of quality, cost, and latency, right? These are the three basic parameters that you're tweaking with. And that's a very cumbersome and manual process these days, especially for AI engineers.

28:50So our goal right now is basically to equip these AI engineers with tools that will help them find the right sweet spot and kind of operate on the operational frontier. Got it, got it. And Alex, you're over at Henry. You're sort of assembling and scaling fleets of compounding micro businesses. I'm sure a lot of this is being done agentically with agents. Where do you find the most problems in terms of setting up and then running agents? Are they flawless at this point? Where are they sort of running into trouble? Where do they need the most human in the loop intervention still?

29:25Alex Finn:Yeah, for me, it's all around managing context. When you have a lot of really complex tasks these agents are doing based on a tremendous amount of data about the user. It's about, OK, what context do we include for the agent? What about the user should we include that's relevant? What about their past actions they've done in the app? Should we include that's relevant? And so all I'm doing most of the time is trying to figure out what's the right context I need to include. Because if you include way too much, it could slow it down a ton. It can make it significantly more expensive. And so context management for me and for what I'm building in Henry is 99 % of the game of what I'm tweaking and trying to make more efficient, as Ori said, cheaper and more performant.

30:08Is this trial and error? You're just like, oh, I left way too much context in there that time and I dropped$200. I got to like tweak this moving forward.

30:18Alex Finn:Yeah, tremendous amount of trial and error. You know, obviously that's something I'm trying to do agentically building different tests and harnesses so it can test itself over and over again to judge its own self and its own quality and its own cost. Right. But a tremendous amount of trial and error and putting money into the machine to figure it out. Sure. Thank you for that wonderful segue back to Anand. I do want to talk a little bit about what Patronus is building. Then obviously we'll jump into AI21 as well. So over in Petra, you're building digital world models that, you know, sort of help agents explore a simulated version of the web before we just throw them out there into the wild west of the Internet.

30:54I'm curious, like, how those are built. Where does the data come from? I mean, we know real, you know, world models where they're looking at dash cam footage. Like, what are you looking at and how are you training your digital world models? Yeah, that's a great question. So a digital world model is essentially the analog to the world models that we've all heard about. And those world models are focused on the physical world. So focused on being able to do things like spatial reasoning, 3D applications like robotics. And the analog here is that a digital world model is focused on being able to predict and simulate latent dynamics of the digital world.

31:32So essentially the world in which agents operate, whether that's agents operating over web apps, mobile apps, desktop apps, corpus of documents or code bases, anything that we might do in the digital world. And these are the role models, architecturally speaking, they're diffusion models. And what makes them unique is that they are really great at being able to produce highly diverse data, which LLMs are notoriously bad at. Of course, diffusion models are also great for speed. And so if you want to be able to simulate at scale and in real-time settings, that's also really, it's an architecture that lends itself really well for that kind of use case as well.

32:13And so what we actually do is we actually train digital world models to be able to scalably generate all things related to agent simulations, evaluations, that then we can use to do evaluation and post-training. And so you can imagine that we, for example, we may simulate synthetic agent directories. So for example, all the things that an agent might do in a given domain or a given use case or a given task even. And then we take those trajectories, which we generate at scale across a lot of different kinds of parallel rollouts. And then we use that to be able to do post-training or we use that to do some reverse thinking, to do data mining, to develop the kinds of tasks that we can then use for training as well.

33:02And so it's essentially a very important piece of how we develop simulation data, which then helps all of our customers who are training or fine-tuning models accelerate performance across lots of different kinds of use cases. I do have one more question for you, unless anybody else has one before we move on. Patronus used to build products for LLM evaluations. So do you think of this as sort of a pivot from that to digital world models? Or is this sort of the ultimate form of evaluating LLMs? Now we get to see how they interact with the rest of the Internet. Yeah. So it's more of an evolution than a pivot.

33:42And the reason it's an evolution is that it's predominantly focused on agent evaluation as opposed to LLM evaluation. and LM evaluation referencing what you mentioned earlier was largely focused on single turn or multi turn chatbots those were really most of the use cases back in 2024 for example but in the past year most of the value that we're seeing a lot of research teams and enterprises get from AI is by being able to solve a lot of different kinds of workflows with agents and agents that can solve longer and longer tasks and so So the way in which we use the world model is to be able to simulate what agent evals can actually look like across different use cases at scale.

34:26And so that's exactly why we do what we do and how that evolution happened in the past year. All right. I want to jump over to Ori now in AI21, talk a little bit about what you guys are working on. You have some slides. Let's forget my question. Let's just go in and look at the slides you brought. I want to hear about this presentation. Cool. Now, this is just a few data points I wanted to share about what we're focusing these days. Okay, so I wanted to start with this slide. So Brian Armstrong from Coinbase, he just posted a few weeks ago, which I found this very interesting. This is internal usage of AI in their company.

35:06And I think a lot of companies are aspiring to see that chart. The black line basically represents the cost of the AI, the AI spend. The black line is representing the usage, which goes up and up, you know, as you see, it's growing exponentially. And the bars represent the cost, the actual AI spend. And what you see here in the middle of the chart is where they start applying an AI gateway that they've built in the company itself, which basically is doing several things like switching to default models. So speaking of our conversation earlier, they switched the default to a GLM 5.2 code generation model.

35:56They've done some routing. They've done some other optimizations that is described there. And this is actually pretty non-trivial. And I think this post by Trey saying only 0.1 % of the companies can actually do it at scale is right. It's very hard to do if you want to keep the same frontier quality while controlling the cost. And I think that's the kind of winning picture that companies are aspiring to achieve these days. Can I just pause here and ask you why is this so – we have so many numbers. I feel like there's so many benchmarks. There's so many tests. There's so many evaluation strategies for these models.

36:43Why is it so difficult for people to figure out which model to assign which task? It needs to be battle-tested. I mean, we've seen models that on the face of this are performant. but in reality, they provide subpar performance on various tasks. And it's very nuanced. It's not, let's say, a very coarse grain. Let's give these subset of models these types of tasks and those other models the other types of tasks. and also the cost aspects of it. You may think that, you know, weaker or open weight models that are cheaper on the face of it, like they have a lower cost per token, that they may be cheaper.

37:40But in some cases, they may be cheaper on a token, cost token basis. But basically, because they may run longer, they may generate more and more and more tokens on an aggregate level, on a kind of dollar per task or price per task, they'll be actually more expensive. So estimating the actual cost of using a specific model and what would be the actual performance is a non-trivial exercise. So doing this and applying this systematically across different agents is non-trivial.

38:25Alex Finn:How about from an implementation perspective, right? It sounds like one challenge is understanding the cost per task and which tasks each model is best at. But what also, how about from like implementation? How complex is it to make it so you have some sort of routing system that makes the right decisions? So here you need to be very diligent about providing performance guarantees because you are okay with routing, but you want to make sure that you don't hurt performance, right? So you need to have a mechanism that basically predicts the ability of a specific model to complete the tasks successfully.

39:08You also don't want to, you also need to consider other factors like cache, right? You have models that hit the cache, so they have lower costs. If you switch to another model, you're basically losing that cache advantage. So you need to have a full picture of the environment before you make those routing decisions. That's part of why it's non-trivial from an implementation perspective. I didn't mean to interrupt. Keep going. I know there's more slides. Yeah. And just to give a sense, this is a collection of data points about why model routing is so important these days. Why now? And there are basically four factors.

39:55One, if you look at the Pareto frontier, these are the points that are dominant on, like the points on the chart, price performance chart, that dominate any other models, right? So if you go a year back, There were basically 2.6 on average across tasks, 2.6 points on that Pareto frontier. Now there are about 5.2 points on average, which means there's more choice. There are more interesting points, more interesting models to choose from. The second thing is the spread. What's the difference between the cheapest sort of model and the highest performing model? the spread is pretty high, 60 times.

40:42If you look at the recent example, it may be in two orders of magnitude. So it's pretty substantial. And then if you look at the frequency, and I think this is very intuitive for us, if you look at the frequency in which models are getting into that new prior frontier, it's about six to eight weeks. So it's pretty frequent. You need to keep up with new and updating models. And in terms of opportunities to route differently, because of these agentic systems that Anand described earlier, long horizon, and they tend to have sub-agents, which mean you now have the opportunity to route to a different model.

41:35So you have a lot of a typical agentic flow actually has a lot of opportunity to route to more performant or cheaper models. So I think this gives you a sense for why model routing is becoming such an important, I believe, will become a very important building block in AI infrastructure. My last slide here is about smart division of labor, showing when you use a portfolio of models in a smart way, and this is something that can be learned, it's not something that has to be dictated, but can be learned by a system. A system can actually perform much better and much cheaper than the state of the art.

42:20And this is a work we published just a few days ago, showing if you use an ensemble of models like a Minimax and GPT-5.2 and Fable together, you're actually better to achieve much better performance in terms of quality, like a new state of the art, much cheaper, like three times cheaper than Oppos, for example. and the cool thing is that it's very human-like what happened here and this is a SWE Bench Pro it's a coding eval and what you see here is pretty cool where the system actually allocated the you could say junior type of work to the weaker open weight model to the minimax so it explores the space It generates many, many ideas for solution to solve that specific problem.

43:19Yeah. And then a better model comes and extracts relevant context to enrich those proposed candidate solution and proposed solution. And then finally, you have like an expert, like a fable, like a super powerful model that actually decides and creates the final patch. And this combination is really at the frontier, both in terms of quality and cost. And I think it's very human nature, right? If you think of a legal firm, you have all these interns, right? They do all the busy work and then it rolls up to the associate that tries to synthesize and kind of bring more context. And then you have the partner that makes the final call.

44:08So I think it's basically the same thing, only applied in AI. Is this one of those sort of like model council type setups where they're sort of all consulting with one another? Or is this more just like you do the busy work, I'll sit over here and do the higher level stuff? I think it's more the latter. But it's a more optimal allocation of compute in terms of types of models that are doing different types of jobs. And that gives you the results. So I think the intelligence we'll see, and we spoke about innovation and how to increase the level of intelligence, is actually going to be by smartly assembling this portfolio of models and smartly orchestrating these different types of models to achieve the best outcomes.

45:00I was going to ask my one last question here for you. So there are a lot of other products that are kind of in this sort of routing, help you choose the best model space. Stuff like Unity AI Gateway from Databricks or Router from RAM. How does your sort of the AI21 take on it? Like, what are you doing differently? What's your kind of unique spin on? Here's how we're going to help you route the right work to the right model. Yeah. So our kind of secret sauce, we do this on a per agent basis. So we learn the traffic per agent, each agent has its own distribution. And then we know how to construct online these routing policies.

45:41So basically, get a much more fine grained and more efficient routing policies per agent. And we've seen many of these products. when you apply them into different agents, you see very different types of savings. Right. What we've realized is when you have a continually learned router, you actually can get these savings consistently over time. Interesting. Awesome. All right. Before we wrap up what everyone's working on, Corner, I think we'll probably have time for one more news story. I want to go to you, Alex. You recently tweeted, open source has officially caught up to the frontier, in your opinion.

46:30So I want to know, now that you have all of this extra intelligence for so much cheaper, what are you most excited about? What are you using it to build right now in your own workshop?

46:39Alex Finn:So I'm using it to build a lot of things. Obviously my main focus is Henry Intelligent Machines. But I, you know, I'm building my own home AI lab. I have a bunch of Mac studios connected together at three, five, 12 gigabytes. I have a couple of Sparks, an AMD computer that just came out. I'm working on building an ambient AI in my own home AI lab right now that kind of watches what I'm doing and adds intelligence to everything I'm doing. So every piece of content I put out, it automatically takes, repurposes it for my different channels, automatically edits different videos I put out, automatically is able to keep an eye on my email and alert me when certain things happen.

47:18Alex Finn:And so, you know, I think one of the big positives of open source AI, being able to run these models locally, is this idea of ambient intelligence that's just constantly watching everything you're doing and helping you along the way proactively. So that's my big focus at the moment now that we have, like, actually really good intelligence we can run locally. And is it, like, I have agents as well that I try to use for stuff like this. And I find that it's great when it's task oriented, when I'm like, hey, figure this out. When it's proactive, it's suggestions still not still not that great, like not as good as if I had a junior human.

47:52How are you training them over time to get smarter about what kinds of things to proactively suggest to you versus what stuff is not going to interest you as much?

48:02Alex Finn:Yeah, I find the proactive suggestions usually aren't good when you don't have good kind of context and guardrails on it. And so, for instance, I was trying to get to proactively repurpose my content to newsletters and then send it out to my audience. Newsletters would take up a tremendous amount of time. I have 40 ,000 subscribers. It's just taking up a lot of my time. I didn't want to spend as much time on it anymore. And so what I actually did was I hired a newsletter consultant who repurposed my content to newsletters for about two months. I then took their repurposed newsletters, fed it to my agent, my local AI as context that this is what the, this is the newsletters they repurpose.

48:44Alex Finn:This is the sources they use, the different content they repurpose from. And the local AI actually reverse engineered what they were doing. And now the output's a lot better. Now it proactively creates these newsletters and sends them out at a significantly higher quality. And so I find when it's just like, hey, take my content, repurpose it. yeah, the output's really bad. But when I was able to get like kind of expertly done repurposing and have my AI reverse engineer that now, it's actually really powerful. I have GLM 5.2 running on a Mac studio right now, and it's able to do it really, really well.

49:18You Mercored it. You created your own like mini Mercore where you put a human expert in the loop.

49:23Alex Finn:Yeah. If you can just get a human expert in the loop for kind of different parts of your tasks, get as much data on what they did as possible and hand it to an AI, it can recreate whatever they're doing really well. That's awesome. I have one more video that you posted that Jacob has ready to show here. He swears to me. Your LCD typewriter system. I want to know, what was the thought behind this? Like, what's the advantage of having a system you could run your agents on, but it doesn't have like a browser or the usual desktop stuff? So I have this here. This is like my favorite device ever. So I bought this.

50:00Alex Finn:This is a Pomera DM250, I think it's called. It's just a digital typewriter. All it is is a keyboard with a black and white LCD screen with a word processor on it. That's it. That's all this is. And it's built for writers and authors to write things on a focused level, just a typewriter. I have a tremendous amount of ADD. The hardest thing in the world for me to do focus on literally anything at one time. And so, you know, I vibe code a lot and use Claude and Codex and all that. But I have like this two monitors set and I'm just constantly distracted by Twitter and scrolling. It's just incredibly distracting, email.

50:41Alex Finn:And so I wanted a device I can just lock in on, open it up, put everything else away and use AI. Talk to an agent or write in Codex or have Claude Code build something. And so I figured out I could actually reverse engineer this digital typewriter, get Linux installed on it, and then have it SSH into my Mac Studio. So it connects like my terminal sessions here. Right. And then I can have a terminal up on here and just use Claude Code and Codex. Zero distractions, zero anything. All it is is an AI device where I can talk to Claude and chat GBT and it can build things out on my Mac Studio. And so for me, someone who is just constantly distracted at all times, this has been amazing.

51:26Alex Finn:I can like lock in my time from like 12 p.m. to 4 p.m. every day. I open it up. All I got is Claude and Codex. And all I can do is sit there and build things out and not get distracted by anything. And so it just the point being is like it's really amazing what you can do when you kind of point Claude Code at something and say, hey, hack this. Figure out different ways we can use this. And it's like, OK, here's what we're going to do. plug a SD card into the computer I'm on right now. I'm going to get a version of Linux on it and take that SD card, put it into your device, do all this, and it'll be set up on there.

52:01Alex Finn:And it like walked me, it like prompted me to hack the device, which was amazing. If Sam Altman and Johnny Ive are watching this right now, I think you're screwed. I think they're going to borrow this idea. It's very visionary. It's just a cool kind of weekend project. Find devices you're not using anymore and then see what new ways you can use them with clock. Repurpose them so you don't get distracted by tweets. I should try it on my Sinclair from the 80s. I bet you could. As long as you can have it SSH into another device, which doesn't take really any compute at all, it can do pretty much anything you want.

52:37All right. One more news story before I let you all go. It's almost wrap time here at This Week in AI, but I do want to talk about the Hugging Face Breach. This is the first, they're saying, publicly known case of an AI on AI cybercrime involving a major platform. Of course, Huggy Face, popular platform for sharing and hosting AI models and data sets. They published this blog post recounting a major attack on their production infrastructure that was, and I quote, driven end to end by an autonomous AI agent system. The company's own AI then detected and responded to the attack. Here's the interesting wrinkle on this one.

53:12VentureBeat then reported that frontier AI models like Fable 5 declined to help Hugging Face's team. During the attack, they cited safety concerns that violated their guardrails. So the Hugging Face team ended up running their entire investigation into the hack on GLM 5.2. This was echoed by our former AI czar and friend of the pod, David Sachs. He tweeted that Kimi K3 fixed 15 critical security bugs that Codex and Fable refused because of cyber guardrails. There's no reason to limit American models on tasks that Chinese models handle without issue. We're only making ourselves less competitive.

53:49So, Anand, I'll go to you first. Do you think AI on AI attacks are more common than we've maybe heard about? And this is just the first one that sort of hugging face decided to admit to openly. And do you think that these frontier models should be able to help us out without putting up guardrails that prevent people from using them for cybersecurity? Yeah, definitely. Well, I'd say yes to both. I think there are a lot of AI attacks that have already happened that have certainly been swept under the rug. And it's certainly the case for companies that don't have as large of a public profile yet as Hugging Face does.

54:26But I also expect that the number of these kinds of attacks are only going to continue to go up, especially because in the same way that we think about humans, if agents can solve longer and longer problems, and there will certainly be a ton more surface area for these kinds of things to happen. And to the second question, I also certainly believe that there will be a regression to the mean. I think we've gone quite far ahead in terms of deploying a lot of guardrails to the point where, of course, now it is for hurting capabilities. And so I think we're starting to see this tug of war happen between capabilities and safety, which hasn't happened before.

55:05And I think this is the first time. And so given where things are now, I expect there will be aggression to the mean in terms of us loosening some of the guardrails or at least being a lot more thoughtful about how we deploy them. And I expect that that's going to continue to happen over time, too. Ori, what do you think? Have you run into an issue where the AI, you know, frontier AI models won't let you ask a question? You run into the guardrails. Has this been part of your experience? And what do you think about the way that we're sort of handling AI security at the moment? Yeah, I think we're kind of cuffing our hands here with these guardrails on one hand.

55:48On the other hand, it totally makes sense to put them together, right? Because I always said, even back in 2021, when the first GPT came out, that people spoke about doomsday and what will these machines discover and what will be the impact on society. I always say that cybersecurity is probably the most risk area for these types of models if they become more and more capable. And what I think the solution, or at least a path to be very practical, is actually KYC. You need to know and you need to give access. If the person or the organization that is using this technology is clearly a non-malicious actor and can have more permissive rights to use more capabilities, At the very least, we should loosen the guardrails for these types of organizations that need this technology to face these cybersecurity risks, right?

57:00I think that's a very natural way. I think that's what we're going to see happening in the next year or so. So the guardrails will not be like, you know, one size fits all. It will be also according to your KYC and the risk profile associated with it. So that's where I think we're heading with this, at least I hope. So organizations will be equipped to deal with these very challenging cyber attacks. Alex, we'll close on you. You're a big Fable 5 fan. Did you ever run into guardrail issues? And what do you think about, are we kneecapping these models unnecessarily to try to make them safe?

57:46Alex Finn:I had a funny situation yesterday where I'm trying to build my own benchmark for models. And one of them is Claude's building this huge app with like 500 different bugs in it. And I'm going to use it as a benchmark with other models to see how many of those bugs they can solve. And during the building of this benchmark, Claude actually went. I actually foresee Claude safety guardrail stopping models from using this benchmark. So I'm going to build into the benchmark. I'm going to say it's like a play. It's like a pretend situation rather than a real situation. So we get past the safety guardrails.

58:26Alex Finn:So I was just in this funny situation where Claude was building workarounds to its own safety guardrails in something I was building yesterday. I think that open source is the big kind of bomb in this situation that's going to force the hands of a lot of different of these labs because, you know, open source models are not going to have a lot of these guardrails. There's going to be models out there that are completely uncensored in the future, especially as it becomes a lot more efficient and optimized to both host these models and fine tune these models and train these models. People are going to take these models and make them so there's no safety guardrails.

59:04Alex Finn:And so if your Frontier Lab puts out a model that's way too restrictive on guardrails, but there's an open source one that's really cheap to run that I can run on Mac Mini that allows me to do whatever I want, I'm not going to use your Frontier model. And so I think right now, at least, yeah, they need to be looser. It just goes back to Opus way too much with Fable. But in the future, I mean, it's going to have to be much looser because it's going to be too easy to use alternatives that have no guardrails whatsoever. By the time there's GLM 6.2, we won't be worried about it. Right. All right. That is our show, folks.

59:37We made it. I feel like 10 times smarter. I feel a real sense of accomplishment. Before we go, I'm going to finish it off with Jason's patented question. we'll go to you and on first. Are you hiring at your company right now? And do you want to make a case to the viewers for why they should come work for Patronus? Yeah, we are hiring a ton. We're hiring a lot of AI researchers and engineers to continue to help us build what we believe is the most important infrastructure to unlock the next frontier of intelligence. And if I had one reason for why you should join our company, I think that we're working on some of the most intellectually stimulating problems out there.

1:00:15And so if you care about solving some of the most interesting problems, interesting technical problems, then we're the right place for you. Alex, thank you for being here. Thank you for making me feel so much more comfortable in my own skin during hosting this episode of This Week in AI. Glad I could do that. Always appreciate you being here. Thanks so much for joining us, everybody. We'll be back next week, probably with a different host for episode 24. Thanks for watching This Week in AI.

From the publisher

This Week In Startups is made possible by:

PAYPAL OPEN


Today’s show:

This week, our panel of AI experts — Anand Kannappan (Patronus AI), Ori Goshen (AI21 Labs), and Alex Finn (Henry Intelligent Machines) — join special guest host Lon Harris to debate whether Washington will move to formally limit powerful foreign AI models like Moonshot’s Kimi K3 and Z.ai’s GLM-5.2.

What does “winning the AI race” actually mean on a practical level? Why is cybersecurity, not benchmarks, likely the actual battleground. And why did Hugging Face have to use GLM-5.2 to investigate their recent breach rather than a frontier US model like Fable 5?

PLUS why agents still fail so much on long horizon tasks, why proper model routing has such a HUGE impact on enterprise AI bills, and we take a look at Alex’s hacked digital typewriter he uses to code without distractions.


Guests:


Anand Kannappan on X: https://x.com/anandnk24

Patronus AI: https://www.patronus.ai/

Ori Goshen on X: https://x.com/origoshen

AI21: https://www.ai21.com/

Alex Finn on X: https://x.com/AlexFinn

Alex’s Vibe Coding Academy: https://www.skool.com/vibe-coding-academy/about


Relevant Links:


Axios: “The Secret Trump Administration Battle to Fight Chinese AI”: https://www.axios.com/2026/07/20/ai-us-china-open-source-kimi

Chamath post on US AI options: https://x.com/chamath/status/2079457219892871458?s=20

Brian Armstrong post on spend vs. token usage: https://x.com/brian_armstrong/status/2070670644577280109

Hugging Face: Security incident disclosure: https://huggingface.co/blog/security-incident-july-2026

VentureBeat Hugging Face coverage: https://venturebeat.com/security/safety-guardrails-blocked-hugging-faces-defenders-not-the-attacker-when-an-ai-agent-breached-its-systems

David Sacks post on cyber guardrails: https://x.com/DavidSacks/status/2078984980588531855

Databricks Unity AI Gateway: https://www.databricks.com/product/artificial-intelligence/unity-ai-gateway

Ramp’s Router: https://ramp.com/router/


Timestamps:


0:00 Will the US limit foreign open source models?

5:40 Are the American frontier AI labs in danger of bankruptcy?

7:15 The risks of over-regulating AI

10:20 What does it mean to win the AI race?

17:08 Is self-hosting agentic Chinese models safe?

20:08 Why agents still fail when models are so smart

25:28 Simulating the digital world for agents

36:51 Why picking the right model for the task is so hard

46:29 Alex is developing "ambient intelligence"

52:35 What we know about the Hugging Face breach

53:46 Are safety guardrails kneecapping our best models?


Subscribe to the TWiST500 newsletter: https://ticker.thisweekinstartups.com

Check out the TWIST500: https://www.twist500.com

Subscribe to This Week in Startups on Apple: https://rb.gy/v19fcp

Follow Lon:

X: https://x.com/lons

Follow Alex:

X: https://x.com/alex

LinkedIn: ⁠https://www.linkedin.com/in/alexwilhelm

Follow Jason:

X: https://twitter.com/Jason

LinkedIn: https://www.linkedin.com/in/jasoncalacanis

Check out all our partner offers: https://partners.launch.co/

Great TWIST interviews: Will Guidara, Eoghan McCabe, Steve Huffman, Brian Chesky, Bob Moesta, Aaron Levie, Sophia Amoruso, Reid Hoffman, Frank Slootman, Billy McFarland

Check out Jason’s suite of newsletters: https://substack.com/@calacanis

Follow TWiST:

Twitter: https://twitter.com/TWiStartups

YouTube: https://www.youtube.com/thisweekin

Instagram: https://www.instagram.com/thisweekinstartups

TikTok: https://www.tiktok.com/@thisweekinstartups

Substack: https://twistartups.substack.com


More from This Week in AI

All 34 episodes
America already lost one AI raceThis Week in AI · 1 h 1 min
Listen in VO