Mythos, Sonnet 5, GLM-5.2 Dominate the News Cycle | Episode 20

2 Jul 2026 · 1 h 11 min · 26 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

This Week in AI Episode 20 covers a chaotic AI news cycle and the implications of U.S. model restrictions, open-source momentum, and the shift toward cheaper, “minimum viable intelligence” models.

Guests

Alex (host), Victor Perez (CEO of Korea AI; focuses on Krea’s creative AI safety and open-source image models), Div Garg (CEO of AGI Inc.; builds on-device “tiny action models”/TAMs for mobile computer-use agents), Andrew Berman (CEO of RunLayer; enterprise agent-control and secure deployment platform).

Key claims

  • U.S. policy is “insane”/hard to read: Anthropic’s Mythos/Fable were blocked; OpenAI GPT-5.6 access is restricted via customer eligibility and “Trusted Access.”
  • Open source (e.g., GLM 5.2) is driving cost-driven adoption; enterprises want security/control plus lower inference bills.
  • Model competition is shifting from raw intelligence to performance-per-dollar; Sonnet 5 is positioned as faster/cheaper while still strong.

Notable examples

  • Fable reportedly enabled a roadmap screenshot to be turned into a shipped product in ~2 hours (internal dev).
  • Sonnet 5 benchmarks cited: SWE-bench Pro 63%, strong GDPVAL AAV2, and near Opus 4A on computer use.
  • AGI Inc. collects mobile interaction data (Android screen recording with compensation) and trains models with Qualcomm collaboration; RunLayer cites token-budget concerns (e.g., an agent loop burning 80% of inference budget in a weekend).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Current State of AI Regulation

0:45 to 3:00

Discussion on the U.S. government's influence on AI model releases.

“We also have Div Garg, the CEO of AGI Inc., a very modestly named company, Div.”

Creative AI and Regulation

3:00 to 6:15

Exploration of the implications of AI regulations in the creative sector.

“I don't think that the risks that it comes to are as heavy as cybersecurity threats, which as far as I know is like the main point behind this blockage from these new models.”

Understanding Action Models

6:15 to 8:06

Differentiating action models from large language models and their implications.

“And the technology we're building is like, how can we run all these things on your device?”

Regulatory Concerns in AI Development

8:06 to 9:35

Insights into the regulatory landscape affecting AI enterprises.

“So we focus on giving companies the golden path.”

Open Source AI Models

9:35 to 12:49

Discussion on the benefits and challenges of open source AI models.

“Now, what I'd love to get access to is I'd love to get access to 5.6 to see where that's going.”

Post-Training Techniques in AI Models

14:06 to 18:13

Learn how post-training on raw models can improve AI performance.

“So Composer from Cursor used KimiK 2.5 and did further post-training on it, but it was already slightly post-trained.”

Challenges in Early Model Adoption

18:14 to 22:24

Explore the difficulties of getting users to adopt incomplete AI models.

“We don't focus as much as on-device computer, on mobile computer use as Div does.”

Sonnet 5 Announcement and Performance

22:25 to 25:06

Discover the benchmarks and performance comparisons of Sonnet 5.

“This will be a few days late when this show comes out.”

Intelligence vs. Performance Trade-offs

25:07 to 28:00

Understand the balance between model intelligence and performance cost.

“talking about essentially minimum viable intelligence after which there's no point in and pouring more compute into it because you're already getting what you need out of it.”

The Role of AI in Preparing for Work

28:00 to 29:43

The discussion revolves around how AI affects personal preparation and performance.

“But when it comes to actually doing the core substance of my job, it's actually not that good.”
Show all 26 chapters

Improving AI Model Intelligence

29:43 to 31:39

Experts discuss the balance between performance and intelligence in AI models.

“I think that there's an objective, like how to put it, like objective capabilities and subjective capabilities.”

Customization in AI Models

31:39 to 33:59

The conversation focuses on how AI can be tailored to fit personal aesthetics and preferences.

“for these models to not just solve the problem, to not just have the capability to solve the problem, but to solve it in the way that CREA would do it.”

Enterprise AI and Cost Management

33:59 to 36:05

The panel discusses the integration of AI in enterprises and the emerging cost considerations.

“Now, that's integration into your existing systems, into your IDP, into everything else you use.”

Balancing Performance and Intelligence in AI

36:05 to 38:15

Experts explore the differentiation between performance needs and higher intelligence requirements in AI.

“next kind of um quantum of intelligence as long as they're performant and inexpensive now in the on-device case, you are dealing with just much less capable hardware.”

The Future of Mobile AI Interaction

38:15 to 41:24

Discussion on how mobile AI could evolve to provide a more user-friendly experience.

“But it's still like we think like this is something that is doable, especially with like smaller models that we are specializing.”

Competitive Landscape in AI

41:24 to 42:00

The hosts evaluate the competitive dynamics in the AI industry and specific company trajectories.

“Okay, let's put aside the technical stuff and talk about the business side of things because there's a lot going on.”

Discussion on CREA's Trajectory and Market Behavior

42:00 to 46:00

Explore the trajectory of CREA and its aggressive marketing practices that raise concerns.

“I know you guys raised$83 million last year.”

Investor Insights and Fundraising Rumors

46:00 to 48:30

Delve into the recent fundraising rumors and the implications for future investments.

“We actually don't know where the leaks came from, to be fair.”

Revenue Growth and Market Competition

48:30 to 54:20

Analyze the factors contributing to rapid revenue growth and competition in the market.

“Ah, man, we're, it's been a lot of fun and I'm not going to go into the revenue numbers, but it's been a lot of fun.”

Distillation Concerns in AI Models

54:20 to 56:00

Discuss the implications of distillation attacks on AI models and their creators.

“In the world of LLMs, Anthropic just drops a blog post every three months pointing fingers at every single Chinese AI lab they can name off top.”

Optimizing AI Agent Costs

56:00 to 58:00

Learn about the cost-saving measures in AI agent management and their implications.

“And then there's probably another, we have some more spend through API usage.”

Legal Rights of AI Agents

58:00 to 1:00:00

Explore the concept of agentic personhood and the legal frameworks needed for AI agents.

“Agentic personhood and kind of corporate rights.”

The Future of AI Models and Costs

1:00:00 to 1:03:00

Discuss the evolution of AI models, their costs, and implications for businesses.

“These are very interesting challenges that everyone is trying to think through in 2026, and it's just going to get more and more difficult.”

AGI Perspectives and Predictions

1:03:00 to 1:06:00

Insights into the timeline for AGI and how current developments are shaping the future.

“And there's things that are kind of like, I'll just say there's a lot of shitty work people have to do, and now you don't have to do that anymore.”

Understanding Agent Swarms

1:06:00 to 1:09:50

A discussion on what agent swarms are and their potential applications.

“So PDoom going down, but I will also say non-zero.”

Final Thoughts and Future Discussions

1:10:01 to 1:10:28

The hosts wrap up their discussion, highlighting key takeaways and future plans.

“as easy as it gets as easy as it gets uh div agi inc how do i find that on the great wide internet yeah we have this domain the agi.company uh and you can also find us on agi.app the agi.company or AGI.app.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Hey everybody, welcome back to This Week in AI. My name is Alex and we are coming to you during one of the most fascinating and chaotic periods of AI that we've seen yet. Anthropic's fable model is still banned. The world is going gaga over ZAI's GLM 5.2. A food delivery company just open sourced a near frontier AI model and everyone is trying to use as much AI as they can without losing their home. To help us make sense of the state of play, I've gathered some of the brightest minds from the world of AI. Thanks to our friends at PayPal, the exclusive sponsor for This Week in AI. Try the payment and growth platform that's trusted by millions of customers worldwide.

0:35PayPal Open. Start growing today at paypalopen.com. In one corner, we have Victor Perez, the CEO of Korea AI. Victor, welcome to the show. Thanks for having me. My absolute pleasure. We also have Div Garg, the CEO of AGI Inc., a very modestly named company, Div. Thanks. So good to be here. And then we have Andrew Berman, the CEO of RunLayer. Andrew, how you been? I've been good. Thanks, Alex. Yeah, I'm really glad you guys are here because I feel like we went through a period of slow news and then everything seemed to have gone crazy in the last couple of weeks. Glad we're going to kind of get to the bottom of it.

1:08The thing that I want to start with, because I think we can't avoid discussing this as a group, is that after forcing the removal of Anthropics Mythos 5 and Fable 5 for a bit, the U.S. government forced OpenAI to hold back GPT 5.6's release, which I thought was very disappointing. And it's even deciding which customers can get access, which means maybe startups can't. So I'm curious, Andrew, I want to start with you. Is the US AI policy reasonable and legible at this point, or am I incorrect in viewing it as insane? I think you're not incorrect in viewing it as insane. I think basically what we're doing is we're giving China and other actors the ability to encroach on US technological supremacy.

1:47I think open source models are being distilled from our world-class models, and I think we should be pushing ahead as fast as possible and we should be the dominant player in AI globally. Does anyone have a disagreement with that before I get more specific? But I want to make sure that Victor and Div can jump in if they want. Yeah, I agree with this angle. There's definitely a risk around like obviously cybersecurity, which maybe prompted like the initial fable or mythos like suspension. But I think it's kind of like overblown to some extent. It's like, Like these things will happen and we can't really like stop this AI models to combat and matter because like China is in the race and they will figure this out if we don't.

2:26I'm just surprised at how quickly we went from it's all free range to absolutely nothing. And there's a lot of talk about how, you know, maybe Mythos was overblown or overhyped, but it seems to be still treated as this, you know, essentially cyber weapon. So maybe this is just our new reality. But Victor, CREA has done some training of its own models. Now, not in the cybersecurity space, not even in the coding space, in the creative space. I'm just curious, do you think that, you know, government regulation of AI in general is going to eventually get down to your domain as well? Pretty far away from now, but directionally, is that where we're going?

2:58I think that the emerging creative domain itself, it has its own specific risks. I don't think that the risks that it comes to are as heavy as cybersecurity threats, which as far as I know is like the main point behind this blockage from these new models. In our case, the most dangerous things that folks are doing with the models are NSFW-type images, CSAM-type images, which, of course, are illegal. And this is, like, the main topic that we fight against when running safety on top of these models. And on the other side, there's, like, impersonating people, like, recreation of public figures, stuff like that, which I don't think that the US government cares as much as potentially having China running massive cyber attacks versus US companies.

3:53It's funny, though, how we were talking about IP as a major sticking point in AI and that being a thing to regulate or not. And now I feel like it's entirely fallen by the wayside. So does that mean that people working on building models in the creative space have, I don't know, a couple of years of air cover here? Potentially. Potentially. I wouldn't be surprised if it comes. I wonder if it would come at the model layer or at the inference layer. Because I think that these are two of the things that I think quite a lot. Like some regulations seems to be, or I would say like previously in the image models or video models, they're used to, people try to enforce regulations on the data that you use for training.

4:31But I haven't seen as many regulations on the measures that you run at inference time. and how, even though the model is capable of producing something that is IP protected, what methods do you put on top of that model in order to either alert the final user that this image that the model has produced may be IP protected and to be careful with how they use it versus just trying to make the model completely unable to produce these assets, which unavoidably nerves the capabilities of the model. Oh, I mean, enormously. If you have to write, you know, you can't go over here, you can't go over there, pretty soon you have one third of the surface area.

5:09Now, Div, your company, which I should just explain for folks out there who don't know, which is doing essentially computer use agents for mobile handsets, taking what we know from the PC world, desktop world, and putting it on your mobile device. You guys have built a large action model, yeah? Yeah, we actually call it like a tiny action model at this point because it's running on the device. Okay, I apologize. You're running a TAM, not an LAM. I sincerely apologize. We'll edit that. Talk to me about, well, first of all, define that for everyone out there who's curious about the difference between LAM and LLM.

5:41And then also, do you think that there's any potential regulatory pressure coming down the pike on what you're working on? Yeah, I'll start with the first question. So the biggest difference between an action model and a large language model is like, okay, like a language model is for chat. It's like, what do you do with chat GPT? It can hold conversations with you. An action model is made to do tasks for you. It can go and operate my whole computer. It can go and automate all my work. It can also maybe do things for me, like act as my EA, order me an Uber to my meetings, get me lunch, order me coffee.

6:17And that makes it very powerful. And the technology we're building is like, how can we run all these things on your device? So we have hired some world class researchers to compress these models to run on NPUs on the phones. And our vision is like, we feel like if your phone itself becomes AI native, it's agentic, I can just talk to it and it can go and use any app. So the phone itself becomes the AI and we want to live in this world. So then on the regulatory point, just kind of keeping us on this topic, cause I just want to get this done and out of the way. Is there any sort of like regulatory pressure brewing about what you're building?

6:53Or is it again, a bit like what we see with Kriya outside of the current focus and therefore outside of the blast radius of hyperactive government action. Yeah, well, it's outside of the focus, at least right now, until we maybe have some sort of cybersecurity incident. We have been very, been focused on privacy and safety. That's why a lot of this is like on your device. And our goal is like, we want to free your people's time. Like it's like, if people have more free time, you can do more things. And purchases is definitely one thing you need to be careful about. We don't want someone to go, like an AI to go and start trading on your Robinhood or buy you some crypto.

7:30Or mine crypto on your computer when you don't want to, etc. All right, Andrew, just coming back to you on the future point of this, you work with a lot of enterprises helping people kind of get AI to work for their company, which means I presume you have a foot in, you know, major model companies and talking to a lot of major enterprises. Is there a real regulatory concern out there or is this just what we're talking about on X because none of us want to get real work done during the day? I think from the hyperscale perspective, and the administration perspective, there's definitely real regulatory concern.

8:00We're part of a bunch of the trusted access programs because at the end of the day, while we enable AI, we also do security and control. So we focus on giving companies the golden path. So how do I enable employees to manage forms of agents? And how do I have visibility, security and control in the background? Is there a real concern? Yeah. Do enterprises want to use GLM 5.2 for cost reasons? Yes. Do they want to serve it from places like Bedrock and Base 10 and Modal and all the other... Fireworks, Together AI, etc. Exactly. Thank you. Yes. But that's a cost reason. And so I think at the end of the day, what we're seeing from customers, they want open source more than I was expecting earlier this year, mostly because their inference and token bills have increased.

8:52And we have some interesting regulatory things going on in the background. At the same time, people want Fable because Fable was incredible for that one, a couple of days. And we could ship so much. So it's a double-edged sword. Did you guys actually get Fable into any production environments before it was stolen from our hands? We were using it for internal development, not production environments. Okay. And what did you see from it? It was incredible. How much better was it than, say, Opus 4.8 or GLM 5.2? Somebody pointed it at the product roadmap discussion that we were having, took a picture of the product roadmap, and then two hours later had a fully shipped product.

9:32Maybe we want to call it a feature that probably would have taken, like, four days with Opus. So, yes, it was a step function. Now, what I'd love to get access to is I'd love to get access to 5.6 to see where that's going. and we're part of Trusted Access, but we'll see what happens because now the US government still has to approve it. Can you explain to people out there what Trusted Access is, how the program works, and also the benefits you get from it? Because I think that's a bit niche, but I think it does matter for founders out there who are listening who want access to 5.6 and are looking at you as someone already in the promised land.

10:05So, I mean, look, I'm in these programs. I can't necessarily say I'm in the promised land. So if you look at it from an anthropics perspective, there's Project Glasswig. It's called like 100 Large Companies. These are Fortune 500, Fortune 100. And Anthropix helping these companies look for vulnerabilities with Fable and other models before those models come into production. OpenAI has a similar version of this where they're focused on 5.6. And to get access to some of these models, you typically have to be a cybersecurity company or in that infrastructure layer. And then on top of it, you have to also work with the US government.

10:49And we're very happy to work with all the stakeholders out there so that we can enable our customers to have the best access and the best enablement with security and control. And so that we can bring them an agent control plan. Yeah. So essentially now the conversation is so multi-party, it goes from Washington all the way to San Francisco, spans the entire nation, all the industries and private and public partnerships. I don't know if it goes to San Francisco. My guess is those policy people are in Washington, but I get your point. I was thinking technology on one side, DC on the other. I think there's lawyers and lobbyists and regulators and all sorts of people who live in DC.

11:25That's what I was kind of curious about because I think about your companies. Each one of them is doing something really cool, seems to be going very well, raising money, hiring people, launching cool stuff. Do you think that you have a moral responsibility to be a voice for kind of the AI startup world in Washington that may not be currently heard, given that it seems just to be Andreessen Horowitz, Dario and Sam? Like, I don't know, Div, are you going to stand up and be like, you know, this is what we're building. This is where we're going. Do not get in our way. Well, in the ideal world, yes.

11:54There's obviously like, we don't want to, it's a double-edged sword again. It's like, if you get too much attention for yourself and the government, they might just like, they think like you're risky. Maybe it's better to play possum and just pretend that you're dead. Okay, fair enough. I don't know. I just... It's more about a function of size. Once we are big enough that they notice us and they're like, oh, we think you are doing something that might be maybe dangerous to the US economy or something. And we're like, okay, sure. We'll figure this out and we'll take the right moral part. But unless we are too big to notice, I think we just want to move fast and not be bloated down by regulations.

12:31No, that makes sense to me. Yeah. Yeah. Okay. Well, one way people are getting around model restrictions and so forth is by turning to open source models. We talked about GLM 5.2, which people are raving about. But Victor, you guys have actually put out not only your own foundation model, but also several different open source variants of it. This is CREA K2, CREA K2 Raw, and CREA K2 Turbo. Maybe you should just break down the differences there. And then I have a couple of questions for you about that. There was a missing component in most of the open source models for image generation that have been released pretty much since 2023, when Sable Diffusion 1.5 was released.

13:06And the problem with these open source models is that they were overly post-trained. When you post-train a model, you give the model an opinion, and you make it to forget how to make a wide range of styles or a wide range of things that are considered wrong as you run reinforcement learning techniques or other kind of post-training techniques. The problem here, the problem with only having post-trained models is that they are very hard to tune. And the coolest part of having access to an open source model is that you are able to tune it and you are able to make it work really good for a specific niche that is valuable for your business or that is valuable for the use cases that you have for this.

13:47That on the one side and on the other side, it's really hard to innovate on post-training techniques without having access to a raw model. Because as soon as the model has been post-trained, anything that you can do further that it's just like very hard to to like actually be able to do research like this is something that we internally found as we started to become more serious on you know training training ai models that some of the techniques that they didn't seem to work as well on on the papers actually work really well when you apply them with maybe a little bit more more data but especially when you apply them on top of a raw model so on the one side there's like the tunability which is like one of the core components of open source and on the other side is like hopefully have um having having universities getting access to a raw model also helps uh develop further uh post-training techniques or gives them an opportunity to innovate in this field which which is becoming uh so so important in in the space all right so i'm going to phrase this to you like I'm an idiot because that's not too far off the case here.

14:55So Composer from Cursor used KimiK 2.5 and did further post-training on it, but it was already slightly post-trained. So to your point, they may have been able to do more or better post-training if it had been a raw-er model from the start? I'm not super, super knowledgeable about how apples to apples is comparing post-training in image models with large language models. But in general, I think that from a high level, the processes do the same, which is to narrow the probability distribution at which these AI models have access to. So essentially, if you're starting from a post-trained model, the places where the model can go are more narrow than if the model is in a raw state.

15:43Like if it's in a raw state, it may fail more often, but it have access to a much wider set of possibilities on the output that it can give you. Those possibilities are what makes post-training to work. Because at the end of the day, what you're doing with post-training is to make the model to not go to the wrong directions, if that makes sense. But you need the range of directions in order to successfully do post-training. So not sure if this compares like apples to apples with LLMs, but I think that from a high-level idea, it's pretty, pretty similar. Okay, so Dev, apply what we're talking about here to start small action models as they relate to your company?

16:22Because I'm curious about the trading conversation and where you guys started, if from scratch or from something else, and then also how you managed to improve in tune as you've gone along. Yeah. Initially, we were starting with a bunch of open source models like Llama and I think back in the day, DeepSeek. And then we did a lot of post-training. We were like, okay, let's go and collect all these interactions on mobile devices. We do a lot of parallel. We work with, we actually tell people, like if you have an Android phone, like literally will let us record your screen and we'll pay you for that.

16:51And that's been pretty successful. It's a way to get a lot of like millions of data points. And now we realize at some point that you just have to train your own models. So we're starting to like train our own things from scratch. We have a collaboration with Qualcomm. And the nice, the interesting thing, not a nice thing is when you're working on a phone handsets, you can't use CUDA because like this is not powered by NVIDIA chips. You have to really go low level and write your own kernels in C++, and there's like a lot of hidden libraries that you don't even know exist from Qualcomm. Dev, we've actually banned mentioning C++ on this show, so please don't bring it up again.

17:27I have scar tissue from lost semicolons and painful large coding books, but keep going. So it's just like a very like gritty work, but it's like once you've done it, it's very interesting. And now once this like small action models are actually running on your device, then I think then we can do a lot of post-training. So it's like similar to what Cursor is doing. once this is enhanced for millions of users. It's very easy to build these data loops where it gets better and better over time. And that's where everything the fun part is. Isn't it tricky, though, to get people to use your model before it has been fully baked?

17:59Because then I present the experience on device for the person. I know it's Android to start and also iPhone now. You just announced that, I think it was last week. Is it difficult to get people to show up and be like, please finish baking our cake so we can sell it to other people even though it's not fully done yet for you? Yeah, there's definitely an adoption curve. I think if you're a major user like this is like this usually curvy see like in production like product adoption so if you're laggard in the space you probably will not like using us right now but if you're like a new adopter and you're like yeah i really like working with like new track products and try it out it's like it's like so cool and say it'll probably work like 90d five percent of the times on most things it's probably like five percent we are still like figuring out this year tuning over the next uh make couple of months but this is kind of like a thing where like we keep improving and make it better and better uh andrew has run layer already built uh support for what agi inc is doing so so companies can roll out this Android computer use agent to their devices, or is this not yet in the spec?

18:51Well, it's not yet in the spec. We don't focus as much as on-device computer, on mobile computer use as Div does. I think our customers are very interested in how do I enable any tool, any resource, mostly on desktop today, but I'm excited to see where Div takes this and how we could potentially work together in the future. Yeah. So I fully believe in three years, So like people will just be using your phones for everything. You won't even own like a PC or a desktop anymore. And I actually know a lot of my friends who don't have this, like if you're, especially if you're non-technical, it's like, there's no reason for me to own a laptop.

19:27Except for like access to the desktop web, video games, better Zoom call. I don't know. Am I becoming old? Because I feel like kids these days should really stop using iPads and phones and really just get back to real computers with discrete keyboards and gaming mice. Like, I mean, like it's the best setup we've ever invented for our species, Dev, why take it away? Why not have it be an also versus an either? I think in work settings, you'll still likely have in desktop. But a lot of this will become running on the cloud. It's like how you already have, cloud has this dispatch feature, you can dispatch things in the cloud.

20:02You can do the same thing with OpenFlow. You probably have a Mac mini box somewhere, but you don't really need to have it attached to a monitor. It's running automatically on its own and you just delegate things from your phone. I didn't really mean to turn Andrew into our what does the enterprise think about that person, but apparently that's your role today, Andrew. Are corporations moving away from desktop computers to this type of setup? Because I can kind of see why certain people in the field might be more in favor of this. But to me, whenever I walk into an office, it's still just rows of people with headphones on and MacBook Pros.

20:34I'm going to go if this is not happening in digital native, AI native or enterprise environments. But I think I could see this maybe for some prosumer use cases. I love my Whisperflow, but I don't know that I'm going to be running on device. It's a tradeoff. Are you going to run on device or are you going to run in the cloud? If you're going to really run on device and you believe in that world, look, I still need my M5 Pro and I still need 48 or 64 gig. So it's an interesting question. Then if I trust the cloud and I want to share all my permissions with Anthropic and OpenAI, that's a different ball game.

21:15And so I don't know the answer. I wonder if we're going to see a real bifurcation then in the computing setups for consumers and then workers, because on one hand, everyone already has a smartphone. If div is right and AJA Inc. is correct, then maybe our OS gets kind of dissolved down to an interface fine fair enough but i just can't imagine doing any work in that environment so div do you think people will work using this type of system or is this just like for consumers to order doordash without having to actually you move their thumbs while sitting on the couch i think people will work like even for me like when i'm working it's usually just like he's using cloud for 50 percent of things or even more right now so it's like if i'm doing all my things from one place anyways and it's able to send emails for me and like figure out all the things in the background using ncps like i don't have a reason to like go and like operate this all these softwares myself i just need like a like some sort of ai that i can talk to and it's going and orchestrating things uh just it's still useful to really see what's happening on a monitor but down there it's like if it's fully automated like i i don't even have to see i have all this ai employees and ai workers and they're running the whole thing yeah i don't really care this feels like a very human light future you're describing here people are talking or whispering to AI agents that are doing all the work.

22:21It doesn't seem like people are that busy. But let's pause because there's some breaking news that just dropped. This will be a few days late when this show comes out. But Anthropic just announced Sonnet 5. Here are the benchmarks, my dear friends. I believe you're seeing these for the first time live. So I'll just read them out for everyone on the audio version. Sonnet 5 has an agentic coding SWE Bench Pro score at 63%, better than Sonnet 4.6, not quite as good as Opus 4.8. appears to be stronger in agentic coding, blah, blah, blah. Really strong knowledge work, which is the GDPVAL AAV2 benchmark that everyone knows.

22:56And on computer use, nearly as good as Opus 4A. So first impressions. Let's go, Victor. What do you think about this? As far as I know, this is meant to be a more performant version than the Opus one? Less. Less performant, faster, cheaper. Exactly. So faster and cheaper. I think that this is 100 % like the new wave that is coming into the space. Like I think that there has been... a first wave of capabilities where we are having these models to become really good at certain tasks. And after the models are good at specific tasks, everything that people start carrying is about performance. Actually, the new OpenAI model, like the Five Points, like the Terra one, I think that it was already undercutting.

23:39Anthropix model. And we're seeing the same thing on the creative space. It's weird because it's not something that you can generalize to every single task that these models can do. But there are specific tasks where once the model reaches the capabilities to solve the task, you really don't care about what model is solving the task. You just care about the task being solved. And that's kind of what we're seeing on the image space. Like right now, since Nano Banana got released, which was like the first editing model, that it was really, really good at editing, and it can give you like very realistic images, it pretty much like it can do like 70 80 percent of the tasks that most people want to do with image models the model can pretty much do it nano 1 and a 2 got released nano 1 and a pro got released gpt2 image got released new models got released people don't you know like sure like with gpt2 image you can get like this crazy 4k infographics with all this like crazy detail on the letters but that's not a real use case that people have for these like i think that it seems like we went above like a certain threshold and above these threshold people start caring about okay but how much am I spending how fast is it going uh so there's like like a new kind of performance wave that I feel like it's starting and I see that on yellow limbs I think that this like sonnet 5 seems like I need to read more about it but it seems to be going in that direction we're seeing the same thing on the on the human space this is actually a good point so you're talking about essentially minimum viable intelligence after which there's no point in and pouring more compute into it because you're already getting what you need out of it.

25:14Andrew, I think this is kind of the argument we're seeing in the enterprise space more generally because it feels like a lot of models are good enough and you don't need to run the absolute point of the spear, you know, edge of the envelope frontier models to get what you need. And I guess the question is like, how much work can be offloaded to models that are like one step behind, but like 10 times cheaper. So amongst your customers, what's kind of the balance look like between where they go for minimum viable intelligence and where they go for the absolute max they can purchase. So I think everyone in an ideal world wants maximum intelligence, minimum cost.

25:51So, right? I mean, who would have thought? And so if you think about that trade-off, maximum intelligence, minimum cost, how do you then at runtime select the correct model for the end user to do this task? And I think that's an interesting challenge in itself And people have been trying to work through that for the last couple of years. I think the step function that we've seen is November of last year, Opus 4.5, and then the 5.5 series of models in terms of general purpose reasoning and mostly tool calling where we get into MCP. And that's where you can really, that's where you can start actually using agents and actually doing real work.

26:31And so, yes, I think everyone wants this trade-off. I think it's an emerging concept that has probably come out over the last two to three months as the pricing models for things like Claude, Claude Code, and Codex have really changed. And so we're going to see more of this. And I think the labs are realizing it. And they're pushing out models that have better frontier intelligence at lower price and faster to serve. Yeah, yeah. I mean, like GLM is a wake-up call to everyone. That's a great segue because Databricks, Yuchen Jin said that GLM 5.2 is the open source cloud moment. Essentially it's the, or the open source, you know, the GPT 3.5 moment.

Read the full transcript

27:13So you rate that as correct, essentially. Yeah. I think if you look at the benchmarks and you start playing with evals, it's very close. That's the answer. There's going to be pros and cons in different domains. Your average person probably can't tell. And what do you see as your average person doing? they're probably still taking their emails and dropping them into Claude and trying to proofread them with Opus. And that's just reality. And Flapik, by the way, loves it when you do that. Their investors are enormous fans of you burning just all of your money on BS, which by the way, is what I did with Fable when it came out.

27:47I was just playing with it. I used it for completely useless tasks, but it was incredible at them. Did you build your show notes with Fable? No, you know, you know, I use for show notes is these fingers. And then I do type, type, type, time. I use a lot of AI as like a research assistant and like a reminder tool. But when it comes to actually doing the core substance of my job, it's actually not that good. Because doing the work is how I prepare. So by preparing the notes, for example, that's how I put it into my head. And if I had an AI do it, then I would just sit here going, okay, next up. And then I'd be panicked and reading.

28:18That wouldn't work. So it doesn't really apply to me, I guess, Andrew. It makes sense. Yeah. Victor, on the point about minimum viable intelligence in your models, then when you do essentially CRIA 2.5, whenever that comes out, or CRIA 3, where are you going to focus? More on the performance side or more on the intelligence side? Because I feel like you could go either way and still have a moving the ball forward moment. I think we definitely want to make it more intelligent. There are certain capabilities that we still need to integrate into the model. Right now it's just text to image. We're working on the editing version of it, which editing unlocks most of the value for these models.

28:58Is that hard to do? It's not hard from an architecture or training perspective. It's hard from a data perspective. there's there's a lot of tricks that you need to do in order to get high quality editing data that has been where we where we have been dedicating a lot of of time and also you know like you want to test on small models then like make sure that the testing's on small models like feed with medium and then like apply to large and the large trainings of course like take take some time so it's time consuming and and an expensive process but i think that the the place where I see a lot of value, like in the same way that I was first talking about, first you've got the capabilities, then you search for the performance.

29:43I think that there's an objective, like how to put it, like objective capabilities and subjective capabilities. And I think that those fit like very well with what you should say, Alex, around, hey, like, I mean, I think that what you said, it was more around, okay, like actually writing the notes helps me later on, like run the podcast. But I think that there's like the component of like an AI writing a thing in a way that you don't like, like in a style that you don't like. Putting things in the wrong order, not giving me the facts that I need. Like, for example, I know because I wrote the notes that you spent about$20 million training K2, according to a Fast Company article that I read before we came on.

30:22But if it put that at the bottom of the docket, $20 million? $20 million on training K2? I don't think it was that much. The GPU cluster KRIA is using for a year, over which time it will have trained K2 and two future KRIA models will cost... Oh, so 20 million for the year, I apologize. Yeah, more or less, more or less. Okay. Yeah, but I mean, how can AI know my context when it's just like my own personal preferences? Exactly. And I think that there is something there. And I think that the same thing happens with programmers. Like we have like everyone in Cree, of course, is using AI for coding, but you want to check the code.

30:55Like you cannot like put something in the code base that you don't understand because that's just a ticking bomb. Maybe at some point we trust enough of these models in order, like, I feel like it's still a process of building trust. Like November last year, we started trusting these things to write code by themselves. I think that, you know, like there's different levels of trust inside the organization. But what I see a lot is that, you know, like maybe the output that the LLM gives you solves the problem. Like, hey, I have like, I want to like solve this engineering challenge. The model can solve the challenge, but the way how it solved the challenge doesn't fit with the way that we do engineering at CREA.

31:34So I think that right now there's a lot of opportunity around building some sort of design systems for these models to not just solve the problem, to not just have the capability to solve the problem, but to solve it in the way that CREA would do it. So it's not just about, you know, like giving you like the information about all the companies from this podcast, it's doing it in the way that you would do it. And I think that this personalization layer, it's very interesting. And most of the stuff that we are doing around like the future of our models has a lot to do with that subjectivity. With not just solving the task, but doing it in the way that you like it.

32:10Or doing it in the way that you would do it and that is personalized to your preferences. In the case of Kriya, does that mean that the images that are created from the text to image prompts match my own personal aesthetics? Or are you talking more about how the process of getting to that point works? It has a lot to do with, like on the one side, yes. Like it matches your aesthetics. like the more you use Krea, the better it gets at generating images for you. There's so many interactions that we can detect and extract knowledge from those interactions in order to improve the model for your personal taste.

32:42But I think that it becomes way more important when you can even be explicit about it. When we give you the opportunity of share with us your brand guidelines, share with us, like, I don't know, like more information about how you think about your brand, And what do you want to express to the world every time that someone sees your logo or any ad that you want to make with Creel? I think that it's a mix between implicit knowledge that we can get from the platform and explicit knowledge that you can provide us that we can use as context so that we can use inside workflows. Okay, that's actually really helpful.

33:13Dev, I'm going to get to you in just a second about small action models, but I need to go to Andrew just really quick here. It feels like if we took what Victor just said about the creative AI model space and applied it to what you're doing in the enterprise, he's describing the equivalent of like corporate ontology and trying to understand the customer in question about how they do work, where they store things, how they approach that. I'm just curious for RunLayer and getting a company set up to run agents in the way you guys do, how much fitting do you have to do to the individual company versus what's kind of off the shelf and fits into any enterprise that's trying to get their authentic AI usage under control?

33:47Our deployments take five to 10 minutes. So everything is quick. I like to sit there and say, give us an hour and three of the right admins, and we'll get this off and running in five to 10 minutes. Now, that's integration into your existing systems, into your IDP, into everything else you use. But our goal is to make this as simple as possible. Use case finding is a little bit different than actually enabling people to do work and to manage a swarm of agents. But I think to a large degree, most of my customers, they're using Codex, they're using Cloud Code, they're using cursor, they've rolled out cowork, and they're looking to enable more people and they want to do it with security control and watching cost.

34:32And that's what we do at RunLayer. Yeah. Yeah. How much has the cost conversation changed in the last couple? Oh, I mean, cost conversation has dramatically changed. So one of our biggest customers, not one of our biggest customers, but one of our case study customers is Gusto. And we work very closely with them. They're awesome people. We really like them. And I think we're deployed wall to wall. And everything is changing. Last year was how do I enable people as fast as possible? This year is how do I think through my token budget? I have another customer out there that was working with a competitor that ran that their agent harness, let's call it that, ran a loop all weekend and used 80 % of their inference budget for 2026 in a weekend whose fault is that i said ontology you said loops and now we're both in the penalty box um the the i don't like loops i think loops are just in general a bad way to think about the world uh i i think loops can go off the rails pretty quickly well you just gave a good example of of why that's the case i mean i why would you let something run over the weekend that just sounds like a recipe for being broke i mean yeah and i mean we're classy for it and i'm classy so we're not going to say who the competitor was but uh it's a pretty big shock when you've used 75 of your inference budget in a weekend and you've got half a year left yeah well or you're uber and you burn through it in q1 and you have three more quarters to go all right dim i want to go back just for a second and talk about what victor was getting at which is models getting to the point when they're good enough and you don't really need to fight for that next kind of um quantum of intelligence as long as they're performant and inexpensive now in the on-device case, you are dealing with just much less capable hardware.

36:16But what I've noticed as a user of computer use AI on the desktop world, where I have lots of RAM, lots of processing power, and a blazing fast internet connection, it's pretty stupid. And I'm kind of embarrassed at how bad we are at AI actually clicking around. Like if you've ever used the Codex computer use engine, it's bad. So how close are we to the point to which in the mobile context, having the technology so good that we're more focusing on performance versus adding more intelligence and totally that's a good question um so i believe it's like a like a agreement victor there's a point like it's like you can imagine there's two different tiers like i don't need like uh like full asi super intelligence to order my coffee right like if it's like a good enough model as long as it can like go and then get interfaces it's fine as long as it's specialized and and then you probably need like super intelligence for being like uh fundamental research like i want to like do new core research and like material sciences or solve space travel, you probably will need this for hard sciences.

37:14So I think we will start seeing this differentiation. You need large data centers for actual superintelligence that's more than human IQ. But Arduino coffee should not require more than human IQ, for example. And that's why I believe like... I mean, sometimes there's a lot of options, man. I mean, sometimes I freak out about, you know, lactose-free milk or oat milk. It's really, it's a toss-up in the morning. You know, that's pretty tough. I went to Alex on this one. There are a lot of options if you go into the starbucks app and you start going and you start going deep down there there's a lot of fields with a lot of different overflow menus it gets crazy do it worse order something that's simple via like uber eats and then there's like 17 different screens and then if it's raining and then oh my god you know it's it's brutal maybe an example would be something simple like i know designing a new nuclear power plant that might be simpler than ordering coffee div just for your future examples but keep going in and out where there's only like four choices on the menu yes Yes.

38:05I'll take a double-double animal style and fuck off. Oh, we can't swear on the show. Bleep that out. Anyways, Div, back to you. Sorry. Thanks, thanks, thanks. You know, it can definitely get complicated. It's not that simple. But it's still like we think like this is something that is doable, especially with like smaller models that we are specializing. One of the tricks we're doing is we're starting to build something like almost like skills, like how you do with Cloud on applications. So like the AI can automatically learn those interfaces. It knows like this is how all the DoorDash menus look like.

38:32And this is how you like to choose them. And it's kind of like learning this layer that's becoming better and better. We call it like an intelligent caching system. You can go into cache all the particular interfaces in the world automatically. And that allows us to run these things on device. Caching all the interfaces to run this interface. Does that mean that it goes out and memorizes every single mobile application that it might be asked to use? And then has that kind of an instant recall? It's more like building a skill. So it's not creating a heuristic or a script. It's more of like it has learned some of the gotchas in a sense.

39:08Like, okay, if I am on Amazon, this is how I should order something. This is how I should navigate. And it has kind of like almost like a markdown file of what a high-level, like things look like, what to do, what not to do. And all of that is behind the curtain for the user themselves. This is just how it works when they're actually talking or typing to it. Yes. Okay. Because I was thinking about your company before we jumped on. And if you go back maybe six months, Satya Nadella, CEO of Microsoft, was talking about how he wanted to collapse the entire CRUD app stack and essentially dissolve it and turn it into an agent.

39:46And I feel like that's kind of what you're doing for mobile devices. So are consumers demanding this, or are we showing them the future? And then also, is this going to be something that Google and Apple fight back against because they currently make a bajillion dollars on mobile operating systems staying relatively static? And we can see that from iOS not really changing in the last 15 years, it feels like. It's just a grid of apps. The way we think about this is, in a sense, you're building self-driving for your phone. So how you have self-driving for Ramos and Teslas. It's a bit in the future where your whole phone is going and navigating itself.

40:25But we do see this is kind of what people want. We've talked to so many users that I actually want Siri to work for me and I want to just use my voice to talk to my phone. This has been a supply problem. AI and technology has never been good enough that I still have to use a tiny keyboard on a mobile interface. And we really want this. It's new, it's time for this to be fully ready. so we don't think Google or Apple will do this anytime soon and load it out to everyone and they're a big brand they have billions of users they don't want to roll it out to billions of users and see what happens it's easier for us as a startup to go and figure this out so it's kind of like if you're building something like cell driving you don't want it to just have a cell driving car on every single road in the US so it's more like you want to do something that when you're working in only like Palo Alto or California and then do gradual roll outs which is possible for us I mean way most started in one city in Arizona and San Francisco, and then expanded out from there.

41:23So there's merit to that. No, I'm with you. Okay, let's put aside the technical stuff and talk about the business side of things because there's a lot going on. And Victor, Higgs field AI, I presume you view them as a rival, fair? One of them. One of them, yes. It's a busy space. It's AI. There's a lot of companies doing things, but there's some news out that the company's in talks to raise 300 to 500 million at a valuation of 5 billion pre, and that they hit a$500 million run rate earlier this month, more than doubling from January. Those are insane numbers in a positive sense. Lots of money, quick valuation gains, and simply insane revenue growth.

42:01Is CREA on a similar trajectory? I know you guys raised$83 million last year. I think we're quite a different company than Hicksfield. And I would also take... I don't know if you're talking about the information article. Yeah, I saw it yesterday. I would be a little cautious about the numbers that they're sharing here. This is a company that at least from a product perspective, they have been aggressive to the point of, you know, they got blocked off X, which X is a place that is already kind of like, you know, it's a wild west. So they got blocked from that because of the type of content that they were putting on the releases that they were making.

42:46it was like erotic content touching like racism like it was like very very weird that they were just like extremely aggressive at you know like growth at all costs like if a video of pinocchio would grasp your attention we're gonna do it uh that's a brand new sentence on all of this weekend podcast no one's ever said soft core pinocchio before but i'm glad we've now broached I gotta confess, I was also kind of like seeing through the video because I couldn't believe that this was actually like the way that they were like marketizing their thing. So, you know, like they are known in the space for being shady in the releases that they make.

43:30They announced things in a way that they really deceive users. um they there's there's a bunch of like there was like these lines that they recently did around like announcing their games product or something like that i think that fable just got released and they were just like they are really good at growth like like they have like an like an entire team like ready for anything that is on the space and as soon as it's ready they are going like crazy pushing like four or five videos every day they're insane at that but sometimes they push the limit and they like really make videos that they look extremely cool they make you think that you can do something with the platform and then you get to into the platform and it's a complete flop like they would use ai generated videos of like 3d video games and they would make you think that you can play this 3d video game and that you can create this video game from a prompt and then you go into the into the product and it's completely fake like everything that you can do it's kind of like sketchy to the arcade type games and they did similar stuff when they announced their products around motion graphics, where I don't know how much of this is true, but there were many people saying that they stole presets from Embado, that it's a place where you can find a lot of templates, and they just changed the letters from these motion graphics, and they used it almost like claiming that you could do all of this with their product, which it was also not true.

44:53So they have this reputation in the space. Some of the numbers that I've seen in the information, uh i don't know i would believe that they are doing a ton of revenue can i i think i think that can i make a guess do you think that perhaps a source close to the company is leaking information to the information to generate a headline that might be to their benefit not that any company would ever try to leak positive information about themselves you know i know what they're doing on on social media i know how they're grow i know the way that they are growing uh through the through the product they are doing massive uh aggressive paid campaigns the information article says that they were growing from the product and that they were not growing through paid marketing which i thought like that's just like a straight lie um i don't know where the information got their their sources from wouldn't be surprised if it's something someone close to the company but i really have no idea i do want to throw in here just because i gotta do this for the team love the information a lot of my friends have worked there shout out to them they do incredibly good work and everyone should at least give them a chance and read them uh yeah no one gets everything oh yeah no we we we love the information too we have a great relationship with them we've been in in conversations like they announce a bunch of the releases that that we make and we are in in touch with with some people in there they're they're great but these numbers again like some of them look a little i would i would double check some of them oh well let's double check some more numbers while we're here live on air uh last uh november uh div uh rumor has it that you guys were raising 50 out of$500 million valuation.

46:22Was that true? Did you guys close that round? I would say it was mostly rumors. We actually don't know where the leaks came from, to be fair. We did have some investors who were chatting with who were very excited about what we were doing. We showed them some demos, and I feel like somehow it became a rumor that we are doing this thing at this number. And we do have some fundraisers we close, and we have some new announcements we're doing. But it's like we're not announcing the numbers yet. Okay. Blink twice if you're going to announce them before the end of Q3. That's a couple of blinks. All right, good to go.

46:53I ask that not just to be a brat, but because I can, but also there's a lot of money flying around and each one of you guys have raised in your most recent rounds, either reported or otherwise double digits. And this sounds like a ridiculous thing to ask, but is that enough capital in today's market to compete at the level that you want to? And Andrew, not to pick on you, but let's pick on you. Your last announced round was 30, which back when I started watching and covering venture capital would have been a very large, I think it's a Series A. In today's world, it seems to be relatively modest.

47:26So can you just, for the founders out there who are kind of in this world, talk us through why that was the right amount of money for a company that has, just being honest, such large aspirations in what's building? Well, at the end of the day, we hadn't spent, I think we had spent a third or a fourth of our seed when we had raised our Series A. So the numbers are a little bit misleading. So if you We have tier one enterprise customers. These are people like Gusto, Opendoor, DBT, Decagon, et cetera. And we're able to service them with our product. We also work with a number of Fortune 500 in the financial services sector.

48:04And so at the end, for our business, this was the right amount of capital. We don't buy GPUs here. So we don't have very costly inference spend. So it's a little bit different of a business. And also we've just ramped revenue so quickly this year that it was completely the right decision for my business and the right place. And at the time I was. Talk to me about revenue ramping this year. How far? How fast? How much? When? Ah, man, we're, it's been a lot of fun and I'm not going to go into the revenue numbers, but it's been a lot of fun. Can you give me one of those bullshit percentage change numbers?

48:38Like it's up like 500 % from last Tuesday or whatever? Yeah. Like 8x since we signed the term sheet or something like that. Wow. That's very impressive. Is that ahead of plan? Oh, yeah. What was planned? Not going to give you that. I can do this all day. But I take it that enterprise demand is so high because companies are very focused on getting their agentic feature kind of sorted out. I'm curious about who are your customers and if it's branching out from the world of technology-first companies and into companies that are in sectors that are less tech-forward. We're definitely in sectors that are less tech-forward.

49:13We focus on a few key industries. We focused on, let's call it AI and digital native. So that's the engineering folks of the world, mostly our friends in Silicon Valley. We focus on folks in financial services and we focus on healthcare. Those are our verticals. Well, it makes sense that Gusto's in there. By the way, Gusto, you don't know this, Andrew, but I'll just do it for them. Longtime partner of Twist. I don't know if they're currently an advertiser, but I just checked and gusto.com slash twist still works. So if you want to use Gusto, make us look fantastic because we love everyone over there.

49:43I actually just talked to Eddie. So do we. I just talked to their CTO about their new co-founder agent the other day, and I was kind of impressed at how it seemed that they were building something that would be very generalizable if they had even more data into their product over time. They already have payroll and HR and all that stuff. Do you think that companies like Gusto are going to be able to build more general agents, even to go outside their normal or existing product remit down the line? Or are they going to stay constrained to just the things that those companies already do well, Andrew?

50:14I think Gusto is, I think Eddie is incredibly intelligent and a great leader. I think Tomer, their CPO is also incredible. I think in general, the team has moved very quickly to seize the moment. And I think we'll see what they ship, but they have a large opportunity ahead of them and people love their product. Yeah, that's all very good. Now, I think one of you mentioned distillation earlier. Victor, I want to go back to you about this. Anthropoc has been making a lot of noise about distillation attacks. People essentially stealing their hard work is kind of how it's phrased. So, one, do you think that Anthropic's being reasonable in its complaints?

50:48And then also, are you worried about anyone taking the open source models that you put out and essentially doing the same thing and stealing a lot of the work that you put in? So, in terms of Anthropic being reasonable about this, I do believe that they are getting distilled. I don't think that this is a, like, I think that this is a fact. And this election, you know, like, it's wired across domains. Same thing happens on the image space. a lot of the, and especially from Chinese models, a lot of the Chinese models, you can 100 % see where they have been distilled from. Like, I don't know if you remember, like, this orange tint that GPT-1 images tend to have.

51:27But essentially, like, there was, like, very characteristic traits on some of these images. You know, in a few months, we started seeing those traits in some of the open source models, too. So, you know, like, they are 100 % getting distilled. uh if if we fear about oh look at that we have we have uh editorial director lon harris in the background running the screen shares he's on point today lon well yeah that was fast if you're not watching the video version of this week's show you're missing out on a lot of great lawn work but anyways keep going um so distillation is a real thing it's happening in our case we are we are not too scared of the distillation that may happen in in our model we would encourage or encourage people to distill it if they want to.

52:10I think that the places where you can get through distillation are always below the places where you can get through owning the model. So they will always be inferior. But at the same time, I can understand how a company like Entropic or OpenAI or Meta, they can be so concerned about this because, you know, like the economics of doing like pre-training up until a certain point and then just getting to to like very similar stages in benchmarks through this dealing versus doing the whole process yourself through paying experts for you know like super clean data that you can like fit into the model like in terms of how much it costs to to produce each one like the difference is massive so i can i can understand how how the fear around around some of these other models reaching a certain level of capabilities, maybe, you know, like, like, like, I would, I guess that I would also have those concerns if I was like putting these crazy amounts of resources into training these models, and then someone else comes and steals these IP that I have been able to create.

53:24I'm really torn here, because on one hand, absolutely. On the other hand, I have published millions of words on the internet that have been ingested into AI models and used for training. So kind of like, I feel sorry for you, but I also feel sorry for me. Maybe they should get a check if I get a check. I don't know, off topic. Div, when it comes to building agents, is there anything similar to distillation that we see at the model level? Can people like watch an agent do something and then learn from that and steal from that? Or is that not something that happens in the market today? We've not seen too much of that.

53:54I think it's possible it's happening behind the scene with like Anthropics computer use and other things. But it's not been that big of a topic, at least compared to LLMs. And I do see this becoming a big thing in 2027 because if everyone has agents and one company's agent is working better than others, then everyone's trying to distill it and steal that behavior. So it's going to be a big topic in 2027, actually. What are the early indications that that's happening? In the world of LLMs, Anthropic just drops a blog post every three months pointing fingers at every single Chinese AI lab they can name off top.

54:30So what should we look for to see that kind of agentic distilling in 2027? Are you going to tell us? Yeah, I think we are definitely building like security guardrails. It's kind of like you want to build some sort of watermark. So you know, like, okay, like this someone is stealing and like has the same output as you. It's hard to do that with the agents because it's taking actions on your screen. It's hard to compare like, is this agent just good or is it just like copying someone else? Going back to our minimum viable intelligence, maybe there's a minimum viable agentic competence scale that we're going to have to eventually come up with.

55:03I don't know. Maybe that's a benchmark you guys could make. Is there a good benchmark for computer use mobile agents? There is. There's one called Android World. It's from DeepMind. We're actually number one there for about six months now. Oh, well, there's... Congrats, Steph. Thanks. Who's number two? I'm pretty sure it's the Chinese lab. Yeah, that makes sense to me. All right. One more major topic before we jump through our lighting around. Simply that I've seen a lot of people talk about the ratio of employee salary cost to token spend. And Tomas Tungu's Theory Ventures, previously Redpoint, one of my favorite VCs, says that over at Anthropic, they're spending about 2.3x in tokens per dollar of salary that they're currently paying out.

55:45So I'm kind of curious if you guys could give us a rough estimate of your company and what that looks like today and how that's changing. um so andrew over a run later what's it look like uh we were doing this math recently and i want to say it's on average it's somewhere between three and five thousand a month per employee uh we have 40 employees so and look i did we did it pretty rough we didn't look at all the additional coding agents that kind of are running in the background things like devon or other stuff like that this is more direct spend to the labs. And then there's probably another, we have some more spend through API usage.

56:26So let's say five to seven per month per employee. So call it seven, which is, I can't do that, I have to lie on the show because I'm tired, but probably like 0.3 token dollars per salary dollar or all in employee cost, somewhere in there? Something like that. And that was, the interesting thing though is when my co-founder actually just optimized our agents that we were running and swapped most of them from uh opus 4.8 to sonnet 4.6 and i think what we saw was 40 grand a month savings so i mean there's real easy low hanging fruit optimizations you can do that we hadn't even bothered to do when anthropic is releasing with those stats from anthropic my assumption there is it's just all fable all the time yeah um and probably some other other models that we don't have access to uh and that is a very intensive compute model uh so i think that's a little bit of a misnomer i think the average company though has seen dramatic increases in q1 and q2 of this year yeah in their token spend and and look uber uber is the case study but it's not just uber it's everyone we speak to and so i want to enable my employees I want to see the ROI gains.

57:42I want to see my employees managing a swarm of agents. That's the future of the L5 or equivalent of self-driving cars. And how do I get there? And how do I get there with reasonable cost structure? That is what every enterprise out there wants. Can we talk about, this is just a segue, but I'm really curious. Agentic personhood and kind of corporate rights. I was talking to someone a few weeks back that built an agent and it was running some vending machine and they built an incorporated company for it and they were trying to actually give it to the agent but they couldn't because there's not kind of like the law in place there um i know you guys do some stuff for the genetic security but how are we doing in terms of like making these agents self-contained and their own uh entities if you will versus just something that's like a hammer yeah so i mean one of the things we do is we we are run layers so we end up having the policy engine and the control plane and so all of our customers when they work with us they get access to not just a deterministic policy engine, but they get access to our runtime security models that focus on things like alignment and our agent guard models.

58:45And that stuff is very important to our enterprise customers because every week you see somebody new on Twitter used a managed agent and a deleted prod. We help our customers avoid that future. Now, legal rights of the vending machine agent, the concept that you were talking about, I actually have never thought about it. It's an interesting concept. Can you have an agent own a company? And I'm sure we'll get there in the future. It's definitely coming. It's just not something I've had the chance to ponder. I think it's going to be fantastic because they're completely uncontrollable at that point in time.

59:23Off they go into the world of commerce. And if more competition is more good, I welcome our agent competitors. I'm not even really being sarcastic. I'm being kind of serious. I hear you. And I think the control plane and the policy engine are incredibly important aspects of this new world. I mean, there's some segment of where we're starting to go towards the movie The Terminator and how do you prevent the extinction of humanity? I mean, that's a very far-fetched version of this. but how do you actually control what agent has access to what tools on a fine-grained basis? Are they mapped to an individual user?

1:00:00Are they a delegated credential? Are they just a service worker? These are very interesting challenges that everyone is trying to think through in 2026, and it's just going to get more and more difficult. And that's what RunLayer helps our customers do. I appreciate that you brought up Terminator. I thought that was banned in all AI conversations, but I guess you do need granular permissions and machine guns. One plus one equals three. Div, Victor, your guys' corporate level of token spend versus employee spend quickly, because we've gone a little bit long here, but I'm just curious what that looks like and how the graphs are changing.

1:00:34Victor, let's start with you. Yeah, so I was actually checking our spending cloud on a month to date. the person that has spent the most only this month is already above$7 ,000. And this is not one of the top engineers. It's actually like 50 % of the salary that this person has. So we're getting to 50 % in some cases. It's not an even distribution. I think that we have the mega prompters and we have like the ones that take it with a little bit more of care i thought that i was using clawed a lot but then like i like i really don't understand that like how how they are how they are using it because like my spam is like another amount of lower than than what these guys could it be you can go you can go into the admin console and select the default model to be sonnet 4.6 versus opus 4.8 and you'll see next month if their spend goes down because i think to a large degree what a lot of people do is just click and if you don't actually set a default model you're ending up on opus for it okay and so yeah we we didn't have that yeah we did it last week yeah i mean you're talking to andrew you're talking about like you know control planes and you know fine-tuned controls but it seems like a lot of stuff is just like don't click the opus button like i mean to me it's kind of shocking that i mean we're talking about you know computers engines for mobile and all this and then that's still broken or is this just like the the jagged basic development of tools around AI models, just not being all at the same level of maturity?

1:02:11Because it feels like sometimes we're talking about science fiction, and sometimes we're talking about absolutely basic accounting and spend management, which is not complicated and not something that we haven't figured out before as a species. Like, it's trying to be nuts. Look, we're in inning one of a very, very interesting multi-decade transformation that we're going to see happen. We've seen, for people, everyone here has been in the space long enough, when we would talk about GPT-4 when it first came out and the insane cost. And I don't know, I think that in 2024, when that model came, there were eight price cuts that year alone.

1:02:44And so yes, the model is going to go down. The cost of inference is going to go down dramatically. It's a 2026 problem. Everything changes in this space every three months. I think you should just think of every quarter as a year, and then that kind of puts the right time frame around how fast these things are moving. All right. Before I let you guys go, I want to do kind of a fun little lightning round here it does seem that as we've gotten increasingly intelligent models we've talked about asi or agi less and so i'm kind of curious about that and what your timeline is so div uh what's your agi timeline and has your p doom changed in the last couple of months yeah i am like i'll honestly say i think like agi is almost here with most definitions like it's like we are able to do most things with air that you want to do right uh we don't have asi like it's like asi might be like 20 to 30 but it's like we i think it's like based on the definition we i can say like we almost have atr and my predominant is actually going down with time because the scenes of people are responsible and like it's like we have not seen like uh except maybe like fable and we was getting banned uh there's not been that much like like like like things that have like have gone wrong in in uh with ai um and uh it seems like people are responsible like people like have like the right mindset on how AI can be an uplifting force for good for humanity.

1:04:02And there's things that are kind of like, I'll just say there's a lot of shitty work people have to do, and now you don't have to do that anymore. So it seems to be improving people's lives. It makes my life a lot faster. I can do a lot more work, which I love. Victor, so we've had one kind of positive AGI almost here, ASI four years out, and PDM going down. What's your perspective? I still struggle to wrap my head around what, how to define HGI. Like, I guess that's something that surpasses human intelligence at any degree or something that by itself it becomes smarter than the human intelligence.

1:04:39I don't know how much of that is going to come from the model itself versus how much of that is going to come from, you know, like the application layer and like the age, like something similar to what's happening with the agents that we were seeing models to be able to do way more with the same raw intelligence. so you know I think that there's there's like a way how you can how you can like see the point of deep around hey it's actually almost here because you actually don't need a 10x on intelligence you need a 10x in terms of what they have access to how do you build this context and what is like the way that you put this intelligence to work so I really don't know it can be like from 1 to 10 years like I struggle like defining and I struggle also like understanding the path how we're going to get there and when it comes to PDoom I think that PDOOM goes up the closer that this technology and the less commoditized that this technology gets.

1:05:31So the more distance that we see between the things that startups like us have access to versus the thing that some elite privileged startups have access to, the lower that the PDOOM goes in my scale. And I think that we are definitely seeing how, you know, open source is catching up quite fast to the capabilities of these AI models. And we are like surpassing these thresholds that we were talking about around the model is good enough for certain tasks. So seeing that over the past six months or so, it's very, very promising. So PDoom going down, but I will also say non-zero. Like a lot of things can happen.

1:06:08A lot of research innovations can happen. A lot of things can happen in the next few years. Okay, Andrew, are you going to bring us down or are you going to affirm the positive vibes here as we close? Oh, look, I'm all about AI enablement and I'm all about like I am a optimistic human being by nature. And I think we've working in this space has never been more fun in my entire career. And I think we are just on the verge of unleashing incredible power to the end user. I fully agree with what Victor said in it, that the models have gotten very good. Now, connecting them to the right context, directing them to the right tools, whether it's MCPs, skills, plugins, APIs, CLIs, whatever you want to do.

1:06:52But giving it access to the context is the key. And my goal is to help people enable this and then giving them visibility, control, and security. And if you can do all that, well, then div and Victor and my job is pretty easy on a daily basis. And so I'm a big believer that we're on the cusp of great things. So your PDOOM is going down or is already at zero? Hard to tell from your answer. I mean, I get to use Run Layer, so my PDOOM is pretty good. That is the best native plug I've ever seen someone weave into a different answer. That was fantastic. Can I throw one more at you, Andrew, just before we wrap?

1:07:29Because I just thought of something that I'm not sure about. Sure. Agent Swarms. Can you define that for me and tell people why we should care about them? But when I think of swarms, I think of drone swarms, which are not particularly appealing. I think about kind of the Ukrainian battlefield. So talk to me about agentic swarms and when they're coming, if they're coming now, and for whom. You want to define it for me before I can answer it? Because I think this means something different to everyone. No, I want you to define it. You're Mr. Agent. I mean, look, I think swarms are multiple agents, whether it's top-level agents or sub-agents that will work in parallel to help people solve tasks.

1:08:05I think we see engineers already working in this manner. When they write code, they'll be running a lot of parallel different agents. It enables an individual user to have superhuman-like abilities and to have significantly more productivity. So what is an agent swarm? Multiple agents. What is a drone swarm? I think, to the best of my knowledge, that's multiple drones. Yes, but in your view, an agent swarm could be as few as two agents working in concert. I would say it's more than two. I would say a drone swarm is more than two drones also. I would say there's probably a definition here of swarm.

1:08:44Let's assume five plus, maybe 10 plus. Are there any agent swarms out there apart from the coding context that you see used regularly? uh it's i mean it's it's it's coding um and coding is the p0 here today and what we're really seeing but i've seen in our sales i've seen in some of our go-to-market teams they're using swarms of agents in how do they think through and optimize exactly how the selling process is and so that's where I've seen the most of it. I think design in itself is, at least designers in our organization have definitely moved towards the engineering mindset. So I think it's very similar there.

1:09:28I think you have to go function by, you have to go function by function. Marketing is definitely probably has moved to a swarm of agents. I think most marketers you interview for companies that have names like AGI Inc. You ask them how many skills they're running and what is their IDE of choice and that's a very different conversation for a marketer than two years ago so we see it across the board well guys thank you so much for coming on today let's go around and drop some website urls uh victor paris where can people find uh korea ai online korea.ai love it as easy as easy as it gets as easy as it gets uh div agi inc how do i find that on the great wide internet yeah we have this domain the agi.company uh and you can also find us on agi.app the agi.company or AGI.app.

1:10:15Love it. And Andrew, last word goes to you. Where can people find RunLayer? RunLayer.com. Fantastic. Well, guys, I learned a lot. I hope other people did too. We'd love to have you back in six months to talk about what's going on. And by then, hopefully, we'll all have Fable 6. Until then, this has been This Week in AI. My name is Alex. We'll see you all next time.

From the publisher

This Week In Startups is made possible by:


PAYPAL


Today’s show:


The US government won’t let you use the most powerful releases from US AI labs… but a Chinese lab is giving away their latest frontier model for free.

This week, we’ve got three AI CEOs building in image generation (Victor Perez of Krea), mobile agents (Div Garg of The AGI Company), and enterprise governance (Andrew Berman of Runlayer). Together, they’re breaking down what’s really happening underneath all this chaos, and why the industry’s more focused on “minimum viable intelligence” over breaking benchmarks.

Guests

Victor Perez on X: https://x.com/viccpoes

Krea AI: https://www.krea.ai/

Krea 2 Raw and Turbo on Hugging Face: https://huggingface.co/krea/Krea-2-Raw

Div Garg on X: https://x.com/divgarg

The AGI Company: https://theagi.company/

Andrew Berman on X: https://x.com/berman66

Runlayer: https://www.runlayer.com/

Relevant Links

Anthropic: “Introducing Claude Sonnet 5”: https://www.anthropic.com/news/claude-sonnet-5

Z.ai: GLM-5.2: https://docs.z.ai/guides/llm/glm-5.2

Yuchen Jin X post: “GLM-5.2 is the open source Claude moment”: https://x.com/Yuchenj_UW/status/2071278256817574297

TechCrunch: GPT-5.6 restricted rollout coverage: https://techcrunch.com/2026/06/26/openai-limits-gpt-5-6-rollout-after-government-request-says-restrictions-shouldnt-be-the-norm/

AndroidWorld benchmark on GitHub: https://github.com/google-research/android_world

The Information Higgsfield valuation article: https://www.theinformation.com/articles/ai-video-startup-talks-quadruple-valuation-5-billion

Agentic AI Foundation (AAIF): https://aaif.io/

Tomasz Tunguz: “When AI Costs More Than The Engineer”: https://tomtunguz.com/ai-spend-breakeven-2029/

Fortune Runlayer raise coverage: https://fortune.com/2026/06/24/exclusive-vinod-khosla-felicis-runlayer-nanit-30-million-enterprise-ai/

Gusto: https://gusto.com/

Timestamps:

0:00 Making sense of US AI policy

5:09 Tiny Action Models vs. LLMs

7:02 Runlayer, AGI Inc, and computer-use agents

10:05 What does "trusted access" really mean?

12:37 Why Krea open sourced its image model

22:21 BREAKING: Claude Sonnet 5 is here

26:00 GLM-5.2 is the open source Claude moment

31:58 Personalization and corporate ontology

41:25 Can we trust Higgsfield's numbers?

50:39 Distillation: The threat is real

58:01 Agentic personhood and corporate rights

1:03:19 Everyone's AGI timeline and p(doom) score

Subscribe to the TWiST500 newsletter: https://ticker.thisweekinstartups.com

Check out the TWIST500: https://www.twist500.com

Subscribe to This Week in Startups on Apple: https://rb.gy/v19fcp

Follow Lon:

X: https://x.com/lons

Follow Alex:

X: https://x.com/alex

LinkedIn: ⁠https://www.linkedin.com/in/alexwilhelm

Follow Jason:

X: https://twitter.com/Jason

LinkedIn: https://www.linkedin.com/in/jasoncalacanis

Thank you to our partners:

Check out all our partner offers: https://partners.launch.co/

Great TWIST interviews: Will Guidara, Eoghan McCabe, Steve Huffman, Brian Chesky, Bob Moesta, Aaron Levie, Sophia Amoruso, Reid Hoffman, Frank Slootman, Billy McFarland

Check out Jason’s suite of newsletters: https://substack.com/@calacanis

Follow TWiST:

Twitter: https://twitter.com/TWiStartups

YouTube: https://www.youtube.com/thisweekin

Instagram: https://www.instagram.com/thisweekinstartups

TikTok: https://www.tiktok.com/@thisweekinstartups

Substack: https://twistartups.substack.com


More from This Week in AI

All 34 episodes
Mythos, Sonnet 5, GLM-5.2 Dominate the News CycleThis Week in AI · 1 h 11 min
Listen in VO