In short
Episode topic: How AI coding and agent tools are evolving, focusing on Anthropic’s Claude Code/Cloud Cowork and OpenAI’s Codex, plus adjacent updates in AI design, “token maxing,” and robotics generalization.
Guest backgrounds
No guests mentioned; the host discusses companies and research.
Key claims
Enterprise AI coding is shifting toward compliance/security and workflow ownership (Factory’s enterprise niche; Anthropic moving up the stack with Claude Design). Raw “tokens generated” doesn’t equal long-term productivity due to high code churn. Robotics “generalist” models (Physical Intelligence Pi 0.7) can follow step-by-step instructions for new tasks.
Notable examples
Factory $150M Series A (Matten Grinberg; customers Morgan Stanley, EY, Palo Alto Networks). Claude Design exports to Canva/PPTX/PDF; reads code/design files for consistent design systems. Token churn stats: AI users 9.4x churn; 861% increase under high adoption. Pi 0.7 operates an air fryer from brief exposure + verbal steps; matches specialized robots on coffee, laundry, box assembly. OpenAI Codex desktop: background Mac control, parallel agents, in-app browser, 111 plugins, memory, image generation, enterprise pay-as-you-go.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOAI Coding Startup: Factory's $150M Funding
1:29 to 2:56
Discussion on the AI coding startup Factory, its funding, and enterprise focus.
“Okay, the first thing I want to talk about is a company called Factory.”
Anthropic's New Tool: Claude Design
2:57 to 5:36
Overview of Anthropic's new tool Claude Design and its capabilities for creating mockups.
“Of course, Claude Code is trying to get in there as well.”
Token Maxing and Productivity Metrics
5:37 to 7:30
Exploration of the concept of token maxing and its implications on coding productivity.
“So tools like Claude code cursor codex, they're all generating way more accepted code at the on the first pass 80 to 90 % of initial acceptance rates.”
Physical Intelligence and Generalization
7:31 to 10:56
Discussion on Physical Intelligence's new model Pi 0.7 and its generalization abilities.
“I mean, a normal developer writes code and read, you know, and works on it and optimizes it and fixes it.”
OpenAI Codex: New Features Overview
10:57 to 13:56
Detailed look at the new features and improvements in OpenAI's Codex application.
“And I think OpenAI sees that and they really want to make a big push to win people back or to have people try OpenAI Codex for the first time.”
Integration Challenges with AI Tools
14:04 to 14:24
Explore the difficulties of integrating various AI tools and platforms.
“you know, my Google Calendar and Chrome like synced up.”
Transcript
Automatic transcript. May contain errors.0:28Welcome to the podcast. massively beefed up codecs for desktop control, memory, and in-app browser, and over a hundred plugin integrations. Basically, this is them swinging directly at Anthropics Cloud Code and Cloud Cowork, and I think it matters a lot for where editing agents is going to be going in the future, so let's get into it. Before we do, I wanted to mention AI Box. The thing I keep hearing from people is that they're paying for ChatGPT, Cloud, Gemini, Perplexity, even MidJourney, and by the time you get all of that added up, you're likely$70 or$80 a month across a bunch of different logins.
1:00AI box gives you access to over 80 different AI models in one interface. All of this is just$8.99 a month. If you want to get access to it, there's a link in the description to AI box.ai. And in addition, we have something called the AI box builder, where you can essentially link together multiple AI models, we build up the entire workflow for you and you vibe by build tools without needing to know any code at all. I'm not a developer, I built this for other people that are not developers. So if you want to check it out, there's a link in the description to AIbox.ai. Okay, the first thing I want to talk about is a company called Factory.
1:33So this is an AI coding startup. They're focused specifically on enterprise engineering teams. They just closed $150 million Series A at a$1.5 billion valuation. Koshla Ventures, Led The Round, Sequoia, Insight Partners, and Blackstone all were participating in this. The founder is named Matten Grinberg. He was a physics PhD student at Berkeley. He basically cold emailed Sequoia partner Sean Maguire in 2023. And they apparently were good friends. They bonded over physics research. Maguire convinced him to drop out and Sequoia seeded the company. Their customer list already included Morgan Stanley, Ernst and Young and Palo Alto Networks.
2:11So obviously, this is a very enterprise focused, you know, they're not targeting individual developers in any way. They're focused on the enterprise. Grinberg's pitch is why factory I think is kind of differentiated their model flexibility. They basically can let you switch between Claude, DeepSeek, whatever makes sense. Although honestly, Cursor does that too, as do I think most of the serious players at this point. What I think this does tell us is that even with Anthropic and Open AI and Cursor already in the market, enterprise AI coding still has room for some category specific players. Morgan Stanley isn't going to let some, you know, random developer tool run inside their network unless it's built with their compliance and security posture in mind.
2:48And I think that's basically the gap that factory is filling. And I think$1.5 billion in their current valuation says that VCs are believing there is a real gap here. Of course, Claude Code is trying to get in there as well. And you can look at things like Cognizant just getting, you know, all of their employees, 350 ,000 of them onto Claude and all the Anthropic tools. So I think there's probably competition from a lot of players, but it's interesting that they're carving out a niche there. Okay, the next thing I want to talk about is a brand new tool from Anthropic called Claude Design. design it's a research preview right now it's available to pro max teams and enterprise subscribers and it's powered by a claude opus 4.7 the model that just came out a day or two ago so this is what anthropic just shipped and basically you can describe what you want a pitch deck a one pager a landing page prototype and claude generates a first draft you've probably seen it kind of make web pages before um so it's interesting because you actually can use claude design to kind of come up with the mockups ahead of time.
3:48And then you can refine it by either directly editing it or just talking to it. I actually appreciate both of these options. As I've used a lot of tools, like I mean, I don't want to throw too much shade at level because I know they do have some direct editing features actually never worked super good for me in the past. Maybe they're better now. But in the past, lovable would have something where you could, you know, describe the website or whatever you're trying to build, it would generate the design. And then you're supposed to be able to click on it and edit directly. It never worked. And what sometimes I do the chat after I do that, it would like undo them.
4:17It just kind of bad. So I think Claude has cracked this a little bit better, maybe levels there as well. But you're able to export as PDFs, URLs, PPTX files, and you can send the outputs straight to Canva. So Canva has a big integration with them. And you can keep all of your collaborations there. It can also read your company's code base and design files to apply a consistent design system across all of your outputs, which is actually I think the more interesting piece if you kind of look at this technically. Anthropic is positioning this as kind of complementary to Canva. They're not, they're saying like, look, we're not going to compete with them.
4:48It's a compliment. The target audience is specifically people who aren't designers. So founders or PM startup operators who need to make something look presentable really fast. What I think about this is that Anthropic is continuing to move up the stack. Earlier this year, they launched Cloud Cowork, then agentic plugins for specific departments. And I think now this, I think where they're going with this is that they're not just trying to be an API company. They want to actually own actual workflows and surface area. It's the same play that OpenAI has been making. And I think you're going to hear more why this matters when we go into the deep dive later on in the episode.
5:19But also just shout out to Google, who's been doing this in basically every vertical. Google has Stitch, which is a very similar design tool as well. So yeah, I think we're going to see a lot of these players get more into the software itself beyond just the models, which is pretty interesting. Okay, the next thing I want to talk about is token maxing. So there's a funny story on on tech crunch recently, where they're talking about token maxing, basically, it's the pattern of companies and developers when they're bragging about how many tokens or AI coding tools burn through, as if you know, the more tokens that they're using means that they're more productive, I think the actual data that they've been crunching is very different.
6:00So tools like Claude code cursor codex, they're all generating way more accepted code at the on the first pass 80 to 90 % of initial acceptance rates. But when you look at the same code two weeks later, the effective acceptance drops to 10 to 30 % because engineers were constantly rewriting it, right? So basically, what's going on here is when they're like, look, we're writing like all of our code, 50 % of our code, a lot of companies are like 50 % of our codes written by like Claude, and it sounds amazing. And again, it's all perfect and great or, you know, high level of it. And maybe that's true.
6:31And I'm not saying that's necessarily bad. I use cloud code heavily as well with my startup. But what's interesting is I think people that pretend it's, you know, basically a marker of productivity, because what they found, get clear, I was doing this specifically, and they found that AI users have 9.4 times higher code churn than non AI users. And Pharaoh AI found code churn increased 861 % under high AI adoption. And then you also had jellyfish, which was looking at 7500 engineers and found that the teams with the biggest token budgets got two times the throughputs at 10 times the token cost.
7:06And I think basically, it's just a reality check, right? The productivity gains from AI coding are real, but they're also a fraction of what the raw output numbers suggest, right? If you're like, I can write a million, you know, lines of code a day now, and I used to only be able to write this amount. Okay, well, the you definitely are getting real productivity gains. It's not an argument. But I just think it's important to check ourselves. And, you know, the million lines of code are not all good, because if you look at it three weeks later, a big chunk of it has to be rewritten or fixed, which is fine.
7:35I mean, a normal developer writes code and read, you know, and works on it and optimizes it and fixes it. I think senior engineers, interestingly, are less accepting and then AI of AI code than juniors and probably because they know which parts are subtly wrong, right? So when there's like a code push, they're less likely to accept that code push. If you're a manager thinking about how to measure AI ROI, I think counting merged and shipped is important and not just how much is generated. Okay, let's get into physical intelligence. This is a robotics foundation model startup. It just published research on a new model called Pi 0.7.
8:14And I think this might be the most novel kind of technical story of the day. But basically, what they're claiming right now is that Pi 0.7 can perform tasks it was never specifically trained on by composing skills it learned in other contexts. So the example that they're kind of highlighting all this, they have like an air fryer and the robot had only briefly seen this air fryer in training. So, you know, wasn't trained to operate the air fryer in any way. It briefly seen it in training. And I think it only seen like two short clips of an air fryer. Then they gave it a step by step verbal instruction and it figure out how to operate it.
8:51And in in like some broader testing, the generalist model actually matched specialized models on jobs like making coffee, folding laundry and assembling boxes. Researchers researchers said that the generalization ability was really surprising to them, right? So basically, you train a special, you train the robot specifically to fold laundry, and yeah, it does good. And then they have a new model where it's just kind of a generalist at everything. It's not trained to fold laundry, but you explain step by step how to fold laundry, and it does it just as good as the robot or almost just as good as the robot that was specifically trained on this.
9:25That is fascinating. And I think especially when you look at, you know, I would say, quote unquote, general models like OpenAI or Anthropic that they're building that can do a lot of different things generally good, it's kind of good news for them because, you know, you may not have to have models specifically trained on just a specific task when it talks about, you know, kind of like physical robotics and stuff. So on the business side, physical intelligence has already raised over a billion dollars. They were last valued at 5.6 billion. They're reportedly in talks to nearly double that to 11 billion.
9:56Their co-founder Lacey or Lachie Groom has a track record of backing Figma, Notion and Ramp. So I think obviously this is why VC dollars are going to Lachie. I think something that is interesting, a caveat that I'll put on this if I'm trying to be honest, like basically Pi 0.7 still can't handle a lot of multi-step tasks, right? And it's not doing this autonomously without any coaching. I think the robotics field doesn't really have a lot of clean benchmarks like LLMs do. You know, we have like humanities last exam and like all these different benchmarks that we give AI models on, you know, like engineering and math and other areas.
10:32And we can tell exactly how good they are at those tasks. There's not a lot of that with robotics. I'm sure there's going to be more as we get more into it. So basically for robotics, so you kind of have to just trust whatever their demo is. But I think if this kind of generalized behavior is going to be something that we're looking at, it's pretty significantly stepping us towards robots that actually work in really messy real world environments. So this is something I'll be closely closely watching over the next six months or so. Okay, OpenAI has just released a new a whole bunch of new features to their desktop app, which is called Codex, something I've tried in the past, but have opted for Anthropics Cloud Code and Cloud Cowork in recent, you know, weeks in the last month.
11:14And I think OpenAI sees that and they really want to make a big push to win people back or to have people try OpenAI Codex for the first time. So huge upgrade to Codex. I'll walk through a couple of things that are new because I think OpenAI is basically swinging directly at Anthropics Cloud Code, which honestly has been really crushing it, right? So this is what they added. First, Codex can now run in the background on your Mac, which is phenomenal, right? It can open up applications. It can click around. It can type into your desktop while you keep working on something else. this is actually something I like Claude sort of does this but I'm going to be honest even with Claude co-work a lot of times if I have like an automated task running like I've got a bunch of things that I'm just like you know every day at 9am do this every day at noon do this and it's like grabs like analytics or grabs data or goes and gets me a report on something so actually the thing that I love using it for is if there's no API for a service I'll just have it go log into the account and go grab the data I need and bring it back to me hopefully those companies offer APIs in the future but for now that's what I do in any case it is annoying with Cloud Cowork that lots of times when those automated tasks start happening, all of a sudden this Chrome browser pops up on my screen if it can't do it in the background.
12:23And all of a sudden it's clicking on things right in front of me and I'm swatting flies, trying to get this thing away while I keep working on something different. Should I have its own computer? Possibly. But a lot of the times I just have it running on the side. In any case, OpenAI is trying to combat that and have it work on things in the background. So it's not just writing code in an editor. It's actually operating your entire machine. So this is what I'm excited about. Computer use from OpenAI, which they sort of have done for a long time. They were the OGs way before Anthropic was shipping things in this, but they really just felt stale and bad.
12:53I've tried a lot in the past. And trust me, if OpenAI agents back in the day were like six months, a year ago, were as good as what Anthropic's doing now, I'd be, you know, shouting them from the rooftops. But it seems like now they're making a comeback. So they can also run multiple agents in parallel without interfering in your desktop, which means that you can have one fixing a bug and you can also have one running tests, one writing docs all at the same time. They also have a new in-app browser so it can hit web applications directly. They have 111 plugin integrations, so CodeRabbit, GitHub or GitLab issues.
13:29They have a bunch of new exciting things with their memory feature so it can remember previous sessions. they have an image generation that is now inside of codex which to be fair claude does not have any sort of image generation they also rolled out a pay-as-you-go pricing specifically for enterprise and business customers so well i think anthropic is definitely ahead right now i would say that the uh plugin ecosystem is probably part of one of the most underrated pieces of this entire announcement because they they have like 111 different plugins at launch they're going to be adding more. And with Cloud Cowork, I mean, it's awesome, but I have like maybe like four things, you know, my Google Calendar and Chrome like synced up.
14:07So a bunch of like Google tools and GitHub. But like beyond that, there's so many different tools I use that don't integrate very well with it. It's got to go and use my Chrome browser to access them. So anyways, I think a lot of these integrations that OpenAI is pulling in are going to be very useful. All right, that is this show. If you're getting value from these episodes, please drop a comment over on Apple Podcasts or leave a couple stars over on Spotify. You hit the About tab on Spotify to drop a review. Basically, the reviews help the show reach way more people. It boosts it in the algorithm.
14:37It helps it out a ton. If you haven't done it already, I would be eternally grateful. Also, if you want to consolidate the AI subscriptions you're already paying for, go check out AIbox.ai. There is a link in the description, 80 plus models, plain English automations,$8.99 a month. I'll catch you guys all in the next episode.
From the publisher
In this episode, we assess the evolution of AI with a specific focus on Codex and Claude. Identify trends that are shaping the future.
See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

