20VC: Codex vs Claude Code vs Cursor: Who Wins, Who Loses | Will All Coding Be Automated - Do We Need PMs | The Real Bottleneck to AGI | The Three Phases of Agents and What You Need to Know with Alex Embiricos, Head of Codex at OpenAI

21 Feb 2026 · 1 h 8 min · 31 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: The Twenty Minute VC (20VC)

Episode Title

20VC: Codex vs Claude Code vs Cursor: Who Wins, Who Loses | Will All Coding Be Automated - Do We Need PMs | The Real Bottleneck to AGI | The Three Phases of Agents and What You Need to Know with Alex Embiricos, Head of Codex at OpenAI

Guest

Alexander Embiricos, Head of Codex at OpenAI

Episode Overview

In this episode, Harry Stebbings interviews Alexander Embiricos regarding the future of coding and AI, particularly focusing on the Codex platform by OpenAI. The discussion covers the automation of coding, the evolving role of product managers, and the dynamics of AI coding tools in the marketplace.

---

Key Topics Discussed

  1. Will Coding Be Automated?
  2. Automation Impact:
  3. Coding is one of the first domains where AI (especially LLMs) can greatly enhance efficiency and productivity.
  4. Automation doesn't eliminate the need for engineers; instead, it's likely to increase the demand for them as coding becomes easier.
  1. Role of Product Managers (PMs)
  2. Definition and Necessity:
  3. PM roles can be undefined, and their necessity fluctuates based on team size and project needs.
  4. In smaller teams, strong engineering leads or designers can fulfill PM roles effectively.
  1. Bottlenecks to AGI
  2. Human Interaction:
  3. Human typing speed and the need for validation are significant bottlenecks to achieving AGI.
  4. There’s a desire for AI interactions to become more intuitive and require less effort from users.
  1. Phases of Agent Development
  2. Three Phases:
  3. Phase 1: From coding to computer use.
  4. Phase 2: Development of a productized workflow.
  5. Phase 3: Establishing secure and efficient agentic browsing in enterprises.
  1. Enterprise Implementations
  2. Challenges:
  3. Security, permissions, and compliance are critical hurdles when implementing AI tools in enterprise settings.
  4. The aim is to build tools that empower users while managing security considerations.
  1. Market Dynamics
  2. Coding Tools Comparison:
  3. Discussion on various AI coding tools like Codex, Cursor, and Claude.
  4. Key features and the importance of user experience were highlighted along with the potential for market shifts.
  1. The Future of AI Coding Tools
  2. User Retention Strategies:
  3. The focus is on building user-friendly tools, fostering user fluency, and creating an ecosystem around coding tools that includes community and shared learning.
  1. Advice for New Engineers
  2. Building Skills:
  3. New engineers should focus on building projects and showing initiative, as real-world demonstrations of capability are more impressive than traditional resumes.
  1. Long-term Vision
  2. AI in Everyday Use:
  3. The vision includes a future where AI interacts seamlessly with users in daily tasks, making it a natural part of work and personal life.

---

Key Takeaways

  • Coding is transforming, and AI will likely create more opportunities for engineers rather than reduce them.
  • The role of PMs is evolving, becoming less critical in smaller teams.
  • Human interaction remains a bottleneck in achieving full AGI capabilities; the goal is to simplify this.
  • Building user-friendly tools and fostering an engaged community around AI tools will be essential for market success.
  • Retention in the AI tool market will be driven by creating intuitive, easy-to-use platforms that integrate well into existing workflows.

---

Conclusion

Alexander Embiricos provides a forward-looking perspective on AI in coding, emphasizing the importance of automation, the continued necessity of software engineers, and the need for user-centric designs in technology. The discussions underscore the evolving landscape of AI, the potential for increased collaboration between humans and machines, and the importance of creating tools that enhance productivity without overwhelming users.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Interview with Alex Embiricos on Coding Automation

4:07 to 9:40

Harry interviews Alex Embiricos on the evolution of coding, AI's role in automation, and industry shifts.

“I told you, I've been at a PE conference, and all I could think was, thank God I've got Alex next, because this is going to be a great one.”

Exploring Human Bottlenecks to AGI

9:40 to 14:02

Discussion on human limitations in utilizing AI effectively and the future of AI in various tasks.

“and I am too uncreative to figure out all the ways that AI can help me.”

The Role of Human Intuition in AI Automation

14:02 to 16:47

Learn how empowering employees with AI can enhance productivity and intuition.

“them credit for, I think, especially in large enterprise.”

The Importance of Speed in Developer Tools

16:48 to 17:57

Discover why speed is a critical factor for developers using Codex.

“How important is speed for developers when using Codex and in the future of AI code?”

AI Inference as the New Sales and Marketing

17:58 to 19:18

Understand the shift in business strategies with AI-driven inference.

“One of my dear friends is Jason Lemkin from Sasta, and he says that actually inference is the new sales and marketing.”

The Evolution of Code Writing with AI

19:19 to 20:38

Explore how AI has transformed the coding process and the role of developers.

“a sudden the model was like way better at running for longer, handling tasks end to end, managing its context and following instructions.”

The Future of Integrated Development Environments (IDEs)

20:39 to 21:45

Examine how the definition and functionality of IDEs are changing in the AI era.

“Or maybe you want to like collaborate on a plan, but then have AI fill it out.”

Revolutionizing Code Reviews with Codex

21:46 to 23:18

Learn about the role of AI in enhancing code review processes and maintaining quality.

“Think like architecturally, like how should this code work?”

User Retention and Stickiness in AI Tools

23:19 to 25:05

Discover the strategies to keep users engaged with Codex and similar AI tools.

“haven't tried Codex yet or didn't try it recently, sometimes the way that people see how good our models are is by asking Codex to review a different model's code.”

The Competitive Landscape of AI Development

25:06 to 28:00

Explore how competition drives innovation and learning in AI technology.

“So kind of like both ends of this are pretty neutral, vendor neutral.”
Show all 31 chapters

Defining Winning Factors in AI

28:00 to 29:10

Discusses key factors that determine success in AI models and products.

“and building our harness to be good at the new models.”

The Importance of Product Execution

29:10 to 30:00

Explores the significance of product execution in driving AI adoption.

“That's maybe the company perspective, right?”

User Engagement Metrics for AI Products

30:00 to 31:20

Examines how active user metrics are measured and their importance.

“Something that I've learned the hard way is if we go to an enterprise and we're just like, hey, we're here, feel free to use the stuff, that doesn't work.”

Chat as the Future UI for AI

31:20 to 32:30

Considers whether chat will be the primary user interface for AI interactions.

“It's like, we should probably just be a daily.”

Balancing Chat and Functional Interfaces

32:30 to 34:10

Discusses the need for combining conversational interfaces with functional UIs.

“The simple answer is yes, but actually I think there's two components here.”

Agent-to-Agent Interaction Design

34:10 to 35:50

Explores the design of agent-to-agent experiences in enterprises.

“And it kind of wrongly assumes on my behalf a consumer interaction at some point in that journey.”

Data Acquisition Strategies for AI

35:50 to 36:50

Looks at strategies for acquiring data for building AI models.

“But from another company, you ask him, how do you think about a coding data moat?”

Consumer Engagement with Codex

36:50 to 38:20

Examines how Codex engages with consumers and its future competitive landscape.

“about the data that doesn't exist, so to speak.”

Responding to User Feedback and Market Changes

38:20 to 39:40

Discusses how Codex adapts to user feedback and market shifts.

“even on free ChatGPT plans or on the Go ChatGPT plan.”

Future Directions for Codex

39:40 to 41:40

Highlights future strategies for improving the Codex product and user experience.

“But the shift that we feel last week is, you know, we, we felt like we had the most intelligent model that was cemented with five free codecs.”

Evaluating AI Models: Metrics and Vibes

41:40 to 42:00

Discusses the importance of both metrics and user experience in evaluating AI models.

“that you trust to own an entire microsystem or internal tool or whatever, and can do the full iterative loop, including feedback from users, without having to go through human review.”

The Importance of Vibes in Model Evaluation

42:00 to 44:05

Learn how the evaluation of AI models is influenced by user experience and relationships.

“Probably, this is an annoying answer for you.”

Cursor vs. Cloud Code: Market Dynamics

44:05 to 46:11

Explore the competitive landscape between Cursor and Cloud Code, including revenue predictions and user experiences.

“Do you think it was the right strategic decision to start building their own models?”

The Future of Agents in Coding and Beyond

46:11 to 47:58

Discuss the evolution of coding agents and their potential to handle various tasks in the future.

“So in that world, I don't think you want like 12 agents at the company and you have to like go, your employees have to go figure out the right one to talk to because then they won't achieve fluency.”

Investing in Durable SaaS Companies

47:58 to 50:16

Understand the factors that influence the long-term viability of SaaS companies in a changing market.

“you, Anthropic, others, are going to come for our lunch, so to speak.”

Advice for Aspiring Engineers in AI

50:16 to 53:06

Gain insights on how to navigate a career in AI as a new engineer and the skills that will be valuable.

“Because it's kind of relatively easier to build good product.”

Learning from Claude Code and Dropbox

53:06 to 55:15

Analyze what can be learned from Claude Code's approaches and Dropbox's successes and challenges.

“and you can ask it to plan out changes that would otherwise take you like days to research maybe.”

Reinvigorating Growth at Dropbox

55:15 to 57:20

Explore strategies for enhancing productivity and growth at Dropbox and the role of desktop software.

“The alumni from Dropbox is incredible and really amazing to see the talent that's come out of Dropbox.”

Software Margins and Inference Costs

57:20 to 58:50

Understand the balance between margins and the cost of AI software usage.

“I've been brought up in a world where margin matters.”

Quickfire Round Insights

58:50 to 1:02:00

Gain insights from rapid-fire questions on AI, product decisions, and future visions.

“Which less unknown capacitor do you respect most and why?”

Agent Guardrails and Future Expectations

1:02:00 to 1:04:08

Explore the importance of guardrails for AI agents and future tech expectations.

“And then maybe you're not actually handholding this very painful CI and deploy process, but you're just having agents do things.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:28Welcome to 20 Product with me, Harry Stebbings. Let me know what you think, harry at 20vc.com. But before we dive into the show today, the early story of Atlassian is probably very similar to your own. Atlassian knows firsthand the challenges that startups face every day and that the right tools are essential to go from MVP to IPO. That's why Atlassian for Startups gives eligible companies up to 50 seats free on the premium edition for products like Jira, Confluence, Loom, Jira Product Discovery, Compass, and Bitbucket, so your team can use the best-in-class tools to plan, track, and collaborate on work, whatever that work may be.

1:05Many of today's most successful startups like Cloudflare, Canva, and Rivian relied on Atlassian for their growth trajectory, and Atlassian wants to give that same opportunity to the next generation of builders and investors. We know how important it is to focus on building the right things early. Whether you're in the sticky note stage or well on your journey, teams at any stage can work smarter together. It's never too early to start with Atlassian. Head on over to atlassian.com forward slash startups forward slash Harry for more details and eligibility. After Atlassian helps your team build and ship great products, Intercom helps you support the customers using them.

1:41If you're looking for a way to transform your customer service, let me introduce you to Finn, baby. Finn is the number one AI agent for customer service, resolving up to 93 % of customer queries automatically. There is no other agent that can do that. Not 93 % of customer queries, okay? No other agent can do that. So why choose Finn? Finn is the best performing AI agent for CS. Finn doesn't just answer questions. It takes actions. It automates the most complex customer queries like refunds, transaction disputes, technical troubleshooting with speed and reliability. I wish my team was speedy and reliable.

2:18Beats every competitor in every head-to-head bake-off. completely configurable and code optional setup. My word, I mean, the benefits just go on and on. It's easy and efficient implementation. It works on any help desk with no tedious migration needs. It's trusted by over 6 ,000 customer service leaders, including top AI companies like Anthropic, Lovable, Synthesia, Clay, Vanta. So if you're ready to transform your customer service team, scale your support and give team members time to focus on the really high level strategic work, Learn more about FIN at fin.ai forward slash 20VC. While FIN scales your support without losing speed, Reforge shows you how to translate that scale into durable product-led growth.

3:01Everyone's shipping faster than ever. Cursor, claw code, codex. AI is making code and writing code faster than ever. But here's the problem. Speed means nothing if nobody uses what you ship. That's where Reforge comes in. Reforge is building the product discovery engine that sits upstream of your coding agents. Not another prototyping tool, research repo, or AI interviewer, but a product that will ingest your customer data, generate variations of product solutions, validate the solutions before code is written, and hand off winning directions to your team. Reforge kills product debt before it starts, because every unused feature you ship isn't just wasted engineering time.

3:44It's a maintenance burden, complexity tax, and surface area that you cannot shrink. Used by product teams at companies like Toast, Vimeo, Klaviyo, and many more, Reforge helps teams ship more features that actually get used. Try Reforge at reforge.com forward slash build and use the code 20VC, that's 20VC, for one month free of pro. You have now arrived at your destination. Alex, I'm so excited for this, dude. I told you, I've been at a PE conference, and all I could think was, thank God I've got Alex next, because this is going to be a great one. So thank you so much for joining me, man. So excited to be here.

4:21Thank you. Now, this is a weird first start, but roll with it. You'll understand my British intricacies. I'm fascinated by people's motivations. Are you motivated more by the fear of losing, or the thrill and excitement of winning? I'm a maximalist. I'm definitely much more motivated by the idea of winning and the fear of losing. But I'll admit to you something, when I was running a startup before joining OpenAI, and one of my darkest moments, and there were many dark moments while I was running the startup, was recognizing that I'd spent the fast few months trying to avoid losing. All of a sudden, I was like, oh my god, that is why I'm so unhappy.

4:56And that's probably why the startup isn't going well. You know, I basically every now and then, I have to catch myself and like flip back into this idea of winning. But really, what motivates me even more than that is I think I just love building things and building things for people. And man, I am so excited for this year because many amazing things that don't exist yet are going to be built and given to a lot of people. I'm diving right in. Elon said that coding is one of the first professions to be largely automated. Do you agree, given your position and what you see day to day? For sure, I would agree that coding is one of the first domains where LLMs are really good.

5:28But what does it mean for coding to be automated? It's like kind of a heavy statement, right? For example, now that we no longer write assembly, like when that change happened and we moved to higher level languages, did we say coding is automated? Not really, right? We were just able to write much more code. And then as a result, actually, there was much more demand for code and there were many more software engineers required. But yeah, part of what they used to do is automated. In the same way that like, do you know the origin of the word computer? No. I might pronounce the location wrong, but I think it was at Bletchley Park.

5:56There were all these machines for decoding German Enigma. And there were humans who would punch out punch cards and put them into the machine and do a bunch of tabulated math. I'm probably butchering this, but basically there was an intensely manual part of work. And even like the first spreadsheet software was kind of loosely based off this idea that you would have an office full of desks arranged in a grid and people doing tabulations and then passing their sheets to the next person. And so all these things, like those specific tasks have become automated. But every time that's happened, there's been an explosion in demand for the output.

6:27And so you need many more people actually to do that kind of work, even if the specific task is changed. So you think we'll have more engineers in five years, not less? Yeah. And, you know, sometimes we change what terms mean, right? Like the term computer now refers to something else, but now we have the term software engineer. And so I definitely think we'll have many more builders. And, you know, something interesting that I'm observing now is like, there's this compression of the talent stack. You know, you still need software engineers today. You still need designers. I'm a PM. Do you need PMs?

6:53You know, you can have some fun jokes about that. I don't think you need them. But maybe when you say engineer, you might be thinking of someone who's like much more full stack than has been true before. Like even if you go back a few years, you had many more places where there was like the backend engineer and the frontend engineer. Whereas like now, at least if I think about the Codex team, like that's much less the case and things are much more full stack. Right. And so I think this talent stack will compress, but we'll still have people building. Why do you think we don't need PMs in this world?

7:20You dangled the carrot. Yeah, it's my fun joke. I think, well, first of all, I think it's incredibly hard to define what a PM is, what a product manager is. I kind of think of the role as like actually explicitly undefined and your goal is just to adapt to whatever the team or business needs. Often, if you have a bunch of people like trying to build as quickly as possible, then what a product manager can do is spend time like taking a few steps back and trying to look around corners and figure out what to do. You know, collaborate with the folks and go to market and maybe be the team's like greatest cheerleader and quality raiser.

7:52But like all of those things I just described, which are maybe my current role, could be done by a really strong eng lead or a designer who thinks a lot about product. And so I think it's like often useful to have product managers, but you probably don't want many of them until the team is really large. I was stalking the shit out of you for the last few days, which was a very fun expedition into your writing, into your tweets, into your prior interviews. And you said that human typing speed and validation work is the key bottleneck to AGI, not model, compute or architecture. And it kind of left there.

8:23And I was like, help me understand why human typing speed and validation work is the key bottleneck and what you really meant by that. For sure. OK, that's a fun one. I think there are multiple bottlenecks, but that's maybe the most sort of clickbaity one. So if you don't mind, I will do this slightly socratically. Like how many times would you say you use AI today? 30 plus times a day. OK, cool. How many times do you think, assuming it was like zero energy expenditure from you, how many times do you think AI could help you per day? I mean, in everything. I think we'll have inference running 24 hours a day across every single thing.

8:58Exactly. And like, I hear things now from engineers like at OpenAI and also outside who are, they're telling me like, you know, I constantly have Codex running. I never close my laptop. And if it's not running while I'm in a meeting, I'm like wasting my time. I need to make sure Codex always has work for me that it's doing. and that's like super cool and super exciting, but that's a lot of work, right? To like manage these agents and make sure they're always working. And going back to the 30 times per day thing, yeah, like when we look at how often codex users are using codex, it's like kind of this like tens of times kind of range.

9:26And I think AI should be helping us tens of thousands of times per day, you know, compute budget permitting, and we'll get there over time. But the problem is like, at least if I think of myself, like I work on this stuff, I know I should be using AI for everything, but I'm too lazy to like type out that many prompts and I am too uncreative to figure out all the ways that AI can help me. And so I end up kind of at a similar number as you. You know, I still am at the point where when I use AI to do something cool, like prep for this conversation with you, I'm like kind of proud of myself. I'm like, oh, cool.

9:55I managed to use AI in this new way. That's fine for people like you and me who are like really interested in this topic, right? But I don't think most people we should expect in order to benefit from AGI should need to like put so much effort into how to use this tool. It should just be effortless for them. I think the world we want to get to is one where to use AI, you don't really need to like figure out the right way to prompt. It's just super easy for you. And you don't even need to recognize that AI could help you. It's just like knows you, connected to your context and chimes in helpfully.

10:26That's where I think like Claude has done well in terms of the packaging they've done like Claude for legal, Claude for Excel, where you can implement it and have a DCF model. I'm not into models, but like better than one could do before. Do you think it is your job then to productize the prompts and the human actions to remove that bottleneck? Yeah, totally. So I think that it is our job to make sure that we have the models with amazing capabilities. And then eventually to get to a world where this is like highly productized. And so you just have this like magic text box or audio input or whatever, or you can just add AI to your like group chat.

11:02And it just starts to help. But I think there's quite an interesting in-between stage. And I think that that is actually where the most value lies right now. So here's what I mean. You could try to productize like a specific feature of AI for a specific market. And, you know, many companies are doing this, but I think it's a little bit hard to know what exactly will work, what is the right form factor. And someone was on your podcast earlier, and they said something that I thought was quite interesting about how you cannot adopt AI at enterprise without FDEs. Yeah, it was Matt Fitzpatrick from Invisible AI.

11:33Yeah. So even though I am literally hiring FDs, and if you're an FDE, please apply for a job with me, I actually disagree with that entirely. So what I think we need to do is build tools for people. Like you can use FDs, as Fitzpatrick said on the podcast, like to automate workflows, right? But then you're limited by like what you from your top down perspective can do and what you from your FDE staffing can staff to be built, right? But for me, the most exciting future with AI is one where everyone just feels like a superhuman or a god just like empowered by AI. And for that, we need tools that are for people, for individual users, and that everyone feels fluent with.

12:11I think the phase that's most interesting that we're at now is building for the kind of people who are interested in figuring out how to use AI. So what we need to ship, and I think this was like the genius of like when Cloud Code first shipped, what they really got right, was they had this tool that was super easy to use in whatever context you want, just in your terminal. And people started experimenting with where to use it. And so I think as we think about AI being used outside of coding work, one of the most important things we can do is not overly build it like, okay, this is AI capabilities, but only specifically for finance, only for specifically for this workflow, but actually build a much more open-ended tool that someone can just use for any given task creatively.

12:48But does that not put the onus or the effort back on the user back to the point of your bottleneck of human action and lack of activity on them if you don't define the task, you put the responsibility on them for defining the task, which humans lack the ability or inclination to do. Yeah. So that's why I think it's the bottleneck. So basically, here are the three phases in my mind. First, let's have agents work really well for software engineering and coding because LLMs happen to be good at that. Next, let's realize that for an agent to be useful more generally, using a computer is super valuable.

13:19And also, we'll realize that all agents are actually coding agents because coding is just the best way for an agent to use a computer. So let's take that same super flexible idea, but make it available to anyone who's excited to explore and tinker. And we're already seeing people start to do this with like the Codex app. Like Codex app is built for builders, but we're seeing builders use it for all sorts of non-coding tasks. Then finally, once we see what's working, let's build that productization that you were talking about, where you have highly specific features that just work immediately out of the box for people.

13:49And I think we're going to speed run this entire like one, two, three journey in the next months. My challenge with what you said about kind of FDs and implementation within enterprise is data security, sensitivity, permissioning, access provisions is really freaking hard. And people are much less intelligent and confident than we give them credit for, I think, especially in large enterprise. Sorry. And I think you actually need an FD to go in and custom fit a lot of the different horizontal solutions to make it work. Am I wrong? I think you're right. If you're trying to go like all the way from zero to one and you have this like, and I said, I don't mean grand negatively here, but if you have like a grand vision for some like ultimate workflow automation system, then yeah, you're going to have to clear through all of these security hurdles, all these like compliance hurdles that are really real, right?

14:34Build connections to all these data systems and like systems of record and action. Yeah. So you're going to need NFDA to do that. What I've seen is that when we do these things top down, we end up like massively under leveraging the potential of AI in helping that company. Whereas you can maybe do that in parallel, right? But if you can just give AI to the people actually doing the work, they can start to get a mental model for how AI can help. And then they can start pulling AI into their workflows at the same time. Here's just an analogy or something here is like, imagine if you work in a customer support role, and AI is being brought into your role and starting to automate meaningful chunks of your work, but you've never heard of Chachapiti, nor are you allowed to use it.

15:18So in that scenario, you have like no intuition for what this thing is. Whereas in a world where actually you've been using Chachapiti for work at the same time as like parts of your work are getting automated by an LLM, you have much more intuition for how this works. And you know, I would argue you feel much more empowered about this idea that it's being accelerated and you have some degree of control to steer like where these automations are built, as opposed to like, it's like this complete like ex machina kind of thing that is quite disempowering. So bringing this back, I think there is a way to do this because the data control issues you mentioned are real.

15:49But at the end of the day, every tool, every feature, every workflow is for a human who is somewhere, an employee somewhere. And that employee is accessing that tooling via their browser or via their file system, like at the end of the day. And so at the end of the day, everything comes to an interface that an agent running locally on your computer can work with. And I think it's quite unusual, like in OpenAI, we're building a browser, Atlas. And you might wonder why. And there are many reasons why. But I think one of the key reasons is that by building a browser and by controlling it tightly end-to-end, we can build safe, agentic browsing for enterprise that is a way to access things agentically that are otherwise not yet built out by FDs.

16:30There are so many questions that I have to ask you. I want to go back before I lose Thread. You mentioned about engineers not closing their laptops because they don't actually want to lose productivity and time with building with codecs. You partnered with Cerebrus, and Cerebrus is the fastest provider, obviously, of inference out there. Amazing win, I think, for both, bluntly. How important is speed for developers when using Codex and in the future of AI code? I mean, these simple answers, it's super important. And so is it like an inference monopoly? Like, you have it now and competitors don't.

17:05This is just my opinion, but I don't think we're going to end up in this kind of monopolistic world. I think there's so much competitive pressure that there'll be like multiple answers to this. But I will say that we have like news coming out about that partnership soon. And I'm very excited for these kinds of things to ship. It's going to be awesome. But even so, like, you know, with GPT 5.3 codex, that model is like significantly more efficient than prior models. And so in the feedback we've heard is that people actually feel like now this is like a very competitively fast model than before.

17:34So there's a lot of things you can do just in terms of the model. There are also things you can do like improving how you do inference. So we recently rolled out a change where in the API, those models are served 40 % faster, and in codecs, they're served 25 % faster. So I think speed matters a lot, and we're kind of approaching it from all angles, like both the hardware, how you do inference, and the model level. You mentioned earlier about putting it in the hands of users, and we talked about inference there. One of my dear friends is Jason Lemkin from Sasta, and he says that actually inference is the new sales and marketing.

18:06Instead of sales and marketing teams, you're paying for inference so users can onboard quickly, easily, see value, and you will actually see the removal of sales and marketing teams. It's kind of like next gen of PLG. I don't know. I think I struggle with that. I think, you know, fundamentally in this new world where anyone can build and it is increasingly easy to build things, what is hard, right? I think having a good relationship with a customer and knowing what they need is as hard as ever, maybe even harder as it's just like there's just more stuff in the market to choose from. You know, the other things that are harder, like building the right thing, having a really high quality thing.

18:40But going back to the sales and marketing thing, like I don't think that goes away because I think that's as like I said, I think that's just gotten harder as the markets. Any given market gets more competitive with more software out there. How much of internal code for you today is produced by Codex? I remember like Claude for work, Boris said was like 100 percent or nearly 100 percent. How much is internal Codex used? So I'll speak for myself and then for the team. I would say like most people that I know are basically not opening editors anymore. And this was a step function change that happened in, it's been happening gradually, but I'd say the key external market touchpoint for this was like GPT 5.2 codecs, where all of a sudden the model was like way better at running for longer, handling tasks end to end, managing its context and following instructions.

19:27And so we kind of saw this inflection point. And that's actually part of why we built the app. So I think before GPT 5.2 Codex, the kinds of AI features we were using to write code were like tab completion or maybe you were pair programming with the model. And in my mind, you still need it to be at your laptop with your hands on the keyboard-ish. And it might go off and do a little bit of work, but you kind of still need to be there and drive. It's just like handling these small things for you. And then at the time of GPT 5.2 Codex in December, we kind of switched to like, actually, I'm just going to fully delegate this task.

19:59It's like, you know, I'm going to do a plan with it, make sure we like the spec that it's going to do. And then I'm just going to go let it cook. And this is quite a different way of working. So it's like it's changing like literally as we speak. And so part of why we built this Codex app that we released last week is because we wanted to build like a form factor or user experience where it felt like very ergonomic to be delegating instead of pairing with an agent. And so like delegating to multiple agents at once. And so even at OpenAI, this is changing massively. I don't have a percentage stat for you, but I would say like the vast majority of code is written by AI.

20:29And I would say that now probably like most people are not even like opening IDEs. Maybe if they are opening IDEs to like, maybe you want to own the interface, right? So you'll like help flesh out like the interface between like two modules and then like AI fills it out. Or maybe you want to like collaborate on a plan, but then have AI fill it out. The code itself is not being written by humans anymore. Will we have IDEs as a part of the stack in 24 months time? Depends how you define IDEs. So the formal definition, right? Integrated development environment. I mean, that phrase is so squishy that like literally anything could be an IDE, right?

21:01So I don't think that's very useful. If that's the answer, then yes, you could even argue the Codex app is an IDE. I don't think it is. Like for me, I think of an IDE as like a really powerful editor. And we explicitly didn't build editing into the Codex app because we wanted it to be really clear how you're meant to use it. So, you know, it has a lot of affordances for managing multiple agents, for delegating, for reviewing changes. It has really prominent skills, which are an open standard that are really useful for doing non-coding work. Stuff like, you know, triaging tasks or monitoring deploys or something, but it doesn't have text editing.

21:32If we assume a large percentage is done by codecs in terms of the code produced, how do you do coding reviews and is AI responsible for internal coding reviews? There are a few things here. First off, the spec for what you want to do or the plan becomes more important than ever. Think like architecturally, like how should this code work? You know, we recently shipped like a very prominent plan mode that works a little differently than others where you have the agent go off and like propose how it's going to do something. It's like quite a long plan. And then it asks you questions about if you agree on how it wants to do it or if you want to have input.

22:04And this is very similar to like if you had a new hire who was new to your code base, you know, they had to present a sort of request for comments to the rest of the team before they started doing the work. So even though that's not formally code review, I would say review of the plan is actually something that's becoming more important because we're entering more of this like delegation phase of working with agents. So that's an underrated thing. Then, okay, there's actual code review. I think a problem that I hear a lot of people talking about, especially in the open source world, is like a lot of AI slop.

22:32Like people will just be submitting PRs to these open source repos and they're trash. And like maybe the user hasn't even, the person submitting the PR hasn't even tested them or definitely hasn't reviewed the code. I think this is a problem. And so a common practice with Codex is to have Codex review its own PR or its own change. And Codex is actually incredibly good at this. We've explicitly trained the model to be good at code review. And that included things like making sure it's really good at creating high signal feedback. So it'll basically have few false positives of criticism, which means you can really trust when it has feedback.

Read the full transcript

23:04And so not only do we encourage people on the team and elsewhere to just ask Codex to review, you can then also set it up to just automatically review. So nearly all code at OpenAI is reviewed by Codex automatically whenever you push it to a good repo. Actually, one fun thing for people who haven't tried Codex yet or didn't try it recently, sometimes the way that people see how good our models are is by asking Codex to review a different model's code. And basically, they're like, oh, shoot, I should probably just be using Codex to write my code in general. You said something really interesting there.

23:32You said for those that maybe haven't tried it yet or coming back to it, how do you think about retention with this category? I remember Tom Blomfield, who's a YC partner, tweeted months and months ago, but it stuck with me, a weird brain, about the ease of transition between different providers, whether it was cursor or ClawCode or Codex. I can't remember which one it was, to be honest. But how sticky are users and how do you think about retention? We've taken this kind of counterintuitive approach with Codex to just build it super openly. So the Codex core harness is open source and we're always trying to make it easier for people to switch.

24:06So for instance, when we first launched Codex last year, we created like created as even a heavy word. It was just we just established a convention, which is called agents.md. This is basically a file that you can put instructions for the agent in. And we didn't call it Codex.md. We just wanted it to be something that all agents can use. And pretty much every agent except Claude uses agents.md, which is awesome. And then just last week, actually, we helped push for putting skills, which are a standard for like giving the agent instructions and scripts. We push for those to be sorted in sort of a neutral named folder called agents instead of in like codex or something.

24:40And again, everyone has jumped on it except the usual suspect. I think it's really great for the developers to have a lot of choice. And we're trying to make it even easier for people to try different things. Now, that said, these coding tasks, right, where you're asking an agent to write some code, they're quite hermetic. And what I mean by this is maybe an analogy in TV would be like episodic, right? You can come in and you've got this open-ended agents file that any agent can read from. You've got these skills that any agent can use. And you can ask the agent to write some code and it produces a patch and that patch goes into Git.

25:10So kind of like both ends of this are pretty neutral, vendor neutral. So very easy to move between for now. As agents start to do work that is not writing code, but more general work, again, for software engineers or beyond for any builder, they're going to need to start interfacing with other systems. So as they start, maybe your agent is talking to Sentry, right? Or it's talking to your Google Docs or something. Then I think these agents become much stickier because actually deciding to connect an agent to that system is a sticky decision. And if you're an enterprise really trusting that the agent is going to have access to these tools, but there are really good secure guardrails and sandbox and like controls over how the agent works with these systems, I think is critically important.

25:50And that's not something that you're going to want to do multiple times. And so, you know, we've been kind of building Codex knowing that this is coming. And so we have like the most conservative sandboxing approach. Sandboxing is kind of like a set of controls, OS level controls over what the agent can do. But I'm a fan of Seven Powers, this brilliant book, which talks about kind of seven ways that businesses accrue value and sustainability. And like, you know, your stickiness or your retention is one. if we're on the same team with Codex, how do we create retentive patterns, behaviors, programs to ensure that people stay with Codex and they don't flip to cursor when there's a better model or claw code when there's a better model?

26:27Yeah. I mean, it's interesting because I think on the one hand, like we think about this, obviously we're running a business, but you know, our mission here is to like ensure that like we safely deliver the benefits of AGI to all humanity. And so something that's like unintuitive to people about like the Codex team. Alex, you actually, I know, but your job is the success of Codex. I get that. Our job is the distribution of intelligence. And so we're obviously building out Codex. And this is really unintuitive to a lot of listeners. But we put all this effort into training these models, and then we serve these models to our competitors.

26:58And from our perspective, this is so difficult for me as a venture capitalist to understand. You are aware of this. Yeah, I'm totally aware of it. It's like where OpenAI is like a really interesting and unusual place to work. But basically, because we're playing such a long game for us, if the competition gets better, we learn. It's actually helpful for us. And so we're pushing really hard at growing Codex. You learn because if they're closed and they improve, you don't learn. I don't think so. For example, there are a bunch of recent launches. Like even today, I literally just like quote tweeted a thing this morning about a launch from Warp.

27:33No particular affiliation, right? And there are a bunch of cool ideas in there about how they like framed up the way that their agent can work in the cloud at the same time as working locally. And for me, that's like inspiring. And I think I see all these things from various companies. And like one of the coolest things about the space is it's like we're all kind of inevitably reaching the same conclusions together and then building things out. And so, you know, on the Codex team, I think we have some massive advantages, right? We have the massive distribution advantage with Chat2PT. We have the massive like capability advantage of training our own models to be good in our harness and building our harness to be good at the new models.

28:05And like no one else has early access to those. And so I think we're playing to win and we have a really big advantage or a number of advantages, but we're also playing this long game where, you know, again, we serve our models to everyone, where we push for open standards so that everyone can use like all the things that we're pushing for as well. Can I ask you, what will be the defining factor of winning? And I know I'm using venture language and you're brilliant and kind of much more free and open, but what was like the defining factor of winning? Again, if I push you, is it like GTM, which is like the biggest enterprises in the world do want to work with open AI.

28:39I have many friends in your sales team. The inbound that you get from the largest brands is incredible. So GTM, because of the incredible brand, product execution and just codex being a freaking awesome product, or compute, inference, speed, actual compute advantage. Which one is the defining winner? Okay. So I think if we're going to talk about it more from an open AI perspective, obviously this way above my pay grade, but I would say it's compute advantage and having the best models. And in order to achieve that, we then need to build businesses that generate revenue. And also that something that's really interesting, we noticed with having the Codex team, which is a sort of combined team of research and product, is also by building these successful products, we create a lot of pressure to improve the model in sort of a faster way.

29:27That's maybe the company perspective, right? If we come to the product perspective, I think the single most important thing we can do is build a really good product that people want to use. And like I was saying earlier, I think we really want to build products for individuals and then allow people to become fluent in those products and then pull in automation. And I think that may be counterintuitive, but will result in way more impact than anyone purely approaching it from the enterprise workflow perspective. I think that's mostly a question of product execution. And then that works for, say, prosumer.

29:59When it comes to enterprise, the go-to-market side is really important. Something that I've learned the hard way is if we go to an enterprise and we're just like, hey, we're here, feel free to use the stuff, that doesn't work. There's actually quite a lot of education that needs to be done. And there's a lot of configuration that we need to support and education of the broader team. So that motion looks much more like coming in, pitching, meeting the head of developer experience or whatever, understanding how they want their team to operate, and then giving them tools to propagate that mechanism of operating to the rest of the team.

30:27You said the word revenue there, which is one metric to measure a business against. When you think about like your metric of success, which you sit down with Sam or Brad or whoever it is and say, hey, this is what we're optimizing for. What is the metric that uses the defining North Star for your progression? It's actually not revenues. The primary is active users. How do you measure active users? Like daily active users? Yeah. So we measure weekly active users. And it's, you know, did this person like actually do a turn in our product? You know, did they send a prompt? Is weekly active a frequent enough metric, do you think?

31:06Sounds nice. But if this is actually replacing the IDE, is daily active not better? I think daily active will be better soon. We just happen to use weekly active. It's like a standard here. And I think as we were getting started, it made sense. But I actually agree with the criticism there. It's like, we should probably just be a daily. Like, I think we need to be getting to a world where for any given task that you have, your first instinct is to ask an agent to help. It's kind of like, you know how like with Google search, it's just like, okay, anything I need to do, I just like go into this text box and I can get navigated to the right location.

31:36Then you had ChatGPT. It's like for any information I need, I can go into this text box, type it out and get information that helps me. And I think the next phase that we'll see this year is like for any task I need to do, as opposed to just get information, I go to this text box or this input and something happens that helps me, even if it's not the full task, even if it's only a small part You said about kind of chat that, again, I jump around, sorry, my brain. My mother has to walk with me around London and she like deals with this manic, episodic brain. But you said about chat and the interface there.

32:06I'm really fascinated by this because it is a seemingly incredibly efficient input function for busy humans. But I spoke to Anish Akaya, who's a GP at Andreessen, and he came out the other day. And he's like, no, no, no, this was created by Sam and Elon and it works for very efficient people. but most of the planet want browser-based discovery interactions, UIs. Do you think that chat will be the enduring UI in the next wave of AI interaction with humanity? The simple answer is yes, but actually I think there's two components here. Like if we just imagine the future, like just like let's think of some sci-fi movie, right?

32:40Like what does AI look like? I believe that sci-fi is a really good predictor of what the future should look like. And usually it's pretty simple because it's a story. And I think simple is usually right, it's going to be some just like entity that I can talk to however I want about whatever I want, right? I felt like I shouldn't have to navigate to a place where I work with like my coding AI. And then I have this like different place for my like sales AI. And I have to like be like, hey, I'm now talking to sales thing and like do that. It's just like, I just going to talk to a thing and it's just going to help.

33:08So I think what we're going to have is that we'll have chat or voice, basically conversational interface will be sort of the pillar of everything that you can talk to about anything and that you can add into any group chat or whatever so it can like discover how to help you. But then if you're like a power user and you're very good at a specific thing, you probably don't wanna be disintermediated by having to talk to another person. It'd be like if you had an executive assistant, but you can only work by talking to them. That's like super annoying, right? So at some point you wanna get to the show notes and like look at them yourself and like edit them yourself, right?

33:39You wanna edit the thing yourself. So I think we'll pair chat with like functional like graphical interfaces that are bespoke to like what someone needs. So like in my case, I will probably chat to like do my, you know, podcast prep. But when it comes to like actually looking at product and code, I probably want like the Codex app that I can go into and get deep in, right? Whereas maybe if we're talking to a marketer, maybe that marketer will like chat to ask questions about the product. They're not gonna download the Codex app just to ask questions about the product. But maybe they'll have like a super custom GUI for like ad analytics or something that they go into.

34:10Totally get that. And it kind of wrongly assumes on my behalf a consumer interaction at some point in that journey. And I want to ask you, how do you think about agent-to-agent experiences and designing experiences for agents? We spoke about, for example, going into large enterprises and how you can be helpful. I'm just using the most boring thing ever, expense approval. You could have agent submission of expenses on my behalf for my trip to San Francisco, and then the agent on the flip side doing approvals for that from OpenAI's compliance department. How do you think about that and that paradigm shift?

34:43That's interesting. You know, to be honest, I'm not sure what that's going to look like. My quickest answer to this is that we've noticed as we build Codex that the best interfaces for Codex to do work also tend to be the best interfaces for humans. So when people ask, oh, how can I make my code base more efficient for the agent to work with? The answer is often like, well, have you looked at it yourself? And is it easy for a human to work with? So a very specific example would be running tests in a code base. Naively, if you just set up most test runners, they just emit all the outputs of all the tests.

35:17And so as a human, it's really annoying because you have to go in and find the one that failed. And it's like, you've got to read hundreds of thousands of lines. Turns out that's terrible for AI as well. But if you filter it down to just only emit the failed test, better for humans, also better for agents. So probably the agent-to-agent interaction points will be very similar to if there was a human in the loop. And that's nice because it means you can kind of atomically replace individual systems. I mentioned our show on LinkedIn and a wonderful investor from a different company. It's like Harry Potter, you know, Voldemort.

35:47And it's like, you know, he who shall not be named. I don't want Sam to kill me. But from another company, you ask him, how do you think about a coding data moat? And does Anthropic have all the data now? I definitely don't think they have a significant advantage in terms of data on coding. I think that from what we've seen, and I would defer to my research team on this, but I feel like we have plenty enough data to build really good coding models. I actually think the place that's more interesting for getting data now is as we get into knowledge work tasks, that's kind of data that's not really available most places on the internet.

36:25And so you start to have really interesting brainstorms for how to help a model be good at it. Maybe you have to pay people to simulate doing tasks so that you can learn these trajectories for the model. Maybe you should acquire startups that are no longer a business but have a lot of data, like say they're Slack or something. Yeah, I think that kind of knowledge work task distribution is much harder than coding. That's so interesting you said there about the data that doesn't exist, so to speak. How do you think about your interactions with the data providers, your McCaws, your Turings, your Invisibles, your da-da-da-da-da-da of the world?

37:00Will you spend 10x there Or will you go, we are spending too much on data. We should do it ourselves and do data acquisition. Yeah, I mean, I think the way that we think about these things is just like, how do we move as quickly as possible? And so becoming able to set these things up in-house is like very expensive in time. And we're a small team. So what I have observed so far is that if we need to run a data campaign at scale, we're usually going to enlist help from one of these companies. On the consumer side for Codex, we've spoken about enterprises and going into them, how to engage in terms of developer experience, developer relations.

37:34Do you compete with a lovable and a rapid on a low-end consumer basis in a year or two's time? Is that a business where you're like, you know what, Codex is not for every person to create an about me or a small business to create their own site? How do you think about consumer in that way? Yeah, I would say that right now it doesn't feel like we're competing super directly. But I don't know if you saw our Super Bowl ad, the tagline of which is just, you can just build things. With the app, we noticed that many people who are less technical are starting to build things. And so the kinds of things they're building are much more hello worldy.

38:07And so I think that we will see some overlap in use cases where you have people just pulling up codecs because they have it as part of their ChatGPT. Actually, a big announcement last week was that we're now offering some codecs to people even on free ChatGPT plans or on the Go ChatGPT plan. So this is massive just in terms of like bringing availability to everyone. And so I think we're definitely going to see people with like a free chapter PT plan coming in and just like building simple things where they otherwise might have gone to a specialized tool. What would you most like to do differently, but for whatever reason you can't?

38:39I feel like it's been a very good few weeks for us. So we're very, I'm pretty jazzed about everything that's happening. Yeah, that's really interesting. You said it's been a very good few weeks for us. And I feel that. does the team feel the changing winds of momentum, both in positive and negative cycles? Absolutely. We are very attuned to it, right? Like if you look at the history of Codex, the first thing we launched last year was like this amazing idea that people were super excited about. It's like, hey, we're going to give the agent its own computer in the cloud. You're going to have as many of them as you want work for you in parallel on tasks.

39:13Super great idea. To be honest, it didn't work as well as what we shipped later. It was not the best. And then since August with GPT-5, we started pushing really hard on interactive coding, which is where most of the competition in the market is. You know, we went on an absolute tear. I feel like the public metric we had was like since August, we grew by like 20x. And then like even like late in the year, we like doubled from December to now. I forget the exact number there. But like that was competing neck and neck. But the shift that we feel last week is, you know, we, we felt like we had the most intelligent model that was cemented with five free codecs.

39:48We had feedback around our model being slower and like maybe less fun to work with. And like being less good at communicating with you while it was working, we address that feedback. And that's true, even compared to like the other competitor model that launched like 20 minutes before us and was like, maybe this is spicy. It was like soda for 20 minutes. So I mean, state of the art. And then we'd always been getting a lot of feedback on like the quality of the user experience in codex our most popular surface was the ide extension and our cli which is a command line interface was less polished but with the app the feedback has been like resounding from the market that this is like a really high quality experience it's like simple like unintuitively simple and people are just loving using even our biggest credits are converted so yeah and then we and then we had the super bowl ad and then we went to free and so going back to your question of like what am i most want to do differently i have two things for you.

40:36The first is I actually want to get back to cloud. When we pivoted our strategy from like focusing on the cloud agent last year to working interactively, the thinking was very simple. It was just, and it's kind of like what I was telling you about FDEs actually. If you go too far ahead to workflow automation before your end user is fluent with the tooling and can get it to work simply, then there's like this disconnect and you just have this pipe dream idea that's not like effective except for the most power users. But once you have this base where people are using your tool every day, and they're configuring it, and every time they use it, it gets better, then the step up to letting it run independently in the cloud is a much smaller step up.

41:10So I think it's time for us to get back to building out the cloud product and making it super tightly integrated with the local product. It already is somewhat integrated. And the other thing I want to do differently is start thinking more about the bottlenecks. Like CodeGen, writing code, has become basically trivial now. But the hard part is what you were talking about with code review, right? How do we know the code quality is good? How do we know we're doing the right things? And those bottlenecks, I think, are underappreciated still and underinvested in. So I think we want to get to a world where you can have an agent that is unbottlenecked, that you trust to own an entire microsystem or internal tool or whatever, and can do the full iterative loop, including feedback from users, without having to go through human review.

41:49And that is a really hard problem to solve, both from an intelligence perspective, but also from a safety perspective and a controls perspective. How much weight should we place on benchmarks and evals? Probably, this is an annoying answer for you. It's like some, right? Like they do tell you, in my mind, they give you a good measure of intelligence. And so you can put weight on those for intelligence. And especially before evals are saturated, I think when you see meaningful progress in those benchmarks, it's like very, very helpful. And then I think you have to pair that though with like what it feels like to use the model.

42:21And that's a vibes thing. Like whenever I talk to any, even internally or even talking to like customers of our models, I'm always surprised by how vibes-based the evaluation of how it feels to work with a model is. How vibes-based life is. People want to work with people they like is the lesson that I give to kids. People want to work with models they like. Relationships matter. Can I ask you, I think that Cursor will lose half of their revenue this year. I think we'll go from a billion to 500 million. It's a bold statement. Agree or disagree? Oof. Can I just like no comment?

42:59Yeah, you totally can. I don't know. I think it's really hard to say. Like, more serious answer here is just like, I think they've built a really successful business. We see them a lot when we're in enterprise. Do you? Yeah. Or is it just Cloud Code? Because I don't know anyone that hasn't. No, I see Cursor a lot more than Cloud Code. And it makes sense to me. My sort of narrative for this is that you have to meet people where they're at. For most people, like they're used to using an IDE. They've been used to using tab completion even before there was AI, right? Like top completion existed pre-AI and then AI just made it better.

43:32And so I think what's like coolest about Cursor from my perspective is that it meets developers exactly where they are. And it's a sort of a switch. It's like you used to be using VS Code or something, switch to Cursor, almost nothing is broken about your workflow. Everything works, just certain aspects got better. And obviously VS Code, I still use VS Code. There's like reasons you might like it more and they're improving rapidly as well. But I think that pitch from Cursor lands well with a lot of people. And so, you know, the bet on Cursor, I think, is that they can like continue meeting people where they are and then like ladder into these more advanced agentic features.

44:03You know, that relationship with the customer is valuable and it's hard to I don't think that goes away. Do you think it was the right strategic decision to start building their own models? It's hard to say, but I feel like there is a bit of a gap in the market right now for that kind of model. Again, if we think about what is the thesis, at least my thesis, you know, I'm not like super close to like working with Cursor or anything, but like my thesis for like how they win is that they meet everyone where they are and they like make it really easy to like step up into using more advanced agentic workflows.

44:32Maybe they noticed that the models that, for example, we were putting out or some of the competition we're putting out were like kind of slow relative to what their customers wanted. You know, my first magic moment in Cursor was like when I hit command K the first time. That's like a feature that lets you select some code and just like edit it in line. And I was like, this is incredible. Right. And so if they noticed a lot of their customers want to be able to kind of pair with the AI and then maybe after pairing with the AI for a while, then they start doing more delegation and then they move it to the cloud.

44:58Then there is a gap for that fast model that they trained. So I think that makes sense in that context. In terms of like market composition, as an investor, I have to think through how do I think about the eventual state of this given market kind of a terminal state? How do you think about that? Is it like Uber and Lyft? And like the majority of the market will be on Codex or Claw Code? Or is it like AWS, Azure, Google Cloud, and a 33, 33, 33? Okay, so I think this might end up with fewer providers that are capturing a lot of value in the long run. And here's why. And maybe this is a bit spicy, but I think that we are kind of in this temporary phase where we have agents that are really good at coding.

45:41right and if you look back last year like maybe more people thought we would have agents that are good at other domains too but that didn't happen last year so we only have pmf for coding agents like in the industry overall i would say right and then there's some like very narrow other use cases like customer support etc but i think that's probably temporary and then over time we're going to end up with agents that kind of can do anything for you this is kind of what i was saying earlier like there's just like a super assistant you talk to it about anything and then there is like specific UI that you can go look at if you happen to be deep in a specific function.

46:12So in that world, I don't think you want like 12 agents at the company and you have to like go, your employees have to go figure out the right one to talk to because then they won't achieve fluency. And if they don't want to achieve fluency, then they will also won't like pull automation into their roles. But if you have this one thing that you can talk to about anything, right? So your onboarding is just like, go talk to this thing about anything you need, then people will develop muscle memory to go to it. It'll become the center of gravity of work and people will pull in automation. So I think that that future makes much more sense.

46:38And I think like as the people building ChatGPT were like really well set up to deliver that. This is kind of a stretch, but an analogy here is I used to work at Dropbox. And for a while, this is before Slack was big. And for a while we thought, we wondered if people should like go comment on like documents in Dropbox or if they should like go talk about the documents in Slack. And it was like obvious that it was like more optimal for people to like put comments on the right timestamp in the video in Dropbox or like comment on the document in Dropbox, right? So it was more optimal. However, what we saw is that Slack is just such a center of gravity of people just like talking to each other.

47:14Like nobody wants to comment on the document. I just want to Slack you. And so we saw that like there was this really big pull towards things happening in Slack, even if it was less efficient. And I think we're going to see something similar at work where if there is a single agent you can use for nearly anything, there will just be this giant pull and everyone will talk about how they use that one agent for things. Teams will share best practices with each other. There'll be hackathons around how to use that best thing. Yeah. And you'll end up with just a handful of these. You said about kind of agents not really proliferating in terms of usage other than coding.

47:41And actually maybe this being the time and customer support is one of the examples. My question to you is, I'm an investor today. I'm looking for companies which will accrue value over time and provide incredible products to customers. There is a belief that the durability of revenue of large SaaS companies today is zero and that SaaS is dead because the model providers, you, Anthropic, others, are going to come for our lunch, so to speak. What would you advise me? Things are built for humans. Otherwise, what's the point? Even SaaS tools are built for humans. So for me, I think my question is, does this SaaS company own a relationship with a human on the other end of things?

48:23And if it does, then I suspect it's not going away. Or does the SaaS company own some really important system of record? It's probably not going away. Maybe both of those two things, the interaction with the human and the system of record, are more important than ever, actually. On the other hand, is the SaaS company kind of a glue layer, but it doesn't own either of those two things? I'm not the expert here, but I'm more nervous about that kind of company. If we take that stance, Salesforce and ServiceNow, they're down 20%, 30%, 40%. They shouldn't be. I don't think they should be. I don't know.

48:54What do you think? I would love to hear your take on this. I think it's massively exaggerated. I think there are some companies that legitimately should be. Respectfully, I think Dropbox is in a very difficult position. And I think your Monday.com's of the world, though, for the majority of SMBs and consumers who use it, which is the large majority of their market, actually, could they vibe code a to-do list? Yes. Would it be cost efficient to do so? Not really, actually, by the time you customize it and perfect it. And to be honest, a to-do list is generally pretty bland in terms of what you need to do.

49:25Add task, complete task, show historical tasks assigned to new members. It's not very difficult. And so actually, I think you just keep it. And so I think it's massively overblown. And I think that's the classic knee-jerk reaction from markets. I completely agree. I mean, if anything, like now that it's so much easier to build. But I do think, sorry, I do think like I think you're going to come for customer support. And I wouldn't want to be in that category. I think this maybe changes what kind of founder you invest in, right? I think there was this maybe temporary phase that I liked personally as a product builder.

50:01There was this phase where you would invest in the person who can just build good product. And you could kind of ignore if they had a good thesis around a customer or go to market or distribution or anything like that. Because it was so hard to build good product. And I think that was an anomaly. If we look at where we are now, maybe that kind of founder is not the founder you should invest in. Because it's kind of relatively easier to build good product. And you need to go back to like investing in the founder who's like thought through distribution, who has a good domain expertise of what to build for a specific customer, etc.

50:29So, again, if you were on my team as an investor, how would you think about interesting areas for us to invest in, in companies that will accrue value and not be threatened by model providers? Because, again, you're going into health. You go into code. Obviously, Codex is very clear. You go into customer support. Where are you not going? Where is Claude Code not going? I'm tempted to just say, like, I don't know. I think it's a hard time to be an investor. The market is so dynamic, it's hard to say. It's a really tough time to be investing today. My answer is kind of twofold, actually, which is number one, I look for things with physical infrastructure.

51:03I don't think you're going into energy supply. And then two is the fintech and banking integrations, gnarly financial products. I don't think OpenAI is going to go into building 500 relationships with banks in Southeast Asia. I tend to agree. It comes back to, are you going into a gnarly, complicated market where customer relationships and knowledge of the market are everything? That still seems great. How bad is the war for talent? From the UK, we look at SF and I say to companies, it's better to build in Europe because it's impossible to acquire talent and it's impossible to retain it. Am I wrong?

51:38I think that the war for talent is incredibly fierce right now. Obviously at OpenAI, we have an incredibly strong brand and so we're able to attract a lot of talent. But even so, we put a ton of effort into like closing candidates that we're really excited about. Even like even we feel it. It's not like you don't just get whoever you want for free. Can I ask the entry price that you get stock at? Is it still attractive for the best talent? I haven't had anyone tell me anything to the contrary. To what extent do you think about like finding the perfect fit versus finding someone who's good enough?

52:09So, you know, earlier I made my joke about like PMs kind of being optional. Yeah, I think that's not actually true. You still need product people. But I do think that they have to be the perfect fit. And if you have someone who's like not the perfect fit, they might just do more harm than good. It's kind of means that like we're way more selective than I might have been in other roles. I'm a CS student. OK, I'm at Stanford. I'm an Imperial. I'm at Cambridge. I'm wherever, ETH, great institution. What would you advise me, knowing all that you know now, that would help me navigate the next five years of my career?

52:42I want to be valuable to the AI ecosystem environment as an engineer entering the workforce in the next year. Basically, there's actually never been a better time to be an engineer because you have incredible tooling available to you to get an incredible amount done. And your ability to ramp into a complex code base that you might be hired into has never been faster because you can go ask AI a ton of questions about the code base. and you can ask it to plan out changes that would otherwise take you like days to research maybe. I think first off, I would say like, you should be like very optimistic.

53:17But then of course, like about your abilities once you're at the job. Then the other question is how do you get the job? Because it's never been like easier to build things. The thing that becomes scarcer is like agency, taste and like quality. I would urge you to like just build things and demonstrate your agency and your taste around what you build and like build things that are of high quality and then share those things. You know, we get a lot of inbound from folks both applying for jobs through the careers page or also on social. And this is just me. But when someone writes to me with like some interesting thoughts and like a link to an interesting project, that gets my attention much more than like a normal resume does.

53:56Final question, so we do a quick fire. What has Claude Code done well that you sit back and you learn from? Number of things. It's like I was saying, I think way back last year, they made something that was really easy to use and just like work with all your tools with Xero setup by running it locally in your terminal. When we started investing much more in the Codex CLI and shipped great models for it, like GPT-5, our growth exploded. And so I think that idea of just meet people where they're at, give them something easy to use, let them ramp from there and figure out how to use it has been awesome.

54:26So that's probably the biggest learning we've had from them. What mistake do you think they made that you've also learned from, having had the benefit of seeing them make it? They over-indexed on their initial success with their command line interface tool. I think at the end of the day, it's not the friendliest UI. And it makes it hard to extend beyond pure builders. And it makes it difficult to truly delegate to agents because effectively to delegate through that kind of interface, you have to be kind of a power user of your terminal or Tmux or something. And so that's why we built the app. And I think the market reception around the app, to me, it was kind of a risk when we started.

55:00But it makes me really feel good about that decision because the Codex app is a much more intuitive, simple interface to get started with. It's like less scary. But then it naturally leads you to this idea of like, I'm going to take my hands off the keyboard and like delegate to the agent. You mentioned Dropbox earlier. The alumni from Dropbox is incredible and really amazing to see the talent that's come out of Dropbox. What's your single biggest lesson from Dropbox that has shaped some of your thinking now with OpenAI? Oh, I don't need to think about that one. That's kind of the thing I was telling you about earlier, right?

55:32Like when you're building tooling for people, like for end users, you have to think about like that tooling as a system of engagement. If people don't want to use your tool, if it doesn't naturally feel like the easiest way to get something done, then people just won't use it. Again, I learned that from watching how Slack just absolutely took off. And so I think about that a lot now when we're building these agents. I'm like, if we build our agent purely as workflow automation, then it's always going to be like pulling teeth to get that thing started. You're going to need to hire Accenture or someone to come in.

56:02They're going to need to deploy FDEs. It's going to be tough. But if you can build a system that people just love using, even if they only use it for partial tasks, over time, they'll get better and better at using it. And then you'll get connected to the tools you want over time. And then you can start laddering in automation. Obviously, these aren't mutually exclusive. How on earth do you reinvigorate growth at Dropbox today? At least from when I was at Dropbox, the thing we were uniquely good at was desktop software. And desktop software, it's funny, it was never not back. But anyways, it's so back.

56:32basically because if you're solving for productivity and knowledge work, yes, there are systems of record everywhere that you need to connect with, but everything at the end of the day happens on the user's computer, either in their browser or just like locally in apps on their computer. I do think that the fastest way we're going to see productivity gains from agents at work is going to be at first meeting users on their computer, working with the stuff that they have available to them, without having deployed FDEs to set anything up. And then over time, you'll connect in these various systems.

57:00And so if I was Dropbox, I'd be thinking about how do we leverage our unique domain expertise in like building really good like desktop software and the sort of collaborative layer on top of your computer? How do we leverage that to enable productivity agents? It's a bit broad, but I think that's the angle you go. No, I love it. And I really appreciate the response. Final one before we do a quick fire promise. I've been brought up in a world where margin matters. Software margins are wonderful. And it's what makes software a brilliant category to invest in. We're seeing margin profiles that are very different in inference heavy plays in particular.

57:32To what extent should I put that out of mind and appreciate that costs will come down, cost of tokens will come down, and actually it's about usage and customer love, margins will come, or no, margins are actually freaking important. Keep that focus. I think both costs are going to come down significantly. And I also think that if this is the year of agents being deployed broadly at work, then this is also the year where they're going to have to be connected to all these various systems. And I think that's going to be very sticky. And so I view this year as a race. And so I think you want to win that race.

58:04And you should be OK taking some hit to margin in the meantime. Dude, quick fire round. So I say a short statement. You give me your immediate thoughts. Does that sound OK? Yeah. What have you changed your mind on most in the last 12 months? When I joined OpenAI, I thought that this is a little longer than 12 months ago. But when I joined OpenAI, I thought that we would all just be hanging out with our computer screen sharing. within a year from there. You know, we'd have this agent that we're just talking to. That was completely wrong. I think the rate of like progress in like multimodal models was like slower than I expected.

58:34Multimodal means, you know, like models that work with like video and audio. So instead, what happened was that we saw that like agents that work with your computer through code are the way. And so for me, that's been a complete rethink in terms of like how we bring the benefits of AI to like just people generally. It's not through video and audio primarily. Which less unknown capacitor do you respect most and why? The first one that came to mind was AMP. I think they're building, yeah, AMP. It's out of the folks at Sourcegraph. Their product has a great reputation of just being like, you know, punching way above its weight.

59:06But I think the other thing that I really respect is that they helped initiate this whole like standardization around like agents.md and like.agents slash skills, which are what I was saying earlier about like making it so it's easier for users to manage all these different agents that they're trying. We obviously put out agents.md, but they put out agent.md. And basically, Quinn started this all by putting out a tweet that said, hey, do you guys buy the domain agents.md? We'll standardize to your spelling. And as small as that was, that initiated this whole standardization that I think has been awesome in the community.

59:36Do you think the response to Anthropics ads was the right response? I mean, there were so many different responses. The one that I heard, obviously, I think was right. The one that I heard was, well, one company is being pretty negative about the future. And the other company, us, OpenAI, is being really positive and just telling people they can build things into dream. I thought that response was brilliant. I mean, Sam wrote an essay. Do you think it was a good response? I think so. I mean, I think as one of the cool things that I love about OpenAI is like people are like very unapologetically and authentically themselves.

1:00:09And so for me, that was just like a very authentic response. And I like that we do that. What's the hardest product decision you've had to make since being at Codex? Well, I can tell you the most painful product decision we had to make. For a while, Codex Cloud was effectively unlimited. Not free, like you needed to pay for ChatGPT, but then you had unlimited usage. Every day that we left it that way, we knew that it would be harder to wind back at being unlimited. But we were just so focused on competing on our other things that had more PMF that we punted that decision out. When we wound back that unlimited use to some more reasonable limit, there was a lot of blowback from users.

1:00:43And it was a very small minority of users who thought everything should be pseudo-free forever. but that blowback affected us everywhere because like the social chatter doesn't really distinguish between these things. I think the lesson I learned the hard way there is like you can't make things unlimited for too long. Dude, it's like pricing, grandfathering pricing is just, it's such a hard thing. What do we do today in engineering or product that in five years time you'll look back on and go, oh my God, can you believe that we did that? Well, one is just editing code by hand. I think probably another one, this is maybe spicier, but another one might even be like actually managing the deployment and monitoring of systems by hand.

1:01:23Like I basically think that probably big companies will take a long time to like deploy this. But many startups might actually kind of start building on a completely new stack that's like fully AI managed. To be clear, the stack doesn't exist yet, but a fully managed AI stack where basically it's been built to give you really strong deterministic guardrails over what the agent can do and control over to whirl back deploys and everything like that. And so we'll get to a world where the way you start a company is you start by getting an agent and just asking it to build things. And then you get more agents in that.

1:01:53And then maybe eventually you add your co-founders to this service that you use to work with agents. And so you end up like maybe your main communication tool is actually your agent communication tool. And then maybe you're not actually handholding this very painful CI and deploy process, but you're just having agents do things. Weird question, but I'm intrigued. Are you the one providing agent guardrails? And what I mean by that is agents can go anywhere within an enterprise. Are you responsible providing those guardrails? Or is there a third-party matter provider who is saying, hey, Alex, you can't go into that.

1:02:26That's human resources. Or you can't go into that. That's marketing. How do you think about guardrail provisioning? And is that the role of the agent provider or a third-party provider? I think we'll probably see both. We are putting a lot of effort into agent guardrails. Like I said, we're basically the only company that cares about OS-level sandboxing for coding agents. For instance, there's none that exists on Windows. We're the ones building that. And we're doing it in open source, so hopefully other people can use it. We think about that a lot. Chachptee supports connectors. So you can talk to your Google Docs or something.

1:02:55And we put a lot of effort into guardrails around what the agent can do with your Google Docs. Those are just two examples, but we think a lot about this. And I think probably, though, the way that we'll do it will not be sufficient. Like there'll be third parties who provide like very bespoke things for very bespoke, you know, company needs. And there'll probably be a mix of both. Final one for you, my friend. What are you most excited about when you look forward 10 years? This is probably going to happen in much less than 10 years. But my mission sort of personally, when I joined the company, was I just felt like even with the models we had a year and a half ago, there was so much just capability overhang or just ability for these things to be useful, but we hadn't built the right products around that.

1:03:33And so people like me were getting more benefit than people like my grandma. What I'm most excited for is to get to a form factor for AI that means that they're just helping everyone regardless of whether they're in tech and especially if they're not in tech or especially if they're older. And so the concrete vision I have is at some point, we'll add an agent to our family WhatsApp or something and it'll just start being useful to the family without anyone having to think harder about it than that. There are many other ways that that could happen, but I think concretely, that's the most obvious thing we can do with like my grandma.

1:04:06Dude, I so appreciate you. I so appreciate you putting up with my wandering questions and my very episodic mind. You've been fantastic, man. Thanks so much. I mean, I appreciate you putting up with my wandering answers. So all good. We're two here. But before we leave you today, the early story of Atlassian is probably very similar to your own. Atlassian knows firsthand the challenges that startups face every day and that the right tools are essential to go from MVP to IPO That's why Atlassian for startups gives eligible companies up to 50 seats free on the premium edition For products like Jira, Confluence, Loom, Jira Product Discovery, Compass, and Bitbucket So your team can use the best-in-class tools to plan, track, and collaborate on work Whatever that work may be, many of today's most successful startups like Cloudflare, Canva, and Rivian relied on Atlassian for their growth trajectory.

1:04:59And Atlassian wants to give that same opportunity to the next generation of builders and investors. We know how important it is to focus on building the right things early. Whether you're in the sticky note stage or well on your journey, teams at any stage can work smarter together. It's never too early to start with Atlassian. Head on over to Atlassian.com forward slash startups forward slash Harry for more details and eligibility. After Atlassian helps your team build and ship great products, Intercom helps you support the customers using them. If you're looking for a way to transform your customer service, let me introduce you to Finn, baby.

1:05:33Finn is the number one AI agent for customer service, resolving up to 93 % of customer queries automatically. There is no other agent that can do that. Not 93 % of customer queries, okay? No other agent can do that. So why choose Finn? Finn is the best performing AI agent for CS. Finn doesn't just answer questions. It takes actions. It automates the most complex customer queries like refunds, transaction disputes, technical troubleshooting with speed and reliability. I wish my team was speedy and reliable. Beats every competitor in every head-to-head bake-off. Completely configurable and code optional setup.

1:06:11My word! I mean, the benefits just go on and on. It's easy and efficient implementation. It works on any help desk with no tedious migration needs. It's trusted by over 6 ,000 customer service leaders, including top AI companies like Anthropic, Lovable, Synthesia, Clay, Vanta. So if you're ready to transform your customer service team, scale your support, and give team members time to focus on the really high-level strategic work, Learn more about FIN at fin.ai forward slash 20VC. While FIN scales your support without losing speed, Reforge shows you how to translate that scale into durable product-led growth.

1:06:48Everyone's shipping faster than ever. Cursor, claw code, codex. AI is making code and writing code faster than ever. But here's the problem. Speed means nothing if nobody uses what you ship. That's where Reforge comes in. Reforge is building the product discovery engine that sits upstream of your coding agents. Not another prototyping tool, research repo, or AI interviewer, but a product that will ingest your customer data, generate variations of product solutions, validate the solutions before code is written, and hand off winning directions to your team. Reforge kills product debt before it starts, because every unused feature you ship isn't just wasted engineering time.

1:07:31It's a maintenance burden, complexity tax, and surface area that you cannot shrink. Used by product teams at companies like Toast, Vimeo, Klaviyo, and many more, Reforge helps teams ship more features that actually get used. Try Reforge at reforge.com forward slash build and use the code 20VC, that's 20VC, for one month free of pro.

From the publisher

Alexander Embiricos is the Head of Codex at OpenAI, leading the development of the company's flagship AI coding systems that power automated software generation, debugging and developer workflows. Under his leadership, Codex has become one of the most widely adopted AI developer platforms. 

AGENDA:

05:13 Will Coding Be Automated? Why AI Could Create More Engineers, Not Fewer

07:17 Do We Need PMs? The "Undefined" Product Role and When It Matters

08:06 The Real AGI Bottleneck: Human Prompting, Validation, and "Too Much Effort"

13:04 Three Phases of Agents: Coding → Computer Use → Productized Workflows

13:52 Enterprise Reality Check: Security, Permissions, and Safe Agentic Browsing

17:57 Is Inference the New Sales and Marketing? 

18:49 What % of Codex Was Written by AI?

21:33 Do OpenAI Use AI for Code Review?

23:31 Is there any stickiness to AI coding tools?

28:22 What Does "Winning" Mean at OpenAI? Mission, Competition, and Moats

32:04 The Future UI: Chat or Voice

34:10 Agent-to-Agent Workflows: Designing for Approvals, Compliance, and Automation

35:39 Do Coding Models Have a Data Moat?

36:50 How does Codex View Data: Will They Build Their Own Mercor and Turing?

37:27 How Does Codex View Consumer: Will They Compete with Lovable?

41:56 Benchmarks vs "Vibes": How People Actually Judge Models

42:43 Cursor's Edge and the Case for Building Your Own Models

47:37 Is SaaS Dead? What Still Defends Value (Humans + Systems of Record)

51:28 Talent Wars and Career Advice for New Engineers in the AI Era

01:01:03 Guardrails, the Fully AI-Managed Stack, and a 10-Year Vision for Everyone

 

 

 

More from The Twenty Minute VC (20VC): Venture Capital | Startup Funding | The Pitch

All 521 episodes
20VC: Codex vs Claude Code vs Cursor: Who Wins, Who LosesThe Twenty Minute VC (20VC): Venture Capital | Startup Funding | The Pitch · 1 h 8 min
Listen in VO