In short
The Neuron: AI Explained - Episode Summary
Episode Title
BONUS: OpenAI Codex Demo, Learn the Absolute Basics of Coding with AI
Podcast Description The Neuron covers the latest AI developments, trends, and research, hosted by Grant Harvey and Corey Noles. This episode includes a live demo of OpenAI's Codex with Alexander Embiricos, showcasing its capabilities in coding tasks.
---
Key Discussion Points
Introduction
- Hosts Corey Noles and Grant Harvey introduce a special live-streamed episode featuring Alexander Embiricos from OpenAI.
- The aim: To teach coding from scratch using OpenAI's Codex, specifically the new GPT-5.1 Codex Max model.
OpenAI Codex Overview
- Codex Model: The latest iteration, GPT-5.1 Codex Max, is designed to handle various software engineering tasks autonomously (e.g., pull requests, code refactoring, frontend builds, debugging).
- Capabilities: Codex can maintain context over extended coding sessions, compacting history for efficient refactoring and multi-hour agent loops.
Live Demo Highlights
- Alexander Embiricos explains Codex's architecture and how to set it up for coding tasks.
- The team encourages audience participation, taking questions and sharing experiences with coding tools.
- Coding Tools and Environments: Discussion includes using IDEs like VS Code and Cursor, emphasizing ease of access and integration with Codex.
Key Features of Codex
- Collaboration: Acts as a coding partner, augmenting the productivity of developers by assisting with mundane tasks and offering code suggestions.
- Feedback Mechanism: Codex can review code changes and offer critiques, enhancing the quality of code submissions.
Development Workflow
- Best Practices: The hosts discuss strategies for getting the most out of Codex, including setting up a structured development environment (inner and outer dev loops).
- Incremental Development: Emphasis on iterative changes and thorough review processes to ensure high-quality outputs.
- Plan.md and Agents.md: Frameworks for defining project goals and workflows, which guide Codex's actions.
Security and Best Practices
- Security Concerns: Importance of running code through security checks (Aardvark) and having robust review processes.
- Education and Empowerment: Discussion on the importance of coding fundamentals alongside AI tools to ensure developers understand and manage their code effectively.
Conclusion
- Emphasis on the potential for AI to significantly enhance productivity in software development.
- Encouragement for listeners to experiment with Codex and other tools, while remaining aware of best practices and security measures.
---
Key Takeaways
- Codex's Autonomy: Codex can independently handle coding tasks but still requires oversight and human input for complex projects.
- Iterative Development: Utilize feedback loops and incremental changes for effective coding practices.
- Use of AI Tools: Embrace AI as a means to accelerate learning and productivity rather than replace foundational coding knowledge.
- Security Measures: Always prioritize security measures when deploying AI-assisted code to avoid vulnerabilities.
Final Notes
- The episode encourages listeners to engage with AI tools in a thoughtful manner, blending traditional coding skills with modern AI capabilities for a more efficient development process.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOAWS Reinvent Insights
0:45 to 2:00
Discussion about experiences and insights from AWS Reinvent.
“So for everyone tuning in, let's just give a quick, I guess, introduction of what we're going to be doing.”
Exploring Robotic Innovations
2:00 to 4:00
Talk about interesting robotic technologies seen at AWS and their implications.
“Someone said, wow, I've never been so alone.”
AI in Fulfillment Centers
4:00 to 6:00
Discussion on Amazon's use of robots and the debate around humanoid robots.
“So they have at least I don't know, but it's and they've had robots on the floors in their fulfillment centers for years.”
Codex Introduction and Goals
6:00 to 8:00
Introduction to OpenAI's Codex and goals for the live demo.
“So I'm excited and anxious to see people freak out about them in the Midwest because they absolutely will for a hot minute.”
Getting Started with Codex
8:00 to 10:00
Overview of Codex's capabilities and how it can assist coders.
“So we'll also have Alexander tell us a little bit about that.”
Understanding Codex's Evolution
10:00 to 12:00
Discussion on how Codex is evolving and its potential use cases.
“And there's a ton of work done before, like ideating and planning.”
Evaluating Coding Capabilities
12:00 to 14:00
Explanation of the evaluations used to measure Codex's coding abilities.
“makes a lot of sense where do you think codex is in that process right now like on the scale of like Yeah.”
Understanding Evaluations in Coding
14:02 to 15:36
Learn how coding evaluations measure model performance in real-life tasks.
“We're still very focused on the coding use case, but we're right about at the time where it's like, interesting to try more.”
The Advantages of Codex Models
15:37 to 17:35
Discover the benefits of using Codex models for coding tasks.
“And so that's where Codex really shines.”
Exploring Open Source Tools
17:36 to 20:06
Learn about open-source tools and how they enhance coding efficiency.
“Yeah, so we shipped it to Codex like the CLI.”
Show all 51 chapters
Introduction to IDEs and Terminals
20:07 to 23:09
Understand the roles of IDEs and terminal windows in coding.
“I don't even know if this is going to work.”
Hands-On Coding with Codex
23:10 to 26:05
Watch a live demo of Codex fixing coding issues and its underlying processes.
“So let's so let's fix some of the things that we wanted.”
Harnessing the Power of Codex
26:06 to 28:00
Explore how Codex models are built to work efficiently and effectively.
“test the jump by the way oh wait I need to refresh you might yeah I need to refresh it.”
Understanding Codex Features and Context Windows
28:00 to 29:15
Learn about Codex's ability to handle long-running tasks and context windows.
“If that's your if that's if you're doing a really hard task that takes longer than like the models context window, you would kind of historically get stuck.”
Exploring the Role of Prompts in Codex
29:15 to 31:00
Discover how prompts are used in Codex and their impact on model behavior.
“But yeah, you could totally just like have codecs run for ages with just a single window.”
Creating Customized Coding Agents
31:00 to 34:20
Understand how to configure agents.md to tailor Codex's responses.
“So, for example, this is actually one of the prompts that the Codex agent uses.”
Getting Started with Codex Installation
34:20 to 35:30
Learn the process of installing Codex and the various options available.
Comparing Manual Coding and AI-Assisted Development
35:30 to 41:45
Explore the differences between traditional coding and utilizing AI tools.
“Like I saw someone in the comments was like, yay for Tmux.”
The Challenge of Manual Coding
42:01 to 42:38
Learn about the difficulties of manual coding in modern times.
“A fun fact is that I used to work on a startup and I started that company with my co-founder obviously, but the point is that I wrote the initial prototype for something we ended up building on an airplane with no wifi.”
The Debate on Coding Tools
42:39 to 44:46
Explore the pros and cons of using coding tools for productivity.
“And then they said they're missing the point.”
Learning to Code with AI Assistance
44:47 to 47:00
Discover how AI can enhance learning and coding efficiency.
“But regularly, if I don't know what something is, I'll have a chat GPT window over on the side and I will just grab a snippet, throw it over and be like, what is this?”
Starting Projects with Codex
47:01 to 48:27
Understand how to initiate coding projects using Codex and AI tools.
“And like you showed earlier, you can actually have a chat inside the terminal or, you know, your code editor and asking questions there.”
Real-World Coding with AI
48:28 to 54:48
Get insights on applying AI to solve real coding problems effectively.
“I've downloaded, I've got, I'm all logged in here.”
Finalizing Your Coding Project
54:49 to 56:00
Learn about the last steps in creating and refining a coding project.
“Built a beautiful, a playful cat coder landing page with animated hero achievements, grid, live claw, soul, honks, and a tongue-in-cheek bug list.”
Introduction to the Demo
56:00 to 56:40
Exploration of initial coding tasks using AI.
Blogging as a Learning Tool
56:40 to 58:00
Discussing the benefits of blogging while learning coding.
“Yeah, no, we definitely need some more cats in it.”
Editing Code with Codex
58:00 to 59:40
Showing how to make changes to code using AI tools.
“And that's what I'm trying to do on this.”
Real-Time Changes and Testing
59:40 to 1:02:30
Demonstration of live coding changes and running tests.
“I don't want a user to have to read this text every time, right?”
Code Writing without Direct Input
1:02:30 to 1:05:20
Exploring how Codex can write code based on simple commands.
“Now what it's doing is it's running cargo test, which is a way to confirm, basically it's running tests to validate its change.”
Creating and Refining Plans
1:05:20 to 1:10:01
Detailed discussion on structuring coding plans for better outcomes.
“So I'm just getting finding I have some links for myself.”
Workflow Optimization with Codex
1:10:01 to 1:12:09
Learn how to integrate Codex into your app development workflow for efficiency.
“So this is a very common workflow we see where you kind of have a generated plan, ask some tweaks and do it.”
Real-World Application of Codex
1:12:10 to 1:14:19
Explore practical examples of using Codex in coding tasks and project management.
“And this 17-minute task that ran before I got in was one of those.”
Debugging and Task Management with AI
1:14:20 to 1:17:06
Discuss strategies for effective debugging and task management using AI tools.
“but for folks who don't know, pull requests is where it then merged the code back into the GitHub folder.”
The Future of Coding with AI
1:17:07 to 1:24:00
Understand how AI is reshaping coding practices and what skills are needed.
“Well, and this is where I think like for me personally, right?”
Understanding the Role of AI in Programming
1:24:00 to 1:25:19
Learn how AI tools can accelerate programming but still require human oversight.
“like what I just showed you, then, you know, it's going to be much more like, yeah, you're making like very careful changes.”
Value of Learning Computer Science
1:25:20 to 1:27:49
Discover why a strong foundation in computer science remains essential even with AI advancements.
“Just by the nature of it being so much more very, very hands-on day one.”
The Dynamics of Building Apps with AI
1:27:50 to 1:31:45
Explore the process of developing applications leveraging AI tools like Codex.
“I usually had a ton to answer questions about the code base so I don't have to bother engineers.”
Enhancing Code Quality with AI
1:31:46 to 1:34:09
Learn how AI can assist in code review and improve software development processes.
“It was, you know, again, it was Vibe-engineered.”
Ensuring Security in AI-driven Code
1:34:10 to 1:37:10
Understand the security measures necessary when using AI for coding.
“So if you go to Codex web, you go to settings, then code review, and you can turn it on super easy.”
Future of AI in Software Engineering
1:37:11 to 1:38:01
Get insights into how AI tools like Aardvark are shaping the future of software engineering.
“So I could imagine we're still figuring this out, but I could imagine like just us making sure that Codex is excellent at security review.”
Identifying and Fixing Issues with Aardvark
1:38:01 to 1:40:09
Learn how Aardvark utilizes Codex to identify and fix issues in real-time.
“And then it will actually send Codex in the Codex Cloud to go fix the issue.”
The Evolution of Codex and AI Adoption
1:40:10 to 1:42:02
Explore the evolution of Codex and its impact on AI adoption among users.
“And I really love that you're taking an approach to that.”
Phases of AI Integration into Daily Work
1:42:03 to 1:44:36
Understand the three phases of integrating AI tools into daily tasks and workflows.
“Like, like, so the AI part is, is already so powerful, but actually there is still a ton of work to just bring those capabilities to all humanity right now, right?”
Enhancing Human-AI Collaboration
1:44:37 to 1:46:46
Discover how AI can enhance collaboration without taking away human agency.
“I can ask us to do something totally non-coding related or powerful in coding, but it's, you know, it's pretty open-ended, right?”
Skill Development and AI Features
1:46:47 to 1:48:51
Learn about the importance of skills and features in maximizing Codex's potential.
“A lot of people were like, hey, I would like a comment that would be more clear.”
Implementing Connectors and Skills in Projects
1:48:52 to 1:51:24
Explore how to effectively implement connectors and skills in AI-driven projects.
“A skill, the way that I think about it is actually more of a set of instructions that create a capability.”
Understanding the Interdev Loop
1:52:00 to 1:53:23
Learn how Codex can adapt to your coding style and workflow.
“So the interdev loop is basically like, does Codex know where you keep your context?”
Exploring the Outer Dev Loop
1:53:24 to 1:54:58
Discover how to manage tasks and review processes with Codex.
“And it kind of takes away the idea that maybe it's a cookie cutter approach for every engineer or vibe coder out there.”
Growth and Experiences with Codex
1:54:59 to 1:56:01
Hear about the rapid growth of Codex and its impact on productivity.
“Is there anything else you want to share with us or demo or just any parting thoughts for the crew?”
Recommended Tools for Developers
1:56:32 to 1:58:35
Find out about IDEs and tools recommended for coding.
“You know, I think this is a really good thing.”
AI Integration in Development
1:58:36 to 2:00:08
Explore how AI can enhance the coding process and specific tools to use.
“I know a lot of folks who do more like Java development and Python really like the JetBrains IDEs, like PyCharm or IntelliJ.”
Transcript
Automatic transcript. May contain errors.0:00.
0:32Hey, what's up, everybody? How you doing? Hey, Corey here with Grant, as always. Got a really exciting live show today that we've been anxious to bring you for a little while and managed to make it happen. And we're super stoked about this. Corey, where are you right now? Where are you recording? Where in the world is Corey knowles where in the world is cory knowles i uh today cory knowles is is in the encore in vegas i've been at aws reinvent all week and how's that been how'd you like it it's wild uh it was a lot uh talked to tons of ai companies we have so many cool people that are going to be coming on the show in the coming months you'll be really excited um uh like enough that we're trying to figure out where to wedge great people.
1:25Fair, yeah. It's a good problem to have. So for everyone tuning in, let's just give a quick, I guess, introduction of what we're going to be doing. So we have Alexander Ambirikos coming on the show. He'll be joining us in just a moment. We're going to be talking about Codex, which is OpenAI's coding agent. And the idea here is that we are going to cover as much Codex as we can to try and take you or myself or anyone else from zero to I can code with agents and I'm not afraid to do it. That's the idea. That's right. That's the goal. We're going to see if we can do that. We have lots of fun questions.
2:06We'll take questions from the chat. Hi to everyone coming in. Someone said, wow, I've never been so alone. Love that. And yeah, Jimmy Roland says, thanks for doing this, guys. Excited for this session. Awesome. So we're just going to ramp a little bit talk about aws until uh alexander's ready to jump on uh yeah yeah what was the coolest thing you saw aws reinvent this week cory i'd like to say that the coolest thing i saw at aws was at replay last night and it was it was this giant like 60 foot maybe even bigger robotic hand and arm and it was very mad max style so like they uh they come in and and they would take people put them in a chair lift them up and uh give them a remote control and they could take this robotic hand reach down grab an actual automobile lift it in the air and slam it on the ground and uh it was it was pretty sick but uh you know on a serious on a more serious note though sure uh lots of talk about physical ai this week lots and lots of it and you know new types of data, new types of models, new types of everything, you know, trying to get, you know, beyond the internet and literature and things that we have and figure out how to teach AI, not just the, I guess you'd say, not just the knowledge of the world, but maybe about the world in a more physical way.
3:39Yeah, that's huge. And I don't know if people know this, but because of the extreme demand of their fulfillment centers, Amazon is one of the biggest employers of robots. So not only do they employ a lot of humans, they employ lots of robots. And I'd actually be curious to know which one they employ more of now. That's a stat I'd like to look up because I think I heard they at least employ a million of each. So they have at least I don't know, but it's and they've had robots on the floors in their fulfillment centers for years. like yeah you know and now you know another interesting discussion this week was are humanoid robots of any actual value i mean it's a big debate like you you know there's a lot of well yeah the argument was like are legs really the best way for a robot to get around you know if you're a factory setting absolutely not you know a lot of people think it's wheels a lot of people think it's like crab legs you know like the six uh i've heard hexapods referred to a lot as the ideal form factor.
4:38I guess the benefit of robots is that, or humanoids, is that they can go everywhere a human can. So I can walk through my door into my house and walk around all of this stuff. So could a human. I don't know if a six-legged robot could be walking around in there. I like, we got a great comment here for Jim Moran. I want that robotic arm in my truck to adjust driver's habits. I literally thought that. We were talking about, man, wouldn't that be cool to have to do something with us? I want my trunk lid to pop and the arm to come out and just slap what needed. That's funny. I don't know how new this law is because somebody was referring to it as if it was coming up new.
5:26But there's a law in California that at least has been proposed that requires – Because at least at a certain point, you couldn't ticket a self-driving car the same way you could a human. I don't know if that has been addressed, but this law was aiming to address that, where it's like somebody has to, you know, from the company has to comply with like police basically say, hey, this robot car is doing something. I'll take the fines. Somebody's got to be that guy. Somebody's got to take the fines, right? Speaking of, while I'm out of town, Waymo came to St. Louis, Grant. really okay which shocked listeners of yep listeners cory's fromst louis famously and now now uh waymo's in your town all right you're gonna be taking it as a as of like this week they said the for the first little bit they'll have humans driving the cars around while us humans get used to seeing them uh but i'm so excited i'm so excited to be able to take one at home now and and uh i tried to take what's the other one zoo uh zooks which is uh amazon actually yeah and and they're here in vegas but there were only like four stops directly on the strip and it never coincided with where i was or where i was going so that's an interesting car because there's a steering wheel sort of like a bus right yes yeah it's um it's neat but it looks like we'll have the little Jaguars in St.
6:48Louis. So I'm excited and anxious to see people freak out about them in the Midwest because they absolutely will for a hot minute. So I'm going to go ahead and just chat a little bit about Codex for people who don't know. Looks like Alexander is ready to come on. So I'll just give a quick intro and then we'll let him tell you about Codex because he'll tell you much better than I will. So just real quick, As a reminder, the idea of this is that we're going to try to go from zero to hero with Codex today. I'm going to see if Alexander can assuage my fears of Codex, help me become a legendary Codex agent.
7:28I've got my computer here, so I'm going to try and follow along with what he shows us. The vibe-coding God you know you can be. Yeah, I want to become the vibe-building god that I know I can be, that I know I have inside me. But just as a question, how many people have used Codex before? Shout off in the chat if you've used it before. It's okay if you haven't. Or I guess you could say what coding agents have you used before. Let's go ahead and share that while we're getting this set up. And one of the also really cool things about today is that yesterday, GPT 5.1 Codex Max, which is OpenAI's new coding model, just became available in the API with the same pricing and rate limits as GPT 5, which is great.
8:11So we'll also have Alexander tell us a little bit about that. All right. So it looks like we're getting some responses. All right.
8:19Alexander Embiricos:I'm going to go ahead and I'll go ahead and bring Alexander in here. Hey, man, how are you? Great to meet you. Nice to meet you. Good to meet you too. Hey, thanks for being here today. For sure. Fun. Yeah, I'm just catching up on the chat. Looks like we've got some people using Codex, people using other stuff. Cool. Yeah, yeah. So this is good. So this is, we're going to, you're going to show us how to become Codex masters today, Alexander. No pressure. All right, all right. Hopefully it goes well, yeah. For people who don't know, Alexander is the product lead on Codex. Alexander, can you tell us a little bit about it and what it is and what people should know about it?
8:56Alexander Embiricos:Yeah, for sure. Okay, so Codex is OpenAI's coding agent. We launched a research preview earlier this year. It worked in the cloud on its own computer. And since then, we've just been like massively improving it. So right now, the most popular place to use Codex is in your IDE. So like in VS Code, a bunch of people pop it up, I'll show you this, got a VS Code extension. And it's a great tool to just use to like, ask questions about your code base, get them answered, and write code together. You can also use Codex in your terminal as a CLI, in the cloud, on your phone. You can use it a bunch of places.
9:30Alexander Embiricos:And like broadly what we're trying to do with it is it's just a teammate that works with you everywhere you work. So, you know, as we go over things today, I would love to show you like, yeah, like the basics, how to do very simple things in Codex. And then I have some screenshots because I was like trying to figure out how can I demo like the coolest stuff we do. And I was like, this is a live stream. If I like accidentally live stream some secret. right right so i just have to show you like okay here's kind of like some of the cool ways that we use codex like as a full teammate at the company at open ai i love it that's awesome i like that you said as a full teammate at open ai i think that's that's a fascinating approach yeah that's the goal we were just talking about waymo and waymo refers to their system as the driver right it's a single system do you look at codex the same way i mean i think we just call it codex this interesting like you're I'm kind of thinking like would we could we would we go around and call it that teammate I don't think so um but definitely that's like it's it's it's like less about personifying him in that way like the goal is still to like just accelerate every human developer um but it's kind of more like we think about like what can codex do today and what what should it be able to do for it to be way more useful so like again we'll show you like today you can ask it to write code that's mostly what people use it for but I think of like what an engineer does like Like, you know, the common stat is like engineers are writing code maybe like 30 % of the day.
10:53Alexander Embiricos:And there's a ton of work done before, like ideating and planning. And I think that's getting even more important nowadays now that it's like possible build so many more things quickly. And then there's a ton of work afterwards, you know, verifying that your code is secure, reviewing it, getting it deployed, monitoring it. And sort of, you know, I think it's quite funny. Most of the coding agents today, even codecs to a large extent, are very good at writing code. but it's kind of like they're this intern that's really good at writing code and refuses to check slack or like refuses to check right and it's like well how much would you trust that that person without you also being in the loop and verifying the answer is well you definitely need to be there but what we're trying to get towards is a world where um more like a teammate uh it's connected to your tools and it's like doing much more of that work including the validation step afterwards uh and really excited about that so that's one side and then the other side is where you don't have to tell it every single thing to do right again like you know imagine an intern first day and then eventually they get smarter and smarter and grow into an amazing senior engineer a large part of the difference is actually that you you just stopped having to give like very specific instructions right and they start telling you yeah cool that makes a lot of sense where do you think codex is in that process right now like on the scale of like Yeah.
12:10Alexander Embiricos:I suspect. Oh, I mean, we're super early. We're super early. But, you know, it feels like we're like at the time where there's a ton of this potential now. Or, you know, if you think about it, we've been like, you know, the research team has been cooking. We've been shipping new models. We shipped a new one to the API yesterday. The new model is way smarter at coding, way more token efficient at coding. we can share more details but it's actually I mean I don't know what I should screen share but I can show you a cool chart there let's go let's jump into it yeah let's yeah let's dive in okay so you know this is a first for me I'm going to share my entire screen I know the anxiety that comes with that make sure you mute notifications you don't want I think we're going to be okay don't want to leak any secrets yeah yeah I mean I would like it if you leaked secrets you would not like it if you leaked secrets i you guys know how we launched um like a new browser yeah atlas yeah i accidentally screen shared it to a bunch of journalists
13:18that's hilarious you just you were using it as your browser and yeah it's just my browser i mean it's great um the chat's offering to sign an nda for you yeah okay you can take this video
13:29Alexander Embiricos:down and make it exclusive. I'm using Chrome here because the new browser has a bunch of features in it and I couldn't figure out how to feature read them off. So, you know, here we are. Cool. So basically what I was getting to is like, if we think of the models we'd be training, like they're getting really good at coding, really intelligent. I think it's the best one out there. But the focus still for now is like very much like working with the model in, you know, a repository of code, basically. And so we're starting to see all these use cases of people using them for other things. But that's not yet our focus to push there.
14:02Alexander Embiricos:We're still very focused on the coding use case, but we're right about at the time where it's like, interesting to try more. So yeah. Could you explain Lancer and TerminalBench for people who don't know what that is? Yeah. I mean, basically these are tasks that relate to like writing code. And so what we do is we have, it's kind of like an exam, if you will. Let me know if this is like too basic of a way to explain it, but basically... Well, it's too basic. Yeah. Yeah. So like we have these things that we call evals, like evaluations. And the idea is that it's kind of like an exam. And when we train a model, we give it this exam, which is like, hey, like, do you know how to do this?
14:33Alexander Embiricos:Like very basic task? Like, can you navigate to a folder in a terminal? Okay, what about a harder task? Like, can you fix this bug? Okay, like, what about an even harder task? Like, can you set up my computer in the right way? Or can you even like perform this task? And so a lot of these evals are, you know, inspired by real life, because we're trying to measure like real life value um and so freelancers thinking about like freelance tasks that you know have been available online for humans to go attempt and like how many of these can like a model achieve terminal bench is like the idea of like you know if you're do i have my terminal up yeah so like if you're using you know terminal um can i like a very basic task would be i don't know could i i don't actually know if this is a terminal bench task but this is my standard use case it's like can i use ffm beg to convert like some png into like some other format right these are the tasks and what we do is like when we have a new model we give it the exam we see what score it gets and then we compare it to the previous model and so like these are some cool things to notice like oh cool like this new model like codex max is like significantly better than codex that's wild yeah that's a big jump actually because when you get into this range it's it's it's clay yeah yeah that's where the jumps start getting smaller yeah i feel like the the reputation that sort of the feedback that we get for the codex models is that they're really smart and so if you're wondering like why should i try codex um i think the simple answer is just you should try it so that um you can uh you can like benefit from the smarter model that can do harder tasks especially in like i mean i'm going to show you a bunch of like vibe coding demos at the beginning of today just because they're easier and more fun but like yeah you know the real work that we do is in like more complicated code bases with harder tasks.
16:16Alexander Embiricos:And so that's where Codex really shines. The other cool thing, though, that I mean, for me is even more exciting is this graph. This is the one I wanted to show you. So we have to like explain a little bit with this graph, how to read it. But basically, the takeaway is the model is smarter, but also faster. So, OK, what does this mean? If you see here, this x-axis, this is like how much time the model spent thinking. So for example, if I ask Codex to do some task, like, I don't know, let's just do this here. I might say, okay, how many, this is a really simple question, right? Folders are in here.
17:01Alexander Embiricos:We're going to see that Codex is going to think a little bit about what my question is and then what it should do. And then it's going to run a command and it's going to think about the answer and then it's going to give me an answer, right? So this, like, it had to think for a little bit, even though this is a simple question, because I didn't tell it, like, run this command. I told it, I asked it this, like, natural language question, right? So it thought for a very short amount of time. And so the more stuff is here, the more thinking, the further along this axis we go in terms of thinking tokens.
17:29Alexander Embiricos:And then if you take this exam, this eval, so we've been verified, there's like, okay, well, how good is it? What score does it get? and basically my analogy would be like if i give you an exam it's quite a hard exam and i say you have 30 minutes total to finish the exam versus like 15 minutes versus you have two hours we can kind of tell the model like hey like think kind of medium hard about it think very hard think extra hard or extra um and so what's really cool um here is that not only is the model like able if it has a lot of time to like way out class prior models in terms of how well it can do in this exam but also it can achieve similar scores with many fewer thinking tokens which means just faster for you as a user and also it means uh it's cheaper for you like you can use yeah the truth is that's still a good score at low reasoning yeah yeah yeah so actually i use codex at low reasoning like most of the time when would you when would you use the high and x high in your put in?
18:29Alexander Embiricos:If I if there's specifically a task that I think is hard or if there's like a specific thing where I know like I'm going to step away from my computer, you know, the bit sliders like the rise of coding agents has led to like terrible hygiene in offices because people just have all you walk around and there's all these computers running just like codecs going. um that's hilarious you know but uh yeah so if i know explicitly that i i'm not gonna look for a while and i just wanted to go off and come up with this best answer like i have a thing running that maybe or maybe not i'll show you that okay i finished but you know 20 minutes question here um we'll see if we get into that uh but that's the kind of thing where i'm like fine like take forever i don't care uh but most of the time when you're just kind of in flow like with what i just did you want the model to be fast yeah yeah definitely um this is awesome this uh this is cool.
19:21So I guess I was a little confused. I didn't know that GPT 5.1 Max was new. And it's really good.
19:27Alexander Embiricos:Yeah, so we shipped it to Codex like the CLI. So our Codex products a couple weeks ago. Okay, that's what I was thinking. Yeah. And then usually like, what's quite interesting about what we're doing as a team is we are we're training models. And we're building the open source harness. I'm going to I want to show you another page. do I risk launching a new tab from here? Risk it. Risk it for the biscuit. Let's do it via the terminal. Let's live a little YOLO. I feel like this will avoid you guys seeing my tab completion. There you go. Is that going to work? Oh, shit. What is it? I don't even know if this is going to work.
20:10Alexander Embiricos:I've never tried to do this specifically. Okay, that worked. Oh, awesome. No, I avoided typing into my address bar. So, you know, this repository here on GitHub is open source. So, like, any of you can go read it. And it actually contains all the tools and prompts that the model gets. So, is it useful for me to explain, like, what a harness? Yeah, yeah. So, we'll just do some broads. This is GitHub. GitHub is where all the code for most of your projects live. So, for people who are completely new to coding, like, you're going to want to get GitHub. You're going to want to use GitHub. Continue from there.
20:47There's definitely great tutorial videos out there.
20:50Alexander Embiricos:Yeah, so let's do this. Let me do a task here, and then while the task is running, we can kind of talk about how this is actually happening under the hood. But you guys tell me what people are doing. I don't know. That'd be great. Okay, so let's see. Good task. Maybe actually... While you're doing this, I'll also explain what he was just showing his VS Code. So if you've never used what's called an IDE, this is an IDE. And this is what it looks like right here. Yeah. So like, you know, maybe we can even show more stuff here just to. And this is the terminal for people who are complete computer noobs.
21:31This is the terminal. So this is where you're talking directly to your computer. I feel like it's important to say this type of stuff, because when I was learning this, I was really confused at the difference between all of these things.
21:40Alexander Embiricos:Yeah. So like here, like, you know, if I'm on my computer, there's just like a bunch of folders. Right. And these folders contain code and various documents. And so there's a thing called an IDE, which is an integrated developer environment. But really, it's just a fancy thing that you can use to like go read the code or stuff. So I happen to have this folder open. It's called 3JS. it's a game 3.js library for making stuff. And let's say I just open this folder and I have no idea what this is or how to use it. That's a place where Codex can already be helpful. So I can go here in GitHub and I could say, how do I play this game?
22:22Alexander Embiricos:Let's just say you told me it was a game. Okay. And so for those of you who are used to using ChatGPT, you're used to doing this kind of thing where you can just ask a question. and then get an answer. So, okay, I got an answer. It happens to be able to read the stuff that's in here, which is the same as the stuff that's in here, just a different view. Cool. So I get an answer and I can go into the terminal here and run the command that it asked me to run. Cool. So I can play the game and just for fun, let's talk about the game briefly. Yeah, I just have a little commander game. I can move it around with WASD.
23:02Alexander Embiricos:I can recruit troops. have sound actually yeah let's play it it's gonna echo i don't think we have that up whatever it is like fun sound effects um yeah and uh looks like there's a button to plant windmills but it doesn't work and i've heard that something's wrong with the jumping animation okay whoa all right right this looks like something i made so that's something that i you know made part of So, you know, so there's a bunch of code in here and sort of in the olden days, we would just go in and we would read the code and like understand all the details. But this is a vibe. Find the problem manually, I assume.
23:42Alexander Embiricos:Yeah. So let's so let's fix some of the things that we wanted. And I actually just so I don't have to type. I can just type them up. So what did we say? The jump animation is way too high, right? And the windmill button doesn't work, but I don't know if you want to do those at the same time. Yeah, okay, so we can do that. So let's say that. Let me open a new chat. I'm going to get to do a different thing at the same time. What was it? The windmill. Okay. The windmill button. The plant windmill button doesn't work. Okay. And then this is just for fun. Make the commander more like it had a halo and a cape, right?
24:18Yeah. Yeah. Ooh, cowboy with a halo. Yeah, I love it. So and I, you know, this is the demo I like to give.
24:26Alexander Embiricos:I can also ask it to add music, but let me not do that because you can't hear anything. So so OK, so we can watch codecs kind of working now. Cool. And so now this is a good time to just like think about what it means to have a coding agent. So looks like OK, so it figured this out. Right. So basically. If we look at what happened over the hood, like so the UI is kind of compressing this because I don't really care about all these steps that the agent took. But under the hood, what happened is a model, GPT 5.1 Max Codex, Codex Max, basically received this query, and then it was told, hey, you're actually running as an agent in this folder, repository with a bunch of files, go figure out how to solve this user request.
25:13Alexander Embiricos:So what it did is first it was like, what files are in here? so it ran ls which is a command that like shows you you know the folders uh from there it uh you know it's like okay let me go search for the jump function so it's like running various commands it's like oh like rg is a command to search for things so it searched for the word jump and then like okay it found the jump it found the command that does it it read the file that contains the jump command and then it saw the like oh like there's this variable called jump impulse which seems very high let me set it to a new value and so then it made that code change so you know normally I don't have to know any of that I just know that like it worked and then it tells me that lowered the value but behind the scenes what it was able to do was to like use my computer actually right and so this is what we call the hardness so when I was showing you all this code well let's test the jump by the way oh wait I need to refresh you might yeah I need to refresh it.
26:10Yeah, there you go.
26:12Alexander Embiricos:We did it. Also, this cowboy hat. That's good. Oh, yeah. Kind of like a fedora, actually. A little bit fedora code. You know, see if the windmill planting works. Sick. There you go. You know, okay, cool. It deleted the cape. That's just me preferring no capes. All right, there we go. No capes. So, you know, under the hood, basically what happens, what we have is we have a bunch of code, where was it? Here, in this repository that everyone can go read. And so when I send this query, the model sees this and then it decides, okay, like, what can I do? And so there's some prompts in here and some tools that it can use.
26:51Alexander Embiricos:And the prompts say like, hey, you're able to execute commands in the user's terminal. You can run these commands in this way. And to make sure that it's safe, we have this thing called a sandbox, which limits what you can do to only things that the user would want you to do unless you can ask for permission to do more stuff. Anyway, so there's a bunch of code around this. And that's what we call the harness. So a bit of a rambly answer to that. No, that's great. It's a good question. It's a good answer. That's a question I've had for a while. I hear a lot of talk about harnesses, especially around benchmarking and stuff and how different companies are doing it.
27:30Alexander Embiricos:Right. And so one of the cool things that we're able to do on Codex actually is we, because we're both training the model and building the harness we're able to do these things together um so to build them to work really well together um that means we're able to like move really fast and like try experiments uh and you know it's kind of like instead of building like a general purpose thing we're able to build something very specific so like when you're when you're using the codex models if you want to like just try using them directly they're a bit more like a sports car that's like very good at like being used a specific way and a bit less good at being used if you want to use them for something like outside of codex they're not as useful um but then you get a sports car um and uh you know we get to build cool features like i think we actually talked about it here so yeah like one of the cool things you can do with codex is you can run these like really long running tasks um models have something called a context window which is basically like how long they can think for before you need to kind of like reset them maybe it's like a human going to sleep you know it can be awake for so many hours and so your task um is a terrible terrible analogy, but whatever.
Read the full transcript
28:33Alexander Embiricos:If that's your if that's if you're doing a really hard task that takes longer than like the models context window, you would kind of historically get stuck. And so this is something we've been working on, like both making codecs faster, which is what we were talking about with this chart, but also making it smarter running through long tasks. So we have this feature here called compaction. Right. That basically we train the model how to like basically prepare itself to run again in a fresh context window. And then we built a harness to like do that in a very specific way that the model is trained for.
29:05Alexander Embiricos:So we're able to ship that feature like super quickly. I have a question about that. So does that mean that I, you know, unless I want to run multiple agents at once, I'd never have to open a new window? Is that what that means? Yeah, I mean, I don't know. I think that's up to you. But yeah, you could totally just like have codecs run for ages with just a single window. There's a lot of different patterns. like actually one of the common patterns at OpenAI, which, you know, I think in the future, you might not need this pattern as models get better and as the harness gets better. But one thing that we commonly see is people just like, so I'm just opening another terminal here.
29:47Alexander Embiricos:You see people like working with the model in one pane and this context they're managing. We're kind of skipping some steps and getting more to the hero side of what you guys are talking. Yeah, well, we'll circle back. Yeah. So this basically, like if I'm working with Codex here, I might want sort of this one chat with the model where we call them threads, where I'm like very opinionated and confident about what I want it to do. So I'd be like, go do this next, you know, like giving very specific instructions. But then maybe the truth is, you know, I'm just a mammal. I don't know what to do exactly.
30:19Alexander Embiricos:So I need to get some questions answered. So I might have a chat on the side where I'm asking questions about it. like is there we just had a good question that's kind of relevant to this uh yeah what you were working on a second ago here and they were wondering if you're leaving the prompts in the code specifically for the ai to read um that was referring to like github right or when what was that referring to yeah that was just a minute ago or maybe three minutes ago um The harness. In the harness. Yeah. There are prompts in the harness. Yeah. Whoops.
31:01Alexander Embiricos:So, for example, this is actually one of the prompts that the Codex agent uses. And, again, this is all open source, so you can just go check this out. Okay. And so, you know, this just explains a bunch about how we want the model to behave. And so, yeah, I don't know if that answers your question. Let me, I'm trying to see. Are you leaving prompts in the code? Yeah, let me know AC main if this answers your question, but there are plenty of prompts in the code. We actually are often updating these. And so it's quite interesting, like if you were to go into our open source resource, maybe we can do this.
31:33Alexander Embiricos:And you were to look at the prompt for, you know, this prompt here, it's quite long, right? There's a lot of description about how codecs should behave. And this is because this is actually the prompt here for GPT-5 model, which are general, right? You can use GPT-5 for a bunch of stuff, right? creative writing all the way to coding um and so because it's like it doesn't necessarily know like by default like how to work in a coding harness we give it quite a lot on the other hand if we go check out the prompt for gp5.1 codex max this model is like trained very specifically to work in this harness and so the prompt is well i don't know if you can tell but the prompt is much shorter um and uh you know we don't have to tell it anywhere near as much and like broadly speaking almost anything that we have in this prompt, we eventually try to delete because the model should, I mean, most things, there's some places where you want to do durability.
32:27Alexander Embiricos:So that's on like the prompts in the R code in the harness. And there also are some valuable ways to prompt codecs. Let's say that every time I prompt, I want codecs to work a certain way. One thing that I can do, so like uh yeah one thing that i can do is i can add a file here called agents.md yeah and this file is actually a file that a bunch of coding agents support uh so like codex is just one of them but also like amp and gemini and this was like a really cool thing that like when we when we created this name agents.md we intentionally made it generic and like plural to like invite like hey like yeah everyone should use this and it was really good community rally around that so like for example this is a silly example but i could say you know hey um i don't know how are you doing this is but you know it's going to give me some answer right and let's say that actually um i don't want it to write like that i always want it to be like sound like a pirate
33:37Alexander Embiricos:this is that you know a better example would be more like telling it about some like specific code style thing that's like hard to infer from the code base right like you know let's say you use like two tabs versus four tabs or four two spaces versus four spaces like codex will just figure that out or like semicolons versus not it should just do the right thing but let's say it's something harder to figure out like i want you to sound like a pirate Now, if I start your chat, then I can say, let's see if this works. You know, now it will like read this prompt every single time. I'd be sailing.
34:11Yeah, this is good. Yeah, this is good. So that's amazing.
34:15Alexander Embiricos:All right. So anyways, that's your answer to the question around prompts. You know, teams eventually build out like fairly like robust agents.mds as they like start configuring codecs to work specifically the way that they want. do you have to be chatting with the agents.md for it to reference it or I can be doing this one thing I was considering could be fun to show you would be we could go up here to one of the demos that we put in the GP5.1 Codex Max blog post so we built the solar system well so you know if I gave it this prompt here for this one I might do medium just because I think it's like writing a whole lot from scratch um you know
35:02Alexander Embiricos:so this task isn't going to finish instantly so we might have to wait but i guess the point is here that you can see the first thing it did is it read the instructions so yeah it knows that it should sound like a pirate even though i wasn't like looking at that file and so when it answers and it's done writing the solar system it'll it'll just sound like a pirate that's awesome yeah where should we go from here i'm curious let's let's start let's start from say I'm a noob say I'm interested in this um how do I download how do I download this program you're using how do I download codex let's like go to the codex website and show the process because I was doing this today and I thought it was kind of interesting there's so many options so I was a little bit curious like what option you recommend uh people people start with yeah so I think so I'm gonna you know you mentioned say I'm a noob right so let's let's say you're just getting started.
35:49Yeah.
35:51Alexander Embiricos:Like I saw someone in the comments was like, yay for Tmux. And like, you know, they were looking at this terminal thing. Like, yeah, look, terminals are actually not that scary, but I don't think like the noob thing is to use a terminal. So, and like Tmux and like all that stuff. So I think really the sort of more basic thing is to use an IDE. So like, I like VS Code. There's plenty of other great IDEs. I also like Crissor and plenty of others. so um what i might do yes code's been around forever tons and tons and tons of people use it yeah so i might go to you know i don't want to risk a web search and like accidentally how to complete some secret project i can uh how about this i can show my i can show my screen which i have pulled up and you can you can show me what to do great oh that's awesome just so you know also if if there's ever something you want to bring up and that is a concern uh we can pull your screen share down from the screen for a second so you can do what you need to do then week.
36:47Oh, perfect.
36:47Alexander Embiricos:Okay. Well, yeah, let's set you up. If you screen share, I can kind of walk you through something. That'll be fun. Yeah, let's do it. Because I was trying to code some stuff earlier and we'll see where I went wrong. So, Monique, can you share my screen? Great. Grant's computer. All right. Oh, I need to present. One second.
37:12So, one of the other things that we want to talk about, Alexander, when we get there, is you released two really awesome guides recently that I want to kind of touch on some of the advice from. One was the building an AI native engineering team, which was really interesting. And also the Codex Max prompting guide, which was really cool as well. Yeah. So we'll touch on that at some point. Yeah, and we'll drop a link to that in the chat real quick here.
37:42All right. So my screen is up. I think. Okay. So.
37:55All right. So I am on, I went to openai.com slash codex, and I see there's a button here to get started, and there's all these options. What should I pick? Okay.
38:08Alexander Embiricos:So it all, like, it starts with how do you, like, how do you write code today? So like, can you show me, it looks like you have an editor open behind. Can you show that to me? Yeah. So this is one. This is. Okay. So that's cursor, right? This is cursor. Yep. Cool. Okay. So now let's say you go to the Codex website and use cursor. So go ahead, get started and click cursor. Okay. And that's an extension. That's our extension. So you would, you know, it would say install there. I mean, you can, we can show the install if you want. I don't know. Yeah. I'll uninstall it and then I'll install it again.
38:44Alexander Embiricos:Super fast. Nice. Okay, sweet. So now you have Codex. And so you're going to want to get the extension. It's a little bit, this is kind of a tricky thing. They tend to get buried. But can you also, can you zoom in? Yes. How about that? Yeah. And then, yeah, go to your extensions pane and launch the thing. So have you already moved Codex somewhere specific? if not no i have not yeah okay so under you see where it says agents and editor then there's like a few icons underneath no yeah yeah yep yep yep yep yep yep yep yep yep yep yep yep yep yep yep yeah and i think that there's a drop down to the right uh of the the extensions icon so like go higher gosh it's painful uh higher how about this this is where you can find it like just go a little higher like you up higher higher zoom out there you go ah there we go you got codex right so that's where climate is okay got it whether the fact that that's hard to find is you know i don't know if cursor people are listening but their extensions are very hard to find in cursive um yeah much easier to find in vs code do you want to open this yeah so i'll open it and then i'll sign yeah so i mean maybe just to call it out for anyone who's watching so what grant did is he clicked the login button really fast so you click that eyeball for me grant sorry right here nothing i was i was had your password up i was like would you click that eyeball for me yeah yeah and you really got to do a lot when you're streaming yeah okay let me uh he's logging into his chat bt account because codex is included in plus and pro plans as well as like um business and enterprise so monique could you take my screen down for a second second long in thank you and you're on pro aren't you grant uh on which one on cursor oh no i was thinking on uh i was thinking yes on this on this account over here um so by the way a good day to all okay i see what you're saying i'm not sure exactly what's going on there but if you just go into curse if you use cursor as your ide that should work if you use a different ide um then you're going to want to click a different one um but no matter what id you use if it's a vs code fork you can um so i if it's vs code or cursor or windsurf you can just go into the extensions marketplace and just search codex and you'll get it awesome okay we had one more here while Grant's downloading, wondering if we can maybe explain the differences between manual coding and using AI.
41:37What's the difference? What's the process going to be like in comparison?
41:41Alexander Embiricos:It's a good question. It's actually pretty interesting. There's like fully manual, and then there's sort of this like accelerated, like using AI native IDEs coding, and then there's like coding with agents. So this is kind of the spectrum. I, you know, I, and I guess all of us used to just code manually. A fun fact is that I used to work on a startup and I started that company with my co-founder obviously, but the point is that I wrote the initial prototype for something we ended up building on an airplane with no wifi. Oh, wow. That's awesome. I actually recently have been on an airplane with no wifi and decided that it would be cool to prototype something and like wrote three keystrokes and i'm just like why am i here this is like there's no point i should just like watch a movie because it is so much slower and maybe i'm worse at manual coding probably but um it's just too hard now so i have a question about that actually but we can get into that a bit later uh okay um but actually i'll ask it now there there was a conversation that i saw on twitter recently and basically it was like oh i was talking to some engineers who said, oh, I don't use coding tools because it gives me like the equivalent of engineer brain rot.
42:58And then they said they're missing the point. It actually gives you 10X capabilities. 10X is your capabilities. And I actually think both are true to an extent, like both have validity to them where if you're not doing the task, you're not gonna learn the task, but at the same time, you can be 10X more productive using the tools. And I wanted to get your take on that, on how you feel about it as you're designing basically the equivalent of an agent that codes for people. Yeah, I mean, I think it's kind of just like abstraction layers, right?
43:28Alexander Embiricos:I think one of the things that makes software so powerful is that over time, we've had more and more layers built by people who are coding before us that make it easier for us to do things. Right. So you've got like HTML and then you have JavaScript and React, right? So now you don't know, very rarely are you like manually manipulating the DOM unless it's some like very special use case. So like if you develop VS code, I, you know, I don't, as far as I know, they don't use react, right. Cause it's like a super high performance use case where they need something like more advanced. Um, similarly, like most people don't need to know assembly today.
44:01Alexander Embiricos:Assembly is like a low level, like machine language. I don't even know if that's the right definition. You know, I took one class in college. Um, and then, you know, on top of assembly, there's like languages like C and like also like most people never need to learn C today. Right. And we just keep getting to easier and easier to use languages that are more and more powerful. And you sometimes need to understand some things about the compiler or the runtime. You definitely want people on a team who understand the details of how the internals of what you're building work. But you don't need everyone to do that.
44:33Alexander Embiricos:And definitely as you're getting started, I don't think you need to do that. You can learn that knowledge over time as you realize you need it because you're trying to solve some bug. And so I think that's totally right. Right. Right. That's kind of what Carpathia was saying is that, you know, it's actually better to take a project, work on it, and then, you know, figure out what you need as you go. Exactly. Yeah. It's so hard. Well, that's kind of how. Oh, sorry. No, no, you go for it. Oh, that's kind of how I learn is like if I'm working on something and trying something and I don't necessarily understand what I'm looking at, like quite often you can kind of understand by just looking through it as you go, even as you're new and you learn more.
45:12But regularly, if I don't know what something is, I'll have a chat GPT window over on the side and I will just grab a snippet, throw it over and be like, what is this? What is this doing? And have it like kind of break a section down into something that makes sense to me.
45:28Alexander Embiricos:Yeah, totally. And I think that's one of the most powerful things about, you know, this age with AI now. It's that almost like any given thing that was impossibly hard before is still hard now, but, or maybe even not that hard now, because you can just like try with AI and like AI will help you understand it. And so, you know, we used to be in a world where like, you know, you had to be very strategic about like what you learned and then you then do things related to what you learned. And you could like slowly expand your frontier of knowledge and capabilities as like a human, but it was like harder.
46:03Alexander Embiricos:Right. And now like, yeah, you still, it's still, you know, you still want to learn as much as possible, but you can aggressively expand your domain of knowledge and you can like lead, like you can almost do more than, you know, now, and then like catch up with your knowledge using AI. Right. And like, whether or not it's worth your time to like deeply understand like the underlying system or deeply understand code that is written is something you can be very strategic about. So if you're building like most production systems that like real like you know tools used by many people like you do want to understand like all the things but if there's a place where the code is like extremely well tested and like modular and compartmentalized maybe you don't need to anymore and you can just trust the boundaries and your tests similarly like if you just are trying to learn something and you're not about to deploy this to like millions of users you could just prototype and you can vibe code that prototype it doesn't matter right so we get to like we have this new lever of like where we spend our time that I think is like really empowering.
46:56Alexander Embiricos:Totally. Yeah, I love that. Yeah, it's an effective way to learn. Yeah. And like you showed earlier, you can actually have a chat inside the terminal or, you know, your code editor and asking questions there. You don't even have to pull up chat dbt if you don't need. What are 32 to 47? You know, I don't understand that. Yeah. So like, right. Yeah. And you could just highlight it. Actually, that's the thing we can show in the ID. You could just like highlight some code and be like, what does this code do? um and then can we put up consulting fives question about like if i need to start writing python because i think it's an interesting one yeah absolutely my take here is that you should just follow your energy right so like some people like reading docs and tutorials i actually i do like reading docs and tutorials like a decent amount not i'm not going to spend days on it so i would just try just like google like you know or use chat gpt actually like search like learn Python.
47:49Alexander Embiricos:But I don't even think about that anymore, because I just use Atlas. So like automatically, you go basically go search it. And then if there's a tutorial, try reading it. And if you get bored in 30 seconds, stop reading it. But if you don't get bored, then just keep going until you get bored, right? So you might as well learn. But similarly, try building some project you have, and see where you get stuck and just like use AI to like build that project. And like, if you try to do these things like some things will just be easier for you based on your personality than other things and you can just go down pull that thread you know you don't have there's no specific way i think i like that i like that a lot it also solves the cold start problem where you're sort of like uh i don't you know i i'm like putting this off because it's going to take a lot of cognitive load or just say like read it for 30 seconds and and if you can read it for 30 seconds you can probably read it for two minutes if it's interesting right yeah um okay so I have my screen up.
48:44Monique, we can pull it back up. I've downloaded, I've got, I'm all logged in here. So it's giving me a little intro that I'm going to click through. It's explaining how it works. So I'm going to click through that and then try the new model. Let's test this out.
49:03Alexander Embiricos:So you're in. So is this, yeah, so you can just say hi or whatever you want here. And then maybe we can zoom in also yeah let's zoom in how do i i'm going to close this one sec
49:19which thing you want oh i wanted to close this yeah close close all there's just a lot of stuff open here okay
49:34Alexander Embiricos:cool okay so like i don't know what this folder i guess it's claude test but like what do you what Do you have a thing you want to do here? I actually had something I wanted to do on the web version. So I can show you that. But I just wanted to show people, like, so you're saying get Cursor, download Cursor,
49:53download the Cloud Code extension into Cursor, and then you can start a project here. Codex, excuse me. Codex. Get the Codex one, put it in Cursor. and then how would I actually start a project here?
50:09Alexander Embiricos:Okay, so first off, I would definitely do the extension in IDE if I was someone who was new to coding. I wouldn't do the web version. That's like a little bit more advanced. So if we wanted to start a new project, then I would just do it the way that my editor does it. So like we would make a new folder somewhere on your computer. We could try this actually if you want. Yeah, let's do it. So it's like... New folder? Yeah, you need to make a folder. So just do that the normal way you make folders on your computer. Call it codex test. This is actually really high pressure that I'm doing this. It'll be good.
50:49Don't worry. You got it, man. You got this. It's in some other folder of some other thing. But we're going to roll. That's fine. It doesn't matter.
50:57Alexander Embiricos:Just a folder. We're going to rock with it. There we go. Folders can live wherever. They're just folders. Yeah. Okay. Git repository. That's your editor talking, so just do whatever you want. Okay. Cool. So now I'm here. How do I start my project? Okay, so let's pull up Codex again. So you click the drop of the arrow because it's buried. The famously difficult to find one. Yeah. Like, yeah, obviously because I'm wearing Codex. I would recommend VS Code because it's easier to find Codex. Oh, I have that open as well. Where is it? I have a lot of things. You have all of the things open, don't you?
51:32I do. Oh, no, they're not all open.
51:34Alexander Embiricos:okay okay yeah but we're chill here i mean chris is fine once we have it open so um yeah so like what do we want this project to be do we want to like i i like to make joke sites about dinosaurs that write code but you know what do you want to do make a joke site about um cats that write code how about that sure there we go there we go so you know that's a way to get started with the project You know, I think for me, if I was sort of knew what I was doing, so this depends on how newbie I am, but if I know what I want to do, I might know what programming language I want it to be in, or I might know what frameworks I want to use, like maybe I want to use React or something.
52:16Alexander Embiricos:So I would probably tell it how I want to start. But this is a way to just simply start a new project with Codex and like we'll have it work and we can open it in a couple of minutes. I will say that like if you are an experienced programmer this is actually not the best way to start using codex how would you start it as an experienced coder because I would go find like think about like some hard problem that I have to solve like today in like my real like normal full-time project or hobby project or whatever and I would ask codex to fix the real problem because that's where it really shines like there's a lot of models like like all the models out there can like buy code a site about cast the right code and I'll even say like you know I don't even know that like our site will be like the best like vibe coded site about its cats you write code like maybe maybe not right but if you if you go to like a real production code base and you have like a real problem and you ask for help there like that's where codex will shine the most yeah we have another good question in the chat uh someone said can you recommend any just enough just in time dev incremental bloggers etc uh diving into python and actually enjoying it an LLM real time write partner code for me and debugging equals chef's kiss.
53:29What's your take
53:30Alexander Embiricos:on that? I so the where I used to read those kinds of blogs was mostly when I was doing Swift and I was like just getting to Swift when they got started. Yeah. So I don't know if that's super applicable to folks here, especially doing Python. But I remember liking the names are kind of hard to come by but like like there was like swift lee it was like bobby lee i don't know that i don't have too many that i can remember off the cuff this is a bit of a previous life but what about you guys what do you recommend uh well i really like so some of them trended to be a bit uh technical but uh um simon willison's really good at working with lom specifically he he's he's really focused on on that in particular um and i also like all the content that swix puts out and then just reading Hacker News and seeing what the most interesting stuff is.
54:23That's usually what I do.
54:24Alexander Embiricos:Yeah. Yeah, I think the forums are really great for if we're talking about AI versus like the language fundamentals. Forums like Hacker News and Reddit, I find really good. I find like being on, you know, Twitter X is there's a bit more hype. It's like it's like more one thing or the other. Whereas these platforms where there's like upvoting of comments, I find like quite useful to like just see what other people are doing. It's good signal to just kind of see, okay, what are other people doing? And then you can kind of adjust. So now we're thinking here. Yeah, we can read what it did. Built a beautiful, a playful cat coder landing page with animated hero achievements, grid, live claw, soul, honks, and a tongue-in-cheek bug list.
55:08And it created three files, index HTML, which is the layout. Then we've got the style.css, which is the style. And then script.js, which renders the scripts, like the functionality of it. Great.
55:25Alexander Embiricos:So it's telling you you can open it. So you could, you know, if you know how to open the file, you could just do so. But if you don't, you could just say run open index.html for me. Run open index.html? Yeah, I mean, you don't need to say whatever. And so this is actually an opportunity to illustrate the sandbox. So before, don't click that just yet. Okay. launching launching your browser is sort of a an action that's like very powerful and so we've decided as like you know when we build codex like we don't want the agent to be able to just like launch your browser to whatever site it wants um you know without you first confirming um so in this case it's asking you can i can i run this command and you can say approve once or whatever you want here um and then it'll open so yeah if you want to click okay and here we go we refactor yarn into shipping features that's pretty good deploy cat i wonder what that does uh probably nothing yet yeah maybe nothing elite developers
56:29Alexander Embiricos:so yeah if you want not so now you have your first thing right so like is there something about this um it's pretty good it's maybe a little overwhelming i don't know go ahead yeah Yeah, no, we definitely need some more cats in it. It needs more cats. I have a question while that's running there, Alexander. Is there a value to over here in your chat before you even start having a conversation about what you want to do and getting like kind of a spec sheet you could go drop in when you start? Yeah. We're doing this the right way, not just vibing, you know? I don't know if we can wait for that to finish or not.
57:13Alexander Embiricos:We'll leave it up. Let's switch back to you and let's see how you do it. Everyone saw the amateur version, the zero version. Now let's see the hero version. Cool. All right. Give me one sec to open stuff.
57:30Okay.
57:34What's there? Who is there? There was one more question we missed that was interesting, I think.
57:47Oh, I can't find it. Nathan Weber, who asked the question about bloggers, suggested I ought to try and start one, fail forward with me. That's a great way to learn. I've actually thought about that. And that's what I'm trying to do on this. I know a little bit more about this type of stuff than I'm letting on. But basically, you should totally write a blog where you are failing in real time and sharing your learnings with people because that is how you like by teaching others, it creates a pathway in your brain that basically allows you to learn it by repeating it. So it's like you can't really, one of the famous lines is if you can't teach it, you don't really know it.
58:31So you like sharing it with others kind of like reinforces the pathway in your brain that teaches you, you know, to learn from your failings. It's a great idea. You should totally do it. Cool.
58:43Alexander Embiricos:Can you put my screen up? There we go. All right. So let's say, so what I'm going to do is I'll show you three things. First, I'm going to show you just like some like slightly more realistic change that I might ask for what that might look like. Then I'll show you how we would do like sort of a longer, harder task. But I mean, we're not gonna have time to watch it run because it'll take longer. But yeah, about how to do that. And then lastly, I'll show you some fun just like screenshots of like us using codecs that I just I did this morning. From shipping linear integration. Yeah, that's awesome.
59:20Alexander Embiricos:So let's say that I am working on the Codex CLI app, and we have this command called slash status.
59:32Alexander Embiricos:And looking at this, it says that you should go here to see more about your up-to-date rate limits. And I kind of wish that this string was down here so that we could just tell people the status earlier. I don't want a user to have to read this text every time, right? Yeah, for sure. I'm going to show this editing the VS code extension just because that's where we've been working. But so I'm going to copy that and I'll go into VS code where I have opened the codex code. So this is I have dragged you can actually this is a fun fact. So you have codex down here. This is where it is in VS code. It's much easier to find.
1:00:09Alexander Embiricos:Yeah, so much better. Yeah. But anyway, so I can I've dragged it here just kind of happened to like it on the right side. So I'm going to ask it to make a change. I'm going to say when I run slash status, the command, the string, and I'll just quote it roughly. Let's clean this up slightly.
1:00:35Shows up at the top of the output.
1:00:42Alexander Embiricos:Let's move it to the bottom. so users you know don't have to read that immediately nice it's extra info and this is this is personal not not everyone i know writes this way yeah but i like to give the model like a lot of information about um what i'm thinking and i often will say things like i don't really know where it should be just like figure it out so like here i did and then like put in a new version where i had moved the string to where i want but the way that i use like oh i was like i don't even think about what i want i just kind of or like what it exactly should be i kind of think about what roughly i want um another weird habit we're on it is like when i'm talking to chat gpt i'll like often thank it or let it know what i did so it'll give me the helpful answer i do too the chat there but like it might give me three options and i use one of them i actually tell which one i used um because chachapiti has memory so like i was on holiday break and i asked it like where should i have dinner and it was like well since you're staying at this hotel you know you go here and since you like this kind of food so that we don't have um like automatic memory like that in codex yet but i'm like very excited for like those kinds of things oh that's awesome you're working on perhaps?
1:02:05Alexander Embiricos:You know, we have plenty. People are trying to. We can see Codex here is like working on this task. And so this is, this is, oh, this is going to take forever. I've been messing around with my git state. So basically what I want to show here is that Codex went and figured out what to do and then it made the change. So it removed this code that rendered that string from the top and then put it below. So that's great. Now what it's doing is it's running cargo test, which is a way to confirm, basically it's running tests to validate its change. Now it still happens that because I've been on like various different versions and be messing around, I think it's going to have to run a long compilation.
1:02:46Alexander Embiricos:So we might not see this finish, or I don't really want to wait. But it's making sure it works before it's telling you it's done, right? Exactly. And so like, you know, this is getting a little bit more real than our Vibe code of demos, where like Codex is actually going to go and do the work about its changes. It's going to take a little longer depending on your setup. But yeah, this would be where, you know, people are at the office. They're taking a break. They're, you know, walking around their computers running. And then you come back and check it later. Yeah. So like what I'm going to do here is I'm just going to let's just stop that.
1:03:14Alexander Embiricos:That's fine. Let's assume the test brand and then just run codex again. Oh, that's not what I want. This is just a, where am I?
1:03:31Alexander Embiricos:okay i've made messing around too much what's going on let me not where i'm oh i need to live this live life yeah there we go there we go yeah all right um so yeah we'll see if this works um so for context this is you're still editing the solar system no this is the codex remember i showed you like the status text in the in the codex This is like a real change that I actually want to merge. Oh, you're actually editing Codex right now. Yeah, yeah, yeah, exactly. So like I had it and I showed in the wrong place. And so I want to move it to somewhere else. That's awesome. Oh, yeah. Okay. The reason it's taking forever is just in case we wanted to demo a long task.
1:04:16Alexander Embiricos:I had Codex do a 20-minute task before this. And so there's a lot of code changes. And actually, this is very YOLO of me because, you know, hopefully that worked. We're taking 20 minutes of risk here. so it's great it's great yeah this should work i'm gonna wait like 10 more seconds and then if not i'm just gonna show something else okay we can also ask you some questions too if we need yeah yay the text is down here right so there you go earlier the text showed up here right yeah yeah and this is like something i want to merge like i i find it annoying that the text is this high it's like sort of polluting the ui for users i think it would be much better if the text is down here.
1:04:56Yep.
1:04:57Alexander Embiricos:Right. So I can send you guys the PR later today if I have time. But yeah, so it made the change. Now, it also so happened to have made changes to a bunch of other files because I had it write an entire Python SDK. I don't have yet. And I forgot that I had done that. So it took a long time. So let me show you how I got it to write the Python SDK. Yeah. Yeah. Okay. So not this. So I'm just getting finding I have some links for myself. Oh, yeah. Okay. Which I think I already have this open. So we actually put out a guide for this in October, which is the idea of using plans.md file. Similar to agents.md, but for planning specifically.
1:05:43Alexander Embiricos:Exactly. So basically the idea here is that if you want someone to, like the same analogy works for like a human teammate, if you want someone to go off and do a lot of work, you might want to agree on what work they're going to do. AC Min, I didn't write any of that code.
1:06:03Pull that, if you could, Vinique, or Corey, pull that quote up just so people know the question. Yeah. Oh, sorry. Yeah.
1:06:09Alexander Embiricos:So how much of the actual code did you write? None. Yeah, none. I mean, but this, you know, there's certain tasks, right? This, like this string change that I, you know, or moving a straight label around is not a hard task, right? So I'm like, yeah, but that, you know, I don't have to write that code. There's many tasks where you're going to have to write a lot um you know i'm i'm on the product team so you know i try not to like write code for complicated systems uh that i'm going to have to maintain so if i am going to like actually use codex to like land them like something in place i usually try to make sure it's like something well tested and then i'll like i'll probably write like the definition of this the tests myself or like work go back and forth with codex on that and then the actual implementation of like some clean thing, maybe that part I won't write.
1:06:52Alexander Embiricos:Yeah, that makes sense. But okay, so going back to plans, basically, the idea is that if you can collaborate with a human teammate or an agent on what should be done, then you can go back and forth together. And at the end of that back and forth, you can you can like, say, yes, this is this is what I want, right. So one way that you do this if you just want to try this is you could go into let's just do something simple for now. Let's say we have this solar system thing that we've built. Let's look at it. Solar index.
1:07:36Alexander Embiricos:Okay, cool. So I don't know. We have this solar system thing. Let's make the timescale faster. okay that's still very slow. But let's say we wanted to sort of upgrade this thing. I could just give a prompt but if I want to have it work on like a harder longer task what I might do instead is I might say let's work on the solar prep together.
1:08:03Alexander Embiricos:together first write a plan in plan.md
1:08:12Alexander Embiricos:explaining how you think we should I don't know what we want to upgrade let's we should add moons right so this would be like a very sort of like if you're just getting started you could just do this you don't have to do anything more complicated all you're doing is asking like write a plan into plan.md. And so this is a thing that a lot of people do. And this one guy at OpenAI, Aaron, he is like a pro at this. And so he no longer just says write a plan. He actually has a lot of things that he wants to see in his plan. It's like a plan template. And so I'm not going to go over all of his template, but he has a lot in here.
1:08:48Alexander Embiricos:He's like, make it self-negotiable, self-contained. I want a certain kind of formatting. I want milestones in the plan. I want design decisions, you know, and he sort of gives an example. I want to do's in the plan. So you can do whatever you want here. Like you could just, so I'll show you how to just copy his, if you just want to use his, the recommendation is just start, start easy. You don't need any of that. Right. Maybe, or maybe what, what someone could do is they could start with the plan. .md like template, and then they could open another window and make their own plan and test both of them.
1:09:22Yeah. Yeah.
1:09:22Alexander Embiricos:His is, his is like really robust. That's the most robust one that I've seen. So, you know, you could start there or not. But anyway, so if I like this plan, I mean, maybe we want to change something. Right. Let me see. Hmm. You know, maybe I don't want any controls here, you know, actually for the moon. It's like, do not list. I'm changing something. Right. Do not add controls for the moon. Moons just populate them by default. Right. So, OK. So I've adjusted. I've gone back and forth with the plan. I might even ask it like, hey, update the plan yourself here. But anyways, okay, great. I made some tweaks.
1:10:04Alexander Embiricos:Do the plan. So this is a very common workflow we see where you kind of have a generated plan, ask some tweaks and do it. And like, again, we're doing this like very quickly here, but we use this a ton. Like the Sora Android app, I don't know if you guys have heard this story just pulling up the number we built the entire Android app in 18 days no joke and it was 4 engineers, 18 days to get a functioning app that we were testing and then 28 days to just fully ship the app and so the way that they did that was we had an iOS app already and so they would run codex where it could see the ios app and the android code base and they were basically figuring out okay here's a chunk of something i want to do let me ask codex for a plan okay go back and forth on the plan and then they would send it off and it would go up with the plan so they would do what we're doing here but just like a little more advanced probably something in between this like very basic thing that we did and the full plan that md version of it um but you know when you when you decide that that's how you want to work so you can do this quick workflow and if you like it you can add an agents.mv file and you might say something like um so here's the thing that folks can do we can let's say we like this version we could make it we could just copy this in i'm just pasting in the the which you guys can share with your viewers and then um what i might do here is i could then say something like this in my agents.mv i'm saying hey when you're writing complex features or a significant refactor use this approach to do the implementation right and so now whenever um whenever i'm working on something i could be like hey uh you know create a plan and show me so i can go back and forth or i can say um just go ahead and use it but i want you to use this sort of structured way of thinking about it because that's how I want you to work.
1:12:13Alexander Embiricos:And this 17-minute task that ran before I got in was one of those. So my prompt at the beginning, just to give you an idea, wow, that's okay. I don't even know if I can go all the way up. What was this task that you were trying to do? Basically, what I said, this is my prompt, actually. Can you read it? Is it too small? Yeah, using execplan.md, write a new Python SDK on the JS SDK for codex. Yeah, so the Python that you were talking about earlier. Yeah. And so this is, this is a bit like sort of abstract. I was in a rush trying to figure out if I was going to show you a long task. I just did this.
1:12:45Alexander Embiricos:If I was doing this in real life, I would first just say, you know, using exec plan.md propose a plan for how we're going to do this. And then we would stop and we would look at it. And then I would say, okay, not right. Yeah. And I would tell people to remember to be patient. You know, some of these tasks do take a little bit, but the fact of the matter is, you know, just sit back, let the magic happen, go drink coffee and watch it work. yeah and learn by watching it even yeah so last thing i want to show you guys we can see if this finishes this uh the solar system change but i just thought it'd be fun to show you some cool cool screenshots of like if you guys get you know going with codex yeah what you might end up with a team so um when you're really going with codex you can start connecting it to your various tools and so for instance uh last night at a i don't know if it's going to show it doesn't show the timestamps good at a slightly toxic time.
1:13:39Alexander Embiricos:We were like looking at this integration and we wanted to make some changes. And so, um, you know, I filed a linear ticket saying like, Hey, we need to update the prompt. Uh, this, I know it's a bit small, it's going to work. Uh, we need to update the prompt for this thing, for this, you know, uh, this, this feature, uh, so that it's better at answering questions. And so what I did is I went to linear file the ticket. Notice I actually wrote out a bit more of a ticket than I would have in the past because I actually want the agent to understand what I want. So I wrote up and then I assigned it to Codex and then Codex made it, got to work on it and then created a PR.
1:14:16And then I merged the PR.
1:14:19Alexander Embiricos:I tweaked it actually slightly, but for folks who don't know, pull requests is where it then merged the code back into the GitHub folder. Yeah, exactly. So like, and linear is a tool for tracking tasks, right? So if you have like a tool to use for tracking tasks you know you can think about okay i just want every task to be first looked at by an agent before i look at it so this is super powerful uh we just launched this integration yesterday um another cool thing is like this is an example workflow of um of an engineer on the team um and it was just cool to see how he works now so we were in slack so i've kind of taken these screenshots out of context just to not show too much here but like We're in Slack and he was trying to debug some problem in the macOS app.
1:15:00So he asked the question and someone shared a helpful idea.
1:15:05Alexander Embiricos:So he said, perfect. Let me give it a try. Notice me. Can you try? That's amazing. Which is amazing, by the way. This is like, this is the future, right? I got to go after that. Yeah. Eight minutes later, Codex had an answer and ran these tests like I was showing you. and then she's like at person with the idea here's the pr right and then later uh that person was like huh i don't know like i think actually we need to sort of do this in this like other way and that's what was like at codex can you change it to do that other way it's sort of like the equivalent of let me google that for you do you remember that website yeah yeah yeah yeah we're like the guy's like okay let me i'll i'll i'll have codex do it for you yeah exactly so you know that's a mindset change i'll bet too getting to that point where you just remember well let's just have this do that let's just have this do that yeah um and then one last screenshot that i took literally this morning but it's kind of like when you really get into this idea of like i'm just going to maximally accelerate myself um working with agents you kind of start doing weird things so like like i was saying like you know i use atlas chashpy atlas my browser so everything i just type in there is just like you know gets ai help um the right thing just happens you know if i want to search or if i want an answer um and so i kind of have started treating even like my terminal this way so this is this is actually me this morning i just typed into my terminal like are there work trees for this repo which that's not a valid terminal command right and then i was like it's like what do you mean so i was like oh i mean codex are there work trees for this repo.
1:16:46Alexander Embiricos:It's like, what do you mean? And then it answered my question, right? But this, this you can see here, this is a kind of task where I'm not actually writing code anymore. I'm just doing stuff on my computer. And like, you know, if I needed to, I could remember that the command is git worktree list, right? Going back to our earlier topic, but there's just like no reason for me to know that command. Well, and this is where I think like for me personally, right? So I started my software engineering experimentation journey back in 2022. And it was right when, right before ChatGPG came out because I was suddenly in charge of building spreadsheet templates at work.
1:17:24So I had to learn all these crazy Excel commands. So it was sort of perfect timing that I was like, as I was learning all of these commands, which was my first foray into really seriously programming, if you could call that programming, the engineering mindset, let's say um like then all of a sudden chat gbt came out and now i could like not i was i was torn whether i memorized the commands or whether i just like work with chat gbt and it's gotten consistently better and better and better at that where now it is you can basically do so much with natural language um but terminals have always scared me because i don't know these like little like get commands here or all these little cds like what like how do i speak computer so what you're doing there is you're saying codex dash dash and you're asking the question and then it
1:18:10Alexander Embiricos:will basically run those commands for you yeah sorry thank you for explaining i totally agree with what you're saying and yeah it's just this is just a shortcut um you know if i let me just close this so if i want to use codex i can type codex yeah um and then i can ask you you can type directly in there right yeah if i just do dash dash hello you know that it's just sort of that works the same way. That's cool. I like that. Yeah. Yeah. So that's all that's happening. But yeah, what you're saying totally makes sense. And actually, I think an interesting thing is that people who are not coding all day, like PMs, or maybe like folks like you when you were early in your journey or whatever, often we are very experimental.
1:18:52Alexander Embiricos:So some of the best AI prompters actually are folks who are like semi technical in their day to day rather than like always coding because they're used to like not having the time to like go super deep and they're used to like figuring out how to explain what they want with just the right amount of ambiguity right so that it's like yeah whoever is actually doing the work will like actually come up with a better answer than you could come up with because they know what to do way better than you right yeah totally totally instead of telling them something specific that may or may not be the right thing yeah yeah I want to ask you a couple of questions specific to basically the problems that I've had with coding with AI.
1:19:37And you might have good answers for these, you might not, but it's all good either way.
1:19:43Alexander Embiricos:Cool. By the way, sorry, I just opened this to check if the moons are appearing. Let's see it. Let's see it. You know, who knows? I did say remove all the controls, which might've been the wrong. I was trying to figure out what I could change in the plan. So we don't have much we can do, but okay, what is that? I don't know what that feature is. A moon just shot off. Whoa. There's a moon rotating around. So, you know, kind of works. All right. Okay. What are your questions? Okay. So, like, let's say we're, you can keep that up. Or, yeah, I feel like the solar system is a good example. So, sometimes what will happen when you're vibe coding something is you'll code yourself into a hole where the AI gets confused and can't backtrack.
1:20:20And, you know, or, you know, worse is when you're using a vibe coding platform and it's like, I just fixed it. And then you look at it and you're like, no, you didn't. You absolutely did not fix it. And then it just keeps lying to you and gaslighting you. What are your strategies for getting out of that situation? Does that happen in Codex? Or what would you do if you were in that situation?
1:20:41Alexander Embiricos:Yeah. So I think one thing is it's like a lot of the code that we tend to be writing, it's more like, if you want to use the term, it's more like vibe engineering. right so like people are people are like reading the code and like looking at it um and so unless we're like just like prototyping or something usually the workflow is much like more human in the loop than just like go off on a prompt so i think oftentimes you'll see things where like people will ask for a plan and then they'll like work on the plan you know like together or maybe you'll ask for the agent to just like build the entire feature that you have in mind actually you know we we tend to like doing something we call best event where like in codex cloud we'll like run codex like maybe say four times on the same idea same task or even locally we might do it in work trees or like multiple repo clones i can show that but i think it's a bit advanced maybe i won't show it but the point is that you'll ask the model to solve the same problem four times and then we might look at it and be like oh wow i didn't even think about doing it that way or now i understand that there's this like other way to do it and then then i might just like start from scratch and be like cool okay here's exactly how i want you to do it right because we've looked at a plan or we've seen something and we'll like do often like and then like create aggressively to create new new chats if you find the agent going getting stuck um and ask for like very specifically what you want personally i have found that i don't need to create new chats nearly as often as i used to like i used to create new chats really aggressively um because i would find that the the model would like not do super well as context got really long but recently actually I found that like the longer the chat goes there's you know off the the model has more context on what we're doing right and so I just haven't had that problem as much do you say that compaction compaction um will help with with that like let's say you get to a point where you're like you do have errors and you're still trying to debug it like will the compact how does the compaction impact that yeah it's interesting at least from most people that I've spoken to and myself like compaction is like very specifically useful for like really hard tasks but like in sort of like day-to-day more like interactive coding or like you know vibe engineering you know more thing then um I don't think you're like super likely going to like run into like that context window I think a lot of us are naturally kind of like compartmentalizing like each cycle of what we want to do yeah yeah my second question I think that's great answer i think that's that's really helpful which is essentially you you want to do multiple versions and you know maybe do it more like incrementally when when you get to a point where you're stuck and test and test things maybe separately and then you can bring that back in would you say that that's right yeah but i think what i would say is like you want to be thoughtful about like basically you know we have these amazing like ai superpowers now and then i think you just have to decide how you want to spend your time.
1:23:39Alexander Embiricos:Right. So like, I think the first question is like, what are you trying to achieve? Right. Are you trying to just like learn something with a prototype? If so, then just go off. Don't read any of the code. But like also like, you know, if the code becomes like hard to maintain, you know, which will happen nowadays with where models are at, it could happen. Then you don't really care. Right. Whereas if you're like, if we're building codecs, like what I just showed you, then, you know, it's going to be much more like, yeah, you're making like very careful changes. you're looking at them, you know, carefully, you're going slowly, maybe you sometimes live code something, but like often you will like, you don't expect to just merge the vibe coded thing.
1:24:15Alexander Embiricos:If it's a hard, complicated time, like I expected to merge, like when I wrote that problem, like, this is for sure going to work. Right. But the heart, the more complex the task is, then the more you're not expecting to just like merge a vibe coded thing you're expecting to like use that to massively accelerate yourself, but you're still going to be involved to make sure that like architecturally it sounds so you don't get stuck. I would say that's true. If you're using it, for writing or other tasks as well where it's like you can't expect it to just go off and write you something like you you do have to be prepared to like incrementally edit and and tweak it um we had a question in the chat that was but on that point how much knowledge do you need to know about programming to get the job done in the end which is part of my second question which is if you are trying to deploy something we've seen a lot of really cool vibe coded demos but how do you actually like how does your team when you actually launch something like say you wanted to launch the solar system app what what do you do what do you do how much knowledge do you need to do it and how do you launch it yeah um so i mean the question is like yeah how much knowledge do you need about programming i mean this is a common question it's like is it worth learning software engineering and computer science and i definitely think it is worth learning computer science still um yeah and i still think if you want to build something like really meaningful then you you still want to understand like the underlying systems that you're going to be working with and you still want to be an expert the difference is that you just are a massively accelerated expert um you know or as i think you know i don't think we're going to reach the point where it's i don't think we even want to reach the point where um you know there's no point learning computer science anymore because learning computer science is fun it's just like right learning computer science was like let's say like let's go back like five years like what learning computer size was five years ago is different than what it was 20 years ago right yeah or 10 years ago and so like well it means it's evolving but like you still you should still definitely be learning it and maybe the answer is trying something like you know if you're working in these and you're you know vibe coding vibe engineering your projects you know setting aside even time each day to to make sure you're really understanding processes and maybe architecture and going back because i I feel like we're going to have a lot of people kind of bootstrapping this over the next few years, you know, and, and learning along the way.
1:26:39And, and, and maybe that's just, maybe that's, I don't know that I would say it's equally valuable, but, but I feel like you might walk away with something very different than what you would walk away from in a traditional university setting. Just by the nature of it being so much more very, very hands-on day one.
1:27:00Alexander Embiricos:yeah I mean that I I think I agree with in the sense that I think that well some people are great at just like reading textbooks and learning I am personally someone who is a little bit in between but I I learn best when I just have a problem that I want to solve and I like get all my knowledge just in time so I'm like a bit outside my comfort zone and what I know that that's the best way for me and I think that like there are many places where that is hard like you know universe I studied both mechanical engineering and then computer science in university and if you want to do mechanical engineering stuff you need like um you know machine shop you need all this stuff it's like really hard and then the machine shop is dangerous if you don't know what you're doing so you need instruction right it's like hard to just start doing things there's a little bit more of a boundary whereas now with software it's always been really easy to start doing things with ai it's like even easier to start doing things um and so i think that that is a real like amazing progress for people who learn by doing yeah so yeah that's a good call yeah i think that's evolving yeah it's i would say it's also not just engineering and software and it's it's you know that could apply to anything i mean you could you could stumble your way through you know understanding literature just as well uh or understanding i don't know sociology for example um and and have a different approach and maybe it's all of those approaches is where success lies for for people is you know being more of a generalist yeah perhaps yeah i mean i think being more of a generalist maybe but even if you want to be an expert in a certain field i think you can also be like very accelerated right like so we yeah we use codex at openai uh you know i i'm a i'm a generalist in my role right so i'm doing like sort of the easier side of of code changes There's like really easy ones for like what's getting deployed, medium ones for what's getting prototyped, you know, just trying stuff.
1:28:58Alexander Embiricos:I usually had a ton to answer questions about the code base so I don't have to bother engineers. It's like me and folks like me on like product teams are using AI and codex a ton for those kinds of use cases. Yeah. They have engineers, right? And they're like more specialized of a function and they're using codex to just be like massively accelerated. Like I gave the example of like Sora, like Atlas was like some of the top codex users internally at OpenAI. I mean, like basically, you know, nearly all of technical staff at OpenAI is using codex. And engineering, it's like a massive accelerant for us.
1:29:31Alexander Embiricos:And there was a time where codex usage was like growing, but it was like went from like a little bit under 50 percent to like over 90 percent. I forget the exact months, but it was like earlier this year, like spring or summer and then like fall. I was going to guess May. There was a moment where it just felt like all of a sudden OpenAI was cooking with gas fast on everything. Yeah. And then all the codex models after that were really big for us. Yeah. And we improved the product a ton as well. And so there was a time where we were just noticing, wow, the people working with codex are making 70 % more PRs per time period.
1:30:05Alexander Embiricos:And that's not like that's the best way to measure engineering productivity, but it was an easy thing that we could just pull. um and so we saw productivity go up a ton there um within engineering and like then projects like atlas which we just launched have been like there's like amazing anecdotes of like engineers saying like yeah what used to take like maybe like three engineers three weeks is now one engineer one week uh wow so it's we've been working so much faster thanks to that um and then you could take like arguably an even more specialized function like research and like researchers are using codex a ton as well you know as part of like how they're doing training runs uh they're doing a lot of interesting things like also like across research and design and other functions to like analyze and like understand what's going on like previously would be quite expensive to write code for like a one-off analysis but now you can or you could write like a custom viewer for like you know to evaluate like how the model is doing and we're even starting to experiment with like using codex like as part of like deciding how to handle like training runs as they're ongoing going oh that's smart i don't think this was all in response to like i don't i think like people like me i like i'm thrilled to have these tools because as a generalist it's very useful for me um but i also think that if you're a specialist you can actually use it to just like go way faster down your specialty as well it makes sense um the two other barriers that i face right the first one we covered which is what do you do when it gets stuck in this loop and you don't know what to do um the second one is um how do you actually take like a a vibe coded thing and then get it out to the world like you mentioned Sora you made that app in 18 days like obviously you guys are professional engineers but if I wanted to make an app and get it on the app store what with Codex what would what would be the process how could I do that I hear people doing it all the time actually um awesome again I think it starts with like what do you want right like Sora we you know is a very serious project from us.
1:32:01Alexander Embiricos:So it was not Vibe-coded, right? It was, you know, again, it was Vibe-engineered. So a ton of the code is written using Codex, but there were humans, like, highly qualified humans in the loop the whole way through, right? Like, really great engineers who, like, know exactly what they're doing and you've been using AI for a while. So I think if that is your goal, then you should be treating Codex as a massive accelerant to your own productivity. But ultimately, like, you are still... treating this like an engineer, you're being very thoughtful about architecture, you're defining that before you get started on what you're building and telling the model what the architecture is, using tools like planning to make sure that it builds in that way, maybe being more modular with code than before so it's easier to test and easier to trust units of code without having to make a ton of changes or even arguably review all the code super carefully if you trust it.
1:32:54Alexander Embiricos:You kind of want to understand the different modules yourself so that you trust it and you know how it works yeah yeah um you know you definitely want to be doing things like writing and like typed languages actually i really think like even compiled languages are kind of almost there's like a potential comeback for them um like codex is in rust um because it's like gives you even more confidence that um the code is going to work if it passes its validation conditions um you know you can you can you enable code review which is actually something I can show as well briefly. More stuff of Codex just being a teammate.
1:33:33Alexander Embiricos:One thing we have is we have Codex set up to work in GitHub. Here's an example. I don't know if you guys know him, but he's head of DevX here. He made his PR. You ready for me to bring that up or you want me to wait a minute? Oh, yeah. Sorry. Bring it up. Okay, cool. I just wanted to make sure before i hit the button before you reveal secrets yeah no thank you so um yeah so he was he was writing this pr for gpt oss and this is the this is the kind of thing that if i was like code reading this i i'm you know depending on how sort of rushed i am like i don't know if i would read this like super carefully you know but you probably should yeah um but you know the beauty is like codex doesn't get tired so you can send codex in uh and they'll just like automatically it's like hey i think you should like you know this is an issue this is a real issue here's another one that was like quite interesting um where uh let's see what was the issue yeah it's like hey i found some issues the streaming code has some issue here i actually haven't read this yet but basically then we ask you know is this useful react thumbs up and so he comes up and you know homosh is like yeah good point codex please fix and then like um and how how did how did you get to this point where you called the the reviewer how do you so yeah so you can set this up in your codex settings.
1:34:50Alexander Embiricos:So if you go to Codex web, you go to settings, then code review, and you can turn it on super easy. You just need to connect to GitHub. You know, a lot of the time, if you're making simple PRs, you'll just get the thumbs up emoji. It is actually one of the most popular, like sort of integrations that we've built for codex, where again, codex feels like a teammate. So bringing it back to like, if you want to deploy codex code, like deploy codex, um lab engineered work like having code review is like just an excellent way to like increase the confidence like when we when we first rolled this out we were actually quite nervous that people would hate it and just be like i don't want bots telling me that my code is bad so like to be the codex models to be great at code review and we train them to um have a very low false positive rate uh so basically what that means is we were like hey for now like you single opinion you have about the code.
1:35:45Alexander Embiricos:Just flag when you think it's a real problem, like a serious one. And yeah, it's turned out to be a really popular feature at OpenAI and also outside of OpenAI. This is related to my third barrier, which is security. Like if I'm a total noob and I know nothing about coding and security, how do I use Codex to make sure that my code is as secure as possible? Would it be through this process? Yeah, I would say connect to GitHub and have Codex code review, look at your PRs. You can also do this in Codex itself. So I could pull up Codex here and I can type slash review and then ask for, you know, basically this will just run the same process.
1:36:27I could be like, yeah, like please review, you know, the most recent commit, right?
1:36:32Alexander Embiricos:And it'll go over to it. And we also, we announced recently a beta of something called Aardvark, which is like custom designed for security and security team. So I'm really excited for Aardvark to come out of beta and just like become a standard part of how we make sure code, whether or not it's written by humans or AI, is like maximally secure. Would that be a separate agent from Codex? Would it be inside Codex or is that still to be determined? Well, I mean, like it's separate right now. And, you know, if I think about who's using it, like it's like mostly for security engineers, But I also think that the capability of like just like a really good security reviewer is probably something that every engineer should have.
1:37:15Right.
1:37:16Alexander Embiricos:It might be a different form factor. So I could imagine we're still figuring this out, but I could imagine like just us making sure that Codex is excellent at security review. But that doesn't mean we don't want a product that's like specifically amazing for security engineers, which is like a smaller population within the company. Right. Yeah. So, you know, that's a population, but still crucial in the world for sure. You know, especially with, you know, when you think of things like we've talked to the guy from Threat Down recently about the patch gap and how much quicker bad actors are able to attack what you're doing.
1:37:53It's there's not a months and months delay anymore like there once was. It's just it just keeps getting quicker. and i guess if you're throwing something out into the world that is we'll call it loose uh you could have real problems yeah it's been pretty cool i think with aardvark we're seeing that like um not only are we
1:38:16Alexander Embiricos:identifying issues uh especially when people newly deploy it like we identify issues like wow i can't believe we didn't know about this but then also the way the way that the the beta works right now is that when it identifies an issue, it attempts to reproduce the issue and can demonstrate that the issue is real. And then it will actually send Codex in the Codex Cloud to go fix the issue. And then it can test if the issue just no longer reproduces. And so I think Aardvark in a way, it's because it's a very specific thing that it actually uses Codex for many things under the hood. Oh, wow. It's building a very specific product for a very specific audience.
1:38:52Alexander Embiricos:So they're speed running it um and uh yeah i think it's a bit of a glimpse of the future oh wow so by the way so all this to say like this is i think for what you should do if you want to um use codex uh for something that you know a production like series project i think if you're using if codex for something that's like you know much more prototype experimental throwaway then i think if it's just like a single player tool that doesn't have any user data in it like you could just put it out there like it's fine you know yeah put it up a lot of like next as projects coded with codex then you can like deploy them on a tool like for sale or something yeah yeah that's the one place where you want to have considerations is just like if you're asking users to upload data then you've got to make sure you do right by making sure that it's secure enough for whatever you're asking for yeah yeah that's that's a can of worms that i'm still figuring out And, you know, some of it, I think, is just a matter of being as someone who is not an engineer, but really enjoys tinkering in here and learning as I go.
1:39:59And I've learned so much over time. There's still so much out there, but there are certain things that I'm just kind of naturally gun shy about. And one of those is security fears, causing some problem for, you know, maybe people who want to use the app or something. And I really love that you're taking an approach to that. And Codex has just come a long way. And what's it been now? Has it been a year?
1:40:33Alexander Embiricos:I forget what month we launched the research preview in, but it's one of the M months. time in ai is a little odd you know it's kind of it all blurs together i still can't believe you just had chat gbt's third birthday yeah yeah i actually had a weird question about that i'd love to get your take on this um so we talk a lot about fast takeoff scenarios with ai where like you know the ai is like recursively improving itself and then that leads to this like intelligence explosion um but i wonder if such a thing could happen with people's skills in ai so a lot of times i feel like trying to keep up with the news like there's so much stuff coming out and people seem like they're accelerating way past like my ability or other people's ability and but then i was also thinking we've seen cases in in history where there's like entire continents or countries that leapfrog certain technology waves do you think that there's like something to this idea that there's like a fast takeoff for people who are really good at using the ai tools and could someone go from like, you know, like almost like leapfrogging the traditional coding, like pathway to if they jump on that train and they're able to like accelerate themselves, they maybe skip the need to learn some of this stuff.
1:41:44Alexander Embiricos:Yeah, like, okay, so this is something I care about a lot, actually. So, you know, our mission at OpenAI, I might not get the exact wording right, but our mission at OpenAI is to ensure that AGI benefits all humanity. And I'm, I believe, you know, we're going to have amazing AI capabilities and AGI and we already do already have amazing AI capabilities though. Like, like, so the AI part is, is already so powerful, but actually there is still a ton of work to just bring those capabilities to all humanity right now, right? Like a lot of new models coming, but like, even if we didn't, I could just as a product person work for years, like to help make sure that people get the benefits.
1:42:26And, um, you know, if you think about
1:42:31Alexander Embiricos:an AI adoption, I think there's like broadly like three phases. The first is where we have these like incredibly powerful technologies exposed to you in very simple ways, right? And so, you know, chat, you can just go do anything and like codex even mostly for coding, but like it turns out coding is a good way to use a computer. So you can actually use codex to do like non-coding work. And the product doesn't help you do that, but you could do it if you know what you're doing or if you were creative with how you prompt. And so that's kind of the first thing, this first phase of like, it's open-ended and it's powerful and you should decide what you want to do.
1:43:07Alexander Embiricos:And this first phase puts all the onus on us to like be online and like read about stuff and like take time, you know, when you're not busy to like try using AI. So it's like, this is a really fun time for tinkerers, which I guess is like you guys and like me. Yeah. But it's also hard, right? So then there's this next phase, which I think we're just entering now, where actually what we want to do is we want to make it so the tinkerers can help the non-tinkerers. Right? People who are too busy to try random prompting ideas should be able to just set Codex up or ChatGPT up or any tool up so that everyone else can use it.
1:43:49Alexander Embiricos:So we're actually working. I can show you a feature that we haven't shipped yet, but it's like we're open source, so everything is in the open. I can show you something that would be amazing. Sounds great. Yeah. Yeah. So that's the second phase, right? Where you don't maybe like Grant figures something out and then Corey just gets it. And then Corey figures something out and then Grant just gets it. You don't have to figure it out. And then the third thing is where you no longer have to think about using the feature or using AI. AI is just like a teammate and it's just helpful for you without you having to do any work.
1:44:23Alexander Embiricos:And I think if we want to have a takeoff in terms of like humans getting value from AI, we need to be at this phase where you don't actually have to think about using AI. It just AI just helps you. Right. Yeah. I actually think that like, you know, you were just showing Cursor, which has an amazing top complete model. right i was using vs code which has a great tap completion model and um you know that tap completion is actually a really basic example of a feature that existed before ai but it's a feature that you don't have to think about ai it just like does stuff and you're just like yes no yes no right so like in my mind like i think about like how do we get to a world where i get to tinker so we're not overly prescriptive but i get to tinker my tinkering is helpful to all my friends and all my team and then we can just turn stuff on that's helpful without anyone having to try so yeah you know to show this over a screen share this is just an illustration but uh you know we can show the pieces of this let me know when you're good no oh yeah you can put it up sorry yeah okay so is it okay that the check comments are here it's fine right yeah yeah that's fine i can see him yeah so like this is this is the open-ended part, right?
1:45:39Alexander Embiricos:This is super powerful. I can ask us to do something totally non-coding related or powerful in coding, but it's, you know, it's pretty open-ended, right? Yeah. You probably have to read Twitter to know what to do with this, you know, if you're not, or, you know, you could, right? So that's the one extreme, right? The other extreme is like the, you don't actually have to think about using AI. It just is helpful automatically. So actually I did show you something like that already, which was here, right? We've decided that when someone at OpenAI submits a PR, Codex just goes and reviews that PR, right?
1:46:19Before it even gets to the person in the queue. Yeah, exactly.
1:46:23Alexander Embiricos:And so at this time, like you could decide to do this. This part is still like, you still have to decide to do this, right? But we're starting to do things of value immediately. And so something I think about is like, where are more places where we can do this? And this isn't like, let's just go plaster AI everywhere. I think we actually need to be really tasteful in terms of how we do it, because we want to be in a world where like, you know, we are paying a lot of attention to like the sanctity of human attention, I guess. So like, you know, we had kind of a funny number of debates about how when their review looks good.
1:46:50Alexander Embiricos:What should we do? Right. A lot of people were like, hey, I would like a comment that would be more clear. The problem is a comment sends an email. So, you know, we ended up with a thumbs up. You could argue if that's right or not. But the idea is like human attention is really important. like don't waste it. And we like emojis. That's the other extreme, right? So then the middle extreme is actually something that actually a bunch of the other CLIs have already, but we're starting to really lean into. And so now we have this experimental feature, enable is how you turn on experimental feature called skills.
1:47:23Alexander Embiricos:Oh yeah, this is cool. And so one skill is the OpenAI job skill. And so finding open worlds at OpenAI. And so partly I wanted to show this because we're hiring and I want you all to know. But partly, this is like a weird thing to use Codex for, right? Like Codex is a coding tool. Like you just told me all this stuff about how to work the repo. Like why would I use it this way? But, you know, contrived example, but I could say like is Codex is the Codex team hiring. um but this is very fresh this is like a few days old but um you know the idea here is that this is that middle ground where like someone on the team can figure out this codex skill this open-air job skill would make sense and then they can make that available to like other people or like the broader hey look at that i love it wait sorry a bit experimental i'm gonna do this make it work yolo um yeah yolo it's a little bit of a pinky name um but yeah okay anyways you get the idea um yeah that's kind of the picture we're trying to paint it's like a world where you can use it open-endedly there's like very deep configuration that you can push to your team um skills just being a small part of that uh and then you can like automate things so that they're like whatever you want like whatever value you want is just provided um this this is a minor thing i was hoping to touch on this but we're getting close to time but um as far as tools and connectors go what's what is the difference between like a connector and a skill in your mind um i know they are different but but yeah so so we'll probably refine this because you know we haven't actually like published like our thinking on it there's just like some prs in the open source repo yeah but my thinking is that like connectors are basically ways to use like other tools or like you know, gather context.
1:49:18Great.
1:49:20Alexander Embiricos:A skill, the way that I think about it is actually more of a set of instructions that create a capability. That capability might include a connector. So an example would be, we might have a connector to say Sentry, right? And so you could use the Sentry to like identify crashes that you want to fix. But like the truth is like, if I navigate to the Sentry website right now, I will have to do a ton of work to click around to get to my team's crashes. And then I'll have to like click around some more to get to which things I care about. Right. So there's actually a bunch beyond the connector that I might want to prompt.
1:49:58Alexander Embiricos:And so you could have a Sentry connector. You could have a Sentry skill like vaguely, but what I might really want on my team is I might want a skill that is like identify century crashes related to the most recent release specifically pertaining to the code that I wrote. And then like plan, investigate changes and plan the fix. And like, for me, all of that thing that I just said is like one skill. Right. Okay. So you're sending it to look specifically for your stuff. Most recent. Yeah. Yeah. For example. Right. Or maybe if there's an on-call skill, there could be a century on-call skill. and that one actually produces like an output um that uh you know uh lists everything by engineer and like sends an email right so um yeah at least that's how we're thinking about it right now um you also in the plan plan.md uh there would you address connectors there like for instance if i was creating a project that's going to talk to like one of the things i do is i uh do video game design on the side.
1:51:01So like I have a Godot connector that I'm working with, which is open source game tool. So would I address whatever project I'm working on? Like, oh, use the Godot connector in my plan.md or then do all of that there? Or how do you do it at OpenAI?
1:51:19Alexander Embiricos:So this is maybe taking one step back, but the way that I would think about making codecs successful, I didn't actually have anything written down for this, but we could just... That's okay.
1:51:32Alexander Embiricos:So, you know, the first step is just like install it in your favorite tool. Right. So like it could be like your IDE or your terminal. So that's step one. Right. Then step two is just test it. Like you don't need to overthink it. And then I think from there you want to configure. And the way that I think about configuration is you want to do first you want to set up what I would call the inner dev loop. Then there's the outer dev loop. and then there's like automation. So the interdev loop is basically like, does Codex know where you keep your context? Does it know how you like to write code? And like, you should just test it because it probably will just infer your style, but there might be things about your style that are hard to infer.
1:52:20Is it able to run tests and so forth?
1:52:24Alexander Embiricos:And so a lot of that is just, you know, you can just tell it in agents.md, okay, this is how we like to work, right? I'm really excited for skills because a lot of skills, there might be like, maybe there's a skill that you use in your interdev loop using Godot, right? And so you might make a skill that says, yeah, this is the, I don't know. Sorry, I'm not familiar with Godot, but like, this is how to use it, blah, blah, blah. This is when, and then in your agents.md, you'd be like, this is our interdev loop. And like this part here where we validate our changes, you need to use the Godot. Right, right.
1:52:57Alexander Embiricos:Excellent. Yeah. That makes good sense. You know, I love the agents MD idea as a whole, just the idea that you can be in there and be like, here's, here's how I like to do things. Here's what I want, you know, whenever possible use, use this repo, use this, the ability to go in and really kind of control that and have it move forward and act taking those steps for you is really awesome. And it kind of takes away the idea that maybe it's a cookie cutter approach for every engineer or vibe coder out there. I noticed in talking to people just this week, one of the things was that a thing engineers and writers have in common is that, you know, they have a style.
1:53:44They have the tools they like. They have the approach they like. They have the libraries they love, the style guides they enjoy. and I think this is a really cool way to be able to hold on to some of that creativity as a developer as well.
1:54:00Alexander Embiricos:Totally, yeah. I mean, the idea is to give you as much control as possible, right? This is just the thing to accelerate you. Awesome. Yeah, so that's the inner dev loop. The outer dev loop is like, where do you track your tasks? Where do you brainstorm your ideas? What do you do after you've written a PR? Yeah. You know, for request, like how do you review it? So you could like set up codecs and today a lot of people do this agents.md but i think a lot of this is actually going to make more sense with skills um you know like okay here's like a code review helper skill that like automatically addresses all code review feedback that you get um wow so that's the outer dev loop and then automation is the idea of like okay i tend to want it to do these things all the time let me set it up so today we have the codex sdk which you can use and you can just build automations around an sdk is basically a really convenient interface to use codecs programmatically.
1:54:51And, you know, that's something we're thinking
1:54:53Alexander Embiricos:about too, just how to make that even more out of the box for people. Excellent. Awesome. Well, we're just about at time here. Is there anything else you want to share with us or demo or just any parting thoughts for the crew? It's funny. I keep remembering things that I can demo, but I I think I feel like I'm tapped out. That's fair. We went two hours. I appreciate you were so down to do this. Yeah. Yeah, for sure. So, yeah, I mean, like, I would just say try out Codex. Try it on your, if you're like a professional in here, like try it on a real hard task that you have. See where it works and where it doesn't work and like set up the interdev loop and then the outer dev loop.
1:55:38Alexander Embiricos:And then do missions. It's been really fun working on it. We've grown 20x since August now. Wow. 20x. You guys grow in exponentials, and that's just got to be such an experience at a workplace. Yeah, absolutely. And it's basically a superpower for us internally. It's a massive accelerant to the rest of the company. So I just can't wait to ship a lot of the new stuff that I couldn't show you. But there's a lot coming and you should come check it out. Excellent. Can't wait to see it, man. Can't wait to see it. Thank you so much for joining us. Oh, yeah. Also, if you are an engineer or a product manager or a designer and you want to work on Codex or a salesperson and you want to come work on Codex, please go check out the job site and apply.
1:56:28Alexander Embiricos:I don't know. I mentioned you found out in this podcast. Yeah. Excellent. That's awesome. Thank you so much. Thank you so much for the time. Thank you, Alexander. Very much. You know, I think this is a really good thing. It's helped me tremendously just listening and watching. And I'm coming away with what is definitely a clearer understanding of how I'm going to go tear into it this weekend as soon as I'm home. I want to see my cat one, right? Yeah. What did you say? Oh, yeah. How did it end? Let's see. Let's see. Might be to bring you up. Yeah. Okay. There is more cats. It's via emoji, but...
1:57:10Alexander Embiricos:Yeah, I have noticed it tends to put emojis instead of images. You kind of have to tell to put images. Hey, Kwin Cottle just asked, do you all have educators on the team? We work closely with a team in an opening night called Developer Experience. That's like a motorcycle mom who was just featured in the screen share or like Dom or people I work closely with. And so there, I guess you could call them educators. Excellent. Grant, is your screen shared? I got kicked out. I'm vlogging. Uh-oh. We can't have that. Well, while Grant's getting back in, just a reminder to everyone, if you're here today, we are so grateful that you've joined us.
1:57:49And we hope you'll continue to come back for a long time to come. Please like, subscribe, do all of the things that help us bring you awesome guests and cool new products to talk about. Also, make sure you stop by. Sign up for the Neuron newsletter at theneuron.ai. join 630-ish thousand people who read it every morning, and we'd love for you to be one of them. It's not letting me rejoin, so I'll have to include the screenshots of the cat with the email that we go out recapping this. There we go. I'll send it to you, Alexander. I'll send it to the cats. One last question as we close. Recommended IDE, would it be VS Code?
1:58:32Alexander Embiricos:There's so many good IDEs. I would say I personally use VS Code. Cursor is awesome. I know a lot of folks who do more like Java development and Python really like the JetBrains IDEs, like PyCharm or IntelliJ. So those would be my off the cuff. Excellent. Do you have any other tools that you use? Oh, tools? I like... It's funny. I was just talking to the founder of a terminal tool that I like, Zach and Warp. I used to use it all the time. Now I use Ghosty. I saw that a little while ago. Yeah. I think don't sleep on Atlas. Yeah. It's there's a whole whole rant there. But I think that a rant about like what the future is there.
1:59:17Alexander Embiricos:But I think it's it's really interesting to sort of just switch to a mode where it's like, OK, I'm going to like make sure that AI can see the sites that I'm on when I want it to so that I can ask questions. it so like earlier you actually mentioned an article that i was like oh yeah i need to like refresh my memory on that you didn't end up asking me about it so i like went into atlas and like i literally launched atlas in order to like load the article and then i was like summarize this so like that kind of stuff is uh is very valuable i'm still anxiously awaiting atlas on windows that's my uh oh we are yeah i mean codex is working on it i heard that's awesome we heard rumors there might be a 12 days of Christmas this year.
2:00:02Can you confirm or deny? It's okay if you can't.
2:00:04Alexander Embiricos:I don't know. Okay. No telling. Fair enough. All right. See you guys. Alexander, we sure appreciate it. Thanks to everyone, and we'll see you next time. Alright, let's see.
From the publisher
In this week's live-stream replay, we go live for a 2-hour, hands-on deep dive into GPT-5.1 Codex Max with Alexander Embiricos, product lead for OpenAI Codex. You’ll walk out feeling like an agentic-coding wizard, even if you’re starting from zero. GPT-5.1 Codex Max is OpenAI’s latest frontier agentic coding model. It’s built on an upgraded reasoning backbone and trained to handle real-world software engineering tasks end to end: PRs, refactors, frontend builds, and deep debugging. It can work independently for hours, compacting its own history so it can refactor entire projects and run multi-hour agent loops without losing context. In this live session, we’ll set it up together, build real agents, and push Codex Max to its limits.
