Scaffolding is coping not scaling, and other lessons from Codex | OpenAI’s Thibault Sottiaux

27 Jan 2026 · 40 min · 18 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Dev Interrupted - Episode: Scaffolding is Coping Not Scaling, and Other Lessons from Codex | OpenAI’s Thibault Sottiaux

Episode Overview In this episode of Dev Interrupted, hosts Andrew Zigler, Ben Lloyd Pearson, and Dan Lines engage with Thibault Sottiaux from OpenAI's Codex team. The discussion revolves around the development of AI agents, the importance of simplicity and scalable primitives, and the impact of open-source models on software engineering.

Key Themes

  • Agentic Autonomy: The concept of creating AI agents that operate with more independence rather than relying heavily on scaffolding.
  • Vertical Integration: Insights on how integrating research and engineering fosters innovation and efficiency.
  • Simplicity vs. Complexity: The importance of keeping designs simple for scalability and performance.
  • Open Source Community: The benefits and challenges of open-sourcing AI tools and fostering community engagement.

Detailed Discussion Points

  1. Understanding Codex as an AI Agent
  2. Codex is described as a state-of-the-art agent capable of performing various coding tasks.
  3. Sottiaux emphasizes the need to conceptualize Codex as an agent first, enabling diverse applications beyond just coding.
  1. The Bitter Lesson of AI Development
  2. Richard Sutton’s essay is referenced, which posits that clever tricks and domain expertise do not scale as effectively as fundamental primitives.
  3. The Codex approach focuses on developing scalable primitives that stand the test of time and improve performance.
  1. Vertical Integration Benefits
  2. Codex’s design benefits from vertical integration as it allows for seamless interaction between research and engineering.
  3. This integration enables the team to address problems more flexibly, deciding whether to fix issues at the model layer or through engineering solutions.
  1. Simplicity in Architecture
  2. Maintaining simplicity is crucial for scaling capabilities as models evolve.
  3. Complexity can lead to biases that hinder the full expression of an AI model's capabilities.
  1. Challenges of Open Sourcing
  2. Open-sourcing Codex helped demystify agent development and encouraged community tinkering and innovation.
  3. Early challenges included managing contributions and maintaining control over the codebase, leading to a shift from TypeScript to Rust to streamline development.
  1. The Role of Agents in Software Engineering
  2. Agents are transforming how teams collaborate and the dynamics of software development, with an emphasis on planning and integration.
  3. Developers are expected to adapt to new roles, often taking on responsibilities similar to tech leads due to increased automation in coding tasks.
  1. Future of Developer Workflows
  2. There’s a shift towards collaborative workflows where developers communicate more effectively, aligning on tasks and sharing insights.
  3. Codex’s introduction of code review models aims to reduce bottlenecks associated with increased code generation.
  1. Emerging Skills for Engineers
  2. New skills are required for engineers to manage and orchestrate AI tools effectively.
  3. The episode discusses the concept of "super bus factor," highlighting the risks of relying too heavily on a single engineer’s expertise.
  1. Advice for Future-Proofing Skills
  2. Sottiaux encourages engineers to build personalized skills and automation tools to enhance their productivity and maintain joy in programming.
  3. Emphasizes the importance of remaining open to iterative development and experimentation in engineering practices.

Conclusion The conversation highlights the transformative impact of AI agents in software engineering, promoting collaboration, scalability, and innovation. As software teams adapt to these changes, the focus will continue to be on simplifying processes, fostering creativity, and ensuring that human engineers remain at the center of development.

Additional Resources

  • OpenAI Codex: [OpenAI Codex](https://openai.com/codex/)
  • Codex Open Source Repo: [Codex GitHub](https://github.com/openai/codex)
  • Agent Skills Open Standard: [Agent Skills](https://github.com/openai/skills)
  • The Bitter Lesson by Richard Sutton: [Read Here](http://www.incompleteideas.net/IncIdeas/BitterLesson.html)

Follow Us

  • Subscribe to our [Substack](https://devinterrupted.substack.com/)
  • Follow us on [LinkedIn](https://www.linkedin.com/company/linearb/)
  • Check out our [YouTube Channel](https://www.youtube.com/@DevInterrupted)

Hosts

  • [Andrew Zigler](https://www.linkedin.com/in/andrewzigler/)
  • [Ben Lloyd Pearson](https://www.linkedin.com/in/benlloydpearson/)
  • [Dan Lines](https://www.linkedin.com/in/dan-lines/)

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Exploring Codex and Agentic Autonomy

0:45 to 3:00

Discussion about OpenAI's Codex as an agent and its implications for software engineering.

“But I want to start at Codex, just in your own words, like learning about OpenAI's flagship coding agent.”

Vertical Integration and Research Influence

3:00 to 6:00

Insights on how research and engineering inform each other in Codex's development.

“It opens the door for like a lot of like, you know, great ideas that you don't necessarily have.”

Simplicity and Capability Overhang

6:00 to 8:30

The importance of simplicity in design to prevent biases and maximize capabilities.

“And that's what we would refer to as a capability overhang.”

The Journey to Open Source

8:30 to 11:00

Reasons for open-sourcing Codex and the importance of community engagement.

“And it's all about figuring out when is the right time to remove like pieces of the scaffold and to build in a way where you can do that.”

Lessons from Community Contributions

11:00 to 14:01

Reflections on community input and the evolution of the Codex project through collaboration.

“And so, you know, it plays this like dual role.”

Adapting AI for Specific Domains

14:01 to 14:59

Learn how developers are customizing AI tools like Codex for unique problems.

“then get a glimpse of how the agent can be adapted to other things.”

Lessons Learned from 2025

15:01 to 16:43

Discover insights on simplifying workflows and the importance of context management in AI.

“I'm curious to know, you talked about earning all of this simplicity and getting back to the primitives.”

Enhancing Performance with Multi-Agent Networks

16:44 to 18:21

Understand the benefits of multi-agent networks and their impact on productivity.

“So that's like something that we just, we removed a whole bunch of complexity just by solving it, you know, in a way that we're uniquely positioned to solve it.”

Customizing AI Personalities

18:22 to 20:08

Explore the need for adaptable AI personalities to enhance user experiences.

“So that's one thing that we're working towards as well.”

Coupling AI Tools with Foundation Models

21:25 to 23:17

Investigate the balance between tightly coupling AI applications and ensuring portability.

“And I'm curious, like, how that affects the simplicity or the primitives that we've been talking about.”
Show all 18 chapters

The New Normal for Developers

23:17 to 26:08

Learn how AI is transforming the daily workflows and collaboration among developers.

“But I want us to take a step back and talk about the developer, the engineer, the person sitting at the computer using the agents and what their new normal looks like.”

Embracing Change with New Grad Developers

26:08 to 28:00

Discover how fresh graduates are leveraging AI tools effectively in their work.

“and it's inevitable that if you're generating so much more code, you're also going to produce some amount of code that you then want to catch.”

Identifying Individual Movers in Teams

28:00 to 29:10

Learn how to recognize and empower key contributors within your organization.

“has never really built these habits in decades of software engineering of like, oh, this is how you do things.”

The Importance of Collaboration in an AI-Driven World

29:10 to 31:08

Explore the challenges of maintaining collaboration amidst increasing automation.

“The idea that like if a single engine, like you said, the biggest bottleneck is just a siding and then somebody fires off an agent and then it gets done.”

Navigating Plans and Implementation in Product Development

31:08 to 32:58

Understand the balance between planning and flexible development in tech.

“So that's like very much like top of mind for us is, you know, solving towards that.”

Career Paths in an Evolving Engineering Landscape

32:58 to 35:08

Discover how career trajectories for engineers are changing with new tools.

“so that the lightning strikes it so we can figure out what's going on, right?”

Building Skills and Automation in Workflows

35:08 to 37:48

Learn how to create and customize skills for better productivity with AI.

“like, you know, building that empathy for the user.”

The Joy of Programming: Balancing Automation and Creativity

37:48 to 38:59

Explore the importance of preserving creativity in programming while using automation.

“like all the other things that you want to automate, actually you can keep like, you know, the most delightful parts, you know, of your day is just like intact.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:05My guest today is Thibault Sotio from OpenAI's Codex team. And Tebow has been working shoulder to shoulder with research and engineering in San Francisco to solve one of the hardest problems in tech, true agentic autonomy. Tebow, it's really great to have you here on Dev Interrupted. Hey, awesome to be here. Really excited to talk about all of this. Great. So today we're going to look at the bitter lesson of AI development. There's so much we can dig into about what you're building and about why clever tricks and domain expertise can sometimes fall short and why scalable primitives are winning.

0:38We're going to explore the Codex team and how it balances that research with exacting the requirements for production and making a happy developer tool that works for everybody. But I want to start at Codex, just in your own words, like learning about OpenAI's flagship coding agent. Can you tell us a bit about what it is and why you describe it as an agent first instead of a product? Yeah, so we think about the two parts. First and foremost, we're building a SOTA agent that is able to act and perform incredible amounts of work on the coding front and help software engineers in this world. And this agent is fairly general.

1:18You can put it to work in many places, and that's also where the products come in. So it's like figuring out what the best way is to leverage an interface with this agent that is going to be ever increasingly more capable. And there is something really interesting when you shift your mindset to building an agent first and then figuring out where to put it to work. It's like you find like a remarkable amount of places where this agent, you know, comes in handy and, you know, can actually do economically valuable work. And we're also thinking about, you know, what does this mean beyond just coding, right?

1:48Even for software engineers, it's not just about code generation. It's about solving many other parts of the day-to-day that are actually the bottlenecks. And so building that general agent is what we're after. Yeah. And we're going to talk more about what those other bottlenecks are. We talk about that a lot on Dev Interrupted, especially in the last year, about how agentic development has really made bare all of those human and communication problems that are affecting teams really deeply. Something that you said there about thinking about it differently from being a product, it actually makes me think of almost like a product and a platform analog.

2:25With a platform, you can put a lot of products, a lot of things in an ecosystem together. An agent, it seems to be emerging as something that operates the same way you can build an agent, like you said, and then figure out what are the use cases, how do we step this forward or backward into the products that make sense for people. Yeah, that's right. Right. So if you think about, you know, building an autonomous, like ever more capable entity and keep your mind sort of flexible about, you know, what the best form factor is, what the best product around it is. And, you know, perhaps we put it to work inside our own products, but we also partner with other companies to put it to work in their products.

3:00It opens the door for like a lot of like, you know, great ideas that you don't necessarily have. You know, we don't have them ahead of time. It's like, you know, we can focus on like building the agent and then, you know, figuring out like where to put it to work later. And what is it like to work in an environment where you're building a coding agent and you're sitting on top of a frontier model? Like you have the vertical integration of that entire process, which is a unique leg up. What is that like? That's right. It's very interesting because we're able to take some of the best ideas from engineering and have them influence research and then do the reverse as well.

3:36where research influences the entire engineering roadmap for how we build out the agent. And one of the things that you can do when you're vertically integrated, you can decide where you actually fix problems. So you don't have to fix everything in your harness. Some of the things we decide to fix downstream by training new models. And we know that by training the model, we will have a jump in the capability that we need three months down the line, six months down the line. And it allows us to do these trade-offs that you can't do without a vertical integration. There's also this thing called, you know, the no free lunch theorem, which is basically like if you're trying to adapt and be intelligent in any possible distribution.

4:17Well, this is going to be strictly less optimal that if you were to actually build something for a very specific distribution. And so by coupling like the harness and the model and like that's what we named the agent, we're able to get a lift in capabilities. And that's, you know, obviously very interesting, like, you know, for something as important as coding. Yeah, it's fascinating that you mentioned too that those two things, they feed into each other, the research and the engineering. It's a big loop. And inside of it, there's lots of little loops. I think that's what we've all been discovering in the last year, especially people who have a lot of success working with agents, is finding those loops, those workflows that work perfectly, the inputs and outputs, feeding into each other and making the system better.

4:58So that is a unique advantage and kind of like a vertical integration front. And I guess in this world where there's so much greenfield, you have the model, you have the coding agent, you can build it and take it any direction, take this almost like agentic first approach and then build towards what individual people and organizations need. But how do you, as the person at the center of that, prevent complexity from trapping your own architecture and keeping it simple? That way it works for everybody. Yeah, so we try to stay grounded in things that have stood the test of time and that we know have been true in the past, where simplicity just often leads to better scaling over time and better overall performance of the system that you're building over time.

5:45And what we're really looking at is the ability of the agent, so it is like the system of the harness and the model to continue to increase in capability over time as our models continue to increase in capabilities. So it would be really unfortunate if we introduced bias in the harness or in our solutions, which means that when we have a capability jump in the model, we're actually not able to express that because we somehow constrained it in a way that prevents it from expressing its true capability. And that's what we would refer to as a capability overhang. And by keeping things simple, you reduce the amount of things that can actually go wrong.

6:30And it's all about choosing the primitives that are proven to scale with model capabilities. And that's actually also closely related to a research problem. We can do the research on that of, hey, this harness scales with different model sizes and different model capabilities. We look at on the smallest model and then a mid model and then a frontier model. How do we get the scaling that we expect? So you can extend the scaling laws on the entire system. Right. And so in order to apply the scaling laws, you have to have simplicity, ultimately, that those things are leveraging, that they're driving.

7:14Otherwise, things are going to clash. They're not going to work. And like you said, you can't take advantage of immediate gains. I think that's really the biggest win of the system is that when you're so lightweight, it's so easy to pick up and put it onto the next new thing, which is what's so important when you're coupled closely with a model. And so that's something that's really fascinating. I think the simplicity, though, you know, it can't be understated. It's earned, right? It's like earned through research and engineering. Like you had to climb a summit of complexity, I'm sure. And at one point, things were very complex.

7:48And then you have to come down from that summit, like you said, to find the primitives, the things that emerged as more clear, right? And so that's like the journey that you're kind of describing. Yeah, it's an ever-continuing search for the right primitives. And once you discover those primitives, they seem delightfully simple. They often seem delightfully simple. But the search for those primitives, it's a complex matter. And you might not find those primitives immediately. And so you have complexity in your system, and you look at that complexity, and you go, we think there's quite a bit of bias here that we introduce and we constrain the model in a way where it will hamper its productivity at some point.

8:33But this is also this delicate balance between the harness and to me it's called a harness because you're scaffolding it in a way where you want to remove the scaffold over time because the model is able to stand on its own. And it's all about figuring out when is the right time to remove like pieces of the scaffold and to build in a way where you can do that. And the advantage of like coupling the two together is like, you know, we have to care about like our model series and our model series only. And so every time, you know, we improve things, we can remove bits of the scaffold without, you know, being afraid of like breaking something that is not under control.

9:10That's a really smart insight that I want to like click on again because, you know, you're calling out that ultimately the harness, that scaffold that you build, it should lean towards that simplicity because like you said, But it's supposed to be kind of there until the agent can stand fully on its own and do what that harness is kind of getting it to do. Some people lean, of course, the opposite approach. And they treat that harness more like a jetpack. And it doesn't go the way of simplicity. It goes of how can I put so many tools and so many things in this box? And so I think it's like two different types of ways that people come at them.

9:44And they obviously achieve different levels of success. It doesn't scale as well as having a lightweight harness like what you're calling out. And I think that's kind of part of what y 'all have earned from your years of research. And, you know, there's also Richard Sutton's Bitter Lesson, which calls out that clever tricks and domain expertise, they don't scale compared to these primitives, right? And when we talk about these primitives, they're compute and power and time. And that's what you're on the mission to find is how can we take best advantage of those so everybody can get it. But, you know, along that journey, you mentioned that you open sourced like the repo, you open sourced the process by which people can build these agents.

10:24What was that like? Like what led you all to that decision? So around the time that we decided to open source a repo, I think there was a lot of mysticism around agents and how they worked and what you needed to build on top of a model in order to have it be able to act safely into the world and perform economical, valuable tasks for you. And so we wanted to sort of show how, you know, delightfully simple it can actually be and show some of the primitives that are important to get right. And, you know, we also wanted to show that if you do these things, you know, you can get incredible performance out of our models.

11:04And so, you know, it plays this like dual role. Additionally, on top of that, sort of like had this idea that, you know, if you solve for code generation, open source is going to change. and we want to deeply understand how open source is going to change by being open source ourselves, at least for part of our technology. And then one additional idea and like reason for why, you know, I thought this might be interesting and like others agreed was that we were going to see a lot of tinkering in the space of agents and a lot of interesting products being built on top of it. And having like an open source core, you know, allows people out there in the world to like, you know, get inspired and tinker and build things that we might have never imagined.

11:49And so it allows to tap into this creativity of the open source community and to let people innovate and invent new things. And I think altogether, it just made so much sense for the agent to be open source. And yeah, it's something that I've been very proud of. Yeah, and as part of building that community, making a place where people can come and be part of the development, Have there been things about that that surprised you or stood out? You're like, this was an earned thing from the community. This is a proven success of why open sourcing this was the way. Yeah, one of the things is initially we got things somewhat wrong.

12:27We accepted too many contributions and we sort of lost control over the repo. We course corrected later. At some point, we migrated to Rust and we had decided that we had conviction on some of the primitives that we needed to build. And, you know, we wanted to build those primitives, right? And, you know, we also had conviction that we were going to see millions, if not billions of agents running concurrently at some point. And we wanted to write this like in an efficient language. So like the Rust migration was a bit of a difficult moment with the community because we had been accepting a lot of PRs through like our TypeScript open source repo at the time.

13:06And then, you know, we just sort of like rewrote things to Rust. we did build like very good partnerships with like awesome contributors over time to the rust core and so that that relationship to the community has evolved over time one of the things is just how great it is to have it all there out in the open it's like we do get a lot of reports and people like sort of like point at you know this is where it is like you know here's like how you should fix it or like you know here is like an attempt at fixing it and then we also get a lot of inspiration from all the forks. I think there's like, I don't know the number, but like over a thousand forks of the repo.

13:42Some of them are actually quite popular. And, you know, they just come up with like new ideas and like we work with the authors of like the forks like to port some of the changes back into, you know, the Codex open source repo. And it's pretty cool like to be able to learn those things from how people change and alter your code. And you can also then get a glimpse of how the agent can be adapted to other things. You're going to get a really good front row seat at how these early adopter developers are picking up these systems and molding them to their specific domain problems. So it's a really great partnership and definitely something I think people should check out because that's kind of rare and it's part of that relationship that even goes back as deep as the model, right?

14:24That's kind of how deeply penetrating this type of technology can be. Every week I hear from a company that, you know, builds their company around like the Codex open source agent where we're like, oh, you know, it's just like it works so well. It works actually like on non-coding things. And we adopted it like, you know, we adapted it like in this and this way. And like now it's able to do like this other economically viable task, such as editing spreadsheets, for example. And or like we embedded it into a browser. And then I get these demos and like it's always very cool to see that. somehow we're contributing to those things as well.

15:01Yeah, totally. I'm curious to know, you talked about earning all of this simplicity and getting back to the primitives. And we're at the top of a new year. It was spent all last year experimenting with these new tools. What is something that you picked up in 2025 that maybe you're not going to pick up again in 2026? How is your workflow and the way that you think about agents getting simpler this year? It's one of the big issues and challenges that we had last year There was very long-running sessions where the agent goes through something called a compaction, where we allow the agent to perform work beyond the context window that is actually available to the model.

15:42And a lot of the complaints were, you know, this is not working well. And this was really part of the scaffolding at the time, where it was just kind of like a heuristic of how you needed to summarize the work done so far, and then reset the context window so that the work could be allowed to continue when given a fresh. And the model was losing context on quite a bit of the work that was performed before. It's really hard to get this heuristic, right? You can try and prompt models. You can try and scaffold your way through it. For a lot of agents out there, this is a very significant part of the complexity of the harness.

16:21And so we decided to solve this at the model level. And, you know, we trained on this like end to end. And now this is something that we receive like almost no complaints on anymore. And, you know, the model is able to, like the agent is able to like work across like 20 windows, 20 context windows, like, you know, for very, very long time horizons without losing track of like, you know, what it was doing before. So that's like something that we just, we removed a whole bunch of complexity just by solving it, you know, in a way that we're uniquely positioned to solve it. And that's the great thing about like, the kind of trifecta of the simplicities of those simple primitives is when one moves and gets a major advantage, when you come back to another one, you realize that you can move it way further because of all these gains are so closely coupled.

17:06Are there any other ways that you're approaching Codex itself for like this year that stand out for that, like will be different, you think, from how you approached it next year? Like what's like the top of the year vision for how we make this like the most popular code agent in the world? Yeah, so there's a couple of things that we're really pushing on. We think last year, single agents became reliable enough that they can perform work. Like this year, we're going to see like reliable multi-agent networks performing significantly more work. What does that mean for the user when you have, you know, maybe an order of magnitude or like two orders of magnitude more, you know, economical, valuable work being produced in, you know, the same time span?

17:51It means that you're able to spend a lot more tokens in the same time, but it also means perhaps you need to review a ton more code or how do you do that exactly. So this is something that we're looking to solve. We're working towards making things significantly faster as well. We feel like we're at the frontier with our models on the intelligence front. We're not at the frontier yet on how fast it can be. We expect these models to get significantly faster this year. And it's almost like a sweet spot that you want to reach of like a level of intelligence at a sufficient speed that, you know, it feels delightful to use in the product.

18:31So that's one thing that we're working towards as well. And the last one I would say is like, you know, it's really this super collaborative personality. Our models, like in our agent, like Codex is known to be a little bit terse, a little bit stubborn sometimes, you know, just really this blunt, pragmatic engineer. that is not the right persona for everyone. Personally, I like to be a little bit more validated in my ideas, not to be told that I'm right when I'm not right, but I want the model to acknowledge that I'm there behind the laptop as well and trying to do something and to be there along the way with me.

19:07And so solving for that is important as well. I know what you mean. I think that's something we've experienced with all of the front foundation models is they all have a different level of politeness, terseness. And like you said, like sometimes too much politeness is just a waste of everyone's time. It's not accurate either. So, but there's like a happy medium there. Like you said, you can't, like what personality works for everybody? No personality. Who's liked by everybody? No one. It's like, so how can you build an agent that's built as liked by everybody? You can't, but you have to make it easy to adapt to what people do like.

19:39That's right. So it's really about making it personal to you and, you know, working in the ways that, you know, you like working. So if you're like highly creative and you like brainstorming and, you know, you don't want to get like pedantic like nits on like the quality of your code is like that should be possible. If you're working on like a super critical code base and, you know, you want every single thing that can go wrong to be flagged to you, then, you know, you should also be able to get that. And that's something that our models are incredibly good at. Codex has been involved in finding some of the most impressive exploits last year, some of the React exploits, which we can talk about.

20:23But for that, you don't need a super friendly personality. You just need Codex to go and figure out what a potential exploit is or fix a very gnarly bug for you and then come back to the solution. And you want to be very confident that it's right. And for others, it's like, you know, you just really want this bubbly, like, you know, super collaborative thing, you know, and it needs to be personal.

20:49AI has changed how we build software, but faster code doesn't always mean faster delivery. That's why Linear B is launching the Essentials Plan, your toolkit for measuring and improving AI productivity. You get the new Linear B MCP server, AI Insights dashboard, and developer surveys, all designed to reveal how AI tools impact delivery speed, code quality, and developer experience. Stop guessing which AI tools work and start leading with data. Visit LinearB.io to learn more and unlock the next chapter of AI productivity. so speaking of like being able to customize and couple things together like you're in an environment where you have the distinct privilege of being like we've talked about connected to the foundation model all the way up to the implementation of the agent and all the research in between so the full stack but for other teams that are building with agentic tools and are in our and are making these kinds of agentic systems like how should they think about like coupling like really tightly coupling to a foundation model versus being more portable in what they build.

21:56And I'm curious, like, how that affects the simplicity or the primitives that we've been talking about. Yes, I think the primitives should roughly be similar, although they might differ in shape a little bit. And I think one of the challenges is, like, if you're building for many foundation models and are completely agnostic to the model provider is you have to find a common ground of all these models. And it's quite inevitable that if you do not adjust at least a little bit, you are going to see the penalty and the overall performance that you can get. And so what we're seeing is we're working with some of the players in the ecosystem and we're advising them on how to make it work very well.

22:45And this is also why our code is open source. is like, it's just all out there. It's an example of how you can get the best performance out of the GPT models. And, you know, we work with them in order to adjust things. So I do think that you need to adjust it to like a certain degree in order to benefit from the models. At some point, if you want to do that for, you know, the thousands of models out there, it becomes like quite prohibitive. And I expect like, you know, major players to like, you know, only do that for like a handful of models. So we've been talking about the agent, But I want us to take a step back and talk about the developer, the engineer, the person sitting at the computer using the agents and what their new normal looks like.

23:27And that's what we've all been exploring in the last year of how this type of agentic coding is impacting developer teams and individual engineers. Different people are getting at this at different speeds, right? And with it, the folks who are experiencing that hyperproductivity, it really blows away all of the code bottlenecks and new bottlenecks emerge. And they're often around like planning and integration, actually taking the time to like sit down and chat with somebody. So, you know, there's a new skill set involved with what makes a developer successful. In a way, each developer almost becomes like their own mini team that they can customize in whatever way works for them specifically.

24:06And they become the representative, almost like the decision manager of this team. You know, what has that been like for your engineering team? How does their day-to-day work transform in that kind of new world? Does everyone just spend less time in an IDE and just lots and lots of time planning and reviewing? Yeah, so interestingly, it actually has brought people closer together. And, you know, we have like more FaceTime. We have more, you know, creative like ideation and planning together because everyone is so accelerated that once you decide on the thing, you can just like almost immediately do it.

24:48And agreeing together and getting organized together, you know, as like a small team, even like of 20 people, it's like it matters a lot because you will be able to achieve like in a week what you traditionally were able to achieve, like, you know, maybe in a month with the team of like, you know, the same size. And so it's kind of like interesting to see that it does bring engineers, to talk more and align more, maybe upfront than before. We're also thinking about, obviously, hey, how can you help and accelerate the planning phase and building consensus and routing things into user feedback and the realities of finding PMF and these things.

25:30One of the bottlenecks that is downstream of that is code review as well. And this is something that we proactively thought about. Last year, we built a custom code review model, which is SOTA. I believe it's still SOTA today. And we deployed it internally across all of OpenAI. And this has been, to my surprise, one of the largest successes within OpenAI for Codex, where it's now pretty much enabled for everyone by default. A lot of teams just require, it's mandatory to have Codex review the PRs because it catches so many bugs. and it's inevitable that if you're generating so much more code, you're also going to produce some amount of code that you then want to catch.

26:18And figuring out what those bottlenecks are starting to become and solving for them one at a time, I think that's the big challenge of 2026. Yeah, I think so too. And those bottlenecks, those are things that are popping up. Those problems, like groundhogs popping out of the ground, right? for these teams to solve because they've already been in motion with each other. They're already working. But there's also this whole other wave of people right now who are finishing up their CS degrees, who have been experimenting with these tools while they've been studying, who are getting their first jobs as engineers, and they're not coming in with all of that baggage.

Read the full transcript

26:54And so they're coming in very, very differently. And from my own personal experience, dealing with folks that are in that skill band, they're absolutely crushing it with these tools. When they use it, when they lean into it, they have a native instinct because they aren't cluttered and bogged down with all the stuff you and I are from just being in the working world. So what do you see with OpenAI? I know you all work with a lot of young developers and folks that are engineers. What has that difference been like in how these agent-first developers are coming on the scene and how do they solve things?

27:29Within Codex, we have the entire spectrum. We have industry veterans who have been working for over 35 years. I think that's the record. It's before I was born. And then we have new grads. And one of the people I trust the most on the team is this new grad, Ahmed. He's just got this incredibly creative ideas, almost was molded by all of this existing, has never really built these habits in decades of software engineering of like, oh, this is how you do things. It's super open to new ways and very, very much adapting every day. And teaching a lot of the rest of the team how to actually be productive.

28:19And it's remarkable to see. We have a couple of others like that on the team as well. and it's essential for us. I would say, like, you know, without that, it's like we would actually be moving, like, way slower as a team. Yeah, I agree. I think you start with those, like, there's individual movers within your org, and I think every org has a responsibility to identify them. They're there. They're doing something. If you don't know about them, I promise you they're lurking somewhere under the surface. And if you find them, there's so much that they can teach and enable to the rest of your org.

28:50Because so often during, like, AI rollouts and people experiment with these tools, things get trapped inside of teams and they don't really get broader than that. Or you get one high-powered individual who is able to just massively ship everything and no one else really knows how any of that, how that guy does anything, right? And so, you know, and speaking of, it kind of makes me think of the emerging like super bus factor problem, right? The idea that like if a single engine, like you said, the biggest bottleneck is just a siding and then somebody fires off an agent and then it gets done. If a single engineer can ship a whole product solo that we're seeing people do with this kind of tool, how do you keep collaboration alive in that kind of world?

29:34Why even bother handing anything else to someone? Me and my army of agents that I can spin up on demand can just do it. Yes. Maybe in the limit, you'll be able to run a company solo. and it's like very much like, you know, maybe like, you know, five, 10 years away in my head, like maybe we get there sooner. And, you know, I think that will be like an interesting time. But right now, what I'm seeing is that disseminating plans is important. So disseminating information and like intent and recording intent of changes is important. And, you know, an agent can help you achieve that as well. So like, you know, as you go through your conversation in a session, you know, you can sort of like get a summary of your intent, attach that to the PR later so that others can understand.

30:20I've started to build tools for myself and for folks in the team as well, to keep track of changes that are happening across our team, across the organization, at different levels of abstractions. Because there's more happening. And you want to offset that by also allowing to give everyone a superpower to understand things faster. understand what is changing, understanding how things are implemented, why things are implemented in a certain way. And so it's not just about, you know, making good generation like 100 times faster. It's really about giving a boost to like how quickly you can find the right ideas to solve for and how quickly you can understand as well, like, you know, as a human, like the state of things.

31:08So that's like very much like top of mind for us is, you know, solving towards that. You mentioned this world, like getting plans out there so people can see them and understand. And I think that's really important. And I've seen a lot of tools and a lot of companies try to tackle this, right? Like everyone has taken a bite out of the whole spec problem, making a plan. And I'm curious, like that solves a lot of things that like it addresses the human to agent handoff, but also agent to agent, agent to human. Like, what do you think about that as like, it's a primitive, just like the rest of these things that you like it's if you have it has to be in there, then how do we make it as simple?

31:42as possible. Yeah, so the downside of like relying only on plans and, you know, building a large spec over time of your product or implementers, like, you know, it gets like a little bit, you know, too wild. And then, you know, you find like contradictions in there as well. And, you know, at some point, it's like so big, that's like, it becomes like inscrutable as well. And maybe there is a mismatch between the plan and like the implementation. But I do, I am a big believer and even before all of this, things like design docs and getting your ideas together for where the product is going and writing down your vision, your strategy.

32:20And so I think that is ever more important. At the same time, iterating on new ideas has never been easier and gaining that signal and that information that you need in order to make a product decision has never been easier. This is really something that's greatly accelerating. And so sometimes you might not know what you actually need to do, but you know the kinds of things that you need to build in order to gain the signal that you need, in order to know what you need to build. And so sometimes the plan is just like, we need to gain the signal. Here's the five things that we're going to do. Yeah, we need to build the lightning rod so that the lightning strikes it so we can figure out what's going on, right?

33:03And that is what's free now, right? It's like being able to just like, you know, you can spend that, like you said, it's easier than ever to investigate something about your product, how people are using it, like all these different types of things in the meta space around your product, right? Things that can be now solved with engineering that before would have taken like deep data analysis and a lot of time and a really like clear thesis up front. Now you can just like freeform explore, like sometimes even in parallel with like different ideas, right? So if that is the new paradigm by which developers build, then what do you think the career path looks like for senior IC or becoming a staff engineer?

33:44Is that seniority defined by how well you can orchestrate those agents and your plans and then, I guess, share it with your colleagues at scale? I don't think it's fundamentally different to before. I think the path to senior staff and beyond is the impact that you have within the team, within your org. This is how I've always thought about it, is how effective are you at building impactful pieces of the product and then making that work in the organization, making that work for the company. And so there is an aspect there of, yes, it is increasingly, you have to scale yourself up. You can do so many more things.

34:29Solo, you can go and look at user feedback. You can look at logs. You can run a few queries in the background so that you understand which database schema is the most appropriate so that the queries that you have in production are actually going to be running performantly. You can run a whole little engineering team by yourself. The skills that are interesting and that you need to build, like really as an engineer, is like more and more towards like everyone, you know, growing to like a tech lead role or like a tech lead manager role together with like, you know, wearing like, you know, an increasingly large like product hat as well, like, you know, building that empathy for the user.

35:11And, you know, models and like the agent can like help you with that as well, like over time, because it should, you know, you should also be able to send like the agent, you know, like interview users or, you know, summarize, you know, what the internet thinks about your product and potential things that you should try. And yeah, so just like maybe this accelerating path towards like, you know, TLM type roles, I would say a lot of people, you know, that's what they've always strived, you know, to do. So I think that's quite compatible. No, I think that's great advice. And it kind of mirrors with a lot of what we've been hearing, especially on the show in the last year, what people have been talking about the skills that matter most of them now.

35:49And just to end our conversation, we've covered a lot of amazing ground. We got a really great look at how you think about Codex and how you're tackling the coding agent problem. But from your perspective, you have an incredible vantage. Any final advice you want to impart to our listeners for how they can future-proof themselves or how they work at the top of a new year? So I'm sure there's lots of new ideas we could tap into. Yeah, so one advice is one thing I've had a lot of fun with and a lot of people on my team at OpenApp at FunWid is like skills. So this is like now an open standard and it's like a little thing that you can teach the model to do like in the way that, you know, you think is like, you know, most effective.

36:33You know, for example, like looking at logs or like, you know, running a performance test or like I have like this QA skill like where Codex is able to like QA itself. So like whenever I build a new feature and, you know, I just send Codex like, you know, to play with a version of itself in the terminal and like make sure that it's implemented up to spec and there's like no regression. Like building, like finding skills and like really my advice is like to make them your own so that, you know, you're building the skills that you need from your agent over time so that they're adapted to your workflow.

37:05And my analogy to it is like, you know, the other day I was like thinking about it. I was like, this is the closest I feel to like having trained like a little Pokemon, you know, where I'm like, oh, you know, this thing is like leveling up. You know, every time I interact with it, I'm like, learn this new thing. And it's just like, oh, my God, it's like doing it like a little bit better every time now. And that is like, you know, it starts to feel like this sort of like reliable, like, you know, you're building this, this almost this bond because it's like more and more reliable over time. And, you know, your work gets more delightful as well because you're automating the parts that, you know, you actually want to automate and, you know, that you don't want to do.

37:41and there's this pitfall as well, you know, which this is why I'm recommending this, like this pitfall of like, you know, just only automating code generation. But like, if you use skills and think about like all the other things that you want to automate, actually you can keep like, you know, the most delightful parts, you know, of your day is just like intact. It is kind of like, you know, get the joy, you know, it's like preserve the joy of like programming. I love that advice. It's like build your own toolbox and think about those tools that you bring. I find a lot of the same luck. I also use skills.

38:12I have a lot of customized things that do little weird, quirky things for me or do it in a specific way I prefer. And I love those. I always pull those out first thing when I jump into something. And I couldn't agree with that advice more. I love the Pokemon analogy. I'll push it even further and say it's more like a chef in a kitchen. You have your knives. You bring your knives to work. You sharpen your knives. You take care of your knives. Your knives are your tools. and developers can think about their skills and how they work with their agents the same way. It's like, that's your bag of dives.

38:41You can fold it up and take it with you. You can make it sharper and better and bring it to the next thing. That's how everyone should be thinking about building the primitives that do plug into the system. So that's really great advice. Tima, this has been a great look at, you know, behind the curtain at OpeningEye. It's been a pleasure to have you on the show. Before we wrap up, is there anywhere you'd like to point our audience to to go learn more about Codex, go follow you and the work you're doing? Yes, you can follow me on Twitter, fairly active there, sharing tips almost daily. A lot of my team is also active.

39:13And we have increasingly amazing developer documentation that is quickly coming together. And then, of course, like the open source repo where you can file issues, contribute to the discussion, and be part of the community. And I really want to thank you for having me on the show today. It's been really a pleasure to talk about all these things. Yeah, basically, it's been great to have you too. And we're going to include the links to all that stuff in our show notes. People can go check it out. And for our listeners, that's it for this week's Dev Interrupted. But if you want to go deeper on how agentic AI is specifically reshaping engineering leadership, definitely check out our LinkedIn and Substack newsletters.

39:52Just search for Dev Interrupted anywhere. If you're listening to this, be sure to check out the included newsletter as well. We share analysis and takeaways from engineering leaders like Tebow each week, and we're going to continue to follow this story. And Tebow, thanks again for joining us. It's been so fun to have OpenAI on the show, and we'll have you back some point soon, I'm sure. Take care. Of course. Bye.

From the publisher

If you rely on complex scaffolding to build AI agents you aren't scaling you are coping. Thibault Sottiaux from OpenAI’s Codex team joins us to explain why they are ruthlessly removing the harness to solve for true agentic autonomy. We discuss the bitter lesson of vertical integration, why scalable primitives beat clever tricks, and how the rise of the super bus factor is reshaping engineering careers.

LinearB: Measure the impact of GitHub Copilot and Cursor

Follow the show:

Follow the hosts:

Follow today's guest:

OFFERS

  • Start Free Trial: Get started with LinearB's AI productivity platform for free.
  • Book a Demo: Learn how you can ship faster, improve DevEx, and lead with confidence in the AI era.

LEARN ABOUT LINEARB

  • AI Code Reviews: Automate reviews to catch bugs, security risks, and performance issues before they hit production.
  • AI & Productivity Insights: Go beyond DORA with AI-powered recommendations and dashboards to measure and improve performance.
  • AI-Powered Workflow Automations: Use AI-generated PR descriptions, smart routing, and other automations to reduce developer toil.
  • MCP Server: Interact with your engineering data using natural language to build custom reports and get answers on the fly.

More from Dev Interrupted

All 208 episodes
Scaffolding is coping not scaling, and other lessons from CodexDev Interrupted · 40 min
Listen in VO