In short
Podcast Episode Notes: VS Code and Agentic Development with Kai Maetzel
Episode Overview
- Podcast Title: Software Engineering Daily
- Episode Title: VS Code and Agentic Development with Kai Maetzel
- Description: Discussing Visual Studio Code (VS Code) as a transformative tool in software development, its evolution, and the impact of AI-assisted programming.
Key Participants
- Kai Maetzel: Engineering Manager leading the VS Code team at Microsoft.
- Kevin Ball (K-Ball): VP of Engineering at Mento, an independent coach, and co-founder of the San Diego JavaScript Meetup.
Summary of Discussions
Introduction to VS Code
- VS Code has evolved into a leading platform for software development with:
- Intuitive User Experience (UX)
- Rich Extension Marketplace
- Deep Integration with modern tooling
Kai Maetzel's Background
- Kai's early career focused on DevTools, with a decade at Microsoft dedicated to VS Code.
- VS Code grew from zero to 44 million users in a decade.
The Evolution of Development Tools
- Traditional editors and IDEs often lacked optimal user experiences.
- VS Code aims to find a middle ground between lightweight editors and full-featured IDEs.
AI's Impact on Programming
- The emergence of AI tools, such as GitHub Copilot, is significantly altering the coding landscape.
- The integration of AI has shifted VS Code’s design philosophy, emphasizing the need for:
- Enhanced user interfaces
- Improved completion suggestions
- Interactive coding assistance
The Concept of Agentic Programming
- "Agentic programming" refers to leveraging AI agents to assist developers in writing, optimizing, and managing code.
- Key discussions centered around:
- How to balance AI suggestions with user control.
- The unsolved nature of effectively integrating AI with human workflows, particularly in coding.
User Interaction Design Challenges
- The balance of showing suggestions without overwhelming users is a significant challenge.
- Different user experiences (e.g., typing speed and familiarity) necessitate adaptive interaction models.
- Continuous evaluation and metrics help fine-tune the integration of AI suggestions in the development process.
The Future of Development Tools
- The idea of running multiple parallel agents to handle complex projects is emerging.
- Discussion around managing interactions between foreground and background agents, with an emphasis on user experience.
- Predictions on how agents might evolve to manage significant projects across multiple repositories.
Collaboration and User Experience
- Future tooling must consider how to enhance collaborative efforts in development.
- Exploring new forms of interaction, such as visual interfaces and voice commands, to make coding more intuitive and enjoyable.
Security and Trust in AI Tools
- Security concerns around AI tools executing code and commands are paramount, especially in enterprise environments.
- The need for robust identity management and permission structures as AI agents gain more autonomy.
Conclusion The conversation explored the dynamic landscape of software development, particularly through the lens of tools like VS Code and the integration of AI. It emphasized the necessity for ongoing research into user interactions, the evolution of agentic programming, and the future of collaborative coding practices.
Key Takeaways
- Rapid Evolution: The coding landscape is changing rapidly due to AI advancements and user experience improvements in tools like VS Code.
- User-Centric Design: Ongoing adjustments are necessary to ensure that AI integrations enhance rather than hinder productivity.
- Collaborative Future: Development tools must evolve to facilitate both individual and collaborative efforts, enriching user experience while maintaining security and trust.
---
For more detailed insights, the full transcript and additional resources can be found on [Software Engineering Daily](https://softwareengineeringdaily.com/2026/01/06/vs-code-and-agentic-development-with-kai-maetzel/).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Journey of VS Code
0:45 to 3:40
Kai Maetzel shares his journey and the evolution of VS Code since its inception.
“and what the future of development might look like.”
AI's Impact on Development
3:40 to 6:40
Discussion on how AI is reshaping the design philosophy of VS Code and programming models.
“So when you think through this one, we easily forget what we knew and what we didn't know.”
AI-Powered IntelliSense
6:40 to 10:00
Exploration of AI-powered features in VS Code and their core integration into the tool.
“But it was super clear that this needs to be a core part of the experience.”
Balancing AI Suggestions
10:00 to 12:40
Kai discusses the challenges of balancing AI suggestions with user experience and acceptance.
“And so I'm curious, as you've layered on each of those pieces, as you've gone down, you know, chat-oriented programming and now agentic programming and all these things.”
Adaptive Systems in Completions
12:40 to 14:01
Exploration of how typing speed and user experience may be encoded in AI suggestions.
“So if you show a little bit more, has that a positive outcome, negative outcome, how often do people hit the escape key?”
Exploring Global and Adaptive Input Models
14:01 to 14:40
Learn how typing speed and user experience are integrated into coding models.
“When you're tuning these knobs, are they global knobs applied to everyone?”
Interactivity Trade-offs in Development Tools
14:41 to 20:52
Discover the trade-offs between speed and interaction in developer tools.
“And looking at that feature space you mentioned, typing speed is like how new a user is or some sort of representation of that also encoded in some way?”
Building Customized Agents in Development
23:07 to 28:00
Understand the framework for creating agent-based development tools.
“one of the things that you talked about, right, for these different modes and how steerable they are, there's obviously model differences, right?”
Understanding Tool Invocation in VS Code
28:00 to 28:36
Learn how tools are invoked and managed in VS Code.
“What is all of the amounts of tokens that you actually use in order to get there?”
Exploring Model Preferences and Customization
28:36 to 29:10
Discover how different models have distinct preferences and instructions.
“If you have a watch task, it should not try to spend, you know, a bunch of turns in order to figure out how to build the project called the watch task.”
Show all 26 chapters
Current and Future States of Model Tooling
29:10 to 30:29
Examine the current and future states of tooling for different AI models.
“So there's a couple pieces I'd like to dig in on that.”
Assembling Context for Agents
30:29 to 31:54
Learn about how context is assembled for agents in different environments.
“We really separate between GPT-5, GPT-5 codecs, GPT-5-1 codecs.”
Tool Representation and Prompt Limitations
31:54 to 33:53
Understand how tools are represented in prompts and their limitations.
“There's people who do different amounts of like pre-injection of context.”
Trade-offs in Tool Usage and Cache Management
33:53 to 35:29
Explore the trade-offs in tool usage and how they affect cache performance.
“So in some cases, right, some models actually support that you put this at the end of the prompt.”
Dynamic Inclusions and User Control
35:29 to 36:54
Discuss the importance of dynamic inclusions for user control in AI interactions.
“You say what the repository is that the user runs in, right?”
Foreground vs Background Agent Interactions
36:54 to 39:46
Analyze the differences in agent interactions when running in foreground and background.
“But when they did a really good job, then they can actually make sure that the agent extremely quickly gets to the right place, knows where to start, knows where to look.”
User Experience and Agent Tooling
39:46 to 42:00
Discover how agent tooling impacts user experience across different environments.
“and is still exposed to me closing the lid, neither a cloud agent is not.”
User Experience in VS Code Extensions
42:00 to 45:45
Explore how background and foreground agents improve user interaction in VS Code.
“It's like, for example, when we come and say, no other agent needs to say, hey, maybe you should install this extension to have a better user experience, right?”
Connecting Interactive and Background Modes
45:45 to 51:21
Learn about the interplay between interactive coding and background processes.
“Where you pretty much, and this is a really interesting point, right?”
Extensibility and Security in VS Code
51:21 to 56:00
Understand the implications of extensibility and security in VS Code's evolving landscape.
“We can do some things with GitHub and github.com, but I think it's not necessarily where the line is.”
Trust and Security in Tool Usage
56:00 to 58:00
Explore the balance between user control and security in executing commands.
“Do you trust everyone who's able to get anything into that?”
The Future of Development Environments
58:00 to 1:00:00
Discuss the implications of using containers or VMs for development and user experience.
“It's like this is chatted in one way, right?”
Managing Code with Background Agents
1:00:00 to 1:02:00
Understand the complexities of managing multiple coding agents and their interactions.
“oh, there's this one or two background agents that I have, right?”
Challenges of Parallel Development
1:02:00 to 1:04:00
Investigate the cognitive limits in managing parallel development processes.
“When we talked about there's something that is really sacred land and then there are other things, right?”
The Role of Mobile in Software Development
1:04:00 to 1:06:00
Examine the potential and limitations of mobile devices in coding workflows.
“But you can see that this is happening, right?”
Collaboration and Creativity in Coding
1:06:00 to 1:08:00
Discuss innovative collaboration tools that enhance creativity in software development.
“And if I have an idea, I just type it in my phone and send it off.”
Transcript
Automatic transcript. May contain errors.0:00Visual Studio Code has become one of the most influential tools in modern software development. The open-source code editor has evolved into a platform used by millions of developers around the world, and it has reshaped expectations for what a modern development environment can be through its intuitive UX, rich extension marketplace, and deep integration with today's tooling landscape. Now, in an era defined by rapid advances in AI-assisted programming, VS Code is at the center of a profound shift in how software is written. Kai Metzel is the engineering manager leading the VS Code team at Microsoft.
0:36He joins the show with Kevin Ball to talk about the origins of VS Code, how AI has reshaped the editor's design philosophy, the rise of agentic programming models, and what the future of development might look like. Kevin Ball, or K-Ball, is the vice president of engineering at Mento and an independent coach for engineers and engineering leaders. He co-founded and served as CTO for two companies, founded the San Diego JavaScript Meetup, and organizes the AI in Action discussion group through Latent Space. Check out the show notes to follow KBall on Twitter or LinkedIn, or visit his website, kball.llc.
1:25Kai, welcome to the show. Hi, Kevin. Thanks for having me. Yeah, I'm excited for this conversation. So let's maybe start a little bit with you and your background and your journey to leading this VS Code team. Oh, so actually it started very, very early on. So my first internship was already with DevTools and I never really left DevTools. so and then you know now 10 years ago i joined microsoft explicitly for the vs code effort so there was you know there was promise that there is something that that could get traction in the market so and that's the moment i joined and we pretty much went from no users to to a whole lot of those 44 million by now yeah i remember when vs code first emerged and i was like another IDE and then it kind of took over the market.
2:17Yeah, that's true. I mean, when you think about this, right, a very well-established market, right? There were editors forever, right, IDEs forever. But all of us somehow lived in this in-between world where it's like, we're not super happy yet, right? It was like, yeah, I can do this there super, super fast and I can do this there in a good way, but I have to wait until it starts up and it has too much stuff in my face and so on, right? It was really finding the sweet spot in the middle. And that's actually also how we talked about this, right? It's really two ends of the spectrum, editor on the left hand, full-fledged IDEs on the right-hand side, where's the spot in between?
3:00And that's really what we tried to find. And I think we had a, you know, we really hit the bullseye. Yeah, absolutely. And I feel like you were winning for a ways. And now we're in this kind of moment in the tech industry where what it means to write code feels like it's shifting very rapidly. And so I'd love to kind of dig in with you about the ways in which you are thinking about this. I think, you know, initially bringing Copilot into VS Code and looking at that. But what has been the sort of VS Code journey to this new agentic coding world we find ourselves in? Yeah. So when you think through this one, we easily forget what we knew and what we didn't know.
3:50I mean, you just go six months back and what our understanding was of how coding should be and what it is today. But finally, I just had a conversation with someone and this person said, oh, 60 days ago. I was like, what? I thought that was in March or so. And we have November right now. So it's like we're working on very compressed timelines. Lots of things are happening. So and I just want to keep this in mind when we talk about all of this. So at the very beginning, we started working actually at the time as a VS Code team with the GitHub Next team. And the GitHub Next team is pretty much the internal research area.
4:31And GitHub Next had good relationships with open AI. And so this is pretty much where the AI-powered IntelliSense suggestions came from. There were already attempts in other areas before, right? This was not new per se, right? And it usually was integrated with Code Assist and so on, right? And then they've got little stars, for example, saying, oh, those are AI suggestions. and the other ones are coming from the language servers and these kinds of things. So that was the first part, right? And now we're like, no, no, that needs to change, right? We need a different UI for this. This is then where we really pushed hard into ghost text and these kinds of things, right?
5:12And at the beginning, we already had multi-line completions, but then we realized no one is using them because now you have to review code rather than to stay in the flow, right? So then we pretty much walked backwards and saying, smaller completions. So that's pretty much this whole journey of completions. Then ChetGPT came along. So we started off with the year right after ChetGPT launched with a hackathon within the VS Code team. We're saying, you know, we have four days here and we just go and build what we think we can actually build with these new models, with this new kind of approach. And that was super interesting because that pretty much immediately made clear that you cannot really just put this on from the site.
6:00It needs to be really part of the tool itself. AI really infuses in every single aspect, right? So you think about the command palette in VS Code, right? You type in there, the very moment you have good AI, you think about that it should be smart enough to figure out what you actually mean rather than what you type, right? But then you have to find the right spot between, no, no, I actually meant what I typed compared to, no, don't guess widely, right? And then there were, of course, performance considerations, right? Everything in Vue's code is about performance. All of a sudden, the AI answers were not that fast and so on.
6:40But it was super clear that this needs to be a core part of the experience. side. Then there was, for us at least, very interesting conversations because GitHub at the time, GitHub Copilot was already an established brand. And GitHub had VS Code didn't have a sign-in that you needed. So you just fired up. But in order to use GitHub functionality, you needed a sign-in. GitHub had already billing and all of these pieces in place. and then there was established brands so how do we now find the balance between what comes in through an extension compared to what is in the core and that journey this duality that really took us a while to get right so then there was a lot with chat there was a lot with what models do you actually have available not just what models, but what capabilities do those models have?
7:42How much context window do you actually get? And so on. And I think there is a difference between if you sit in a startup and you think about those problems, or you come from a world that is already profitable. And so Microsoft thinks about this in very different ways. And like, for example, it took us a really while to convince others no, we need larger context windows. You cannot work with a 4K context window very efficiently. right so there was these kinds of of challenge in the beginning was an extremely steep learning curve i think for an organization as a whole right and then i and i think over time we kind of figured this out right we're still learning so i'm not done here right so but i think we figured this out yeah and then you know when you should not just think about the last year We came with edits.
8:35We came with what we call NES, so the tap-tap-tap model. We have the agentic loop. We integrate the cloud agent that GitHub has a co-pilot coding agent. Right there is now the co-pilot CLI. We integrate all of those now into the VS Code interface. We use the agent sessions view in order to make that. We're now actively working on improving the agent session views because it's still somewhat rough in usage. So we're actually improving this. So there is a lot of these kinds of things that happen. And at the same point in time, the competitive landscape has shifted. The capabilities of the models have shifted.
9:19It's now not only about capabilities. Now it's about how long can it run? How fast does it respond? like time to first token what model mix are you actually using at the right time while people actually learn how to use it right so it's an extremely dynamic area it absolutely is yeah so an area i'd love to to kind of dig in with you a little bit more you mentioned how even from the beginning as you started to look at more advanced tab completions ai enabled tab completions not just language server you had to kind of find this balance between how much were you showing? How much were you asking to allow the developer to course correct, right?
10:01If you infer, I think one of the beautiful things about these AI models is you can do this sort of intent-based UI where you kind of try to guess what the user is doing and lead them there faster, but you can get it wrong. And so I'm curious, as you've layered on each of those pieces, as you've gone down, you know, chat-oriented programming and now agentic programming and all these things. How do you think about that balancing act of, well, we can do a lot for you, but how do we make sure we're doing the right things? So I would actually start with the sentence of saying that this is an unsolved problem, right?
10:37And it seems maybe surprising because we have been doing this. But like, for example, right, you look at the tab completion, So NES, so next edit suggestions, right? There it's always between what do you show to the user? How often do you show something to the user? How often does the user actually accept what is being shown to them? And then also how often do they explicitly dismiss it? And that is the space in which you operate, right? And you try to find, as we discussed before, right? What's the right thing between an edit on an IDE? you try to find this spot, right? And there are extremes where you can go.
11:24Like, for example, you show everything that the model proposes, right, immediately to the user. And of course, the acceptance, the absolute number of acceptance goes up and up, right? But you also annoy the user at the same time more and more, right? So you really have to find pretty much how often do you show something, how many opportunities are there that you actually show to the user. and how many of those does the user not explicitly dismiss but accepts, right? So you really have to find this explicit acceptance, explicit dismissal, and so on, right? And that is an ongoing kind of fine calibration because people also learn that that is the next part, right?
12:03A person who actually uses NES for the very first time has different expectations from a person who is actually much better, right? It also has to do, like, for example, the typing speed that a user has. So how long do you wait, for example, until you show something? A slow typer, for them, it might be way more annoying if the model is actually very fast. While a fast typer is annoyed that they don't have the proposal. They pretty much want to type in without stopping, just hit the tap key in order to accept because they anticipate what the model actually will bring. So it's a really ongoing effort.
12:42And we have quite elaborate dashboards with metrics on this where we really go back and forth and adjust those pieces and then see and run a 5 % flight and seeing does that actually change how people actually interact with it. So if you show a little bit more, has that a positive outcome, negative outcome, how often do people hit the escape key? And it's really interesting. Like, for example, the escape key hit rates are not that high. right they're around three percent the last time i checked but when you ask someone they're actually saying i hate escape all the time and then you look at the data i was like no it's actually you don't it's like so this is really but really really interesting because in the end you have to you have to get a happy developer and and happiness is is is a combination of how productive you feel of how well you actually thought you could go through your thought processes right how focused You could be how little annoyed you were, all of these kinds of things.
13:42And that is an ongoing kind of process to really get this right. Absolutely. I will say as a longtime Vim user, no amount of escape in VS Code ever feels like a lot of escape. Yeah, absolutely. I'm curious, you talked about how this can vary across, for example, different typing speeds or experiences. When you're tuning these knobs, are they global knobs applied to everyone? Do you have some sort of adaptive system in there such that, for example, if I'm a faster typer, I get more rapid completions? Like, how does that end up working? So it's mostly global right now, right? But for example, I had made rocking on how is typing speed encoded actually in the input that the model gets, right?
14:30So that it actually can take that into consideration. Got it. So it would still be a global model, but this would now become an input with features developed based on it. Interesting. Okay. And looking at that feature space you mentioned, typing speed is like how new a user is or some sort of representation of that also encoded in some way? We don't have a good way to really encode that as an experience level, right? Because you could argue from, you look at the workspace, right? And I think the Rookspace is a good indication, but it doesn't necessarily tell you if the user is new to that particular Rookspace or not.
15:14So it's quite complicated to have a good profile of a user because each of us actually goes through those different stages depending on what repo they are looking at. If I open a Rust repo, I might be more intimidated than when I'm looking at a TypeScript. repo, right? And these kinds of things also should, in a perfect world, play into what we're doing. They're not right now, but we're thinking about those. Then looking at some of the other interaction modes beyond the next edit suggestions, as you start looking at chat-oriented development or even this increasingly agent loop types of development, what are the interactivity trade-offs that you're exploring there?
16:00I mean, let me put it this way. The original chat interfaces across all of the different tools, ours included, they were interesting. You looked at those and saying, oh, this is really amazing what it can do. And at the same point in time, oh my God, this is bad. This sucks so much. I feel like this is my experience with all of AI. All right, because what I mean by this one is we spend years optimizing to go from, let's say, three seconds for a particular interaction to two seconds for a particular interaction, right? And we do all of this in order to keep you in flow state. And then we put you into a chat.
16:43And now the answer takes 20 seconds. In a good case, it might take longer. you know, a couple of minutes and some other cases and so on, right? So, and you only tolerate this because you still think that the outcome at the end is quicker, right? It's better than what you have done on yourself. So you torture yourself a little bit in order to accept the better outcome, right? And that is pretty much the baseline where we started with these kinds of chat interactions, And since then, I think you see that the world actually changed a bit. So first of all, we have much faster models. People have also developed different styles of interacting.
17:30Like, for example, one style of interacting is you use a really fast model in order to do your research. So you go and interactively, you go and try to figure out what you want to do. but this is something where you actively research and then you kind of know what you want to do that you are able to actually put this in a reasonable prompt that you then delegate and let run in a background agent for example right that's one thing there's another school of thought or another behavior that is almost the inverse of this where people go and say no i run multiple exploratory, asynchronous agents, right?
18:14I roughly tell them what I want. They're responsible for creating a plan for this. And it can take a long, right? They use a large, slow model for this, right? Until you get your plan, you work a little bit on the plan, but then you pretty much use a dumber, fast implementation model, right? And you do this because you kind of know that AI doesn't get it right all the way to the very end. And so you are actually helping along, right? You're perfectly fine to go only to 90 % and do the other 10 % manually, right? And there's not such a big difference between going to 92 % and doing the last 8%, right?
18:54That's why people accept a model that is not that sophisticated sometimes for this, right? But those are really different work styles, right? And one, you do all of the synchronicity and speed and exploring and thinking. You do this interactively and then delegate and then review. But because you actually had the thought process at the beginning what to do, the review is kind of easier compared to that the models think, that the agents think. Then I'm helping with the implementation and that also makes the whole review much, much less. So this is quite interesting, right? And there is really this kind of trade-off, right?
19:38And what you're saying, right? The interactivity versus not, right? We have different, we implemented different custom agents, right? And so like, for example, we have the one that only the ask mode where you really just go and make no edits, right? We have one where you can define yourself. What is the scope of the modifications that actually can happen, right? That's called edit mode. Then we have the agentic mode. I were rocking on or playing around with something that's called interactive mode, where the model becomes, or the agent becomes exceedingly steerable. So when you go and say, make this change in this file, it will make this change in this file and not go off and fix five other files that actually now have compiled errors, right?
20:26Because it's you who actually steers it. but again that needs to be super super fast like we have a playing mode that we ship a planning agent right so there's all these different kinds of trade-offs and we're learning at any given point in time right but it is like what it always was with developer tools there are different breeds of developers with different different interests and different preferences right and you've got to be giving the right tools the right combination of tools to each of them so that they can find a place that they're happy.
21:29for teams of all shapes and sizes, from startups and side hustles to SMEs and enterprise, and is especially great for teams that build with Ruby on Rails, Elixir, Node.js, and Python. Start your free 30-day trial and get 10 % off a yearly plan with code SCD10. Go to www.appsignal.com slash SED. That's www.appsignal.com slash SED and use code SCD10. If you're an engineering leader, you know this cycle. Your team's focused on building product, but someone in ops needs a dashboard. Marketing needs an admin panel. Finance needs a custom workflow. The requests pile up. You can't get to them all. So people start building their own solutions.
22:16Shadow IT spreads. And eventually, you're the one stuck cleaning up tools that were built with duct tape and good intentions. Retool breaks that cycle. Their AI AppGen platform gives teams a governed place to build the tools they need, so everything stays secure and under your control. Someone could type, build me a customer admin panel that manages accounts from Postgres, and they'd get a real, production-ready app with proper permissions built in. Your teams get unblocked, and you don't inherit a pile of technical debt down the road. So if you're tired of being the cleanup crew for Shadow IT, head to retool.com slash SEDaily and see how other engineering teams are democratizing app building without creating chaos.
22:59Because honestly, we could all use a better way to handle internal tools. Sometimes you just need Retool. So let's dig in a little bit because one of the things that you talked about, right, for these different modes and how steerable they are, there's obviously model differences, right? Like Sonnet loves to edit all the files. It just likes to talk, whereas some of the other models don't. But there's a lot that you're doing in the agentic harness and how you're defining these agents. Can we maybe dig in a little bit? I think coding tools are probably some of the most advanced agent software pieces we have out there.
23:36How do you build it? What's the stack for defining one of these agents, an ask agent or what have you? So, I mean, at the very core, right, and I'm pretty sure you have heard that answer several times, right? An agentic loop is not that particular complicated, right? It's like you give it a bunch of tools. You give instructions how to use those tools. Most of those instructions are actually with the tool description. Sometimes they are outside, right? Each model has certain kinds of preferences, right? And there are prompt guidance for each of those. Some, for example, like that you tell it, oh, give the user an update from time to time.
24:16Others actually stop the agent loop when you give such instructions. Like, for example, Codex is one of the models that stops and it wants to give an update to the user. So there are all these kind of differences, but that is the basics. And then you've got to pretty much instruct the agent. And that's actually one of the more interesting problems is when is it done? At which point in time should it actually consider to be done? And so that is the basics of all of this, right? So what we then, when you ask a custom agent on top, right, that's actually something where we say, okay, in a custom agent, you can define what tool set is available to that agent.
25:01So out of all available tools, and they actually can come from different sources, right? There are built-in tools in VS Code. Extensions can actually define tools. and on top of this you can install MCP servers. So you can have a quite large set of tools. So then you can specify pretty much in a custom agent file, you can say, oh, here's the tools that you should make available. And then after that you have pretty much the normal syntax that we use for everything else, which is a Markdown-inspired syntax where you go and say here's that you can say how an agent should actually operate. And then you have seen, you know, many people have seen what Claude's skills look like and so on, right?
25:46All of the definitions are pretty much comparable to each other, right? That kind of aspect, right? It's a markdown file where you give references to other instructions files, where you can say what tools to use under what circumstances, right? Like, for example, you would go and say something like, hey, because I'm in a workspace that actually I know I have defined it all. I really like that you use the test runner tool rather than go and do NPM run tests or cargo test or something. I was like, you say this explicitly, right? Or sometimes a model goes and kind of assumes that, let's say, that your tests are wrong, right?
26:28Right. By the funnily, I think that was the first moment I thought models really get intelligent and they pretty much rewrote tests to a search room in one way or another. It was somewhat obfuscated, but that was pretty much the bottom line. I was like, oh, all my tests passed. So it's good. Yeah, you did. So sometimes you just go in very explicitly and say, never touch test files. everything is good there. You might make, if you do a refactor, you can adapt them, but asserts are untouchable, for example. So it's pretty straightforward like this. And then a lot of work then actually goes in and saying, okay, how much instructions do you need for tool usage?
27:15How much guidance do you need to give? And that's where pretty much this whole machinery comes into play of what emails you run, how many of them you run, how often you run them, right? How do you really think about actually assessing slash evaluating an outcome, right? And that is quite different, right? Like, for example, we run like the rest of the industry, you know, Sweebench, for example, right, as one of the benchmarks. And we're not just looking at the resolution rate because the resolution radius you get from A to B, right? That can be super, super messy, right? We look at how fast do you get from A to B?
27:59How many tools did you actually call? What is all of the amounts of tokens that you actually use in order to get there? Did you call the tools that we think you should use, right? Like, for example, if you have a terminal tool, you can use that tool and you can get to the end of it. And that's perfectly fine if you run as a background agent. But if you run as a foreground agent in VS Code, right, then the user actually expect that when you say that the tests failed, that they can look at the test explorer and see the failing tests and click on there, right? So you wanted to use that tool then.
28:37If you have a watch task, it should not try to spend, you know, a bunch of turns in order to figure out how to build the project called the watch task. It's right there, right? So these kinds of evaluations we actually then run and compare. And then there's the fine tuning going on, work changes going on, how you group tools in different categories that actually makes a difference. There is a lot of that work that actually goes in. Yeah, it's deceptively complex inside of this very simple wrapper. So there's a couple pieces I'd like to dig in on that. So one, as you mentioned, different models tend to have different preferences, I guess we'll call them.
29:18in terms of how they invoke tools, how they check in with the user, things like that. In the UI, I have this very simple model switcher. I'm just changing models. All of that is opaque to me. Are you customizing tool descriptions, the core agent instructions, all these different things by model to help them behave consistently? Or how is that all functioning? So there is a current state and then there is the near future state. The current state, And you can, I mean, we're all open source. So you can take the repository, you can look at this and you actually see. I mean, when you run inside Vue's Code, right, we have a log view where you can see every single call that is actually being made to the model.
Read the full transcript
30:03You see every single detail of this. So you can actually verify my words here. So in your own day-to-day experience. So the current state is we actually have specific prompt for, I would say, roughly every model family, sometimes more detailed. Sometimes we go down and say, oh, it's not just a GPT model. We really separate between GPT-5, GPT-5 codecs, GPT-5-1 codecs. We move so fast that I say 5 rather than 5-1 as an industry. So different entry points pretty much. in our main profile generation, right? And then we'll pretty much pick what tools are available for that particular model, what are additional instructions that need to be given and so on.
30:53We customize the instructions that are outside of the actual tool descriptions. But we don't have a model, we don't have it yet in code where we actually say, no, we know exactly that this is the tool description that works better in this particular kind of model. That's actually something that we discussed several times, never quite made it to the point of, yeah, now it's coming. The last iteration plan, we again had the same conversation. And I was saying, oh, we do this right when we have shipped at the beginning of December. We'll have model-specific tool descriptions, right, where pretty much the prompt file can overwrite and saying, oh, if this tool shows up, right, And here's actually the tool description that you should use.
31:43So another kind of detailed topic in here is how you assemble context for the agent in terms of you're operating in this type of repo. You have these things beyond tools. There's people who do different amounts of like pre-injection of context. Maybe it's not just system prompt, but it's got a whole bunch of additional things. how do you think about the right ways to sort of present things to the agent so it kind of starts out going the right direction versus everything's in on-demand tool calling so there are a couple of things here that and let me actually start with the tools first right so most models these days have been trained with a particular tool set right so out of the box they already know a certain set of tools, like for example, apply patch, right, for GPT models, string replace for sonnet models, those kinds of things that they are must have, right, in those individual tools.
32:46And then the next question is beyond this, how well do models actually generalize, right? And so you will, you want the tools and then again, right, the context, as I said, in what kind of environment is that particular prompt now executing, right? Is that foreground agents, background agents, and so on? So that's the first kind of question, right? Those tools, how are they actually represented in your prompt? Then the next one is, let's say you have a couple of MCP servers installed. So most of, some MCP servers have only one or two tools, right? But others come in with dozens and dozens of tools.
33:31Most models have a limit how many tools you can actually put in a prompt, so 128. But then on top of this, there's a lot of tokens that you actually put in there. So how many tokens do you actually want to spend on tools that are rarely used or only in particular specific situations are being used? So a technique we're using there is, like for example we go and take all of the tools that an mcp server gives us and we actually now create pretty much virtual categories of tools right and in these kinds of virtual categories they are represented as tools in their own right we give this to the model at the very moment the model decides to call one of those virtual tools then we pretty much expand it so but now you have immediately this kind of trade-off discussion, which is the very moment you do this, you actually have...
34:30You've blocked a KV cache. Exactly. Exactly. Right. So in some cases, right, some models actually support that you put this at the end of the prompt. Others actually don't. Right. So now you immediately have to make this trade-off. And that's pretty much where a lot then also of evals come in. Right. You run in all of those different configurations. You compare, this is optimizations functions that you have to hit here, which is like, oh, if I blow my cache once or twice over such a long time, I'm still good, right? Or you're saying, no, actually, I can run with a slightly larger prompt, right?
35:10That is fine because I have a cache hit rate of 87 % or whatever, right? And I was like, this is okay, right? There's no big advantage here. So it's a constant kind of trade-off. And that is also true for all of the other context part, right? And there is not a real stable. I mean, there are some stables. Like you say who the user is. You say what the repository is that the user runs in, right? But the very moment already, like how much information do you give about the project itself? Like, for example, we put this kind of prefix in where we're saying, This is what pretty much the top level of the project looks like to the user.
35:50And we still believe that this is actually reasonable token spent. But then on top of this comes what we dynamically include. And dynamic inclusions is clearly like if you have an agents MD file, we have custom instructions that we actually do support. And custom instructions actually can be tailored in different ways. They can be just in a certain location, and that's fine. They can, in the front metal of those custom instructions files, you can say apply to, and then you can actually give glob patterns and saying, oh, you know, in this particular test folder, in a TypeScript file, this is a file that actually applies.
36:31And then there's yet another mechanism where you can actually give a natural language description under which circumstances that's custom instruction implies. So, and then we actually start collecting these and actually putting them also in the prompt, right? And that is actually a process that I think is the one that is the most valuable, right? Because you make sure that the user is in control of how much they pretty much AI prepared their code base, right? But when they did a really good job, then they can actually make sure that the agent extremely quickly gets to the right place, knows where to start, knows where to look.
37:14Let's maybe talk a little bit about interactions between agentic pieces and the IDE itself. You mentioned a couple examples of this, of if it's running in the foreground, use tools that connect to parts of the IDE so that they're running in the right place rather than using the terminal. But what is the surface area that you expose to the agent, to the IDE? And how do you think about changes coming from the agent versus coming from a human? So, I mean, it's all about how you interact with it, right? So again, the most straightforward form is you actually have a foreground agent running in VS Code.
37:59And that foreground agent, we give it actually quite an interesting set of things that it can do. It can look at terminals, it can read selections in the terminal, all of these concerns. It can run tools, watch tasks, all of these specialized edit tools that we then actually, where we actually are able to run pretty much snapshots at a given point in time, so that we can show you, oh, here's all the appropriate diffs and so on. There's a good chunk of tools that we actually give a foreground agent. In a background agent, that's quite different. In a background agent, we give significantly less.
38:40Why? The first thing is if you run the agent in the foreground, you have this kind of expectation that the agent actually is reasonably quick. If you think about this more, that is the interactive part. You don't want to sit there and wait two minutes and twiddle your thumbs. You want to get the answers relatively quickly. And then, again, you want to make sure that this all kind of is like the extension of what you would do anyways. But when you go and move something into the background, then you clearly don't want that it touches your UI state at any given point in time. You don't want it to, like we mentioned the example of the test runner a couple of times, right?
39:25You don't want it to mess around with your test runner. You don't want it to open up a terminal on you, right, so that your mouse all of a sudden clicks a different place and so on, right? So there are different tool sets that you're actually giving. A cloud agent is yet different. So while a background agent is still running on my local box and is still exposed to me closing the lid, neither a cloud agent is not. And there it's about in what containerized environment is that agent actually running? What is the project that you actually have? Can it build successfully in that container? or can it execute and test run in this container?
40:06Yes or no, and so on, right? And the agent. But again, the cloud agents have significantly less tools in order to do so. Remind me of your question again. Well, so my question was kind of how you think about those interactions. And you've given me a fair amount. This actually leads to something that, or a curiosity I had as I was listening to you, is how much does this differential exposure of different types of tools end up influencing how well the agent does. I'm imagining the same prompt in a IDE context versus a background agent versus a cloud environment might result in quite different coding behaviors.
40:49The model choice, I think, has a much bigger impact on the actual outcome. So what we're trying to do is really straddle the line between the user experience and the success of the agent, right? Because as we said, right, if you have a terminal tool, so execute terminal commands tool, you can get really far. You don't need an edit push, you know, and cat command with input redirection, right? And you see this is your edit and then it writes it to the file system and so on, right? So you don't really need a whole bunch of tools in order to make an agent successful going from A to B. There are some differences, like, for example, the industry introduced the to-do tools in order to have longer running agents, self-organizing, and so on.
41:48But in big parts, when you think through this, you don't need a huge amount of tools in order to make that successful. So that's one. So when we actually bring it in the foreground and give it more tools, then that's very specific to the environment, right? It's like, for example, when we come and say, no other agent needs to say, hey, maybe you should install this extension to have a better user experience, right? But inside VS Code, that clearly is a tool that is available and is particularly interesting. And you scaffold, for example, a new project, right? You go, you say, oh, I want to do this, right?
42:22And now create this workspace for me. are like, oh, and you go because you told it to go, but you don't have the Go extension installed. So it makes sense that the agent actually goes and saying, by the way, go install. Should I install the Go extension for you? So it's really more about the user experience that we try to give to folks in the appropriate environments they are in. That's really the biggest difference. And then it's also coming back to how you interact actually with an agent, right? I think for a background agent, I don't want to have a lot of interaction. There I want to have context isolation.
42:59I want to make sure that it's even running in a sandbox environment so that I'm not bothered by tool calls, right? So that I have to approve tool calls and these kinds of things, right? But in a foreground agent, right, that is more like, well, I'm not quite sure yet exactly what I'm doing, right? At least that is the use case I see primarily. people are talking about code, right? They kind of go and make a selection and say, hey, change this, right? Short prompts, right? They are not particularly long, right? Sometimes people go and actually use NES in order to start a change, but then they don't finish it and just go and forget everything, you know, hey, finish this up, right?
43:45And it should then be very quickly just, you know, in this particular file, just do the rest. or when people create new test cases, right? So test case generation usually is not something that takes particularly long, right? I mean, it depends, but in most cases, right, it's pretty straightforward. Pushing this in the background, right, and then coming back after a while to review it and all of this, it's usually more, it's like, no, no, I do it right now, right? And then immediately run it and then let me review it, right? So that there is no cheating going on, right? That the tests are not already playing to how the actual behavior is.
44:26So very different styles of interacting, right? One is that the foreground agent is really like short interactive behavior, talking about code, pointing at code, collecting pretty much the context that you want. That's all very code specific, right? Two background agents you don't really do, right? It's like you try to be precise the moment you started, right? And then at the end, yeah, you can follow up a little bit if you want, right? But it's different expectations, different levels of preparation and so on, right? And what I'm actually saying here is, right, there's more to this, which is when I say you talk about code, then I'm more like you actually are a person who cares about code and you are actually really rocking on something where you need to guide in regards to software architecture, certain patterns that you want to enforce, et cetera.
45:23There's this whole other world where you don't care, right? So you don't care what the code looks like. It's really just outcome-oriented, et cetera. And those lines, they also shift back and forth, right? They're good within the same project, by the way, right? 100%. I care about my core architecture. This tool, just vibe it. I don't care. Yes, exactly. Right. So it's exactly this, right? Where you pretty much, and this is a really interesting point, right? And I think as an industry, we might not, or we maybe don't talk about this enough, right? Which is that, how do you actually AI ready your code bases?
46:01That is exactly right. If you have a project, like, I mean, our code base, right? The initial commits and all of this are more than 10 years old, right? And since then, we built on top of this, right? In order to make our code base AI ready, we really have to think about what are the core abstractions that we really, right? Each and never go and change those things, right? If you should change anything here, we tell you, right? But then there are other parts that are a little bit more peripheral, as you said, right? some tool or so that you just want to have on the site now tool in a more generic way right that's like yeah just just go do it right and you might even just check in pretty much the prompt file that you use in order to generate this right so it's a very very different right so you think about what is untouchable what kind of lives at the periphery right where you care where you don't care most people really love using test-driven development for a bunch of the same tests are pretty much my prompts that I use for the implementation side.
47:08So there's really this great flexibility in people operating in quite different ways. I'm curious, when we talk about these different modes of operating, and the fact that we kind of flow between them, how do we connect the dots? So an example that I'm going to bring forward, I'm very interested to how you would think about this. I'm often working on something kind of interactively in that interactive mode, I'm thinking about it, and an idea comes, oh, it would be great if we do this. And I have a set of sort of predefined research style prompts that I can just kick off. So I'll like kick off a background agent and say, okay, go and research in my code base what it would look like if I were to do something like this, write me an analysis doc and go.
47:48It'll go off and do it as I continue on my main line. And at some point, I want to come back and pull almost suck that into interactive mode. Now, I can do this right now with like branches or doing things like that. But I'm curious if there's something in the IDE that lets me kind of, it's almost like I'm pushing ideas into the stack, and then I want to pop them down into my interactive world. There are different ways of thinking through this, right? So when you actually, it's interesting, because we really just discussed that. And we have a mockup, we have not implemented this yet, where a similar discussion came up, but more about at which point in time do I go back to a background agent or to a cloud agent, right?
48:36And yours is similar because it depends on the output of what the cloud agent generated for you. In your case, you want to see an analysis, right? You want to look at this and so on. So the way you think through this is whether you kick them off And at some point, you've got to go back. So you need an indication that it's telling you that it's ready for review. But then the interesting thing is that this is not necessarily just looking at something does not necessarily mean that you did all of your due diligence. So in a way, you need interaction saying, yeah, it's ready. You can go there. But then at some point saying, I took action on this is actually good.
49:19right so you want to have this awareness of those right and when you think about this i mean there's really really how should i say prior art right i mean when you think about email management tools and so on they're quite similar kind of characteristics so one thing that we had was pretty much at any given point in time right when you interact with this chat and so on clearly you can make these things disappear but you pretty much have awareness about where your background agents are and which ones are ready to review, right? You don't see the, if you don't want to, the running ones, you don't see the ones you took action on, but really just those that you haven't acted on yet, right?
50:00But they are done, right? They have produced what you asked them for. And then one of the mock-ups that I just talked about, right, is pretty much where we bring this right into pretty much the very top of the title bar, right? Where you have pretty much something that just can come down, right, as an overlay. And you just, like, super quick, you just see it, right? And it needs to be a first-class citizen. If this UI, what I describe is what it will look in the end, is a different question, right? But it is absolutely clear that you need this kind of peripheral awareness. No matter what you do, you need this kind of peripheral awareness that something else is right, right?
50:42I mean, there are workarounds for this, right? And that's actually something that sometimes we're maybe too focused on the one tool that we own and that they operate in, right? But again, right, I mean, if you want something from me, you select me, right? And I get a notification that you want something from me, right? So integrations and these kinds of other work environments that actually tell you, right, that you have a Slack channel with your agent, right? And just comes up and actually says, hey, I'm done, right? It's good, and you have the notification, and it fits in your other workflows and so on.
51:16So there is a lot to explore here. That's where I'm going. We can do some things in the IDE. We can do some things with GitHub and github.com, but I think it's not necessarily where the line is. If an agent is an actor, there's a lot of other kind of tools that already are custom-built for actors. right and so i think we we've got a broad analogy of thinking through these problems these problems i love that and it's a good segue to another topic here which is i mean vs code has always been very extensible very plug-in centric very open uh how are you thinking about that within this new world are there things that need to change i i know you mentioned mcp servers that's definitely one way of interacting, but like, what does, are there changes going on in that landscape as well?
52:12There's an interesting duality here, right? One is that in order to get new functionality into VS code, you had to write an extension. If you actually have direct access to LLMs who can actually operate some of that, right? You said, oh, I have a custom prompt file that does X for me, right? So to kick off and research background aging. Wow. But you can also have custom prompt files and do actually things in interactive mode in the IDE. And now if you give you a capability to keyboard shortcut this, all of a sudden there's extensibility right there without writing any extension and so on, right?
52:54So interesting kind of duality here, because some things you still want this extension, but some others, you have a lot of flexibility already without even required to write an extension. So that already changes extensibility, right? Then MCP services in your newish, new in these days, right? And a newish concept, right? And that is interesting, right? Because it has MCP spec covers a lot, right? But the most interesting one, the most used one is the one that you actually can make tool calls, right? You get pretty much a tool host, right? So, and that is clearly, I mean, there we're still very early, right?
53:40When you think about this, kind of obvious some of the aspects, right? But now you actually have people who kind of say, like Anthropic just posted this a couple of days ago, right? the whole part of a couple of weeks ago about Procramatic MCP tool calling. And then you're like, oh, I can clearly see where the cups are, but now we're really just making different APIs. It's like, why are we not calling them the real APIs? Why is there a differentiation between the normal APIs and the MCP servers, right? You can see that they then start fusing together, right? And that's just, or MCPs is just the API that you publish, right, and nothing else.
54:20And so that would make a lot of sense, right? But now you put this together with autonomy of an agent, and you end up in a potentially scary world, because now you need to think through the security implications. You need to think about how do you actually control this, right? What do you allow? What do you not allow? identity management permissions for agents, all of these kinds of things. So what we are doing today is we create a sandbox for some of this. But it already starts like people go and say, oh, context seven. It should be quite careful how I phrase this so that your takeaway is not, oh, this is context seven is an issue.
55:08But I give this as an example. right so you register your website there right it's crawling in markdown files i'm sure there's some sanitization going on but you know wherever there's sanitization you can actually play it now you have an mcp server that actually finds those pieces of documentation puts it in your prompt and now what now it starts building and executing code and so on right so it's this poison the well kind of problem right so you you need in a way control all entry points into this but this creates the most awful user experience yeah no this is this is a fascinating domain because in essence all of these large language models they're another form of running computation it's word programmed rather than formally programmed and so any mcp server is injecting code that's running on your box.
56:06Do you trust it? Do you trust everyone who's able to get anything into that? Yes, it's exactly the question, right? And what we just, in VS Code, for example, we had a, so the way you actually do tool approvals, right? First of all, this is just running the tool, right? So again, input outputs, right? So what you just said is the part where saying, oh, are you okay with this command being executed? And it can be local. If the MCP server runs local, it might install packages, right? UVX or NPX or SI if it didn't run yet, right? So interesting, right? But there's a number one, right? Are you okay with being this one executed?
56:53Saying, okay. So how do you want to do this for this session, for this particular call, right? Is there certain patterns in a command that you actually want to allow, right? There's a lot of room in order to get this kind of configuration, right? But then you make the call, right? And let's assume this is a remote MCP server. Now, what you pretty much said is, I'm okay that this server is being called, maybe with my authorization token, right? But now there is a response being computed and that response also goes into your either in a summarized form or in its actual form into the history of your chat, right?
57:35And I was like, oh, what now? Now you need to pretty much review everything that comes back. But again, right? That is awful. So you want to use then specialized security models to actually do this kind of monitoring, right? But now you're pretty much in this kind of, you know, who wins, right? Head-to-head wins, right? So that's what I haven't touched on. It's like this is chatted in one way, right? But now you go to terminal in terminal commands, and do you want every terminal command to ask for permission? Now you've got to go and say, no, no, let us actually read and understand what that terminal command is.
58:18And if that terminal command actually feels safe, let's do it. But then you also need to give the user control. I mean, a user in an enterprise setup might think about what tool should be called without permission, explicit, every single time permission, differently than if I'm running on a VM in the cloud that I just leased for this particular kind of use case, for example. I do wonder if it leads towards a world where essentially all development is actually happening inside of a container or VM. I think if you think this all logically to the end, that is the, I think, the part you're getting to, right?
59:03But then still, you need to control the inputs and outputs to this container, right? So fetching a web page, right? When you go and say, hey, I need the latest version of Node, right? You get an install command that runs in the terminal and this moment something comes into your box, right? So yes, you want this to be safe, right? You want to be able to close the doors and say you cannot get out of it, right? My point here is the problem doesn't go away, even if you put this into containers, right? But you can control the environment in a better way. But we still need to think about how to make this a good user experience, how to make that understandable for you and so on.
59:49And when I say understandable, is like, we're now talking about, we didn't say this explicitly, but I think our conceptual model here when we talked was, oh, there's this one or two background agents that I have, right? Now multiply this with 100, right? And all of a sudden, that is a very, very different problem, right? And we need to rock through all of this, right? I mean, that's already starting, right? When you think about this, right? So with cloud agents, for example, let's say you're on GitHub, right? You groom, you assign a bunch of issues to co-pilot, or you actually have auto-triaging enabled, right?
1:00:28So you go and say, oh, we auto-triaged certain ones. Now you get those. You get those PRs for this. You review them. Reviewing five or 10, depending on size might be good. Reviewing 200 a day, you're out of luck. I've reviewed more code in the last six months than I can remember, right? It's ridiculous. It's wild. So I think this gets to kind of where I want to take us towards the end. And we're getting closer to the end of our time here, which is where do you see this going over the next year or two? I hesitate to go too much farther out because as you've highlighted, things are moving so fast.
1:01:07But like, how is VS Code and this whole world of how we're managing the writing of code? I say it that way because maybe we're not actually writing the code, but we're managing the generation or writing of code. where do you see it going and what's coming down the pipe? I'm not sure I look two years down the pipe, right? Because we might surprise ourselves how quickly we end up in a different place. But when I think about, so we're still learning about the interactivity models, right? And that's active research where we go in and say we implemented one way, we implement the other way, we look how people actually accept it, right?
1:01:48Where does the percentage go? In the beginning, it was a lot of tap, tap, tap. Now it's more like, oh, the percentage is lower, but depending on the experience level of people and what part of the coding, right? When we talked about there's something that is really sacred land and then there are other things, right? So all of these kinds of things are influencing this, how those interaction models are. And that will change and we'll figure this out, right? I mean, different ideas will come from different areas and so on. But I think then there's this whole point about how do we use agents effectively?
1:02:26And I think that is also a very hard problem because, again, people, we run many agents parallel. Yes. But now, in what circumstances? If I have a project like VS Code where we go through 3 ,000 issues a day, sorry, months, not a day. but that's projecting forward two years right probably right but uh 3 000 issues a month right and that was just based on human activity right you you put now ai into the mix that number needs to go up right then how much of you can do more parallel work right because there are different boxes that you can execute, right? You own, let's say, as a team member, you own a bunch of tags, right?
1:03:21Those are yours. And you can go and have a couple of agents running on each of those tags. And that's kind of fine. But if I go and think about creating something new, right, then actually running multiple things in the background, that's way more complicated, right? Because it's easier to think step one, two, three, where one, two, three built on top of each other rather than, oh, it's one, and then it's two A, B, C, D, right? It's like it's way, way more complicated. We hit cognitive limits. Yeah, absolutely, right? As a human, we are pretty much the weakest link in the chain, assuming that we really get to a place where that code that comes out is in a good shape.
1:04:02But we're not there yet, right? But you can see that this is happening, right? So how do you actually really work with this level of parallelism and so on? So I think there's work to be done. I think where we clearly will end up, right? And that is now, just think about real world large scale operations, right? Where you go and say, hey, I have to make a change here. I want this new vertical feature to go in. but it now actually touches dozens of repositories, different service deployments, all of these kinds of things, right? So now you end up in a world where you pretty much need to create a plan, almost like a project plan, right?
1:04:49So agents need to go in, right? One solution is you have a monorepo and you just have one agent running around, right? But what is more likely is that it stays a distributed world, at least for many people out there. And then you need different instances of agents that actually went to manage and delegating to other agents. They are running and they need to communicate to each other, right? They need to report back where they are. The reporting back cannot be a markdown document anymore. They need to potentially go back and say, no, here is really the change tracker, right? Here are the different issues.
1:05:26They're linked to each other and so on, right? Maybe you need to see this on the planning board in order to understand this. It has a lot to do with, again, the human is the weakest link in the chain. It's about transparency, what is actually happening and so on. So I think there is a lot that will happen in this particular area where we need to go. Then there is one other aspect, but that's the one I personally struggle a little bit the most with, which is what is the, it is very easy to say, oh, I can code wherever I want, right? And if I have an idea, I just type it in my phone and send it off.
1:06:06And so I'm certainly true, right? There are some use cases for this, but I'm not quite sure how, I mean, I want to work on my iPad. That's clear, but this is just a replacement for, you know, I just sit in a different place. I don't want to use my laptop. So these kinds of things. But really, what's the role of mobile? Of smaller, smaller form devices. Really, how much do I want to do on my phone? I can see maybe voice that plays a big role in this one. But other than that, still, I don't want to review code. I was going to say, kicking things off, great. Reviewing code, miserable. That's right.
1:06:51That's right. And then I think that last part, and again, it's actually not that surprising when you think about this, right? We like to be creative, and in what environments are we creative? And I really could see that we're still in very traditional kind of collaboration forms. And as we said, you could have a Slack channel with your AI agent, for example, right? So these kinds of interactions. But I think the other one is, what is it that we really like as humans, right? We like to stand on a dashboard or together, you know, sit harder together and do something together, right? There's materials on the table that we shuffle around in order to talk, right?
1:07:36So these kinds of things. How would you replicate these kinds of things, right? And I think sometimes Microsoft had this studio PC, right? It's a large screen that you could flat down, right? And you now take something like this, right, where you can draw on the screen, particularly when you do UI development, for example, right? You go, you draw on the screen, you say, this is what I want, right? And so on. And then you can actually talk to it at the same time while you're drawing and saying, no, this here, right, should be a little bit more over here. What do you think about this? Give me two alternatives, right, and so on.
1:08:15And then all of a sudden you have a very, very interactive rock star that makes us happy as human. There's a lot of dopamine in this, right? And it can be very well AI supported by voice, by how it's actually multimodal inputs and so on, right? And there might be or there might be not code involved in this. There's code involved in it behind the scenes, right? But again, right, this is a rock star I can clearly see, right? And I think we will see a lot of this coming forward. I think that's a great cut point.
From the publisher
Visual Studio Code has become one of the most influential tools in modern software development. The open-source code editor has evolved into a platform used by millions of developers around the world, and it has reshaped expectations for what a modern development environment can be through its intuitive UX, rich extension marketplace, and deep integration
The post VS Code and Agentic Development with Kai Maetzel appeared first on Software Engineering Daily.
