Making Sense of Agentic AI | ThoughtWorks Birgitta Boeckeler

12 Nov 2024 · 48 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Dev Interrupted Podcast Episode Notes

Episode Overview

Title

Making Sense of Agentic AI | ThoughtWorks Birgitta Boeckeler Hosts: Andrew Zigler, Ben Lloyd Pearson, Dan Lines Guest: Birgitta Boeckeler, Global Lead for AI Assisted Software Delivery at ThoughtWorks

Description: This episode discusses the impact of AI in software delivery, exploring both its practical applications and shortcomings. Birgitta Boeckeler shares insights from her research on AI agents and tools, while Dan Lines provides perspectives on how engineering leaders can leverage these insights for effective AI integration within their teams.

---

Key Themes and Discussions

AI and Business Impact

  • Understanding AI's Role: The conversation begins with questioning whether AI agents and tools drive tangible business impact or merely add to developers' workloads.
  • Practical Applications: Boeckeler emphasizes the need for AI to enhance software delivery effectively, highlighting her role as a practitioner focused on applying AI to real-world software challenges.

Birgitta Boeckeler's Research

  • Agentic AI: Focus on utilizing AI agents to assist in software delivery. Boeckeler experiments with current tools and discusses their effectiveness.
  • Experimentation: Detailed insights into experimenting with legacy systems, where AI was used to improve code search and understanding of domain-specific terminologies.
  • Challenges Found: Despite positive outcomes in some areas, limitations were identified, particularly in debugging, testing, and managing legacy code complexities.

Example Case Study

Legacy Code Application

  • Open Source Application (BAMNI): Boeckeler used this medical records application to evaluate how AI could assist in implementing features effectively.
  • Results:
  • AI helped in understanding domain terms and improving code search beyond simple string searches.
  • Struggled with legacy code complexities, particularly in reproducing behavior and creating tests due to the absence of unit tests.

The Future of Agentic AI

  • Defining Agentic AI: Boeckeler defines agentic AI as models equipped with tools to perform specific tasks autonomously, sparking a discussion about the definition and scope of what constitutes an AI agent.
  • Promising Tools: She identifies several emerging tools in the agentic AI space, including GitHub Copilot and others, noting that many currently overpromise their capabilities.
  • User Experience Concerns: Current AI tools still require human oversight for maintenance, and the user experience needs improvement for effective integration into developers' workflows.

Engineering Leaders' Takeaways

  • Measurement and Impact: Dan Lines emphasizes the importance of measuring the impact of AI tools within development teams and understanding their business value.
  • Standardization: Encourages leaders to implement AI tools that standardize processes, such as AI-based pull request reviews.
  • Context Awareness: Future tools should integrate data from various sources (e.g., Jira, support tickets) to provide a comprehensive understanding of projects and enhance productivity.

Key Takeaways

  • Current Limitations of AI: AI tools can assist in code generation but often struggle with complex tasks, particularly in legacy systems. Developers still play a critical role in oversight and maintenance.
  • Adoption Strategies: Engineering leaders should focus on specific, realistic use cases for AI implementation, ensuring that these tools contribute to measurable business outcomes.
  • Continuous Experimentation: Organizations should foster a culture of experimentation with AI tools, balancing their implementation with an understanding of their limitations and potential risks.

Resources Mentioned

  • [Birgitta's LinkedIn](https://www.linkedin.com/in/birgittaboeckeler/?originalSubdomain=de)
  • [Birgitta’s Website](http://birgitta.info)
  • [Martin Fowler Memos](https://martinfowler.com/articles/exploring-gen-ai.html)
  • [2025 Engineering Benchmarks Insights Webinar](https://linearb.io/event/2025-benchmarks-report?utm_source=Substack&utm_medium=referral&utm_campaign=202410-Dev-Productivity-Insights-IMC)

Support and Follow

  • Subscribe to the [Dev Interrupted Substack](https://devinterrupted.substack.com/)
  • [Leave a review](https://ratethispodcast.com/devinterrupted)
  • Follow on [YouTube](https://www.youtube.com/c/DevInterrupted), [Twitter](https://twitter.com/DevInterrupted), or [LinkedIn](https://www.linkedin.com/showcase/dev-interrupted/)

---

This structured markdown file organizes the podcast's content effectively, summarizing the discussions while emphasizing key insights and takeaways that are relevant for software engineering leaders and teams navigating the evolving landscape of AI in software delivery.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00And I haven't really ever in a demo video or for myself seen a tool really solve the original problem. There's something weird going on with these marketing videos where even the use cases that they choose for the video often don't make sense.

0:22Hey everyone, I'm your host, Ben Lloyd Pearson, Director of Developer Experience at Linear B. And today I'm delighted to be joined by Brigida Bokela. She's the Global Lead for AI-assisted software delivery at ThoughtWorks. Vegeta, thank you so much for joining me today. No, thanks for having me, Hyben. So we're going to spend a lot of time talking about agentic AI today because, you know, it's been a big part of your research in recent months. And one of the things that I sort of stumbled upon while I was learning about some of the work you've been doing is some of these articles or memos, as you've been calling them, that you've been publishing over on Martin Fowler's website.

1:01and I've seen like it looks like a series of experiments that you've been doing to to test out some of the cutting edge of AI so maybe let's just start there like tell me about like what's going on there and like the types of experiments that you're doing and what you're finding out from them yeah I would say I mean first of all like I'm so I'm a developer by trade right developer architect you know practitioner right so I'm not an AI or machine learning expert and the way that I see my role at the moment is that I'm a domain expert at effective software delivery because I've been doing it for over 20 years.

1:37And, you know, and it's not just like coding, but it's also like effective teamwork and stuff like that. Right. And so that I see as my domain expertise. And now I'm trying to apply AI to that domain. Right. So how can we use AI to be better at coding, to be better at software delivery, teamwork and so on. Right. So and but of course, to do that, I have to understand the technology below the hood, like to a certain extent to understand what are the possibilities and understand if a tool is claiming a certain thing that it can do you know do I think that's like viable like do I want to try that or not like or where do I see this going where do I see this potential right and that's why one of the things I was trying of course like one of the hot topics right now is like how do you use agents or like agentic applications to help with software delivery as well right so that was something I was trying but also a lot of the things that I'm trying is not even me myself trying to build a tool, but just trying to use the tools out there and seeing if I can come up with an example where it's like a, as realistic as possible workflow that I would usually see on a team and just like try to stitch those things together and see if it actually gives me value.

2:44Right. So another one of the things that I wrote about was I used like an issue and like an open source tool that is actually like a business application like tool. And it's like very old code base so you could call it a legacy code base right and I was trying to see okay how would I usually try to find out how to implement this issue this ticket and can AI help me with that right so that's kind of like the the stuff I'm doing so I'm experimenting but I'm also talking to our team so I work for a consultancy right so we have teams in a lot of different domains in a lot of different situations a lot of different tech stacks so I talk to them and you know what what are you using what's working what's not working then I try things I tell them about it I talk to our clients.

3:24So that's kind of like my role right now. It's a lot of different things. And the fire hose of AI, like of change in the space is actually still going. So even with the full-time role, it's very hard to keep on top of everything. So nobody has to feel bad when they feel overwhelmed by what's happening. Yeah, and I love the approach of, you know, you mentioned how you're sort of focused on figuring out what's viable, right, in this day and age. Because, you know, I think we talk to a lot of companies right now that are just, they're like, I don't really know that anyone has figured out how to adopt Gen.

3:55AI yet. They're all like, everyone is experimenting at this point and trying to just learn, like, what are the tools that actually were today? How can we implement them in our organization? And I love the, you know, sort of the legacy case that you brought up, because that actually is one that I think has potentially a lot of promise, right? but also a lot of interest because there's just a really high demand for something that can help manage that. So yeah, so what have you learned? Maybe do you want to dig into that one a little bit? What did you learn about as you were exploring generative AI for a legacy system?

4:30What did you uncover from that? Yeah, so this memo that I wrote basically, so what's always a challenge in this space is as long as I'm not on a team, I need examples. And I need examples that are as realistic as possible. And so we have a team in ThoughtWorks that, like we've had this for years, that is working on an application called BAMNI, which sits on top of an open source application called OpenMRS. And BAMNI is open source as well. And it's like a hospital medical record system, open source. So this is used in countries a lot or in hospitals that can't afford to buy like expensive medical record systems and so on.

5:10Right. So this is like a much more realistic example for me, for my use case, because the business application. So it's not like a Python library or something like that. Right. Because what I'm interested in, like for our clients is like, how does this work for a team building an actual application? Right. Not for like a library or something. So I use that as an example. And what's nice is that all of their JIRA tickets, their Confluence documentation, everything is public and open. Right. So I also don't have to worry about data confidentiality. And I'm not familiar with this application. So I'm actually a user exactly in that use case, right?

5:42So there's like an application that has gone through like, I don't know how many Java versions. I don't know how many different developers, you know, and different approaches, right? And I'm new to this application and there's a ticket and I need to figure out how to implement it. And I don't know the domain terms, right? So it was mentioning things like HI types and FHIR. So all things that I didn't know, right? And I wanted to see, of course, I can search their wiki for this, but can I also ask the code base if it can explain these things to me together with like what an AI model, a large English model would know about these things, right?

6:18So I used a bunch of like different tools to like actually ask the code base, like or based on the code base. And actually, you know, it actually helped me explain some of the domain terms. And also what it's quite useful for is it's like a new form of more powerful code search, right? Where I can go beyond just searching for a string in the code, but I can actually ask for concepts or feature descriptions. And in some cases, I think maybe I would get the same result as a string search if it's like a very specific term. But I can now ask like, where is the, you know, where is, where do we have a list of HI types?

6:54These were like health information types or something, right? And it was actually pointing me to like an enumeration or something like that. So I could find those types of things. That has a lot of promise, I think. Like it wasn't perfect, but I also know that technically there's still a lot of potential how to make these searches better, right? Like some of my colleagues, for example, are working with clients on loading code bases, abstract syntax trees into knowledge graphs and then enriching them. And in this way, making the retrieval augmented generation that is doing the code search, even better, right?

7:27So I know there's a path. I can see a technical path how these things can get better. But then there are other parts of this like journey when I have a code base that I don't know and I need to implement a ticket, right? So now I found the place where it's supposed to be implemented, but now I have to reproduce the current behavior so I can implement the new behavior, right? And for that, usually I would have to do things like either run the application and kind of like find the click path to where I need to do this, or maybe it's about an HTTP endpoint, so I want to run that and try it out, or I want to write a test, or I want to find the right test to change, right?

8:06And this is where it got a lot trickier. So in this case, this is a very old code base. So this was Java 8. This was still using Vagrant. I don't know if anybody remembers that. I think it's not that long ago that we used Vagrant, but a lot of maybe younger developers today wouldn't even know that. So it's still like AI can only help me there only so far, right? And like a complex combination of different old technologies. So I would still have to rely on the documentation. You know, maybe AI can help me debug some things, but I didn't quite go down that rabbit hole. But it was like a realization I had that was like, oh, you know, like even running like older code bases isn't that straightforward, right?

8:49And there are other problems there and setting up my machine and so on that maybe AI is more in the way than helping me. And then the next problem I ran into was this test problem. So for this particular part of the code base, there actually was no test yet. So this is also quite common in these older code bases. Or I hear from clients a lot, oh, we have this old code base that doesn't have any unit tests. Can we just add tests to it now with AI so it's safer for us to change this when we have to, right? And so in this case, I did not have an existing unit test. So the problem I ran into was that the code wasn't very testable, very unit testable, right?

9:31So the AI tools were suggesting test code to me that wasn't really useful. Like it was stuff like the AI suggestions were mocking a lot. So even mocking to the point that it was mocking the things that I actually wanted to test. And then when I was trying to actually use real data inputs, real data structures, I noticed that the data structures used in the code were very convoluted and like a deep hierarchy of objects. So I couldn't just mock that away. hey, I wanted the real data, but AI couldn't help me set up my test data properly because I think it was like the network of these objects was too deep.

10:10It just wasn't quite figuring out the full path. So I kept running into null pointer exceptions and I never succeeded like to write a test with the help of AI that would actually help me. So it was like a great example of, you know, if you don't have unit tests in a code base, probably AI cannot help you with that unless you first refactor the code base so it's easier to test. I've found that a lot of these models just don't work very well if you don't have a concrete example for them to reference as like the baseline. You know, so if you have zero tests, it's like you really kind of have to go from zero to one before you could even expect like most of these models.

10:47Exactly. Test generation works a lot better when you have some examples already and you have like maybe a setup, you already know how you want to do your mocking and all of those things, right? Yeah, but I think one theme that I think I've picked up here is that, you know, particularly when you were talking about asking the code base, it's like there seems to be a lot of value in going from like knowledge acquisition to some sort of practical application of that knowledge, right? You know, because you think about conventional search, like text search, that's how you just find knowledge. But then the human element still had to like bring together multiple knowledge sources to craft like a solution for whatever situation they're in.

11:25So it sounds like that's been successful at the very least. And I also think that this, exactly this, like bringing knowledge together, the right knowledge for the context and kind of like putting it together for the prompt to the large language model, right? The retrieval augmented generation. I think that's where a lot of the improvements in the direct future will come from in these tools and not from like bigger models or like if only we trained the model to even better at Java or something like that, right? but these tools and how good they can be at that. So there are some products also coming out now where they really try some much more sophisticated approaches to pull in your documentation and your JIRA tickets and your code and put it into these knowledge graphs and really categorize that knowledge and then try to figure out the intent of the person working in the tool, right?

12:15What do they want to do and what would be useful information? And the degree to which these tools can do that for us effectively I think that's the big potential, right? Like when we don't have to do all of that anymore. Yeah. And, you know, I've found myself thinking about, so let's maybe use this to segue into like the core of what we want to talk about today. And that's agentic AI. And I think part of the challenge of this is that everyone sort of comes into this discussion with like a different interpretation of what this means. So, you know, a lot of days when I think about agents or AI agents, you know, I'm thinking about having specific GPTs that I've prompted to behave a specific way so that, you know, I do a lot of content, for example.

13:01So I've got the content strategist that gives me ideas. I've got the writer that produces the copy. I've got an editor prompt that will help me edit it. So, you know, as a, like from the non-developer side of me, because I kind of viewed it that way. But then there's also this like agentic AI tools, like developer tools category that seems to be emerging. So maybe we can just start with like, how do you define like agentic AI? Yeah, as a non-data person, right? Like you said in the beginning, yeah. So for me, like in my head, the simplest definition of like what I think of when I hear agent is that you have a model and you tell the model that it has some tools available, some actions that it can take.

13:46And then the model will tell my application, oh, I think we should now run some tests. Please run tests for me. You told me that you can run tests. That makes sense for it. If you think of it as a conversation between my application that is the agent and the large language model, right? So that's like the minimum that you'd give it some kind of tool, right? And then, of course, people talk about multi-agent systems or where you actually have multiple of these applications that then talk to each other. Almost like one agent is a tool that you provide to the other one, right? And that's when it gets a lot more complicated, right?

14:21And I think, yeah. So I think when people say agent, we always have to investigate what do they actually mean. It's one of those terms like service that is very quickly becoming very overloaded. So we always have to ask each other. There's no wrong or right definition, I think, at this transitionary time also when we're still figuring out the terminology. We always just have to ask each other, what do you mean in this context when you say agent, right? The same we have to do that with service or unit test for that matter. Yeah. So given that we're now sort of on the same page about how we view this technology, what do you think are the most promising agentic services or tools that are hitting the market right now?

15:00I mean, there are all of these tools that call themselves like software engineering agents or developer agents. And a lot of the coding assistants are starting to build these in. So like GitHub Copilot has GitHub Copilot Workspace, which I think is still in private beta. I'm not, I keep losing track of what is generally available and not. There's, I don't know that Amazon's product has something built in. There's a tool that is very focused on test generation called Codo. They are also experimenting with agents like that. There's open source things that do that. This is Benchmark, SWE Bench. I think it's from Princeton University, if I'm not wrong, where they're trying to benchmark these tools.

15:40But in the Gen AI space, benchmarks should always be taken with a grain of salt as well, because you always have to investigate, like, what are they actually testing? What are they comparing? So we also can really rely on that, right? So I've watched a few demo videos. I've tried a few of these myself, right? And a code base that I have, and I know, oh, I have to do this little extension, and it actually uses a lot of the existing components I already have. So it needs to create a new React component, it needs to create a backend endpoint, and then it needs to stick those two together, let's say.

16:12So I always think about, okay, what is something where I already have code, but it's kind of relatively straightforward? Because that feels like a good sweet spot, maybe, I think. So it's not quite totally simple. So something that I could do myself, but it's also not like super complex and new and there's no example. Right. So these types of sweet spot things, I was looking for my code basis and then tried these tools. And at the moment, I am not very like excited about them, to be honest. Like one problem that I have is that they seem to often be advertised as like, oh, these things are going to be our agents that can solve like any type of problem for software.

16:50Right. Which is maybe a marketing problem, you could say. Right. Because I don't think that that promise is fulfillable. It's just too much, the scope, right? But what I do find interesting is, are there any specific types of problems that maybe, yes, in the next few years, we could actually get to a point where these can help us, right? Like what I was just saying, like maybe things where we already have code, it's relatively straightforward. Or, you know, what I was writing in the memo is about like a tech stack migration that is not quite as straightforward as a Java upgrade, but has kind of like this fuzziness.

17:24So it's still like quite a lot of work manually. Like a good example today is like migrating enzyme tests to React testing library because it's a very common use case of something that's deprecated right now. So maybe there are these areas where this will work better sooner, right? But then at the same time, when I look at it, I've had these experiences of like, I was talking about this standard case of I want a React component that calls a backend endpoint. I have lots of examples, right? So, and then like an agent is like, I see an agent kind of like go through, you see the text come through the text generation, right?

18:00Here's a plan. This is what you should do. And then I see these one or two things where I'm like, oh, it shouldn't name it that way. That's not a good term how to call this variable or this concept or something like that, right? And we still have to edit this code in the future, right? So we still don't have AI that also takes care of the maintenance for us later, right? So humans still need to be able to maintain this and change this. So it's important to me that certain namings are right. So now I have this big plan that I waited for two minutes for the AIs to spit out, and I have these two little comments.

18:31I want to change the name of that and the name of that. So then when you say, please change those names, it goes and does the same thing again for two minutes, right? It's just like an example of how it seems very tedious at the moment in the user experience for me as a developer to work with the AI to change the plan. So when I look at what Copilot Workspace is doing, it looks like a relatively okay user experience of how I can edit things, right? But this is like one of the things that I'm wondering about. When we have the AI solve these larger problems, how can we still like tweak the little screws and how is that still a good developer experience for us?

19:06So that's one of the things I'm wondering about. And I haven't really ever in a demo video or for myself seen a tool really solve the original problem. There's something weird going on with these marketing videos where even the use cases that they choose for the video often don't make sense. There's a bunch of videos on YouTube of people ranting about that. So I don't know. It still seems very immature. So I'm holding my breath. I think the pie-in-the-sky vision of this is being able to go from a Jira ticket to code that's in production using completely an AI workflow. But just from the experience I've had working with these tools, It's like really the more focused you have each model on a specific task, the better.

19:52And you really kind of have to have at this point in time, you still have to have a human in between each one of those points to validate and make sure that, you know, the input for the next part of the workflow is actually going to be successful. And, you know, I think, yeah, I agree. The marketing around this stuff has been wild, probably because it's so hyped. But I remember one statement I saw a company have on their website. I think it just said, unleash an army of junior developers under a code. And I'm like, on the surface, that sounds amazing. Because you could do so much if you had an army of developers.

20:25But then you think about how much supervision those junior developers are going to need. It's like I said, a promise of a threat. Exactly. No shade to the junior developers out there. Like you're not in an enviable situation right now. So is there anything like today that you think people should be paying attention to as like, like it's here now or if it's not, it's close enough that you should start figuring out how to adopt it into your offer delivery? Yeah, I mean, the coding assistant products that are out there today are actually quite useful already, right? Like for things that we can do today, right?

21:04So GitHub Copilot is a good product. There's this IDE called cursor, right, that I always call it like on my slides, like the developer favorite, you know, because a lot of developers really like it. And it's also like a great tool to look at what are the next upcoming ideas, because they put a lot of interesting things in that, you know, maybe sometimes don't quite work yet, but that gives you an idea of, oh, what if we could make that work? That's really interesting. And they have interesting ideas about user experience. All of these coding assistants now at least have awareness of your local code base that you have open.

21:34So they index the whole code base and you can ask questions about that code base. So that's super helpful. There's more and more things emerging in terms of like what is often called context providers. So that you can say, for example, like a bot, right? So that you can say at Atlassian in your coding assistant chat and it will, and then ask questions about what's in Atlassian, right? So there's like all of these, it's almost like an ecosystem emerging now with like all of these extensions and these tools integrating into each other. So again, that you can have more context and like better context orchestration and better curation of what's relevant in the moment.

22:10So that's really interesting. And what I also would like to see in more products is this idea of like sharing prompts with the team. So there's this open source coding extension called Continue, which is, by the way, also really nice if you want to try coding systems with different model services because you can plug in local models, you can plug in, I don't know, your Google Gemini subscription, your Anthropic subscription, all of these different models. So it's kind of nice for that as well. But one, like maybe not so like flashy looking feature that I really like that I haven't seen in that many of the other products is that you can create a folder and put in prompts that maybe on your team you're using a lot, right?

22:51So just the other day I was trying this thing where, let's say you say, summarize all of my local changes as like a change summary, right? Like let's say it's a basis for release notes or something like that, right? Because you can point it at your local git div. or you can ask it for code review, of course, like these models are not that great yet at code review, right? But you can do that, right? And then you can have these shared like prompts where maybe you have some of your coding conventions that you always want to check, or this is the format of the change log that we want to have, and you can have that as custom commands in the chat, right?

23:24And this way you can kind of codify some of the practices or conventions that you have on your team in these prompts and like share them with the model when it's helping you, right? So I think that's quite promising. Also, such a simple thing, right? And it gives you, as a team and as developers, it gives you a lot more transparency about what's actually going on. Because these tools often do all of this orchestration under the hood. And sometimes it feels like magic, but you don't really know what's going on, right? And sometimes it's very easy. You just say, okay, I want to point at my local Git div.

23:54Here are my instructions. Go. And you know exactly what's happening, right? Yeah, I mean, you know, a year and a half ago, none of us knew anything about writing prompts for these things, right? And today, it's like, we're all sort of collectively trying to figure out what it takes to do it the right way, you know? And it's actually a practice that my team has adopted as well. Like, when we come up with interesting prompts that help us be more productive, that's something we've started sharing them with the team and commenting on how we can get it better, you know? And if something only gets us 80 % of the way, how can we get it 85 % the next time?

24:28I think that's a really great practice for teams to take away. And this term prompt engineering is also kind of annoying, right? Because it's another one of those where we're not quite sure yet what do we mean by that. And the way it works changes every day, it seems like. You know, the models update. Yeah, and the word engineering, again. And, you know, so some people, like sometimes when you say prompt engineering, you just mean the way you write your prompt when you type into chat to PT or something. Right. But then there's like a much broader definition where it's like all of the retrieval augmented generation engineering that you do around it.

Read the full transcript

25:01Like when a coding assistant, for example, puts together the prompt. Right. That is maybe like actual engineering more than that. But because it intimidates a lot of people away from just writing a prompt. Right. And we all now have to because I don't think this is going away. So we all have to like, we're not going to be able to ignore this. It's just too tempting. So we all better figure out how to use this in a responsible way so that our quality stays good and that our junior developers use it in a responsible way. Right. And to figure that out, we have to use it. So we have to get over this like, oh, but first I have to like do four hours of prompt engineering training.

25:35Right. It's very accessible. So we just have to start using it. So you brought up a really great point, you know, about prompt engineering, about how it can feel like intimidating to some people. And, you know, my practical advice to people who might be in that situation, honestly, is just to ask whatever model you're using to help you write prompts. And suddenly you'll get so much better at it. Yeah, yeah, that's a good hack, actually. Yeah, yeah, yeah. So, you know, I think with all of these new technologies that are coming out, you know, I think this happens with just about every wave, whether it's like containerization, microservices, mobile, like every wave of new technology.

26:16I think there's a moment where we hit like tool fatigue, where there's just so much developers have already adopted multiple tools in this space. They might be wondering, like, do we need to actually cut out some of the tools we have? Do you think tool fatigue is something that we're approaching? Like, do you think it's going to get worse before it gets better in the near future? Or how do you think that's playing out? I also think there's like such an explosion of tools in the space right now. And also of just like open source experiments and stuff like that. Because the technology is very accessible in the first steps, right?

26:53Like, you know, you build a little application that sends prompts to a model. Maybe you even provide it with like your first little tool. So you have your bigger little mini agent, right? And it's actually very quick to set this up, right? But then, you know, and we're seeing that outside of the AI for software space as well, where it's for like building Gen AI into your products, right? That there are so many POCs out there in organizations, like with our clients as well, and only so little actually goes to production. And then the things that are in production also still sometimes have to prove like the value that they're actually bringing, right?

27:24I mean, there are some domains where it's more obvious than in others, but there's so many POCs out there. And the same, like so many little tools on GitHub or on people announcing next model that's supposedly better than GPT-4 and once more it is not, right? So I think maybe one way to navigate that is like, as I said, like these coding assistants, some of those products are actually already quite good. I mentioned a few, right? Maybe I should mention a few more, like there's Tab9, which has been around for even longer than GitHub Copilot. So it's also quite a mature product. There's Codium. I mentioned Codo.

27:57So there's like a bunch of them out there, right? I'm trying to not be biased to any of them. But so if you pick one of those in your organization and actually start using it, and then you have a chat available in your IDE. And this chat is like this like Swiss army knife that you can use to get used to using this in your IDE with your code and so on, right? So, and I think it's like a good way to make sure that you learn how to use this today and don't have to like delay that and also do it in like a environment where, you know, the, you know, the, the organization, the company has actually said, yeah, we're okay with using this for our code.

28:34And so it's like a good, how do you say, like springboard? I don't know if that's the right word, but, you know, to, to use it today. And then maybe, yeah, monitor a little bit like the, the madness that is going on around it with, with all of that popping up. and every now and then maybe try one of those tools, right? Something that I always try to look for is, you know, all of these tools are making a lot of claims and I'm always looking, is there some write-up or documentation about how they're saying that they're fulfilling this claim, right? Like, what is the thinking behind this? And I always look for, technically, do I understand what's going on here?

29:07Like I was talking about the knowledge graphs before, for example, right? So I'm like, oh, okay, they're explaining to me here, they're using a knowledge graph and the way they're doing that is actually making the following thing better and that's why it supposedly works, right? And I'm like, oh, okay, I have an explanation. Maybe I give this a try. And the other thing I look for is in like how a company or a framework is talking about something. Do I feel like they understand the real life situation for a developer? Or are they just looking into like how can we build a glorified artifact generator, right?

29:37So those are like some things I look for before trying like every single tool that comes my way. Wonderful, wonderful. So we've covered a lot of the good and bad about these tools. So I don't know necessarily we need to dig a whole lot more into that, but I do want to focus in on one thing that you have said in the past. And that is you said that AI is great at adding code. So what do you mean by that? So this is actually like the best way to demonstrate this is actually to reference a study that I've referenced over the past few months a lot and that got a lot of attention when it came out. But they actually looked at like a whole bunch of like code repositories, both public and from their customers, I think, and found that it seems like the size of those code bases is growing faster than it was growing before.

30:23So they looked at things like, you know, the number of lines being added versus number of code lines being moved from one place to another versus number of lines being changed, right? And they saw an uptick in lines of code being added and kind of like downtick in a number of lines of code being changed, right? And so I think this is like an indicator for us, right? And it's like a very strong hypothesis. I think that if it's so easy for us with the coding assistant to add new lines of code instead of refactoring something where maybe we shouldn't just add that same function almost the same because the AI picks up on this other function, copies it, changes it a little bit, and I'm done, right?

31:04It's very tempting. but then we have all of this code duplication that you know in a lot of cases we don't want right so these tools are often it's it's more tedious to use them to change our existing code than it is to add these new lines because then we always have to i don't know how to explain it it's just more more tedious and it's also like when i ask it for code review or like what refactoring should i do i often get very like basic advice it's also not always the best advice it's like always a good source of inspiration for me and ideas. But when it comes to like more complicated refactorings, usually I already have to know what I want to refactor, but because it doesn't like explain, it doesn't suggest a more sophisticated refactoring to me, right?

31:46I can only tell it, I want to do these things. And then I already know what I want to do, right? So it wouldn't help a junior developer who doesn't have that much experience with refactoring, right? So I don't know if those are like some good examples. There's also a really good paper by this company called CodeScene, who have a product in the code health, code health improvement space. And they sent like a bunch of code smells to LLMs and asked them to refactor them and found that often the large language model would change the behavior or would not actually remove the code smell. And I think they had a success rate of like 36%, which is not great.

32:22But they also go on to talk about like ideas, how we can actually filter out the bad ones, the bad suggestions, and then actually get to a really high success rate. So in a lot of these spaces, there's like ideas how to make this better, but it feels like it will still take like some more time for companies and for products to figure this out. So it's not going quite as fast as things were going last year, but you can see kind of avenues how it's going to get. So I mean, if you have more code coming into your software base and that code is getting churned out faster because the models are just constantly iterating it, Like that's, I mean, that's a lot of risk that you're potentially introducing.

32:59Yeah, yeah, yeah. I think it's also worth like always remembering that these tools are not like other software that we use. So in their non-determinism, right, which some people say it's a bad thing, right? But because there's usefulness here in certain situations, you know, we have to kind of like we think when we've used it two or three times and it didn't help us, maybe it wasn't the right situation. so we can't immediately dismiss them because that's what we would do with another tool that's deterministic, right? We would say, oh, it's not working. Yeah. You know, let's move on. But with this also fuzzy and sometimes it works and sometimes it doesn't, so we have to give it like a bit more time to adjust our expectations of it and to really find out when we can use it and when we cannot use it.

33:42And there are a lot of situations where it does not help, you know? Yeah, yeah. And that may not even be a problem with the tool. It might just be a problem with the code base that it's being run on. right? Like in order for it to be effective there, you may have to make, you know, as we were talking about testing earlier, changes like that. So, well, that's all we've got time for today. So I want to thank you again for joining me today, Brigitte. It's been wonderful to have you here today. If someone wants to learn more about you or follow your work, where's the best place for them to head to?

34:11Yeah, I guess like that memo series on martinfodder.com. So if you go to his website and then there's a generative AI tag, then you should find it. And I also try to keep my website up to date with the articles and podcast episodes and stuff like that. And that is beagita is my first name, dot info. Stick around. After the break, Linear B co-founder Dan Lyons will be joining me to discuss today's episode. Want to know the secrets of high-performing engineering teams? Mark your calendars because Linear B is back with the 2025 Engineering Benchmarks Report, and they're hosting an exclusive webinar to unveil the latest insights.

34:50This year's report is bigger and better than ever, analyzing over 6 million poll requests from 3 ,000 organizations worldwide. You'll get the latest benchmarks on DORA metrics, poll request workflows, and predictability, critical data to shape your 2025 engineering strategies. Join the roundtable discussion with industry leaders and be one of the first to dive deep into this landmark report. Plus, just for registering, you'll receive a free pre-release copy of the 2025 Software Engineering Benchmarks Report. Don't miss out. The webinar is later this month on November 20th. You can register today at linearb.io or use the link in our show notes.

35:30Welcome back. Dan, thanks for being here with me today. I wanted to try out a new segment on the show where we bring you in to dive deep into a few key points from the interview. And we'll use this as sort of a way to discuss some of the bigger themes from the episode and get your take on how engineering leaders can apply the lessons that we heard about in this episode to their own team or organization. So I think my conversation with Brigida is a really great place to start this new segment because, you know, AI tooling has exploded, but it's really hard right now to know, like, what's marketing versus, like, what's actually relevant products that you can use today.

36:12So, you know, I first want to just start by asking you a little bit about some of her research. Specifically, she mentioned how, like, generative AI has been pretty useful for some use cases, but then it still struggled with others. Particularly, you know, she mentioned, like, tasks that required knowledge of, like, how multiple older technologies sort of work together. So, Dan, how do Brigida's findings align with what you hear in conversations with engineering leaders about generative AI? Yeah, awesome, Ben. Thanks for having me. And the interview with Brigida was great. She's really smart, really insightful.

36:48But I think what I can provide is I'm working with a lot of engineering leaders and they're actually running into the same stuff. So the first thing that I would say is the question that I get asked the most, especially when it comes to like the co-pilot, generative AI for code specifically, is is it making an impact? That's what our community is asking. That's what they're asking me. That's what they're asking Linear B. That's the product that we provide. So what I would say is first, make sure that you have a measurement in place, especially for not the usage, not just usage, but what is the impact of Copilot.

37:30That's something if you're going to go experiment there, your business is going to ask you. So that's kind of my first tip. And that's what I'm hearing from the community. Now, when we think about the different types of, let's call them use cases, because I think it's really important to distinguish, like, what is this AI stuff doing? And I can just tell you what I'm seeing. When it comes to things like generative text, I'll call it. So things like I need to generate a PR description. I need to generate documentation based on the code. I need to do, I think Brigida was talking about code search.

38:07Like that's more like tech space. To be honest with you, I think this is where AI is shining. I would, you know, I'm not like an inventor of it, but I would say from the community, a lot of positive results. Even like the stuff we're doing at Linear B right now, like we're providing an iteration retro summary for all team leaders. That's generative text. Like that's doing really, really well. All of that is great. then I think we need to open up the conversation a little bit more when it comes to let's say either test creation that's what she was diving into or co-generation itself what is the impact of that and what is like the accuracy of it I think everyone's in experimentation measurement and trying to understand the impact so maybe two different ways to think about it Yeah.

39:01So, and Brigida had, you know, she had a lot to say about agentic AI in particular, which we're hearing more and more from a variety of sources. And, you know, of course, she outlines a pretty wide range of tools in this space, you know, and all the big players are out there like GitHub, Amazon, there's this newer company that's merged called Kodo, but it almost feels like there's probably an over-proliferation of tools that are emerging in this space. So what recommendations do you have for engineering leaders who are in a situation where they're trying to navigate this new experimental technology space that is potentially over-proliferated?

39:36Yeah. I mean, the other thing about it is there's a ton of money from investors going into this agentic AI, like the amount of startup companies, the billions of dollars that's being invested. So you're going to see like a long tail of companies that are doing something, let's say, in the space. But I thought Brigida made a really good point that it's more about the use case. What does Ingentic AI really mean? It's more about the use case. But here's what I can tell you. Especially if you're a midsize or maybe an enterprise-type engineering team, it's totally worth deploying people, so resources, I believe in being up-to-date on the latest technology.

40:20This technology is changing all the time, but I do think it's worth saying, okay, let's pick a few. And I have a few examples of these, Ben, but let's pick a few maybe down to earth use cases that we want to be really, really good at. It could be, again, like I'll say something as easy as I want to make sure every pull request has the right and accurate description standardized every single time. Okay, we want to go make that happen. I think we can go make that happen. Then I think there's a long tail of maybe other experimentation, like test creation or something with code creation. I think Brigida was saying even in her situation when she was doing the test creation, it kind of depends on the code being used in the architecture.

41:09Okay, maybe you do some experimentation there. But the first thing that I would suggest is I do believe it's worth it to have a few people on your team looking into this stuff. And I would pick a very down to earth use case, something around standardization. That's what I would do. And then maybe start dabbling with a long tail of companies that are coming out with the more innovative use cases. But specifically for you, like what are you most excited about for the, these generative AI capabilities and integrations that are coming out? Like, what do you think is the coolest new stuff that's hitting the market?

41:42I'm going to give you like two versions of this. One is going to be more down to earth. And then one is just something that I've been thinking about specifically at Linear B and like what we're doing with our customers. Like I said, throughout like all of this, what your business is going to ask you for is, okay, you're dabbling, you're spending money on all of this AI stuff. What is the business impact for us? Otherwise, I'm not really sure why we're doing all of this. I believe that standardization is a really good place to start. Again, if you're an enterprise or more like a mid-market customer, meaning you have a good amount of developers, standardization end-to-end is something that's been, I think, pretty difficult.

42:26And I think AI is very good at helping with standardization. So when I say that, start with the pull request review, AI-based review. Make sure every PR is going through at least an AI review. so you have a baseline of review. Make sure every pull request has a good description. You might have developers working in all different languages from different. Yeah, okay, make sure every single PR has a good description that goes through a review. Pick a few things that are standard that you know you can deploy to all of your developers and you guarantee business impact. That's the first thing that I would do.

43:03Now, the second thing is a little more out there and some of the things that we're starting to experiment within Linear B. And I think the next step, actually, Birgitta might have actually mentioned it. It's having like the right context. And I'll try to explain what that means. Think about having AI be able to look at not just code, like your local code, but also have it being able to look at like Jira, like your project management situation, and also have it being able to look at your deployment information, also have it being able to look at your customer, like bug requests, your support queue.

43:44What we're experimenting with, because right now, a lot of the things that I see with AI, they might be looking like at just the pull requests or just your local code base or something like that. But I think the next step is, let me give you full business context end to end. I can see everything that's going on in code, everything that's going on in projects, everything that's going on in your releases, everything that's going on in support. Now I can provide you with something that I think is a little more intelligent. So for example, if I have an AI code review, let me confirm that it's actually meeting the story requirements.

44:20That's pretty cool. I'm moving it from just code to business value. Imagine a review saying, okay, yeah, we found some code base inaccuracies, but also, hey, I'm not sure that you actually address, you know, bullet point number five within this story. Take a look at that. Okay, now we're talking about business value. Now think of it if the AI review came back and said, and you know what? You also touched an area that we just had five support tickets within the last week. So when I assign a human reviewer, I'd like them to focus on this area because it's very sensitive based on your support queue.

45:00So what I'm excited about, I think the next step is actually looking at the whole data set of the business or the end, like product delivery organization. I think that's where it's going next so that we can actually make the outputs of AI more intelligent. That full business context, you know, I think that's really critical. You know, I'm imagining a world where like a developer can come in And, you know, not only do they have this AI assistant for their code base, but they even have like these AI coaches that help them with determining like what's the best, what's the best thing for them to focus on right now.

45:36And not just like how to write good code, but like, you know, what's the task that you need to be focused on right now? Or what's the biggest challenge that your team has today that you can help solve with? And when you get that full context, like that's when, yeah, a lot really starts to happen. That's right. Yeah. And I think, you know, if we can summarize this, it really feels like, you know, what engineering leaders need to be thinking about right now is, you know, there are some use cases out there that are clearly impacting productivity in a positive way and giving your team space to both adopt those use cases, but then also start to experiment with some of the other stuff that's emerging.

46:12You know, that's how you can both get some of these productivity gains today while also like setting yourself up for the next wave of tools. Yeah, that's right. I mean, I really think like pick what I call a down to earth use case. Make sure you're delivering business value, whatever it is. I suggested standardization. It might be something else for your business. That's going to give you a little more leverage to start experimenting with more on the edge stuff, multiple agents working together to solve a story. There's a lot of cool things out there, but I think it starts with measuring the impact and the business value.

46:45And that's going to give you the leeway for more of that experimentation. Well, that's a wrap on this week's episode. Thank you, Dan, for joining me for this session here today. If you're hungry for more content like this, join the over 18 ,000 listeners who have already subscribed to the Dev Interrupted sub stack. Each week, we curate our favorite newsworthy tech articles and do deep dives of the podcast. You can also find us on YouTube where we post all of our favorite moments from each episode. And hit us up on Twitter or LinkedIn at DevInterrupted to let us know how your teams are using AI tooling.

47:21You know, what's working? What isn't? And let us know what you think of this new show format. We'd love your feedback. We think this is a fun experiment, so we hope you agree. Thanks, everyone. We'll see you next week.

47:37Get prepared for it!

From the publisher

There’s AI agents. There’s AI tooling. Do either drive business impact or are they just more things your dev team is supposed to stay on top of?

Birgitta Boeckeler, Global Lead for AI Assisted Software Delivery at ThoughtWorks, joins the show to discuss the practical applications of AI in software delivery. She shares her research on AI agents, highlights areas where AI hasn't lived up to the hype, and offers concrete examples of useful AI tools for development teams.

Dan Lines then joins the conversation to provide his perspective on how engineering leaders can leverage these insights to effectively implement AI within their own teams. He also discusses LinearB's efforts in helping software teams measure the business impact of AI.

Show Notes:

Support the show:

Offers:

More from Dev Interrupted

All 208 episodes
Making Sense of Agentic AIDev Interrupted · 48 min
Listen in VO