Will AI Finally Make TDD Practical? | Diffblue’s Animesh Mishra

18 Mar 2025 · 46 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Dev Interrupted Podcast Episode Notes

Episode Title

Will AI Finally Make TDD Practical? | Diffblue’s Animesh Mishra

Episode Overview In this episode of *Dev Interrupted*, host Andrew Zigler talks with Animesh Mishra, Senior Solutions Engineer at Diffblue. They discuss the gap between the theoretical appeal of Test Driven Development (TDD) and its practical challenges. Animesh shares insights on how deterministic AI could enhance the application of TDD, making it more feasible for software engineering teams.

Key Concepts

  1. Test Driven Development (TDD)
  2. Definition: TDD is a software development practice where tests are written prior to writing the actual code. The goal is to ensure that no piece of software is deployed without being tested.
  3. Challenges:
  4. Developers often find writing tests tedious and unmotivating.
  5. The practice of TDD is not widely adopted despite its potential benefits.
  6. The "fun" aspect of software development often lies in problem-solving rather than in writing tests.
  1. Role of AI in TDD
  2. Deterministic AI: Unlike LLMs (Large Language Models), deterministic AI can systematically generate tests without the stochastic variability.
  3. Automation: AI can automate the mundane aspects of testing, allowing developers to focus on more complex problems.
  4. Productivity Gains: With AI, developers could recover a significant portion of their time usually spent on TDD, potentially allowing a focus on high-value tasks.

Discussion Highlights

A. The Sandwich Generation in Tech

  • Concept: Refers to professionals who take care of both younger and older dependents, often leading to lifestyle and financial compromises.
  • Impact on Work: This demographic often experiences pressure due to balancing work and caretaking responsibilities, especially in remote work environments.

B. Google's Privacy Concerns

  • Discussion on Google's increasing control over privacy settings in Chrome, potentially leading to a more AOL-like browsing experience.
  • Implications of ad blockers and user privacy in the context of Chrome's evolving features.

C. AI in Hiring

  • A recent incident involving an AI impersonator during interviews highlights the potential risks of AI in the hiring process.
  • The need for organizations to develop methods to authenticate candidates effectively.

Insights from Animesh Mishra

  1. Building Trust in AI for Testing
  2. Trust can be established by being transparent about AI capabilities and limitations.
  3. Developers need to experiment with different AI tools to find which ones suit their needs.
  1. Evolution of Testing Tools
  2. Mutation testing is emerging as a realistic measure of test effectiveness, surpassing traditional line coverage metrics.
  3. AI can help modernize legacy systems by automatically generating comprehensive test suites.
  1. Future of AI in Development
  2. Companies should consider whether they truly need AI and what specific problems they are trying to solve with it.
  3. Not all software engineering tasks require AI, and some may be better served by traditional methods or simpler automation solutions.

Key Takeaways

  • Flexibility and Support: Organizations need to accommodate the realities faced by modern software developers, especially in balancing work and personal responsibilities.
  • AI's Role in TDD: The potential of AI to make TDD practical could revolutionize the way testing is approached in software development.
  • Strategic Implementation: Before adopting AI solutions, teams should assess their specific needs and the problems they aim to address, ensuring the right fit.

Resources and Links

  • [Diffblue](http://www.diffblue.com)
  • [Animesh Mishra LinkedIn](https://www.linkedin.com/in/siranimesh/)
  • [Dev Interrupted Substack](https://devinterrupted.substack.com)

Closing Thoughts The conversation emphasized the importance of innovative thinking in software engineering practices, particularly regarding testing and the integration of AI technologies. This episode serves as a call to action for leaders to rethink how they approach development challenges in the modern tech landscape.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:07So, welcome to Dev Interrupted. I'm your host, Andrew Ziegler. And I'm your host, Ben Lloyd Pearson. In today's news, we're talking about a few things. The sandwich generation in tech, Google striking back against ad blockers, and an AI-powered deepfake coding interview applicant that almost got hired and the entire internet is talking about it. Ben, what do you want to talk about first? Yeah, I've seen that last one a ton, but maybe we'll save that for the end. And since I'm feeling like half a sandwich, maybe let's start with the sandwich generation. The sandwich generation in tech is something that maybe you're very already familiar with, but you didn't know the term for it.

0:47It's the folks that maybe you are one of them, who every day when you go to your job, you have people younger than you that you take care of, and people older than you than you take care of. Oftentimes, you're in a remote setting. Your time is split between being a caretaker for your kids, but increasingly, people in the workforce are also taking care of their parents. So you end up in a situation where there's a large identity of folks who are dual caretakers. In fact, there's 11 million of them in the U.S. as of last year, makes up millennials and Gen Z are about a third of them. Most of them work full-time or part-time predominantly.

1:25But overwhelmingly, 90 % of them or more in this article talked about how they have to make lifestyle or financial sacrifices. They have compromise on their own financial security and their future, it's not a temporary issue that goes away. And anyone who's a caretaker is very familiar with this. When someone depends on you, it's there and you have to always keep it front of mind for yourself. And these people don't just work in tech, of course. They work all across different fields, but they also have to wear a bunch of different hats at home when they're solving problems for the older and younger generations.

1:59Maybe they're the power of attorney for someone less capable of making decisions for themselves. Maybe they're tech support for their parents and their kids, because it turns out that millennials and Gen Zs are the only generation that really learn how to become really good tech support. Everyone else older than us and younger than us, you know, they just Google it or they don't know how to do it at all. Imagine having to do all of this context switching and wearing all of these different roles while still doing their job. So when I was reading this article about the sandwich generation, you know, I saw a lot of myself in it.

2:33I'm not fully in the sandwich. I'm more like a pita. You know, I only have one piece of bread, I guess, but I definitely feel the pressures of having to be responsible for them. Yeah, I've got kids, so I'm definitely feeling one half of the sandwich very strongly. And, you know, I guess the real question is what kind of sandwich, you know, because if it's like a panini that's been pressed, like, man, that's a that's a pretty rough life, you know? Yeah. Our listeners can probably tell we have not eaten lunch yet when this topic came across, but you know, you too, Ben, you know, you, you have kids and you work remotely and I'm sure that that takes a lot of your time.

3:06I actually think it points out how much remote work can help mitigate the effects of this because my mom worked in tech and for a while she was also getting her master's degree. So I remember her working really late nights and not getting to see her all the time. And I often think about how, you know, as a remote worker, the flexibility that brings me and the fact that I see I'm way more involved in my kids' everyday life just because I'm around a lot more. I don't have a long commute and I don't have to sit in an office all day. It's definitely an interesting trend or an interesting phenomenon.

3:41And I'm kind of asking myself, is this an HR problem? Is it an engineering leadership problem? Are we simply not hiring enough developers for some companies or are we pressuring developers like to do too much? You know, there's always this like endless drive for productivity. But when do we actually start to see a drive for more flexibility? The return to office movement is definitely not a positive trend for this. No, not not for the sandwich generation. It puts them in that hard position like what you described. It makes me think of even myself. You know, growing up, I was a latchkey kid and most of my parents worked and, you know, they weren't able to work at home.

4:19Remote work wasn't a reality when I was growing up and that wasn't an accessible life for them. And so it really changes how, you know, you become your own independent person, but also just having that person available for you to support you is really, really great and helpful. So the fact that a generation now, like my generation, is able to work remotely and help raise their kids and be there for them when they come home from school. Just those small things, I know that can make a lot of difference. But at the same time, it puts a lot of pressure on them. Like what you said, at work, they're strained between leadership decisions, between project deadlines, between things that have to get out the door and everything right now.

4:56And every industry is crunched and accelerated. There doesn't really seem to be much relief. And so that's why it's always important to stop and highlight these human moments and how we're all working together and having to support each other. I think it really calls out that, you know, the flexibility of remote work is really key for this generation to not only survive, but thrive and be able to set up a future generation for that same kind of success. Otherwise, you're going to end up in this scenario where it's like kind of like a brain drain situation. If you can't accommodate these very talented, but obviously responsible and in some scenarios like strapped leaders and professionals that are just in between life's situations.

5:40If we can't account for them, then they can't be in our industries. They won't be represented. and then that means that the products that we make, the businesses that we build, are less understanding of them and are less able to serve their needs and the situation becomes worse. And that's always why it's important to make sure everyone has a seat at the table, that way that we can resolve those problems as a society. So my favorite part of this article is the very end where they brought up people who were responding to this by building technology to help them caretake or take care of their elders.

6:12So I'm asking myself, Like how long until the pressure builds to a point where developers just start leaving their jobs and using AI to build the next generation of elder tech, you know? Well, the next story is about privacy, a little bit of a different turn. But everyone's heard something about Google Chrome and privacy and how different browsers that we use to surf the Internet all handle our personal data very differently. There's been a recent shakeup, and it seems like Google is getting closer and closer to making Chrome something more like AOL of the modern day. We're talking about getting rid of ad blockers, getting in the way of how users want to use and configure their browsers and ultimately protect their privacy online.

6:58AOL 2.0. I'm already hearing the dial-up tones in my head. My goodness. This was sparked by a conversation we saw on Hacker News, a very heated conversation that was about Ublock Origin being removed from the Chrome store. And the real story here isn't anything to do with Ublock because we all kind of know that Ublock, it's a good app. People really like using it. It's not malicious. It's not trying to harm you in any way. But it's being removed because of these new restrictions that Chrome is bringing into their browser. but also sort of like using their position within the market to force standards to go in a direction that they want.

7:37There's a lot of ways to get around this stuff, like use a different browser, for example. I personally am a massive fan of Piehole, which if you've never used it is a DNS service that you can run yourself and will, and instead of serving you up ads, we'll just give you blank, nothing, but convince the websites that you actually did see the ad. And it removes ads actually on your entire network, which is pretty amazing. The real story here, I think, is that Chrome has this long history and Google through Chrome has this long history of pushing standards that maybe other browsers don't really want to adopt those standards or maybe users, certain users don't feel comfortable with them.

8:16And we're just kind of like going further down this path. And of course, always in the background is things like the antitrust rulings that are coming down on Google that forcing them actually to make Chrome its own company. So I really feel like we're in the middle of this story right now. Like we haven't seen the end yet. I feel like we've been in the middle of the story for a long time. You know, Chrome has made lots of changes over the years against and for privacy that, you know, its users, its consumers have not been happy about. I think you make a very salient point about how the antitrust plays into this.

8:52Because ultimately, you have to remember at the end of the day, Google is selling you ads. And that's where Google gets most of its income from. And so they have an incentive to make sure that they serve you ads. And as the creator of Google Chrome, you know, one of the largest and most adopted browsers, it puts them in a right position to not only serve you the ads, but to serve you the place where you're looking at those ads. That amount of control over our browsing situation is a lot for many people to kind of tolerate. And I can certainly see people going their own ways using different browsers like what you said, maybe even using fanciful acts like the spy hole situation you're describing.

9:27DNS engineering aside, I do think that it highlights the importance of understanding how you're browsing the web, the websites that you're going to and what they know about you. But really, now that we've been talking about AOL, I just really want us to have like AOL keyword dev interrupted. How cool would that be to go back to the AOL keyword era? I'm not on board with that. But I do kind of wonder, are we nostalgic for a web that never actually existed? Analytics is something that's been sort of built into the web from the very early days. And at this point, it's almost like a fundamental component of the very fabric of the internet.

10:04Nothing can really be optimized without data and analytics. And that requires some degree of tracking that, you know, I think is really always existed. I certainly do miss the profound lack of advertisements everywhere. So, Andrew, let's talk about these AI fakers that are trying to enter your organization. So if you haven't heard about this article, maybe you've been under a rock. Maybe you haven't been on LinkedIn. But this has been blowing up the tech industry in the last week, week and a half about AI. I've seen at least 20 posts on LinkedIn about it. Me too. I've seen so many, and so we had to talk about it here because if you've been listening, you know, a few weeks ago, we talked about how AI is disrupting the interview process in tech and about how it changes what we need to have in mind and what we need to be considering when we bring in candidates.

10:51This goes the whole other direction about the dangers that AI can pose within your hiring process. It's about this whole scandal, just to back up and give you some context. It happened for a company. It's a Polish company called Vidoc, and they were hiring for engineering roles. And they were doing this remotely. This is a remote work position. And so they were hiring candidates. And as they were evaluating one of them, you know, they passed the initial interviews. Their resume hit the marks for what they needed. They even sailed through the technical interview. But then at the final stages, there were some questions about their background, about some gaps on their resume, and ultimately an offer wasn't extended.

11:34A few weeks later, or during that same hiring process, another candidate came into the queue who managed to pass all of those same checks, talked to those same people, but was posing as someone else completely. So both of these were actors. These AI powered actors were actually the same person applying for this role, or at least we assume they are the same person applying for this role. And they were using AI technology to basically fake their way through all of the hiring process, even getting all the way to the end. And the only reason that they were debunked is from a now viral video, which is probably the reason you know this story of the of the final interviewer asking the candidate to put their hand in front of their face in the interview.

12:17And as we all know, when you have photo filters, you can't do this because your photo filters are going to pop on and off. The candidate refused. And this is because it would break his filter. and imagine that he made it all the way through that process, not just once, but twice, and almost got the offer, and the only reason he didn't is because he used the slightly same voice. He sounded slightly similar to that first candidate, and this hiring team, they were on it, right? Really props to them for being in charge and on top of their hiring process. The whole time I've read about their post-mortem on it, I think they did everything right.

12:52I think this is something that could happen to any company, And it's a really big wake up call about how hiring is changing. Yeah, I mean, we just covered this recently in another episode about how AI is just fundamentally shifting the interview process for software developers. And this is one of the darker ways that that's happening. But, you know, it makes me think, like, what is the way that you validate that somebody is real? I've been talking to some people I know that, like, in my personal life, it's like, there's going to be a day where AI impersonates me. So we need a code word that only we know that will validate that I am who I am when you hear me talking to you.

13:29And we almost need something for the interview process as well. It's like in my next interview, I might have to request that like the candidate turn off their background filters, put their left hand and their right foot in the frame of the camera and then spin around in their chair three times or something just to validate that they're actually a human being. You know, these impersonations are only going to get better. So we've got to continue to evolve our way of detecting and responding to them. We have to sharpen our evaluative skills for candidates, but we can't let the fear of these AI deepfakes harm our hiring process for real humans either.

14:08We can't make it oppressively difficult for people to apply for and to get roles or to prove that they're real. We need to come up with intuitive and simple ways for people to prove that they're real virtually. without having to add a bunch of Byzantine process. You know, I really commend the engineer who interviewed this person who had the idea in the moment to put their hand in front of their face to try to get the filter to fall off. I think that's one step in the right direction. I think there might be easier answers out there too that maybe a brighter mind in HR and tech is already thinking about.

14:43And if that's you, if you're hiring and you're kind of experiencing this and you already have some interesting stories from the AI world for the jobs you're hiring. We'd love to hear about them. Yeah, I would love to hear what our audience thinks about this. If you've encountered this, head over to our Substack, head over to LinkedIn, wherever you prefer, and leave us a comment that describes what you're seeing. You know, are you experiencing this? And have you had to address this? So, Andrew, let's talk about our guest today. Who do we have? I'm very excited for today's guest. After the break, we're bringing Animesh Mishra on the pod.

15:15He's a senior solutions engineer at DiffBlue. And DifBlue is an enterprise-grade solution for automating test suite generation. They make it really easy to make a massive amount of tests at scale. And when you stick around for this discussion, you're going to learn about the realities of AI and how it can be used to do this. And how this is finally enabling test-driven development for engineering organizations at scale. You really don't want to miss this one.

15:45Are you struggling to explain developer experience to non-technical leadership? Join Linear B's upcoming workshop and learn how to translate DevX into language the business cares about. We'll show you how to present data on developer productivity, AI performance, and engineering health in ways that drive alignment and investment. Plus, you'll get an early access to our CTO board deck template, making it easy to connect engineering metrics to outcomes like faster time to market and cost savings. The link to sign up is in the show notes. We hope to see you there. Today, we're talking about one of the biggest paradoxes in software engineering.

16:23Developers love the idea of test-driven development or TDD, but in practice, almost no one actually does it. And here's what we're going to unpack. Can AI finally make TDD practical? Developers that struggle to trust AI in the last mile, such as for testing, how do we build that trust for them? And AI isn't just LLMs. So if you're a software org that's investing in AI right now, how can you expand your approach beyond mainstream AI solutions? Animesh, welcome to the show. Thank you, Andrew. I'm a software sales engineer at DiffBloot, and I'm looking forward to the conversation. Yes, we're really excited to have your expertise here.

17:06You have a lot of exposure to how people use testing out in the wilds, and a lot of knowledge about how teams can be unpacking this kind of stuff today. But starting with the first topic at hand about test-driven development, and it's one of those things that, you know, people say they do, but, you know, let's be real, it doesn't really happen all that often. Maybe we could start by getting the definition of what is TDD? What is the goal of test-driven development? So the goal of test-driven development, which kind of relies on the name, is to make sure that no piece of software goes out the door untested.

17:45Now, that has always been the case with software engineering. But historically, what people found out was that the fun part of software engineering is in coming up with the code that solves the problem. And all software engineers are good problem solvers. But once you've solved the problem, writing all of the testing code and all the harnessing to make sure that it is delivering every single piece of functionality required of it is not that interesting. And so historically, people saw that the model of you write your code, you see how it works, then you write the test that validates all of the assumptions made about the code.

18:24And when all of the functional assumptions pass, then you can say that this code is working perfectly and no bug or fault has sneaked in. That didn't work because, like I said, the human nature is to solve a problem and then move on to the next dopamine hit. And there's no dopamine hits to be had in writing tests. So test-driven development emerged with the emergence of extreme programming. So there's new ways of writing software, which not only improved the productivity of the team's writing software, but also had lots of techniques on how to do existing things better. And one of the focus was on doing testing.

19:05So they figured that a lot of the time software is badly written or has bugs because the functional requirements are poorly understood by the developers. So they assume that this is how it should work. Then they write the code. Then they write some tests. And then off they go. And because of that poor understanding of requirements, you don't get the right kind of code to go with that. So that's the first challenge that was spotted. So to address that, the idea of test-driven development was created. So the idea being, instead of writing the code first, you would write unit test that define the exact spec of the code you would be writing to satisfy that.

19:52And in the process of defining that test script, your assumptions will be challenged about how this piece of code or method should work. You will refine those with your product team. And once you've understood right down through the detail of a unit test how this thing should work, then you move on to writing code. Because by that time, having thought through all of the eventualities and all of the logic that this code needs to do, you would have a good design in your head anyways. So then you go from something on a piece of paper to something running, working, and deployed much faster. It does work.

20:27so it's not that this is all you know it does it does work it does work and it's maybe just that like to your point it takes the unfun part of it the part that's not the dopamine hit and it front loads it onto what is now what you do every day when you're building things you want to go start building you don't want to start by validating what the thing you build will do so it it kind of becomes a showstopper in a way. Not that it's not effective as much as it is just, it's unattractive to work that way, right? That's correct, yeah. When the problem is fun, then it's actually really cool and challenging to figure out, okay, how many ways this needs to, can go wrong or how many different permutations do I need to account for and enable?

21:13And then that's pretty cool. Unfortunately, a lot of software is quite mundane. And then it just becomes dull work Because then you're just writing tests for the sake so that you can move on to writing the code for it. That's where the rubber meets the road, where you get teams which aren't fully committed or motivated to do TDD. They start cutting corners. So they say, OK, we'll do TDD on the business logic bit. Well, we'll leave the others. We'll see as we go. When we talk about that work and it's not sexy, it's not fun to do. how does AI now fit into that conversation? Because we've talked traditionally about AI being, you know, it fills in and does that mundane work.

21:57It takes me back to the first thing I said. The purpose of TDD was to make sure that no piece of code goes untested. Because what you're trying to reduce production incident and what you're trying to improve is your code quality and your delivery rate. So the fewer faults there are in the code that you're pushing out, the faster your entire team works. Now, if that is the goal, the path you take to reach that goal can be many, right? So TDD was invented, I think, 25 years ago, before my time for sure, because I always worked in a TDD environment. And that was the time when the only way to write a unit test was to get a developer to write a unit test.

22:41There was no other options, which means there was only a singular path from, I have something and I want something tested. Now we have options. Now we have AI tools which can do software testing. And I mean, it's all pretty new. The whole field is new. So we're still figuring out what works and what doesn't and where to use which specific tool. But now more parts have opened up. I mean, the reason we're having the discussion today as well, and I'm going to be in California next week talking about this during developer week too, is the fact that now that we have more parts, it will be foolish not to try to traverse them.

23:17So let's at least walk down the part, see what we find, right? So I understand lots of people swear by TDD because it has delivered. And if you can stick to it, it does deliver. The challenge is it takes a lot of time for development time that you could utilize working on better features, better design. Because the thing that's offering you, there's always a lot more to do than the developers have time to do. So with AI, what you can do is take away this mundane work, automate it away. Give the developer back their 20-30 % of the time they spent doing TDD at the moment and let them focus on other more higher value targets.

23:56And I think that's the core takeaway here is that there's an opportunity to free up the time that you might spend doing other things or, you know, writing unit tests and spend them in time doing more productive and impactful things. It frees up developer time to focus on those higher level problems that you can't automate. Definitely. And the other benefit, which is kind of unsaid that people are only now starting to see, and you get that with certain tools but not others, is the human mind is brilliant, but one thing it's not is consistent. This is why you always have these arguments in development teams about standardization.

24:32This is how we write unit tests. You'll get linters involved, so on and so forth. One of the ways you can, so the benefit you might get with an AI tool, which is deterministic, which always produces the same kind of code and same kind of style, is that you can standardize the style of tests you write. So regardless of whether it's my code or your code or someone else's code, the testing, it follows a standard pattern that you have approved and adopted. it. And that takes away a lot of the cognitive load that you're put under as a developer. The first time you see a piece of code you've never seen, that you go to the test script to see, okay, what it is doing.

25:14And every test script is written in a different way. All of those problems go away. And when we talk about the way that developers can use this tool, you know, part of it is that consistency so they know what to expect. And I think consistency is the core part of trust. There's a recent Stack Overflow developer survey that found that only 43 % of developers right now at work feel that AI is accurate. And when we talk about a process like the last mile, which is testing, you know, doing coverage on an application, or in some cases, the first mile, if you're doing that up front before you develop anything, how can we trust or how can we build that trust in developers for using a tool like this?

Read the full transcript

25:56It has to come from being transparent about what the technology can and can't do and how it works. What we're seeing currently, there's a cambrian explosion of AI tools. Everyone's making it, right? Everyone and their mother and grandmother. But a lot of the time, they're just what's sometimes pejoratively called LLM wrappers, right? I think they add a lot more value than just LLM wrappers, but that's what people call them. What's important from a developer's perspective, So I work with a lot of teams which are evaluating different AI tools for different problems. I don't think the software development lifecycle has a lot of tooling currently, and there's a good reason for that.

26:37It is very hard to use one tool that does it all. I don't think we're ever going to get to a point in software engineering where you have one tool that's taking care of every single problem. The reason for that is there are multiple people involved in the chain and they have different problems to solve. If you just focus on the testing part of it, even in testing, often you have two teams involved in many companies. You have the developer writing the code and there's a testing team that's completely separate, often working off a Cucumber script or just, you know, a Jira ticket. The developer will just give them a code and they'll test it.

27:11Right. So because of these reasons, I don't think people should try to go and find one tool that does everything. And luckily, we haven't seen a company that's trying to do that. So then we come to the trust part. So you can only trust something once you know how it works and whether it delivers, right? So you trust your toaster not to burn your bread every morning because you did some trial and error when you first bought it. You set it to three, then it burnt it. So then you set it to three. But yeah, that's about what I want. But there will be, every new technology is going to be some trial and error.

27:45People need to be willing to experiment and see what works for them. because the company I work for, Difblue, for example, we have a tool that works extremely well for certain companies but are completely useless to other teams in solving other types of problems. So there's no, unfortunately, very roundabout way of saying there's no standard one answer to how to build trust. It's going to have to be built slowly with trial and error, people finding out which tools and which technologies solve a given problem better. within testing i think there is a room for llms and other technologies as well i also think in testing what you want is less creativity and more determinism because you don't want your test to work differently on a tuesday than they do on a friday right that is something where testing is very different from when you're actually writing code trying to you know deliver a ticket as part of that too there's a connection you know between how much code coverage i mean you might get and and and with how much like quality or software it can ultimately reach but is that an actual connection like if you have the if you have an ai that's able to provide that full test coverage what do you gain that you maybe don't if you had a human that was making a non-standard approach to doing that same task?

29:14Well, the first thing you get is predictability, right? So once you've seen how a tool works, particularly if it's using the kind of AIs which are more deterministic, so once you've seen it working on one application and the kind of test suite it's writing, its style, how it creates its assertions, the strategies it uses to test a given piece of business logic, If you like that, and that too is more deterministic than probabilistic, then you can trust it to work elsewhere as well. And this is where we see a lot of POCs happening currently across the technology industry. And that's what they're trying to find out.

29:52Which part of a POC result can we deem to be repeatable? And which was a flu? I see. So that's how engineering leaders are really evaluating that. They're looking for that repeatability. And is there anything that they measure in particular when they're evaluating a tool like this to see if it's successful? So I can speak for testing because I'm very close to that currently. So in testing, we've long used unit test line coverage as a measure of how good a test read is and how well your code is tested. I think we need to get a bit beyond that, particularly in an AI world. I'll give you a very good example for that.

30:29So when developers write tests, they tend to write tests about eight to ten tests at a time. They'll create a pull request. You can review it. Somebody else can come and review it. AI has the capability to go into an application and write 50 ,000 tests in one go. Now, nobody's sitting through and validating those 50 ,000 tests, right? We do POCs with companies who have these large legacy applications which haven't been touched in, say, 10 years. Because nobody knows how they work. They don't want to break it. They do want to modernize them. And to be able to modernize them, they need a test suite in place so they can validate what the new is the same as the old.

31:07For these kind of problems, AI is excellent. Because what it can do is say, okay, fine, give me your 10 million lines of code and I will go and write all the tests you need. So far, so good. It's done two years worth of job in two hours. But how do I know that it's good? That's a question engineering leaders ask. That's a question developers ask, right? Okay, there's 50 ,000 tests you've written. So the first answer usually is to say, well, let's look at line coverage. So you can look at line coverage, but developers listening to this will know, and I've done this in my life as well as a developer, you can gain line coverage.

31:40Line coverage is the easiest thing to gain. And if we know it, then AI knows it too. And I've seen examples of AI just gaming line coverage, writing an excellent test that does nothing, but does give you very good coverage. We need to get one step beyond. Now, when humans are writing unit tests and their tests are being reviewed and the review is really good, then you do get these problems. These problems do get caught. With AI, we're going to have to get smarter. and there's this new technique that I think is gaining ground, which is good. It's not a silver bullet, but it's better than line coverage alone, which is called mutation testing.

32:16And the way that works, it says, okay, you've written a test suite. Whether you've written it or AI written it doesn't matter. The purpose of a unit test is to catch unwanted changes, regressions in the test suite, right? And any business logic falls. So what I'm going to do is I'm going to take your test suite that you've written, keep it as is. I'm going to then jump into your code and start making changes. Complete chaos. So if you have a check in there which says if age is less than 50, I'm going to flip it. I'll say if age is greater than 50 and run your test against it. If your test still passes, and I hope that's not somewhere in some pension calculator app because that is not looking at the age at all, right?

32:59So that's the sort of thing mutation testing tries to do. Logic flipping is only one of the many techniques it uses. But the idea is that I'm going to change your code six ways to Sunday, and I'm going to see how many of your test cases catch those faults. If your test cases make 100 changes and your tests catch all of them, then your mutation coverage is 100%. Your tests are actually really good because even the smallest of changes gets picked up. If you're only catching 20 of the faults I'm introducing in the code, and some of the faults are logic changes like I told you about. So instead of checking for equality, I start checking for not equality and then you're still passing, then that test is not preventing anything out there, right?

33:42So retelling testing is brilliant in scoring your unit test suite at scale. It gives you two numbers. It gives you rotation coverage, which tells you, which is a measure of, it gives you test strength, which is a measure of how good a specific unit test could be. and then it gives you mutation coverage, which tells you how much of this goodness is spread across your application. So you really get all these different actual metrics for figuring out if your code is good to go and if changes to it have modified it downstream. And that's actually really impactful, I can imagine, for modernizing, like you said, having those tests in place beforehand.

34:21Imagine a really, really old legacy system within the government or an enterprise somewhere and they have to turn into something modern. and think of how impactful that project could be. But think of how disastrous it could go if they don't have a way of knowing if the new machine does the same as the old one. So that's like a really big unlock for modernizing stuff. When you talked earlier about how it worked and this is something that stuck with me about it being deterministic and about it doing the same thing over and over and over again with great success, you know, that really works against the narrative for me of how I envision AI and LLMs working today because we all know that they're quite random and they're stochastic.

34:59You know, they want to be different every time. They don't like to repeat themselves. How do you approach using AI to solve something like testing if you want a standardized approach? And what's the unlock there? It's a great question because one of the things I've... So we've been in business for about six years now. We are a company based in Oxford, England. We spun out of the university. and when we were starting out there were no LLMs so they hadn't sucked all the mindshare at the industry so AI and machine learning used to mean more than just LLMs. Now thanks to the success of ChatTPT and OpenAI, AI has now been completely consumed by LLM so if you speak AI you must mean LLM and this is something I actually struggle with in my job because when we are selling into the large companies, banks, et cetera.

35:54They have these model risk reviews. And they will ask you questions like, okay, what kind of LLM model are you using? And then an answer, we don't. Say, oh, because you don't have a local LLM. Do you have it in the back end? It's like, we don't have an LLM. All right, so where is the LLM? We don't have it. It's the challenge that, because ChatTPT is the most successful AI product, right? So that's how I would approach this problem, that there are many ways to do AI. So people looking at doing testing often start with Copilot because they already have Copilot, right? They're already using Copilot for testing and Microsoft as a stellar sales force, credit to them.

36:35They already pushed it everywhere. So everyone else is, like we say in DiffBlue internally, we need to basically treat Copilot like weather. It's going to be there. We are going to have to prove that there are other ways to do testing and there are better ways to do testing. And so recently, towards that, we did this benchmarking study against Copilot. And just to give a quick brief overview, use the reinforcement learning model to understand how your code works and then write unit tests for it. We don't look at just the plain text source code. We also look at the built by code of the application, which gives us the computational understanding of what's going on.

37:14So when we're not able to write a test, we will also leave behind testability insights like, hey, do you know what? This thing is missing a getter. And because this property is missing a getter, I can't write a good test for it. Go add a package to a private getter, come back, run it again, and you're going to get a better test. And all of this works completely deterministically. And so compared to Copilot. So what we wanted to understand was, okay, how do we compare against all of these LLM-based tools? So we did a study that found that unit test generation agent, so it's basically what we're comparing is an agentic system like DiffBlue, which is completely hands-off.

37:51There's no LLM, so there are no prompts to be done. There's no, you know, back and forth. You just click a button, walk away, make yourself a cup of tea, come back, and you've got your test suite, right? So we're comparing this agentic system versus something that is more collaborative, more prompt-based thing like Copilot, but other LLM tools are similar. And what we found was that when using DivBlue, a developer was 26 times more productive than using Copilot, which is clearly easy to understand because you don't have to engage with it. You run a command, you turn around, you do other things.

38:21Then you come back, you've got your job done. Whereas with the Copilot, you know, you're fiddling with the prompt. there's this whole practice of prompt engineering coming up which i think will be quite short-lived because companies will get better at doing prompts so then you don't need to become an expert and it's basically like the transition from command line to gui you can learn all the commands but then somebody builds a gui and you just click your buttons and exactly so uh so that that's the one thing that i found uh which was actually that was expected that wouldn't surprise me what did surprised me was that we achieved significantly higher test coverage than Copilot every single time.

39:02And so when you put those two together, the fact that it's more productive, 26 times more productive, and it's giving you better coverage, over a year, this translates into covering exponentially more code without breaks than something like Copilot. So the challenge in testing will be for companies to figure out, does this testing require a human in the middle or can this be completely automated, truly autonomous operation? We believe, because we've done studies, that unit testing is a problem that can be completely automated because it's a very well-defined, specific problem that you can train an AI engine to do predictably in a deterministic fashion.

39:50So as long as you have the same code and use the same version of DivBlue, you'll get the same tests, but more importantly, at a higher quality. Because again, you take the probabilisticness away from it and then that allows you to train the model better and do better tests with every release as well. When I hear you talk about the way that teams can be using AI and thinking about it beyond LLMs. If you're a team right now and you're building AI resources, AI enablement within your org and you're trying out AI, what's like a good practice or a good habit that you would tell them to be successful?

40:28I would ask them to ask themselves if they need AI. I don't think everybody needs AI. I think, for example, if you're a small team startup writing some microservices, Do you need AI? Because the problem is going to be, and this is actually going to become a bigger problem. I've noticed this myself. AI written code is harder to debug. And it's not because it's AI written code. If you give me a job to do, I write code, and then I ask you to debug it, it'll be harder for you to debug because you've not written it, right? So this is, as a developer, we know this. We don't like debugging other people's code.

41:04AI is just other people's code, right? Somebody else writing the code. So a lot of developers, this is your co-pilot. This is your co-pilot who's doing job for you, but it's not you. So it is adding friction into the process that when things break and when there are bugs, I am seeing and I'm hearing from companies as well, which is why they come to us, right? This is both ways, because most people come to us after having tried co-pilot and not liking it for testing. That it's taking more time. So the whole promise of AI making you more productive goes out the window. Like it's definitely not making you productive.

41:38It's maybe making your job more interesting. You know, you're not just writing code in an idea. You're having experiments. But that's the first thing I would ask people looking at AI in software engineering. Identify a good problem and then ask yourself, do you really need AI there? And do you really need large-scale language models there? Because there are techniques out there where you don't even need AI, right? So simplest example ever, if you want just some way to make it easy for your developers to create microservices, the two ways you can go about it, there's the cowboy way, which is to roll out some kind of an LLM CLI tool, which will create this for you.

42:26And the second way is to create a GitHub template. The first one is a one-off effort, but it's predictable. Every single repository created from that will always come out the same. So then it's easy to debug and you solve one. There's a problem. You solve it once, you solve it everywhere. That sort of thing. I'm seeing people using AI at these sort of things. I think it's actually going to crash and burn. It's going to cause a lot of pain because what we're going from is... And we moved away from doing crazy things to more of a standard DevOps model in software engineering. and now we're then going again to doing some crazy things and then it's going to iterate and get us to a better place.

43:09I think if we're in this middle where there's a lot of churn, people figuring out what to do, without a good problem, you are not going to find AI useful. So my only recommendation and the only first thing I ask people is like, what problem are you trying to solve? Do you need AI? And if you do, make sure you understand how AI works and what types of AI need to be used in what circumstances. Today, we learned about a more traditional version of AI that's beyond LLMs, deterministic, and could be really useful for engineering leaders trying to standardize and scale up solutions within their organizations, especially if they're already investing and putting mindshare into an AI-driven world.

43:50It's definitely a card to play and something to keep in mind. It's also something that allows you to unlock what test-driven development was always meant to give us and maybe bring us better to having secure software that runs our world. So this has been a really insightful one for me. Animesh, it's been great having you on the show. It's really been fascinating to get to learn about your expertise. Before we wrap up, though, where can our audience go to learn more about you and to follow your work? To learn more about my company, DiffBlue, you can go to www.diffblue.com. We are based, like I said, in the UK, but we have customers all over the world.

44:28Our sales pitch is pretty straightforward. We believe that balancing quality and speed is crucial for sustaining reliable and maintainable software products. And the way to do that is to have reliable, maintainable, and predictable development tools. And DivBlue is one such tool that we believe you should have in your arsenal to take advantage of the advancements that AI has produced for us. You can also follow us on X. Our handle is at DivBlueHQ. I think I got that right. And you can connect with us on LinkedIn as well. If you'd like to follow me, I would love to connect with you personally. You'll find me on LinkedIn.

45:08My name is Animesh Mishra. You can search by username. I'm Sir Animesh on LinkedIn. Oh, we'll definitely get your links in the show notes. Be sure to subscribe if you haven't already and share if you found this insightful with your teammates. And also be sure to check out our Substack. Our Substack is a weekly newsletter where we release our podcast as well as a roundup of some of the stuff we've discussed today. And I'll also be including the notes from today's guests in the Substack newsletter and on the show notes. Like Animesh said as well, we'd love to hear from you on socials. So please come find us on LinkedIn.

45:40I'll make sure we're both linked. We'd love to hear your thoughts on test-driven development. Is your organization doing it? And what do you think of this kind of solution? And that's it for this week's Dev Interrupted. See you next time.

46:00I'll see you next time.

From the publisher

The promise of Test Driven Development (or TDD) remains unfulfilled. Like many other forms of aspirational development, the practice has fallen victim to countless buzzword cycles. What if the answer is already in our toolbox?

This week, host Andrew Zigler sits down with Animesh Mishra, Senior Solutions Engineer at Diffblue, to unpack the gap between TDD's theoretical appeal and its practical challenges. 

Animesh draws from his extensive experience to explain how deterministic AI can address the key challenges of building trust in AI for testing. These aren’t LLMs of today, but foundational machine learning models that can evaluate all possible branches of a piece of code to write test coverage for it. Imagine writing two years worth of tests for a legacy codebase… in two hours… with no errors!

If you enjoyed this conversation about the gaps between theory and execution in engineering culture, be sure to check out last week's chat with David Mytton about shift left adoption by engineering teams.

Check out:

Follow the hosts:

Follow today's guest(s):

Support the show:

Offers:

More from Dev Interrupted

All 208 episodes
Will AI Finally Make TDD Practical?Dev Interrupted · 46 min
Listen in VO