In short
Dan Faulkner (SmartBear CEO) argues that AI coding accelerates “clean code” but not “application integrity,” creating systemic digital risk. He says SDLC pillars like application testing, security, monitoring, and requirements must keep pace, with continuous end-to-end validation in real environments (not just unit tests). He highlights agentic-era risks like slop squatting, prompt injection, instruction inversion, cascading errors, and unreliable model benchmarks (e.g., SWE-bench false positives). SmartBear’s approach: API governance via Swagger/OpenAPI-based lifecycle management to prevent spec/test/code/doc drift, plus agentic application testing (“BearQ”) where agents explore apps, build a knowledge graph, generate/run tests, and report results with human oversight.
Guests
Dan Faulkner is the sole guest/host in the transcript (CEO of SmartBear; previously CTO at Planner; earlier career at Nuance Communications). No other guests appear.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOSmartBear's Role in Application Quality
0:46 to 2:33
Discussion on how SmartBear enhances application quality through testing and API management.
“So that means testing applications above the waterline of the code, the actual real application in its real deployment environment, and in API lifecycle management.”
Integrity of Applications in AI Development
2:34 to 4:24
Exploring the integrity of applications amidst rapid AI-driven software production.
“This is something I've spoken to a few people about.”
Probabilistic AI and Code Challenges
6:45 to 9:13
Examining the challenges that arise from using probabilistic AI in deterministic software environments.
“is a registered broker-dealer and member of FINLA, NFA, and SIPC.”
Evolving Issues in AI Software Development
9:14 to 12:18
Addressing new challenges such as cascading errors and human error in AI-generated code.
“How much am I willing to give up the velocity gains of the coding process in order to rush this out into production because my testing practices can't keep up.”
Company Solutions for Unit Testing
12:19 to 14:00
Discussion on a company providing solutions for unit testing in JavaScript, highlighting industry innovations.
“So we need to be incredibly excited about the advances in the state of the art.”
Hidden Issues in AI Development
14:00 to 14:51
Explore the hidden issues and cascading errors in AI coding.
“there are issues that are being hidden, not intentionally, but they're being hidden and we will have to reckon with those.”
The Challenge of Large Code Bases
14:52 to 16:47
Discuss the difficulties of managing massive code bases generated by AI.
“getting traction with it and i don't know where they are now but since then there's been this tsunami of other code test writing solutions.”
The Need for Human Supervision in AI Testing
16:48 to 18:58
Learn about the critical role of human oversight in AI-generated code testing.
“but it's a complicated environment that we live in.”
SmartBear's Approach to Application Testing
18:59 to 21:09
Understand how SmartBear is innovating in application testing technologies.
“It's certainly what we're designing for.”
Continuous Testing and Agile Development
21:10 to 23:01
Discover the importance of continuous testing in modern software development.
“we've created, including ones that contain agentic workflows, they're geared towards a human as the primary doer, the tester.”
Show all 20 chapters
The Role of In-House Development Teams
23:02 to 28:00
Discuss the dynamics of in-house development and the need for effective testing.
“Does that mean the one system that's testing another or one tool that's testing another has to be based on a different foundation model?”
Testing Application Integrity in Software Development
28:00 to 29:52
Learn about the challenges of ensuring application integrity given the abstractions in modern software development.
“Our fundamental kind of simplifying point is, whichever one of those you're doing, you should be testing that your applications actually work.”
Innovation vs. Adoption in Technology
29:52 to 31:31
Discover the differences between rapid innovation and slow adoption in the tech industry.
“And if that's continuous, then it does that, is there one person in the organization that's, uh, doing this application testing with SmartBear or, or is it distributed across everyone that touches the application?”
AI's Impact on Software Infrastructure
31:31 to 33:53
Explore the implications of AI becoming critical infrastructure and the need for audit trails.
“As AI-generated software proliferates, it's already in some areas becoming infrastructure, really, rather than just a productivity tool.”
Risks of Automation Bias in Testing
33:53 to 36:29
Understand the concept of automation bias and its potential dangers in application testing.
“If we choose, the more we choose to allow AI and agents to operate in a black box, the more systemic risk we will introduce into our infrastructure.”
The Future of Quality Assurance
36:29 to 37:48
Discuss how quality assurance is evolving in the face of new AI technologies and automation.
“on automation and it will require significant errors and their consequences to cause people to change.”
SmartBear's Role in API Governance
37:48 to 42:00
Learn about SmartBear's approach to API governance and the importance of avoiding drift.
“Where do you see it in, say, five years?”
Agentic Tools for Software Development
42:00 to 46:00
Learn how agentic coding is transforming the software development landscape and the role of human oversight.
“you're in a kind of heavily people-driven process great we can help you do that if you're in a heavily agent-driven process, we can help you there as well.”
Market Dynamics and Competition in AI
46:00 to 50:40
Explore the current state of the AI testing and validation market and strategies for differentiation.
“right end user whether that's a person primarily agents primarily or a combination of both and And SmartBear with BearQ, is it?”
The Evolving Role of Human Talent in Tech
50:40 to 52:30
Understand how human roles in software development are changing as automation increases.
“I think that we may see that the agent workers fill in the gaps that are already there and the gaps that will be growing very quickly as everything accelerates.”
Transcript
Automatic transcript. May contain errors.0:00My name is Dan Faulkner. I'm the CEO of SmartBear. I joined SmartBear about five years ago, initially as the Chief Product and Technology Officer. I've been the CEO since the beginning of 2025. Prior to joining SmartBear, I was at a small Martech startup called Planner, which we sold where I was the CTO. So prior to that, I spent most of my career working for a company called Nuance Communications, which was one of the pioneers in really field hardening speech technology and making that work well for real people in real world environments. Now, what we do at SmartBear is we provide tools for development teams to improve the quality of their applications.
0:43We do that in two primary domains. The first is in application testing. So that means testing applications above the waterline of the code, the actual real application in its real deployment environment, and in API lifecycle management. So APIs obviously are now the vocabulary of agents, so they've become even more important than they were, and they were critical anyway. So they're the two main areas that we focus our product investment into. Yeah. And did you say testing APIs or configuring APIs? It's the full lifecycle management. So our product is built on top of a very well-known open source product called Swagger, which is essentially the codification of the open API standard.
1:35And the commercial product that sits on top of that is an API catalog that manages all the artifacts across the lifecycle. So the specification, its documentation, and tests all of them. So the contract testing, the functional testing, performance testing, and so on of APIs. So you can kind of keep the API catalog well governed and up to speed. Yeah. And as you said, API is getting more important because of the explosion or what's anticipated as an explosion of agentic AI that's connecting and calling APIs for tools or databases or things like that. But as AI increases software production, the velocity, the speed of software production, if application level validation doesn't scale alongside it, are we increasing systemic digital risk?
2:43This is something I've spoken to a few people about. Yeah, I mean, it's certainly what we believe. that there are pillars throughout the SDLC are resilient and meaning they will still be there in the AI disrupted SDLC. And those are failing to keep pace with the acceleration in coding. One of those pillars is application testing. There are others like security and monitoring, you know, even strategy and requirements, creation and definition, those things all still need to be done. And coding specifically has accelerated way past the capabilities of what those pillars have been able to do. And the risk that that creates kind of the systemic risk is that the vocabulary we use to describe it is you lose the integrity, you risk losing the integrity of your applications.
3:38And what that means is, do you really know what your application does? Is it doing everything you wanted it to do and nothing you don't want it to do. And one way that we're addressing that is through the application testing space. So yeah, it's definitely an issue. Yeah, yeah. And yeah, so application integrity is, we'll describe integrity again. So application integrity is, you can kind of view it as the assurance that your application does everything that it was intended to do and that it will just work for your end users. And you need that to be continuous assurance because the applications themselves are evolving so quickly.
4:24So what's the flip side? How would that break down? You would start to see that there are missing capabilities in your application that were specified up front. It was not behaving as you anticipated. It had some unspecified behavior in it, or even though the code is clean, it's not actually solving the needs of your end user. Those are all ways in which the integrity of an application can break down as we rely more and more on these kind of coding agents to do the code creation for us. This episode is brought to you by Tasty Trade. On Eye on AI, we talk a lot about how artificial intelligence is changing how people analyze information, spot patterns, and make more informed decisions.
5:10Markets are no different. The edge increasingly comes from having the right tools, the right data, and the ability to understand risk clearly. That's one of the reasons I like what TastyTrade is building. With TastyTrade, you can trade stocks, options, futures, and crypto all in one platform with low commissions, including zero commissions on stocks and crypto, so you keep more of what you earn. The platform is packed with advanced charting tools, backtesting, strategy selection, and risk analysis tools that help you think in probabilities rather than guesses. They've also introduced an AI-powered search feature that can help you discover symbols aligned with your interests, which is a smart way to explore markets more intentionally.
6:08For active traders, there are tools like Active Trader Mode, One-Click Trading, and Smart Order Tracking. And if you're still learning, Tasty Trade offers dozens of free educational courses, plus live support from their trade desk reps during trading hours. If you're serious about trading in a world increasingly shaped by technology, check out Tasty Trade. Visit tastytrade.com to start your trading journey today. I'm going to myself. Tasty Trade Inc. is a registered broker-dealer and member of FINLA, NFA, and SIPC. Yeah, you know, the AI is probabilistic and traditionally software has been deterministic.
7:03So when you're using a probabilistic system to write a deterministic code base, I would imagine that on the margins, maybe, you know, maybe not, maybe there's 97 % accuracy in the code writing AI, but that 3 % can slip into the code And these code bases get so large because they're being written so quickly, it becomes a real challenge for humans to go through line by line and figure out whether there's anything wrong. I mean, that's basically what you're talking about, right? Yeah, I mean, I think there's kind of two angles to it. One is in the code. So let's start with that one. There's just a giant amount of code that's being, the phrase I've been using is like fire hosed into the repos of the world, right, at record speed.
7:58And we do need to stay current. In the last few months, it's very clear that Claude Code and OpenAI's latest version of Codex, they've kind of crossed some very important threshold where the code is really good. It's clean. And it's pretty good at testing itself. And it's really now become something that I think people feel confident in generating code that's quite reliable. Our point is that you can have reliable code and you can have very hygienic, clean code that passes all of its unit tests. That doesn't mean the application works as intended. They're very different things. And you need to actually take the compiled application in its end-to-end environment.
8:43So it's distributed AWS architecture, for example, then you need to figure out, does that work end to end for your end users? Does it solve the business problems that really need to be solved? And does it do it across all the operating systems and browsers and disparate devices that your users are employing to use your application? And that has been something that wasn't keeping pace with development before AI. and it kind of gets left in the dust without a different approach in the AI disrupted SDLC and it's a real problem because any business that wants to release an application and they have material consequences for that application failing has to reckon with how much do I test this?
9:34How much am I willing to give up the velocity gains of the coding process in order to rush this out into production because my testing practices can't keep up. So that's where they, if they make that trade-off in the wrong direction, then the integrity of their application is at risk because they will have issues that they're unaware of in production. And if you're relying on AI tools, they can introduce patterns or vulnerabilities at scale, right? Across thousands of applications simultaneously. Is that an issue? It's a real thing. and certainly for the industry there's new categories of issues that are emerging and they're being I would say exacerbated by the emergence of agentic workflows so now we're thinking about there's the coding but then there's the agents who are actually you know using capabilities and there's pretty interesting sets of issues emerging some of those lie in security There's a pretty common issue that's emerged, has a delightful name.
10:36It's called slop squatting. And it's where agents will import or refer to libraries, third party libraries that don't exist. And a bad actor can create that library and put it into the blockchain. And suddenly that's deployed at scale. There's the well-known issues around prompt injection to get agents to create malicious code. But there are also more subtle issues where if you've interacted with anything like Claude Cowork or ChatGPT or Gemini, you'll have had the situation where it does something that you don't want it to do. This could be as simple as if you're writing a document and you can say, hey, can you can you stop doing that?
11:16And for me, one of my pet peeves is they all seem to have this real tick where they like to say, hey, this phenomenon isn't just A, it's B. And it creates this ludicrous kind of straw man that makes cringe and it just screams of AI. And I'll say, hey, can you stop that? And he goes, oh, absolutely, I'm going to stop that. Creates a new draft and does it again. And that's called instruction inversion. And it happens in coding too. So you can say, can you please not do this behavior when you create code? And it says, I will absolutely not do that, I promise. And then does it straight away. And so there's this whole raft of systemic issues that are appearing in code bases, some of which are clear and some of which aren't.
12:02And it even permeates how the frontier models are tested. There's a very, very famous and well-known thing called SWE Bench, which is used to assess the quality of the latest frontier models. And that has also been shown to mark things as being correct when they're not and to mark tests as passed when they're not. So we need to be incredibly excited about the advances in the state of the art. We also need to be realistic that there's a whole set of new challenges that need to be thought through and wrestled with and addressed. So it's an interesting time. Yeah, I mean, it's fascinating that now there are terms that are being developed for all of these various behaviors, whereas at the beginning, it's just you don't know what to call it, but it's doing these crazy things.
12:55I mean, one of the things about the probabilistic nature of AI is that when it makes a mistake, which it's going to do because it's never 100 % accurate, and you try and correct the mistake, if it makes a mistake in the correction, you end up magnifying the problem and you just can never get ahead of it. because every time you, I mean, maybe I haven't used the latest iteration of Cloud Code, but there was a time when that was very frustrating. You know, you'd ask it to fix that error and it would just say it had fixed it, but it had actually created more problems. I mean, I do think Cloud Code and Codex and maybe other models too, but those two specifically, literally in the last two or three months, have made a leap in improvement and I think addressed many of the frustrations.
13:51but they're fitting into a system of other agents and other tools. And it's unquestionable that as they're going at such velocity and such a volume, there are issues that are being hidden, not intentionally, but they're being hidden and we will have to reckon with those. And you also get what's called cascading errors. So agents working with other agents, one makes a small error, that gets amplified. We get errors in the context files that people are providing to guide their agents, striking how many people don't really understand how what their software processes are what their company's rules are so it's great that we can give you know an agents.md file a ton of context and say this is how we want you to build but of course if that's wrong you're building a whole bunch of software with baked in errors that's not the agent's fault that's that's the human's fault but the humans in the loop yeah i know a company uh called diff blue out of the uk out of the uk actually i don't know if you know them out of oxford uh and they have a solution very specifically for writing uh unit tests for javascript i believe it's javascript and they had a real hard time getting traction with it and i don't know where they are now but since then there's been this tsunami of other code test writing solutions.
15:14But what was interesting is we would talk about how as AI coding takes hold, the world is going to end up with these massive code bases that no human has ever read. And maybe, I don't know, is it possible for a human to read these code bases once they're installed? The code itself is readable. It's much harder to read someone else's code than your own. But that's okay. It's doable. And the human code review has been a real thing for a long time. But code review itself is being automated heavily. And we start to get to a little bit of a black box problem. We also end up, and I think the market is cottoning onto this right now, we have a not even long-term, but a short to medium-term expertise problem.
16:03Because to the extent that people are starting to say, hey, I'm not going to hire junior developers, I'm going to rely on my agents, you're going to lose the human resource pool of people who can actually open and understand and think deeply about code. So there's real debt problems, technical debt, knowledge debt, design debt, all sorts of debt issues that will need to be addressed. And we're in this kind of very giddy environment at the moment where we're behaving as if all of these issues get solved by one tool. And they don't. There are some really important things being very well solved by AI.
16:46And I'm actually very bullish on it. but it's a complicated environment that we live in. And people are going to need to be in the loop for the foreseeable future in existing ways and in new ways. Yeah. And when I said that the code bases that no one has read or can read, when I engineer developers would be incapable of reading it, it's just if you have a code base of a couple million lines of code, I mean, that was written by AI and it has to be gone over by a human. That's a pretty daunting task. And you mentioned then using AI to test or validate the code. But that in itself, as you said, the black box problem, you don't know whether it's validating systems that shouldn't be validated.
17:43I mean, it just compounds the problem. So how do you avoid self-validating systems that simply confirm their own assumptions? Yeah, I mean, you need systems that are independent of each other to validate, and you're going to need humans supervising those processes. So, you know, the next iteration of our application testing products are agentic. They have to be in order to keep pace with agentic coding so they can explore an application and create a test strategy and create tests and run tests. And it's good that they're doing that outside of the environment that the code is being generated in.
18:30That alone should give you some kind of third party confidence. But you absolutely must have a human who's looking over that system's shoulder and say, well, show me what you're doing and show me your thinking and show me the results. And are you testing the things I want you to focus on? So we have to design agentic systems, I think, to work in partnership with people. And it's very, it's both provocative and seductive for folks at the moment to be thinking that the agents are working independent of people. I think it has to be a partnership. And that's what we're going to see. It's certainly what we're designing for.
19:08Right. Yeah. So can you talk to me about how SmartBear is approaching this specifically? We have this sort of simplifying analogy that we stole from automotive driving that we call the ladders of autonomy. And many of us will have heard of levels of autonomy in cars, with Waymo is the fully autonomous and, you know, your old fashioned manual car with no bells and whistles is a level one. And if you think about that, from a coding perspective, there's kind of a level one idea where there's a human who's kind of typing all of their code, and then they get an ID that does code completion, and they're bringing in libraries, and then it gets more and more intelligent up to kind of a vibe coding idea where you're saying, hey, build me an app that does X, and it's off to the races and does most of the work for you.
19:55As you go up that autonomy ladder, think about the tools that exist to help you. So you might have a tool that a manual tester uses to help them explore an application and do their testing. Then you might have a test automation tool or an open source framework that will allow you to record an interaction and it turns that into an automated test. But that still requires a person kind of at a GUI pointing and clicking and doing stuff. And what we have tried to do at SmartBear is build tools that will help people wherever they are on that autonomy ladder. Because there's tons of industries that are very cautious about adopting agentic systems and autonomous systems.
20:38And we've got to make sure that they've got what they need. But they all want to eventually when the regulations allow them and their lawyers allow them they'd like to get to this kind of highly automated reliable world the new capability that we've created is a platform called bear queue and what that does i'm sorry bear q q yes bear nope the letter q what that uh does and how that's different is whereas pretty much every other tool we've created, including ones that contain agentic workflows, they're geared towards a human as the primary doer, the tester. And BearQ is designed with the agent as the primary doer.
21:26It's a team of agents that work together and they have different skills. And it's able to explore an application, figure out what it's doing, it builds a knowledge graph and says, okay, I understand what this is all about. And then it can build a testing strategy, author the tests, run the tests, it can introspect if a test fails, and say, well, did I write this test badly? Or, or is this really an issue? And as the application evolves, and there are, you know, each time it's iterated by the development team, it will update its its application exploration, its knowledge graph, its testing strategy, and its tests accordingly.
22:04Very importantly, though, we have not designed this to be independent people. We've designed it to click into the testing, to fit into the testing teams that exist today. We view it as an augmentation to the organizations that exist today. We don't think people want to abandon human oversight of application quality at all, but we do know that they can't keep up. So what we're trying to do is think of it as, If there's a pie chart, there's a slice that's being addressed by the tools and the people to date. We're basically saying, hey, let us fill up the rest of that pie chart and help you keep up.
22:43And that's kind of our approach on the application testing realm. So they get the benefit of the pace. It's independent of the coding agents. It's above the code level. So it's testing, will this really work for real end users? And we think that's solving an important problem. Yeah. And to avoid the self-function, you said you need very independent systems. Does that mean the one system that's testing another or one tool that's testing another has to be based on a different foundation model? Not necessarily. I mean, we exploit the frontier models. But what you're really doing is testing an output of the model, not the model itself.
23:30And so when you present a compiled application to our product, it doesn't know what frontier model was used to generate the code. It doesn't care. It doesn't care if it was a person or a model or a combination of people and models. All it's doing is essentially using the application to learn what the application is intended to do. Yeah. And things move so quickly. Does testing have to move from pre-deployment to sort of continuous testing? Yes, I do believe that you will need continuous testing. I think it's not so much a pre-deployment and post-deployment. it's both. You should be continuously testing pre-deployment and you should be continuously testing in production because there are new factors come into play when you go into production and real world end users that you don't control in your testing environment come into play.
24:28And so you absolutely need to be testing in production as well, but it's got to be continuous because the updates to applications will become continuous. It's in the name of CICD. It's arguable whether it's really been continuous up until now but it's going to get more continuous um when you have coding agents creating updates and new applications much much much more quickly yeah and in the old days not very long ago programmers wrote unit tests or and tested as they went along it was kind of a tandem exercise you would write code you would test you would write you would test at this point in smart bearers tools, do you envision, or maybe it's already happened, that's a role within a separate role within the software development team, somebody who understands the tools for testing and just focuses on that while the programmer continues writing, you know, with a co-pilot of some sort.
25:36Yeah, I mean, I think that unit testing and other types of code level testing have pretty much been subsumed into the frontier models that the coding specific frontier models clawed code creates the code creates the unit test creates better unit tests than most developers would more fleshed out more complete more counter example types of tests. So all of that that that creates good, clean code. What will be needed is the continuous ability, whenever the application is actually compiled, to say, okay, now let's actually use that and see if that works in the end environment, in the deployed environment.
26:21And I do think that that will need to become much more integrated into the SDLC. So I can easily imagine a scenario, if you can imagine continuous code updates continuous builds every time that happens having an immediate testing of the built application feeding back to the the developer um i i think that that kind of it's not so much shift left it's kind of smeared everywhere it's got to be left and it's got to be right as well it's got to be in production as well yeah and and that uh as you say above the code line uh application testing. Even there, I mean, it used to be there were third party firms, UAT firms that did that for you.
27:09I mean, is this, who are SmartBear's customers, I should say? Are they separate testing firms that are contracted to test applications? Or is this intended to be in-house application testing? Yeah, for the vast, vast, vast majority, we sell into in-house development teams within companies. And how they're constructed, whether they have kind of integrated quality and development or separate quality and development tends to depend on the tech stack, the industry, the age of the application, their tool chain. And really, our goal has always been to say, yep, we'll kind of, we'll meet you where you are, depending on, you could be super progressive, you could be doing something once a quarter.
28:00And any of those are fine with us. Our fundamental kind of simplifying point is, whichever one of those you're doing, you should be testing that your applications actually work. And we know that you don't have enough people, you don't have enough seats of any software to be able to do that sufficiently quickly and sufficiently well. And you need to do it more acutely even than you previously needed to do it because you've allowed consciously, you've made the choice to allow a degree of abstraction to come into your development processes. I listened to an interview just the other day where the head of product at Codex said that none of their developers open their IDEs anymore.
28:51They're just orchestrating teams of agents. So by definition, you are further away from the code. And I think people are saying that's okay, as long as the coding agents are reliable and create clean code. but it it creates a distance and an uncertainty and the likelihood that you're going to get more drift between what you really intended the application to do and what it really does goes up and there's going to be a need to validate that what you really intended is really happening and it doesn't matter how well your context file is defined it you're going it doesn't matter how well your user requirements are defined, you've got to validate that and understand what the drift is between, um, those upfront plans, that upfront context, and the thing that has been built at the end of your, of your tool chain.
29:54And if that's continuous, then it does that, is there one person in the organization that's, uh, doing this application testing with SmartBear or, or is it distributed across everyone that touches the application? I don't think we really have the right to be prescriptive about that. In the same way, every company, you can go to 10 different companies, you're going to find 10 different tech stacks, 10 different tool chains, 10 different organizational models, responsibility models. And I don't think that's going to change. There's not going to be one uniform approach to software development to rule them all.
30:34So that's one factor. There's just an incredibly diverse ecosystem that we live in. The second is that while innovation happens quickly, adoption happens slowly. And we see that time and time and time again in technology. It's the adoption of tools and agents and AI is going to move at different paces, company by company, industry segment by industry segment for diverse reasons. Sometimes it will be mandated by regulatory requirements. Sometimes it will be mandated by a huge cost of failure, either financial, reputational, legal. Sometimes it will just be mandated by conservatism. We need to assume that that diversity of approach will persist, even as AI is probably adopted more quickly than any technology that's happened before.
31:30Yeah. As AI-generated software proliferates, it's already in some areas becoming infrastructure, really, rather than just a productivity tool. How do you see that affecting systemic resilience? There's a lot of fear in the market at the moment. We can see it in the public market and in private markets about what this all means for the future of software that we've previously thought of as critical infrastructure. I would actually say that it's uncertainty about how this will all shake out rather than a conclusion that the existing infrastructure itself is going to be replaced. we will continue to need audit trails and proof and therefore those systems can't be replaced by by black boxes we cannot have ecosystems of technology that say here's what i'm going to do i i'm doing it now i've done it and it's perfect and we just say great i think infrastructure will evolve, I think that we will see portions of it be hugely impacted positively by AI.
32:48But if we fundamentally look at AI as a new type of automation, and we think of agents as workers, I think what we're going to see is that the overall population of workers grows, and the overall proportion of workers will increasingly be agents. There's going to be a whole set of new tasks that we need to do to oversee the agents, to validate their work, to validate intent, to make sure that things are as intended. And I think that's actually going to create new opportunities for humans and they will need their own infrastructure and their own tools. So we're not going to see this kind of unidirectional change where everything becomes agentic and everything stays the same for humans and it's a zero sum game.
33:41I just don't see how that's possible without people abdicating responsibility in a way that they don't and won't and can't from any business or government perspective. If we choose, the more we choose to allow AI and agents to operate in a black box, the more systemic risk we will introduce into our infrastructure. There was just two days ago a story about the, I think it was the head of AI security at Facebook or Meta, who was working with Clawbot, this kind of open claw, Maltbook. It's changed its name sometimes, the thing that OpenAI just acquired. and it was an experiment i think so you know i'm but the fact is it she had given it very clear instructions and said do not take any actions until you've checked them with me and i give you permission to take the actions and uh it deleted her email inbox and it it nuked all of her emails and there's actually this published conversation where she's going no stop don't do that don't why did you do that i you know and the agent the agentic system says yeah i there was a rule i understand that you gave me a rule and i broke it i'm sorry about that it won't happen again and it's like how do you how do you know um like it seemed to arbitrarily decide not to follow the rules what's technically feasible and what we can imagine one day needs a tremendous amount of hardening and scaffolding and process and insight that doesn't exist today.
Read the full transcript
35:16There's huge opportunity for us in there. Yeah. And so you're not concerned that, you know, there's something called automation bias, where I can imagine this job of testing applications must be exhausting. But if you're testing something with a tool that's automating, or you're watching the automation and it does it perfectly 10 times or 100 times pretty soon you just thought you start trusting the automation is that a concern of even if smart tool i mean oh you were saying a smart bear i'm sorry the smart bear uh requires a human to be in the loop it's designed for that or where is there a risk that uh people will just start hitting that button over and over and letting run well there's the we we can see what's happening in the market today that people are trusting ai codegen without checking it properly and it's going to create a it has already created a giant technical debt that will need to be unwound um so it's 100 going to happen people will over depend on automation and it will require significant errors and their consequences to cause people to change.
36:42And I think that that is completely consistent with the adoption of any automation I've experienced. For me, that was kind of the migration from on-prem to cloud, the emergence of mobile and those caused huge change some people rushed out early almost kind of proved what's possible and what will be possible but didn't survive because they um they got over their skis and catastrophic failures and we will definitely see that definitely see that play out here um and i think that you know we're not a business to consumer product we're a we're a b2b company and what we're trying to do is look out for businesses and say hey you know let's anticipate environments where we can get the benefits of the velocity and the quality that these AI solutions can bring us now are tremendous benefits that we should be excited about let's also make sure we can do that safely yeah so QA as it's been traditionally is sort of entering this period you know, structural transition.
37:49Where do you see it in, say, five years? Or is at the pace things are changing? Is that too distant or rising? Five years is too distant. I see tight partnerships between people and agents and a much more fluid SDLC where the people have much more generalist skill sets. There's going to, so people who actually know what they want to be built, what the business constraints and priorities are, what they want the end user experience to be. If you can describe those effectively, then I think we'll see more and more reliance on agents to do that work and to do it quickly. I really struggle to see a world where people are not in that loop in multiple ways.
38:41I also think you're going to see a need for new roles. You're going to need to have people who deeply understand software architecture, who deeply understand security, who deeply understand quality, deeply understand code, because things will continue to go wrong. They will continue to break, and you have to have people who understand that and they may too be using specialized agents and specialized tools i'm not saying that that doesn't include agents but they're different roles and and so it's going to be a different organizational model a different balance of human and agentic work um and i think that we will see a huge amount of experimentation in companies as they try and find the thing that works best for them.
39:35But as I said earlier, today and for as long as I've been in software at all, I've never seen two companies that approach software development the same way. I don't see that changing. Talk about some of how someone uses smart various tools because you have a fairly broad platform now from APIs, testing, observability. what what sort of the vision what the company is how it fits into the the application development necessarily just the code and how is that about how are you keeping up with the market yeah i mean we do have quite a few products but you can really boil down our focus to being around api governance so api life cycle management we think is increasingly going to become a governance question your catalog of APIs, which will sometimes be used by humans and sometimes be used by agents and probably mainly be used by agents.
40:35You need those agents to really understand what the API does. We've mentioned that word drift before. You need to make sure there's no drift between the spec and the tests and the code and the doc. They all need to be absolutely lock tight, which means they need to be managed and governed together. All the artifacts need to be governed together. And then that catalog also needs to be able to keep pace with the agents. So it needs to be agentic. That's our strategy there with APIs. To the extent that APIs are kind of the lingua franca of agents and software in general, that just becomes more important.
41:13So avoiding drift, providing great governance and audit trails, and keeping pace with agentic coding. That's what we'll do in that family of products. On the application testing side, we already address the critical application level testing tasks. So whether it's web or desktop or mobile, cross-browser, any endpoint, load testing, functional testing, all of those capabilities are capabilities that SmartBear supports. our goal and our strategy and you know when we launched bear queue i would say our reality is that we're bringing those capabilities to people wherever they are on that autonomy ladder those are all things you must do they're all if you're not doing them you should be doing them but if you're in a kind of heavily people-driven process great we can help you do that if you're in a heavily agent-driven process, we can help you there as well.
42:13It's really our mission, and we're not looking to get out of those zones. We feel that they are going to be needed and needed more as agentic coding really takes off and becomes truly reliable. Those questions or the questions that we answer are going to become kind of more urgent and more acute, even more than they are today yeah and and how does somebody use uh your primary tools i mean is is there a chat interface do you do oh i see give it access to your code base and yeah and so so and it's on the product we have some products that are like systems of record so the api lifecycle platform swagger basically will ingest the APIs from your gateway or from your systems, and then we will organize the assets around that.
43:10And so today, most of our users for that product today are people. So it's a GUI interface, or you can get kind of command line interfaces, but it's all also accessible via MCP server if you do have agent workflows. workflows. We've integrated agentic workflows into the product as well. So literally kind of like with a push of a button, we can read your code and generate the spec. We can read the spec and generate the documentation. So we can do all of that clever stuff. And what we'll be doing over the course of the year is introducing more autonomous agents to do a bunch of those tasks for you.
43:49But for the API catalog, I anticipate that those are being overseen always by people. I think that's just a key requirement. What is the role of the human that's managing this? I mean, you said they would be in charge of governance of the API library. But for example, in testing of an application, so much of this is agentic. What is the human doing? Just watching it or reading summaries or, yeah. Yeah, so what we do with BearQ, the new agentic and highly autonomous application testing product, is we show our work. So we create daily reports. We have great analytics and insights. We literally say, here's what we did for you today.
44:46Here are the tests we created. Here are the tests we run. Here's what failed. Here's what didn't. People can click into that as much as they want to and inspect it down to a fairly atomic level. We mentioned the word trust earlier. As their trust increases, they may prefer to just stay more at the dashboard level. But they can also directly intervene. So they can say to BearQ, hey, I would like you to go and focus a bit more on this portion of the application. Go deeper down there. And that can be done through a chat interface. Or they can even just say, you know what? I'm going to come in, I'm just going to write a test because there's a very specific test that I personally want and they can inject that into the process as well.
45:23We're trying to accommodate the ways in which people will want to work with agents and with an autonomous system. And for some people, that's going to be more hands off, more summary driven. And for others, they're going to say, no, I've got team members who I actually want in there working alongside the agents. so so it's incumbent on us to try and accommodate those use cases use cases as well but in general when you ask how do people use our products we're trying to tailor the user experience to the appropriate you know step on the autonomy ladder and make sure that it's optimized for for the right end user whether that's a person primarily agents primarily or a combination of both and And SmartBear with BearQ, is it?
46:10Yeah, BearQ. Yep. How do you see the market? Is it kind of an endless market right now as Agenic enters the production processes? Yeah, I mean, I think the demand, the need is uncapped. I feel like there's another shoe or two that need to drop in the world's evolving understanding of how software is not just developed, but developed, deployed, and maintained. And right now, we're incredibly focused on how software is generated. But there's an awful lot, like 80 or 90 % of your work is ahead of you once you've built an application. That's where we need to shift our focus next, or else we'll find that the benefits of just creating applications or updating applications very quickly will end up being somewhat diluted.
47:07Yeah. And do you feel like the market that you're focusing on or the, you know, testing and validation, it's traditionally been a crowded market. Is it thinning out as it gets into a Gentic because people aren't there yet? Or do you feel that it's, that there's a lot of competition. And if there is a lot of content competition, what's the strategic bet that you're making to differentiate smart pair? What we're trying to do is create clarity for the market and clarity for our customers and prospects about what precisely we and others mean when they talk about having AI powered solutions. There is a world of difference between pointy clicky SaaS app that has a couple of AI features or has a chatbot on the front and a product that's genuinely is a team of agents.
48:07They're wildly different. And that's why we've been leaning so hard on these concepts of application integrity and also on this autonomy ladder is so that we can create a framework to explain where we are. We think that what we have with BearQ is unique. We don't know what other folks are working on. And it's unlikely that we'll be the only company who has a product that does that, you know, forever. I think there's a disservice being done to portions of the market at the moment by a number of vendors saying, hey, we have this agentic product, we have this autonomous product, but we look at it.
48:44Not really. It's not really. you're just kind of jumping onto a word and sadly diluting its meaning, which will actually end up holding the market back a little bit. We will also see some new companies emerge. There's a lot of Y Combinator startups who are addressing the kinds of problems that BearQ does. And there will be smart startups who address this issue and the types of problems that we care about in novel ways. And so it's always incumbent on us to be innovating as well. It's something that we prioritize is a kind of continuous pressure testing of our own ideas and thinking about what's next.
49:23Yeah. And we're coming up to an hour, but how do you see this developing? I mean, as you said, five years is too distant horizon, but generally, I mean, the volume of code that's being written, the volume of software, the volume of security vulnerabilities that are being created And the demand for human QA that's being eroded to some extent by people over-relying on AI. You said that you need people who have a deep understanding of security or different aspects of an application. I mean, I imagine that the workforce is changing around these things. As tools like SmartPare become available, it allows teams to sort of balkanize in their expertise so that they can go deeper or not.
50:25I don't know. How do you see it developing? I mean, it's a huge question and it's like, I think, inherently subjective. And the only thing I would say in your summary there, which I generally agree with, is I do not think that there will be an erosion of human talent. I think that we may see that the agent workers fill in the gaps that are already there and the gaps that will be growing very quickly as everything accelerates. And I think that the humans are there. I think their jobs will change. There will be things that they do today that can be automated away. But to the extent that you have expertise, whether that's in quality, security, architecture, operations, there's going to continue to be a demand for those skills.
51:17And kind of in a, you know, we went from assembly to C and everyone thought C was ridiculously high level and you weren't a real programmer. And then we went to Python. It's like, well, you're basically writing a story. you have no idea what's going on under there and now we are literally writing natural language and still most people can't write functioning software where we've answered the question like you literally write down what you want what you want the software to do most people don't know what to do and by the way most people don't have a good idea that they can express well so what we're kind of exposing is that there was a fallacy that software engineering was typing and it's not we are going to need people who understand business problems, business processes, consumer problems, consumer needs, then be able to express those in a way that creates a compelling application.
52:11And the application itself is going to have to be performant, well-architected, secure, private, and all of the things that we've come to recognize. And we're still going to need people giving those instructions. Otherwise, you're going to end up with unbearable convergence to the mean. people or businesses will accept. This episode is brought to you by Tasty Trade. On Eye on AI, we talk a lot about how artificial intelligence is changing how people analyze information, spot patterns, and make more informed decisions. Markets are no different. The edge increasingly comes from having the right tools, the right data, and the ability to understand risk clearly.
52:58That's one of the reasons I like what TastyTrade is building. With TastyTrade, you can trade stocks, options, futures, and crypto all in one platform with low commissions, including zero commissions on stocks and crypto, so you keep more of what you earn. The platform is packed with advanced charting tools, backtesting, strategy selection, and risk analysis tools that help you think in probabilities rather than guesses. They've also introduced an AI-powered search feature that can help you discover symbols aligned with your interests, which is a smart way to explore markets more intentionally.
53:44For active traders, there are tools like Active Trader Mode, One-Click Trading, and Smart Order tracking. And if you're still learning, Tasty Trade offers dozens of free educational courses, plus live support from their trade desk reps during trading hours. If you're serious about trading in a world increasingly shaped by technology, check out Tasty Trade. Visit tastytrade.com to start your trading journey today. I'm going to myself. Tasty Trade Inc. is a registered broker-dealer and member of FINLA, NFA, and SIPC.
From the publisher
What happens when AI writes code faster than anyone can test it?
In this episode of Eye on AI, Craig Smith sits down with Dan Faulkner, CEO of SmartBear, to explore one of the most underappreciated risks of the AI coding boom. As tools like Claude Code and Codex push software development to unprecedented speed, the systems built to validate that software are being left behind. Dan makes a distinction that every engineering leader needs to hear: clean code passing unit tests is not the same as an application that actually works.
Dan introduces the concept of application integrity, continuous and measurable assurance that your software does everything it was intended to do and nothing it was not. He explains why the gap between what AI builds and what teams actually validate is already creating hidden risk in production, and why that risk compounds the faster you ship.
We also get into the new failure modes that agentic AI is introducing. Slop squatting, instruction inversion, cascading errors. These are not theoretical. They are happening now, at scale, in codebases that no human has fully read.
Dan also walks through SmartBear's autonomy ladder framework and their newest product BearQ, a team of AI agents that explores your application, builds a knowledge graph, authors tests, runs them, and updates everything as your app evolves. The key distinction: it is built to augment human teams, not replace them.
Finally, Dan shares his honest take on the future of software engineering. The fallacy was always that coding was the hard part. The hard part is knowing what to build. That skill is not going anywhere.
Subscribe for more conversations with the people shaping the future of AI and emerging technology.
Stay Updated:
Craig Smith on X: https://x.com/craigss
Eye on A.I. on X: https://x.com/EyeOn_AI
(00:00) Introduction and Dan Faulkner's Background
(01:05) What SmartBear Does: Testing and API Lifecycle Management
(03:27) AI Is Outpacing Application Testing
(07:51) Slop Squatting, Instruction Inversion and New AI Failure Modes
(17:31) Black Boxes, Technical Debt and the Expertise Crisis
(22:00) How to Avoid Self-Validating AI Systems
(24:11) The Autonomy Ladder and BearQ
(31:30) Why Testing Must Be Continuous and Everywhere
(36:31) Infrastructure Risk and Automation Bias
(44:11) The Future of QA and New Specialist Roles
(50:44) How Teams Use SmartBear Tools Today
(58:57) The Future of Software Engineering and Human Roles




