In short
Software Engineering Daily - Episode Summary
Episode Title
Amazon’s IDE for Spec-Driven Development with David Yanacek Hosts: Kevin Ball Guest: David Yanacek, Senior Principal Engineer at AWS
Episode Description This episode discusses the challenges of transforming AI-generated prototypes into reliable production-grade systems. It introduces Kiro, an AI-powered Integrated Development Environment (IDE) designed around a spec-driven development workflow that aims to streamline the development process while retaining the creative aspects of AI-assisted coding.
---
Key Concepts and Discussions
- Challenges of AI-Assisted Development
- AI-assisted coding tools facilitate rapid prototyping but struggle with creating production-ready systems.
- Large Language Models (LLMs) tend to be non-deterministic, may drift, and can lose track of developer intent over extended coding sessions.
- Introduction to Kiro
- Kiro is an IDE designed for spec-driven development, intended to help developers:
- Capture intent early in the development process.
- Translate this intent into precise requirements and designs.
- Validate implementations through systematic tasks, testing, and guardrails.
- Spec-Driven Development
- Definition: A structured approach that maintains the agility of coding while ensuring concrete outcomes.
- Components of a Spec:
- Requirements Document: Initial prompt expanded into detailed requirements.
- Design Document: Outlines technology choices, architecture, and implementation strategies.
- Task List: Breaks down implementation into manageable tasks, with both mandatory and optional tasks.
- Kiro's Unique Features
- Flexible Specs: Allows for iterative modifications to requirements and designs throughout the development process.
- Documentation: Specs are stored as markdown files and can be committed to the codebase for future reference.
- Guardrails and Testing: Kiro uses property-based testing to ensure that implementations adhere to specified invariants, enhancing reliability.
- Learning and Adaptation
- Kiro incorporates a learning mechanism where:
- It analyzes past interactions to improve future performance.
- Developers can influence Kiro's learning by providing feedback and clarifying requirements.
- Multi-Agent Coordination
- Kiro leverages multiple autonomous agents capable of performing different tasks (e.g., coding, DevOps).
- Each agent can operate independently but also communicate and coordinate through shared knowledge and goals.
- Future Directions
- The emergence of Frontier Agents: Autonomous agents that manage longer-term tasks and automate operations beyond simple coding.
- Emphasis on continuous learning and adaptation of AI models to keep pace with rapidly evolving technology stacks.
---
Key Takeaways
- AI Coding Tools: While powerful, require a learning curve for effective utilization.
- Iterative Development: Kiro promotes iterative discussions and adjustments to ensure that all aspects of a project align with developer intent.
- Progressive Disclosure: Tools should provide contextual information gradually to avoid cognitive overload and facilitate targeted learning.
- Observability and Learning: Continuous evaluation and feedback loops are essential for improving agent effectiveness and understanding user satisfaction.
Conclusion Kiro represents a significant advancement in AI-assisted software development, combining the creativity of AI with structured methodologies to produce reliable and useful software solutions. As AI and development practices continue to evolve, tools like Kiro will play a crucial role in shaping the future of software engineering.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOChallenges in AI-Assisted Development
0:00 to 0:36
Explore the complexities of turning prototypes into production-grade systems.
“AI-assisted coding tools have made it easier than ever to spin up prototypes, but turning those prototypes into reliable, production-grade systems remains a major challenge.”
David's Journey and Philosophy
1:18 to 3:24
David shares his career journey at Amazon and the philosophy behind his work.
“He co-founded and served as CTO for two companies, founded the San Diego JavaScript Meetup, and organizes the AI in Action discussion group through Latent Space.”
Introduction to Spec-Driven Development
3:24 to 3:58
Learn about the concept of spec-driven development and its significance.
“And so that kind of rinse and repeat with that pattern.”
How Kiro Transforms Development Workflows
3:58 to 9:32
Understand how Kiro streamlines the development process using specifications.
“So I am very grateful for all of the work and like where we are today.”
Managing Specifications and Code Base Integration
9:32 to 11:49
Discover how to manage and integrate project specifications within the code base.
“is essentially what my team does, but we do it by hand in the sense of like, we have a bunch of processes.”
Testing with Property-Based Techniques
11:49 to 14:03
Learn about property-based testing and its application in Kiro for enhanced validation.
“So I tend to use specs as these start to finish and then that's like archive for posterity.”
Understanding Property-Based Testing
14:03 to 16:55
Learn about property-based testing and how it ensures system invariants through thorough test generation.
“thorough test generations using this technique.”
Challenges and Benefits of Thorough Testing
16:55 to 17:46
Explore the advantages of thorough testing and the complexities it introduces in debugging.
“So property-based testing, very powerful way of having thorough tests to make sure that the system is what we agreed up front with the agent of what the implementation should do.”
Kiro's Approach to Spec-Based Workflows
17:46 to 19:11
Discover Kiro's methodology for managing specifications and ensuring clarity in requirements.
“So one of the things I love about your description of the spec-based workflow is Kiro's taking you through this best practice, right?”
Customizing Kiro for Development Environments
20:24 to 25:09
Learn how Kiro allows developers to customize their workflows through steering files and integrated features.
“I think this is a really important thing to get Kiro to understand how you work and your team works and everything.”
Show all 29 chapters
The Power of Hooks in Kiro
25:09 to 28:00
Understand how hooks enhance the functionality of Kiro by automating tasks triggered by specific events.
“help package up expertise so that the agent will be good at everything.”
Exploring Full Agent Loops
28:00 to 28:30
Learn about the concept of full agent loops in software development.
“joins back with the main agent that you spawned it from.”
Multi-Agent Coordination
28:30 to 29:50
Discover how coordination among agents works in an IDE environment.
“Now we're talking about essentially multi-agent patterns, right?”
Introducing Frontier Agents
29:50 to 30:50
Understand the capabilities of new autonomous agents in software development.
“It's just, you assign it, it kind of meets you where you are, is part of your team.”
DevOps Agent Functionality
30:50 to 32:30
Explore the functionalities of the AWS DevOps agent and its role in incident response.
“opportunities to optimize your infrastructure or the way that you even deploy your software.”
Agent Learning and Feedback
32:30 to 33:40
Learn how autonomous agents adapt based on team preferences and feedback.
“So the frontier agent working off of my backlog.”
Understanding Learning Mechanisms
33:40 to 35:30
Dive into the learning mechanisms used by agents for continuous improvement.
“So it realizes, okay, I have too much ambiguity in this case to be able to have that kind of optimism that LLMs so often have about moving things forward.”
Topology and Knowledge Representation
35:30 to 37:50
Discover how agents visualize and represent their learned knowledge.
“interfaces kind of overlays on top of that infrastructure, how your CICD pipelines push to it and how you, you as operators interact with it, how you observe it, or given this part of my application, where are the logs?”
Organizing Knowledge for Agents
37:50 to 40:00
Learn about the techniques for organizing and managing agent knowledge.
“That's like a real rinse and repeat thing.”
Agent Reflection and Improvement
40:00 to 41:40
Explore how agents reflect on their learning to improve future performance.
“There are a lot of agent, like other agent loops, like having agents that are responsible for doing that reflection.”
Evaluating Different Perspectives
41:40 to 42:09
Understand the importance of having multiple perspectives in agent evaluation.
“And it's also important to do that reflection during a task because the nice thing about background agents, like frontier agents, as we call them, is that they have time.”
Exploring Perspectives in LLM Agents
42:09 to 44:51
Learn how different perspectives can enhance the evaluation of LLM agents.
“All these different elements of what we've learned.”
Knowledge Sharing at Amazon
44:52 to 47:14
Discover how Amazon facilitates knowledge sharing among its developers.
“And it's funny, but it's also one of the big challenges of our current LLM coding era is things change so incredibly rapidly.”
Building Effective Learning Systems
47:15 to 49:52
Understand how Amazon develops tools to support model training and learning.
“I can then share that directly with everybody using Kiro by providing a power that packages up all that experience that knows all the power user tricks and tips and everything.”
Evaluating Agent Performance
49:53 to 52:52
Examine the complexities of assessing agent success and user satisfaction.
“Everybody's going to have their own slant on this from their prior expertise.”
Leveraging AI Coding Tools
52:53 to 56:00
Learn how to effectively use AI coding tools to enhance software development.
“I said, one of our goals for this product is it should be delightful.”
Accelerating Coding Practices with Frontier Agents
56:00 to 56:49
Discover how new tools can change coding practices and eliminate bottlenecks.
“Because these tools can accelerate coding so much that they force the need to change the rest of the practices around the coding.”
Innovative Testing Approaches and LLM Integration
56:50 to 57:31
Learn about property-based testing and its integration with language models.
“Well, and I think the stuff we've talked about today, like all these things, Cure was baking in some of the practices that six months ago you had to learn, right?”
The Future of Development Speed and Efficiency
57:32 to 58:08
Explore the potential for increased development speed and ease from new technologies.
“Sounds like Kiro should be able to just generate tests in whatever environment.”
Transcript
Automatic transcript. May contain errors.0:00David Yanacek:AI-assisted coding tools have made it easier than ever to spin up prototypes, but turning those prototypes into reliable, production-grade systems remains a major challenge. Large language models are non-deterministic, prone to drift, and often lose track of intent over long development sessions. Kiro is an AI-powered IDE that's built around a spec-driven development workflow. It's focused on helping developers capture intent up front, translate it into concrete requirements and designs, and systematically validate implementations through tasks, testing, and guardrails. It aims to preserve the creativity of AI-assisted development while producing software that is ready for real-world use.
0:43David Yanacek:David Janacek is a Senior Principal Engineer and a Lead Advisor on the Agentic AI team at AWS. Today, his work focuses on Kiro, Frontier Agents, Amazon Bedrock Agent Corps, and AWS's Operational Agents. He joins the show with Kevin Ball to discuss the design of Kiro, how spec-driven development changes the way teams work with AI coding agents, and what the next generation of agentic software development might look like. Kevin Ball, or KBall, is the Vice President of Engineering at Mento and an independent coach for engineers and engineering leaders. He co-founded and served as CTO for two companies, founded the San Diego JavaScript Meetup, and organizes the AI in Action discussion group through Latent Space.
1:30David Yanacek:Check out the show notes to follow KBall on Twitter or LinkedIn, or visit his website, kball.llc.
1:47Jeff Meyerson:David, welcome to the show. Oh, thanks. Great to be here. Very excited to chat today. Yeah.
1:53David Yanacek:So let's start out with a little bit about you. Can you give me the quick rundown of who you are and how you got to where you are today working on Kuro?
2:01Jeff Meyerson:Sure. I'm a senior principal engineer who has spent my coming up on 20-year career exclusively at Amazon with a singular purpose in mind, and that is to make developers' lives easier. I've been just focused on that. It all comes from, I guess, how we build and operate software here. We do DevOps. And so to us, DevOps means this, of course, takes different meanings depending on where you are and how you use the term. That's how language evolves, of course. But to us, that means developers do the ops. There is no DevOps separate thing. It's just a state of how to do dev more than it is to be a separate thing.
2:39Jeff Meyerson:Anyway, so because of that, that obviously puts a lot of work, responsibility on the shoulders of me, the developer. And so I've been moving from team to team over the years at Amazon, mostly AWS, trying to build the next thing that's going to help life as a developer be easier. So that means I found it tedious on the first team I was on to operate databases, especially when they scale and when they need to be highly available. And on one hand, I didn't like doing database ops because it distracts from the thing that I'm actually trying to do. But on the other hand, I kind of loved it. And so when I heard that we're going to make a highly scalable, highly available database called DynamoDB, I signed up and I was like, OK, I'll help join that and build that.
3:22Jeff Meyerson:So the thing that attracted me to that was that it's going to be large scale and just that I'll never have to do database operations again because it's just managed and everything. And so that kind of rinse and repeat with that pattern. And I worked on Lambda, API Gateway, all the serverless stuff. Operations is a big thing. So I also worked on CloudWatch, which is the observability tool that we use broadly and a lot of people do. So anyway, that's been my whole thing. And so most recently, it's been with the advent of LLMs that has opened a whole new way of making developers' lives easier. And that's what brings us to Kiro.
3:57David Yanacek:Yeah, I have to say, I still remember manually sharding databases and having to bring up systems. So I am very grateful for all of the work and like where we are today.
4:07Jeff Meyerson:I've just like, it's like conceptually easy, but then tedious. Yes. That's one in the morning.
4:13David Yanacek:Which is somehow always when things go down or when you're having to run those migrations because yeah, no, it's such a pain and less of a pain now. So let's talk about Kiro then. Cause like we've been doing a whole sequence. I've been really diving into this for probably the last year and a half of like, how do we use these tools to code? They're incredible tools. They are non-deterministic. They have all these challenges. They're also wrought with a lot of learning that we have to do. So what is Kiro? What is the take and overview of how it works?
4:42Jeff Meyerson:Sure. Kiro is an AI development environment that helps developers go from prototype to production grade code using a technique approach we call spec-driven development. Spec-driven development, it keeps all of the agility and fun and iteration that you get with vibe coding, but adds just enough structure to make it so that you can produce the actual result that you want with where it runs as autonomously as possible. So it's something where it keeps the IDE focused with an end goal in mind and keeps focused on tasks that are in service of that goal rather than having it wander off. So the spec-driven development that's kind of baked into Kiro is what makes it so people can produce production quality code that's thoroughly tested with accurate tests and produce the right thing that they need for production without the frustrations that you can get of like a wandering vibe coding experience.
5:40David Yanacek:So I love this. I mean, I've been long an advocate of things like document-driven development and this sort of thing. So spec can have a lot of different meanings with different levels of formality to it and different levels of restriction. So how do you, within the context of Kuro, define what a spec consists of?
5:57Jeff Meyerson:I'm hearing you kind of say maybe also, well, specs can be, they can be overly formal. When I've done kind of documentation driven, like they can be potentially a little bit kind of in the way, but with Kiro, I find it actually doesn't. It actually speeds me up and helps me. I found I just like it and it's adapts to whatever your way of thinking and coding is. So a spec to Kiro is just, it's three parts, but you just start with pretty much the same prompt as you would have otherwise. You say, Hey, I want to build this thing like in here and let me describe this thing for you and let's build it from then it expands on that let's say okay well you said you want to make a traffic light control system so i think that means that you're going to want it to keep track of when cars enter and exit the intersection it kind of expands on your prompt and produces more detailed requirements but then you you can you read this doc it's a markdown file and you read it and say yeah these requirements that's in this ears format, which just has some more like doubt and shall kind of wording, but it's just very clear then.
7:00Jeff Meyerson:It's easy to read and skim to say, okay, this is what I want or not. And then you can chat and say, okay, no, actually I don't want it to do that thing. I want it to do this other thing. And you just chat with it. You can either modify the doc or you can just go back and forth, just like you would any kind of chat based LLM interaction. So you chat with it about the requirements to agree on what we're going to build, whether it's something new or a feature, Like a spec can be a whole new project. It can be a feature. It can be an upgrade. I'll say, hey, let's go upgrade this code to the latest Node.js version because that's some maintenance tasks that I need to do.
7:33Jeff Meyerson:So it'll come up with requirements and that includes acceptance criteria to say, well, okay, I'm going to make sure that in the end that these are the properties that the system needs to have. That's kind of all part of the requirements. So requirements doc from there, once you say, yep, this is what I want, it produces a design, which is another markdown file that includes all of the technology framework choices that it's going to have all the kind of non-functional some more of non-functional requirements if it wasn't requirements doc class diagrams architecture diagrams the little snippets of code to say here's how i plan on doing this or that so then i can skim that and see if i agree with that approach it's nice that i get to see that those kind of code snippets are nice because i might otherwise be like 20 minutes into this project and see, oh, I don't want you to go about it that way.
8:22Jeff Meyerson:I want to do it this other way. And then you have to throw all that work away. So the design kind of helps me make sure that it's going to take an approach that fits with my mental model for how this code base should be evolving or start from. And then once I agree with that, it breaks the project into tasks. And the task, it's actually a separate markdown file, a third markdown file that just says, okay, let's build the framework. Let's put here. Then now we'll do some infrastructure as code to set up some stuff. Let's set up a test environment. Now here's, it just breaks down the implementation.
8:54Jeff Meyerson:Maybe let's create a database schema. Okay. Now let's create it. It just keeps going. And then it even splits out optional tasks, which are nicer with tasks that you can follow up with later. Let's say, okay, then let's, it focuses on, let's get something tangible for you to see working. And then let's go do the thorough, all the tests, including property-based testing. Essentially, a spec is just these three things that you're just chatting about and seeing what Kiro is going to be doing, plans on doing, and how it plans on going about it. And I really just like doing that all upfront because then I can just say, okay, go do the tasks.
9:29David Yanacek:Well, this is fascinating to me because the process you've just described is one that is essentially what my team does, but we do it by hand in the sense of like, we have a bunch of processes. It was like, okay, go and analyze what's going on or let's write a spec together. Let's iterate kind of going through this. And it sounds like Kiro bakes that into the flow as you go. Now, are those markdown documents, are they committed as a part of your code base? Like where do those live? How are they managed over time?
9:57Jeff Meyerson:Yeah, ultimately it's up to you, but the intent, I guess, and what I like to do and what I see everybody do is commit them to the code base. these are they're actually pretty flexible into how you use them in the future like one approach is that you use the specs to add a feature start a project and they're kind of one shot then you're done now i should actually add that during this process i actually go back and forth a lot it's not just a waterfall it's like let's do the requirements and then the design then the tasks and like sometimes i'll go back and forth i'll say when i see the design i said oh this design kind of don't agree with because it's actually missing a couple of requirements.
10:36Jeff Meyerson:Like I forgot to mention that I actually want to use this framework or something, or that I actually, yeah, well, I don't want users to be able to do that or something. So it's nice to be able to go back and forth. Even once it starts implementing, I can go back in terms of the spec or the tasks or like any part of the spec. And so this is also true with when you talk about where do I commit the spec? You can treat them as these one-off things where, okay, now that implementation task is done, I'll put it there for posterity so that later on, if the agent or I have questions on how we got here, I can see that or ask questions about it and the agent can read the specs that exist for that project, consult them when it deems necessary or when I ask it to.
11:16Jeff Meyerson:Some people like to keep a spec to be up to date with the project where they say, okay, this spec represents the overall architecture. Once I make a change, I want to go update some spec, but that's definitely a fine way to do it. Sometimes you might actually just say, hey, I'm going to have a design document also checked in that I can ask Kiro to say, update my overall design that's like separate from a project specific spec. Just once you're done, now update my authoritative design once we're done and just keep it up to date with what we just implemented there. So that's another nice way to use it.
11:51Jeff Meyerson:So I tend to use specs as these start to finish and then that's like archive for posterity. And then I'll use a separate document to keep that overall idea of the current state of the architecture and the system and how it's built.
12:06David Yanacek:That makes sense. So let's now move into something you kind of mentioned a little bit, which is this idea of, I think you said property-based tests or things like that. I think at this point, pretty much everybody has experience using these tools. And depending on which ones you use, they are better or worse at actually sticking to the guidance, the spec that we've agreed on things like this. And so how do you think about building in layers of validations, guardrails, and actually validating those requirements?
12:35Jeff Meyerson:Well, I think it's really important. That's actually what the spec is super great at, is that it captures the actual intent of what you're trying to get done, and whether or not it has done that yet. So that's just the breakdown of tasks. Okay, I haven't written these tests yet. Okay. So that's just one to make sure it actually does the things that you wanted it to do. But in terms of the quality of the tests, a nice thing about capturing the intent in the spec of the requirements and the design is that that gives a bunch of extra information that you wouldn't have otherwise gotten by just doing some prompt that you mentioned like 20 minutes ago that's now kind of floated off.
13:13Jeff Meyerson:To me, when I'm doing coding using an AI agent, that's where whenever I type something into like steering or anywhere that's the value that I'm adding and so that I want to make sure that that's saved and consulted and that's where it's really nice to have that all summarized into the spec where I can see that kind of work that I've been doing with the agent and so from that because it describes the actual intent of what the system is supposed to do and how it's supposed to work we can write more thorough tests or Kuro can write more thorough tests than if it were just looking at some code and saying, okay, I need to write some unit tests now or some integration tests.
13:52Jeff Meyerson:Cura uses something called property-based testing. It's not a Cura-specific thing. It's something that's been around in the industry for some time, but it's something that we realized that by having this spec, we could actually do thorough test generations using this technique. Property-based testing tests invariants, and then it writes tests to make sure that all those invariants are held during any kind of sequence of input. Rather than saying, writing a test for a specific scenario, this tries to generate many scenarios and make sure that those invariants hold true. So let's go back to that traffic example.
14:26Jeff Meyerson:Let's say you're building a traffic light system. A really important invariant is that at most one direction has a green light at a time, right? It's obviously very important for a system to never have more than one green light. It's okay for it to have no green lights, but to have the most one. And so that's, if I think about how to test such a system, I want to make sure that every sequence of things where somebody hits the walk button or emergency vehicle goes through, that it's always like all these different things can happen. These are the inputs to the system can happen, different timings, power outages, power restores.
15:03Jeff Meyerson:And I want to make sure that it always holds. And so property-based testing is where you describe a test in terms of these invariants, which can be generated directly from the spec. If the spec says that, hey, we want to make sure there's only one green light. Okay. So now we can generate property-based tests based on that. So property-based testing uses different frameworks that help drive these different inputs. So you have the test that has the invariant. Okay. How do you drive a bunch of input at it? Well, that's where a input generator comes into play. We use one called hypothesis. It's a Python framework for property-based testing.
15:35Jeff Meyerson:And that just generate from the spec that describes the types of input that you want to a part of your system. It generates a bunch of permutations of that and then feeds that into the test or into all the other tests that might also want to have that input and test different to make sure that different invariants are upheld. And then when a test fails, ultimately, then the agent can go and test its implementation against these and then keep adjusting the implementation until those and variance hold during all the tests. And one nice thing in property-based testing is they're very thorough with all the permutations of input, but that thoroughness can make it tricky to understand why a test failed because it's like, okay, it's not just, oh, it's straightforward as a unit test where you fed it this specific input and then this assertion failed, right?
16:24Jeff Meyerson:That's very easy to understand. It's okay, let's run it again with that input. With property-based testing, you had like a whole series of inputs in a sequence. And so to replay all of those can be a little confusing to see which of those triggered it. So property-based test frameworks use a technique called shrinking to take those different permutations of inputs and it finds the large sequence that caused the failure. And then it just kind of tries to remove those states until it arrives at the most compact explanation for why the code isn't holding the invariance to be true. So property-based testing, very powerful way of having thorough tests to make sure that the system is what we agreed up front with the agent of what the implementation should do.
17:05Jeff Meyerson:Because it just tests so many boundary cases rather than having to just fish for one at a time. Because I've seen agents just will, for those listening, you're smiling here because you know what I'm about to say. You'll see an agent kind of, I can't get the test to work, so I just comped it out the body of it. So now it passes. That's great. Okay, let's move on. And then it forgets that it never did that. Having these thorough tests make it so that it keeps the agent honest where it has to prove its correctness.
17:32David Yanacek:Yeah, there's multiple things. Agents tend to be, or LLMs, I guess, tend to be confirmation biasing and self-confirmation biasing. So they'll try a thing and they'll get into this rut. And by mapping the space out with the full property range, you can help them from getting in that rut. So one of the things I love about your description of the spec-based workflow is Kiro's taking you through this best practice, right? Like I've seen everyone's trying to figure out their set of things and Kuro's like, here's our approach. We're going to be opinionated. Let's go. Does it do the same thing there for spec?
18:00David Yanacek:So it's like, okay, I've done my implementation. Now it's time to do property based testing. Here we go. Or are you prompting it to do that?
18:05Jeff Meyerson:It actually thinks about it. It says, okay, here are the properties that are going to be verifying in the end. But you just decide, yeah, no, it actually has a step where it's coming up with correctness properties. And whether or not I have it actually implement those tests right now is the optional part. Yeah, it's kind of also reflecting on whether during that spec flow, it's saying, do I actually have enough clarity from you? This is actually an interesting thing that Curo is doing to decide how much more it needs from you. I mean, you can obviously weigh in at any time with it, but it kind of checks, it reflects on, are these requirements clear or conflicting?
18:45Jeff Meyerson:So that's kind of an interesting behind the scenes that it's doing, but it's also doing that around what properties then should I test? What are the key tests to make sure that these requirements are upheld? So it's just doing that reflecting during that spec flow. So that's great.
18:59David Yanacek:I bet that leads to substantially better specs. But then that also means, hey, whatever environment I'm in, I just need to find a test framework that can validate these correctness criteria. Right.
19:11Jeff Meyerson:It's pretty powerful.
19:12David Yanacek:In mobile application security, good enough is a risk. GuardSquare uses advanced, multilayered code hardening techniques and automated runtime application self-protection and mobile application security testing. combined with real-time threat monitoring to deliver the highest level of mobile app security. Discover how GuardSquare brings all these together to provide mobile app security for your Android and iOS apps without compromise at www.guardsquare.com. Why is there always a meeting bot in your Zoom call? Blame Recall.ai. Recall.ai powers the meeting bots and desktop recording apps behind products like Cluely, HubSpot, and ClickUp.
19:56David Yanacek:They handle the hard infrastructure work, capturing clean recordings, transcripts, and metadata across Zoom, Google Meet, Microsoft Teams, in-person meetings, and more, so developers don't have to build it themselves. If you're building a meeting note-taker or anything involving conversation data, Recall.ai is the API for meeting recording. Get started today with$100 in free credits at recall.ai slash software. So feeding into this concept, we've talked about what Kiro is doing in an opinionated way, but what mechanisms, hooks, skills, other form factors does Kiro offer for developers to customize it to their environments, to their particular preferences, to their team practices, etc.?
Read the full transcript
20:42Jeff Meyerson:I think this is a really important thing to get Kiro to understand how you work and your team works and everything. And there are essentially three features in Kiro that help with this. First, I can write a steering file or a series of steering files where that kind of just describes my development environment. Maybe I say, hey, I'm always using this. My team is always using this for continuous deployment. We use these frameworks. Just kind of set that, the stuff to always be keeping in mind. And these are potentially, you can have many steering files. They do take up your context window. So you want to keep them relatively small with pointers to where to go for more information about a particular thing.
21:21Jeff Meyerson:And that's really useful because by having multiple of them, you can have maybe a company wide one if you have many teams that try to do certain things the same way, but then you can have your own for your own team to do things the way that your team does it that's maybe different than other teams. And then your own steering file that says, here are my own preferences. So just the steering files help it stay focused and be able to do things the way that you are used to having them without having to repeat yourself all the time. Then there are powers, which Kiro powers is a feature we added relatively recently that just kind of bundles up MCP servers with steering and with hooks, which I'll describe in a second.
22:01But these powers are things that are loaded dynamically depending on what you're doing.
22:07Jeff Meyerson:So Supabase, for example, provided a Kiro power on Supabase. So if I'm using, if I say that I'm using Supabase or my project clearly is, Kiro will load the Supabase power, Kiro power. And it suddenly shows up with these MCP servers and steering files and hooks that would help it when it comes to using that platform.
22:28David Yanacek:That's really interesting. I want to dig into that. because like one of the big things I think people are trying to grapple with right now is what I've been calling progressive disclosure of context, right? It's like, I don't want everything in my context window upfront. One of the challenges with MCP servers is it's hard to do, like they've got a whole bunch of different stuff. And so skills were sort of a step into like, okay, we'll give you a little description and then you can load more if you want it. But this sounds like potentially even more powerful. So how do powers work? Like what are the knobs and levers I have to say, okay, these are the situations in which this context is going to be relevant.
23:02Jeff Meyerson:Yeah, to use a power, it's relatively simple. I mean, on the surface, you go to kiro.dev and list the powers that the website has is just a starting point and say, add it. And then it'll load up the ID and download that whole bundle of stuff. The Kiro powers, I think that you'll find that they're pretty similar to what you're describing with skills of that they kind of have a smaller amount of data that will entice the agent to use it in certain cases, is just enough to pique its interest at the right time. And then when it is time to do a thing using that, it'll load up the MCP servers and steering files and everything for that part of the task.
23:40Jeff Meyerson:I'd say it's not magic, but it is extremely convenient. And so keeping the, I'd say keeping the context window small, it's convenient for that, but it's really convenient for just bringing in the expertise around a particular technology when it's time for that. But I found these features that we build into Kiro, we're building from our own experience and other customers' experience. It's the nice thing about building tools for developers is that we are also developers. And so it's a thing where it's a little more intuitive to imagine what other customers might want. But just as we were, that's kind of just where Kiro came from in the first place, actually, is we were using LLM-based tools to do development and we found that, okay, it would wander off.
24:23Jeff Meyerson:And so we were like, well, let's build specs. or you'd want to do something and it wouldn't know how to do that. Say, hey, I want to use this new Bedrock agent core feature to build a new agent or use strands. It's a framework that we created for open source framework for making agents. I want to use that. Well, at the day of launch of that new service or feature, the agent has no idea what that is. And I can either go and give it a bunch of links to say, hey, go read this documentation. Read this documentation. This is what I'm talking about. Here's how to find it. it points it in the right direction and saves a bunch of back and forth of like, no, I meant this, I meant this, like the power just keeps it focused.
25:00Jeff Meyerson:So we just from like this experience of having to repeat ourselves to the agent say, hey, no, this is what I'm trying to use right now. And here's how to find out more about it. We found that this powers concept would just help package up expertise so that the agent will be good at everything.
25:16David Yanacek:Absolutely. Well, and I think that is what that kind of progressive disclosure allows you to do, is you can say, there's all these things that are going to be relevant at some point. Let me give you access, but only when you actually need it. Let's talk a little bit about hooks, because that was another thing that I saw in Kiro that seemed like it was potentially very powerful.
25:36Jeff Meyerson:Oh yeah, hooks. It's actually another part of this packaging of Kiro power, but it also makes them on their own. A hook is something that will run a prompt or another spin-off and agentic loop in reaction to a thing. So if I say, let's say I have an API, some kind of web service API, and I want to, every time I update my API definition, I might want to generate some things off of that. Just like specs, you can generate things like property-based tests off the specs, API definitions. Also, you can generate a lot of interesting things like SDKs, API documentation, really nice. And so maybe that's a nice time when I save my API.
26:15Jeff Meyerson:If I may ever make a change to my API definition, and I want to go do these things as a result. So I would write that as a hook. And it's pretty simple to make them, actually. You just write a prompt that says, update my API documentation. Simple, like just really that's it. And you would say trigger on whenever this file is saved or changed. So the different triggers that you can kick off these different hooks, like maybe run this code scanner, run this dependency. If every time I update my dependencies file, whatever framework I'm using, I want to check for security vulnerabilities or something using this tool or out of date if there are more recent versions available that I could grab.
26:53Jeff Meyerson:So it's a nice way to just do those things that I need to remember to do or that I'm just making it a little more convenient. You can also manually trigger the hooks. I find that's actually how I mostly do it, just personally. I just sometimes it's not exactly when a file gets saved that I want to do a thing. I just easily, oh yeah, I click this button and it's going to go off and do that thing for me.
27:13David Yanacek:Now, does that, whatever it does, so say you have an agent running on a thing and it's kind of going, it's got its own context window, it's writing things and it touches one of these files that has a hook attached and the hook runs. Does the output of that hook, well, first of all, if it's a prompt snippet that's getting injected, does that run in the main agent context window? It's a separate context window.
27:33Jeff Meyerson:Yeah, it pops out into a separate context window, yeah.
27:35David Yanacek:Okay, awesome. And then does anything from that get fed back into the original context window so that you could create like one of the things that I think is emerging is a pattern with a lot of these is you want to create feedback loops for your agent so that it can self-correct and linters and tests and all these things give these opportunities. So does that get kind of piped back in some way or can it?
27:54Jeff Meyerson:I don't think it does. I'm pretty sure these are one-off independent tasks that just fork. I haven't done it that way, so I don't think it joins back with the main agent that you spawned it from. I think it could be wrong, but it's a neat idea.
28:06David Yanacek:Does it have a full agent loop or is it like a single inference, do a thing and come back?
28:11Jeff Meyerson:No, it's a full agent loop. It's going to just keep working on a task just like any other task.
28:16David Yanacek:So it doesn't necessarily have to go back to the core one because it could go and just do the fix itself.
28:21Jeff Meyerson:That's right. That's right. Yeah. They just all branch off and go kick off these other tasks. So I think hooks are an interesting kind of nice convenience for remembering to go and do other things.
28:31David Yanacek:That kind of pulls us into a world. Now we're talking about essentially multi-agent patterns, right? Because a hook at its core, it sounds like, is a small agent or maybe not so small agent. It could be a large agent. Who knows? So how do you and Kiro think about coordination across those different agents, making sure they don't stomp on each other, et cetera?
28:49Jeff Meyerson:I think an interesting place that I see more of the agent coordination. So when you're in an IDE, that agent coordination, you're seeing them do the things. And so it's not a huge amount of cognitive load to see what one is doing and make sure that they're not overlapping or anything. When you get into this other world of agents that aren't running right in front of you, that this coordination becomes very interesting. This world is actually already here. Recently, we launched a set of what we call frontier agents, which are, they connect into Cura in this interesting way. So with Cura, we found that with Spectre and Development, it could run for longer.
29:31Jeff Meyerson:It could run independently. You can give it a larger, more ambiguous task and have it go do that without having to like pester you. We decided to push that as far as we could into a set of new software development agents that we call frontier agents. One is we call Kiro Autonomous Agent. And so this is, it's a non-IDE agent. It's just, you assign it, it kind of meets you where you are, is part of your team. And you would say, if I have like a, whatever I'm using for my backlog for my team, you assign it a task and it'll go do that task and do its own loop and test and test and refine and refine and test and understand, expand on your kind of more ambiguous tasks that you gave it to make sure it fits with your team's working patterns and implement that for you and then produce a code review, like a pull request.
30:21Jeff Meyerson:Here is the results. Do you want to merge this? The second of these frontier agents is a DevOps agent. So we noticed that when you code, you also need to run that code. That's like what I was saying with DevOps. So we've kind of encoded that into an autonomous agent that we call AWS DevOps agent. And so this does incident response. It'll triage and root cause issues and recommend how to fix these issues in production. Over time, it'll actually look at your whole environment, your whole setup to say, well, actually, I found these opportunities to optimize your infrastructure or the way that you even deploy your software.
30:57Jeff Meyerson:Say, hey, this alarm keeps going off. I noticed you're getting this alarm that goes off because you're doing bad deployments all the time. Let's update your CICD pipeline to add better tests and automatic rollback and better alarms because maybe alarms aren't even catching this early enough. It just tries to prevent future issues by just working all the time in the background to look for opportunities on what to improve. And then we have a security agent. It'll make sure that the code that you and agents write adhere to security standards, and it'll even do penetration testing on its own. So the coordination, you asked about coordination, I think that really comes into play across these agents where they're not running in front of your face.
31:37Jeff Meyerson:These are running kind of all the time. Ideally, they run when you're sleeping or something, right? So getting deeper into the backlog than you would. And so the coordination between these tends to actually be the coordination that you use across your team. Like we make it So that these agents are where you are and your team is. So these agents, you interact with them in Slack or whatever team communication tool you're using in whatever backlog tool you're using, whether that's Jira or whatever. And in the case of the DevOps agent and whatever incident response tool you're using, like ServiceNow or something, whatever observability tool you're using, Dynatrace or Datadog.
32:17Jeff Meyerson:So these are all about, when it comes to coordinating agents, we find it's best to do that coordination where you're already doing coordination with your teammates.
32:26David Yanacek:So that raises some interesting questions for me. So let's take one of those examples. So the frontier agent working off of my backlog. So say it's in JIRA or linear or something like that, it sees a ticket. Does it then follow the Curo process of, okay, I'm going to write a spec? Do I then as a human have to take a look at that spec? or like how does the in the loop piece of this happen or does it, is it completely autonomous and then it's gonna come back and I'm gonna look at it only when it's got a complete working set of code that may or may not match my intent if I had a very poorly specified ticket.
33:00Jeff Meyerson:Right, I mean, it has to have that judgment and learn. Like, so these agents learn about what are your team preferences are and how they work. So they learn based on your feedback. So you might realize, just like I mentioned in the Kuro spec flow, The Curo is asking itself whether or not it has enough information to have a well-formed spec or whether it has conflicting instructions that it needs your help resolving. Similarly, with property-based testing, when it's generating tests and sees a test failure, if it thinks, if Curo, with the kind of reasoning that we've given it, if it thinks it has the right implementation, but it also thinks it has the right kind of requirement, but yet the property-based test is failing, it needs help resolving that.
33:42Jeff Meyerson:So it realizes, okay, I have too much ambiguity in this case to be able to have that kind of optimism that LLMs so often have about moving things forward. So similarly with the autonomous agent, a big part of them is realizing that they have a ticket assigned to them that maybe has some instructions in it that conflict with team practices that it's already learned. This ticket says to use this logging library, but the team is using this other logging library. Like you might need to ask for clarification before it continues. And so that's a big part of it is deciding when to re-engage somebody before just burning a bunch of cycles, doing something that is the best guess.
34:20David Yanacek:Can we peel back the cover a little bit and talk about how that learning works? A couple of different things that I'm curious about. So one is like pure guts implementation, right? Everybody's trying to solve. This is like the frontier problem right now with LLM tooling is how do you make it kind of continuously learning rather than train once and go. And so like, I'm just, there's various approaches floating, I would love to know, at least at a high level, like what approach you all are taking to it. And then I guess the other thing that I'm always fascinated with around this is like, how is that made legible to humans, because LMs, as we know, get things wrong.
34:57David Yanacek:And so like, how do you take whatever mechanism that you're using to derive, like, these are the practices we use, or what have you, like, what form is that then bubbled back up to people to say, you know what, that you learned that wrong, like, that's not correct or, yeah, this is great. Can we expand on this or what have you?
35:13Jeff Meyerson:Great. I'll give you a couple of examples about how these agents do their learning and how you see it as a customer of them, a user of them. One is in the DevOps agent, the AWS DevOps agent. It needs to understand what we call topology, your whole system in all of your test environments, your production environments, what your infrastructure is, how your code and interfaces kind of overlays on top of that infrastructure, how your CICD pipelines push to it and how you, you as operators interact with it, how you observe it, or given this part of my application, where are the logs? What provider do I use to keep the traces, whatever.
35:51Jeff Meyerson:And so that topology is a thing that you can see when you create AWS DevOps agent, we call agent space, visualize that topology for you. So you can see here's the universe that we have learned about and discovered so far. I kind of mentioned this, that part of the DevOps agent runs all the time looking for things to improve. And while it's doing that, it's also discovering more and telling you more about, hey, I was trying to figure out something. I was trying to access this part of, figure out where your logs are for this thing and I can't find it. You've told me I should be able to access this, but I can't.
36:26Jeff Meyerson:And so it kind of is producing these things that kind of got in the way of it learning more. It's like, hey, you might want to resolve this. And so as a result of that, we're asking, how does somebody correct that or do that? Well, they can either, if we are right, and then we found a misconfiguration, they can reconfigure it. Or if we were looking in the wrong place, they can write in the case of DevOps agent, we call runbook, which just it's essentially a set of steering files that are loaded kind of in the way that like progressively loaded. So they include like short names and short descriptions that can entice the agent on when to look at them and when to consult with them.
37:00Jeff Meyerson:But these runbooks can help it say, okay, no, this is actually the observability tool that I'm using for this part of my system. You shouldn't be finding logs there anyway. You should be finding logs in this other place instead. So that's one way that the agents learn is the topology in the DevOps agent case. And you can see that visually in the application. The other one I guess I could talk about is a similar agent that we have called AWS Transform. AWS Transform, especially this custom transforms, it's a service that makes it so that you can, it's an agent that helps you do a longer kind of upgrade or transformation project, or maybe a repetitive one.
37:40Jeff Meyerson:Let's say you have the same kind of code transformation or upgrade or migration that you need to do, like rinse and repeat a lot of times. Maybe you're trying to replatform or move to a new framework or upgrade a new language version. That's like a real rinse and repeat thing. And sure, you could give that to Kiro and have it do it every time and you give it some content, the best practices and everything on how to do that. Give it a Kiro power specifically for it. Definitely works. But we've made transform to be where it learns every time it does a transform of a particular kind and kind of writes down, okay, tips on, oh, this worked really well.
38:16Jeff Meyerson:This didn't work. And so you can write your own kind of custom transform that's just for you. and it learns just, it's not learning off of others. It's learning off of every time you run the transform, your custom transform. And then the way that it shows you what it's learned is it just shows you a bunch of essentially learning document, knowledge items, it calls them, which are just essentially markdown files. So it's showing you what you learn. And then you say, yes, that is correct or no, that is not correct. So you're accepting or not accepting those knowledge items. So that's kind of how that's shown to you and how you can control whether or not it's learning useful stuff.
38:54David Yanacek:And then digging into implementation pieces then, are these markdown files that are globally available in the context for this agent? Are they exposed progressively like powers are? How do you think about as you accumulate learnings over time? And the transform sounds like it's honestly very simple. This will have one type of transformation it's doing in learning, but maybe not. Maybe it's got a whole bunch of different things. or if we're talking about like a just frontier coding agent that's learning all of my team's practices like it could accumulate a lot of documents over time so how do you control which the agent is looking at when or like you had this great example that i loved of like oh you asked me to use this logging library but i know that my team uses a different logging library like that's a detail probably one of many details about coding best practices on this team How did it find the right one?
39:46Jeff Meyerson:It's certainly not a load it all up every time into the context when everything we've learned, you know, progressive disclosure or resurrection of knowledge is super important. Aging out of knowledge is important. It's a mix of so many techniques and around like from rag to simple files that are with like summaries. There are a lot of agent, like other agent loops, like having agents that are responsible for doing that reflection. I was describing in the DevOps agent, for example, how it's actually going back every day and looking at the last weeks of issues that it investigated. And it's just reflecting on those.
40:24Jeff Meyerson:It's reflecting on those for things that maybe are like the way that you see it most as a user of it is you're seeing recommendations on how to prevent this same recurring issue from happening or patterns of issues. But it's also going and reflecting on whether or not it took the right path in that investigation. Like, where did it waste time in the investigation? The trick is sometimes it's good to waste that time. Like, you don't know. Like, next time it might be a different issue. And if you don't explore that branch of an incident response, like, then you miss out on the thing that actually was the problem this time.
40:59Jeff Meyerson:So I guess to your question of how do we organize knowledge, it's a mix of techniques. We're always learning ourselves on the best way to do it. So it's moving so fast that by the time I'm done with this sentence, we'll probably have a slightly different take on it. But it involves a lot of agents behind the scenes. Ultimately, agents with jobs of learning are pretty good at then reflecting on what might be useful in the future and coming up with the right applicability. Yeah.
41:26David Yanacek:No, and I love that example of just like having a set of things that are, their job is to look back over particular time periods, maybe look at particular types of aggregates, what have you, infer what they can, and then expose it or make it useful when need be.
41:41Jeff Meyerson:And it's also important to do that reflection during a task because the nice thing about background agents, like frontier agents, as we call them, is that they have time. Now, during an incident response, we don't have the ideas to get that figured out as quickly as possible. But when working on a coding task, we have time. And so we can have a bunch of just introspection points in the agent where if it's working on something, it can say, okay, let's go look at, let's examine one at a time. We're in parallel, actually. All these different elements of what we've learned. And so that's a neat observation is, sure, we don't have time to go evaluate every knowledge item we've accumulated, but we can kind of really ask ourselves from a different subject areas, whether we've considered everything.
42:26Jeff Meyerson:I guess having the background agent learn, but also having the agent that's doing a thing, kind of know what questions it should be asking of the knowledge and when. And when it has more time to do that, it has more time to reflect on learners.
42:40David Yanacek:Yeah, I like that. Can you share some of the different perspectives you might take? Because I think that is another very interesting thing with these LLM agents. It's like you can have like, oh, I have five different lenses on which I wish you to evaluate this thing, and then we'll come back and compare.
42:56Jeff Meyerson:There's so many different dimensions on this. One that we do make sure that we're always reflecting on is HBS has this thing we call a well-architected framework to make sure that we're building things that follow the learnings over time of what's a good way to build on AWS or just build systems and operate them in general. So we kind of reflect on security, for example. Am I opening up anything that we shouldn't be opening up or following the best practices that you've established as a company? Or even things like reflecting on whether or not we're following the testing practices, the observability practices.
43:31Jeff Meyerson:Just everything you can kind of imagine of, particularly if somebody says that something is important. So if they say that this is important to them, like make sure you're doing this, like then we should reflect on that to make sure we're doing that thing.
43:46David Yanacek:And are there ways for teams, for example, to set up one of these agents to say like, okay, here's a perspective I always care about. And in fact, I'm going to give you some additional resources or things particularly for that perspective.
43:57Jeff Meyerson:Yeah, exactly. Like connecting knowledge bases, steering files, even writing powers. One thing we've done within Amazon in our use of Kiro, but we're back to the IDE at this point, is we have a steering file that gets kind of bundled up. We have a separate VS Code plugin that connects to our internal build system anyway. And so we've had that for decades in the different generations of IDEs over the years. but when it's installed in Cura, we'll install a default steering file. We've just said, okay, this is something that we always need to have. So make sure you're always following these things.
44:32Jeff Meyerson:Again, unless a team has overridden it with their own kind of little bit separate way of going about it or whatever. One thing that you kind of said in passing, but I think is actually interesting
44:43David Yanacek:to dig into, you mentioned, oh, the best practices for how to do learning, they're changing. By the a time the sentence has changed or you're finished, they may have changed. And it's funny, but it's also one of the big challenges of our current LLM coding era is things change so incredibly rapidly. It's hard to keep up with what's going on. So I'm curious, is there anything you're doing either that Kira was doing or that you found effective at Amazon for helping us, like limited human brains, keep up with the changes that are happening in our software stacks?
45:16Jeff Meyerson:A lot of what we do is educating each other, sharing what we do. We have long running just ways that we propagate and share in practices and stories. We do a lot of storytelling of when I was building this thing and solving this problem, here were the important things that I learned and maybe you should too. So a lot of these internal talk series, we do a lot of them externally too at our learning developer conferences like reInvent, do a lot of describing how we build things and how others can learn from what we found that works well or doesn't work well. A ton of just sharing information with each other in different talk series that we make sure we advertise and market well.
45:54Jeff Meyerson:And people learn different ways. Some people like talk, some people like reading, some people like podcasts. And so we try to do a lot of different mediums to do that. One thing, kind of pre-AI era, the most recent AI era of generative AI, we released one thing we call the Amazon Builder's Library. It's a set of long form articles that describe how we at Amazon do certain things when it comes to writing software, operating it or designing, scaling big, large distributed systems and operating them at scale. And so just sharing knowledge like that, getting into the nitty gritty details on like, what's the right way to implement a health check when you have a server behind a load balancer?
46:36Jeff Meyerson:Something that seems super easy, like, oh, yeah, just respond healthy when you're healthy. Okay, but what does that mean? And what are the downsides? I actually, for that one, that's one of the articles I wrote for the Amazon Builders Library. It's a long article because when you health check is and something responding to that health check is an automatic system that's going to can have surprising behavior when it's running unattended. And anyway, so just to your question of how do we keep up with all of the changes in AI? We talk about it a lot. We talk about what we do that works. We try to do that externally for everybody so that you all can learn what we learn.
47:12Jeff Meyerson:And so we can learn from you all. We try to just participate in community of practice about it. And then we try to think of just kind of shortcuts for it around things like hero powers that are going to package up everything that somebody has learned and an opinion that they, especially when a company or a tool owner or a framework owner, platform owner has an opinion about here are the best ways to use my framework or platform to be successful. I can then share that directly with everybody using Kiro by providing a power that packages up all that experience that knows all the power user tricks and tips and everything.
47:52Jeff Meyerson:So just ways to share knowledge and then ways to share packaging of tools. You find that those are the best we can do so far. And then around the learning systems and training in general, in particular, I mean, this is part of why we've been building things like Amazon Nova Forge, which is a service that helps you do model development from early model checkpoints. So it's just, there are so many services that we've been launching recently that help you do that. Like when it comes to different techniques for whether it's building a RAG or knowledge base or actually training models. We do not know that that's super important for agents.
48:31Jeff Meyerson:And so we're doing everything we can to provide services that you can use to just get started with something that learns like agent core memory. So Bedrock Agent Core is a service that helps you run right agents. It's something that we use internally. AWS DevOps agent uses AgentCore because it's a handy way to build and operate and scale agents in a secure way. Isolation between tenants and everything. And so one of the features of AgentCore is memory. And so it will look at the traces of all the agent and tool interactions and has different strategies that are available out of the box to just compress that knowledge.
49:09Jeff Meyerson:to say, okay, here's what I should store for later for the session wise of just how do I reflect on a particular agent run? And then how do I promote stuff to long-term memory where I can have long-term lessons distilled and available afterwards? So agent core memory is one place that we've been trying to build that, like how do agents learn? I know your question was more, how do humans learn about how to use these things, but agents need to learn too. So we need to be able to be giving as many building blocks as we can. So that basically that complex task of training and learning can at least give everybody a headstart who hasn't tried it before.
49:50David Yanacek:Kind of on that note, you're trying to push forward this frontier and make it easier for folks. What do you see, like where the frontier is in these agentic coding tools and building agents and this whole space and what's coming over the next, I don't know how far we can project now, but three months, six months, something like that.
50:08Jeff Meyerson:I guess it's funny. Everybody's going to have their own slant on this from their prior expertise. But one thing that I see of challenges that kind of fit with my own background, it's tricky to tell whether agents are doing things successfully. How effective is an agent at the thing that you've built the agent to do? It's something we sink a lot of time into when we build our own kind of agents. And it's something that we've been trying to build primitives to make even easier. Like we've been trying to make observability and evaluation of agent trajectories were automatic and easier. It's funny how some techniques that I'd say when it comes to agentic applications, some of the techniques that were maybe not as exciting because they're more subjective for say like website, are your website users happy?
50:56Jeff Meyerson:Actually kind of tricky, a little more annoying than measuring is my API returning success or failure? Okay, like pretty, pretty easier signal, still complex. But my website customers happy is a tricky question to answer. You can infer it. And there are a bunch of techniques around, say like, okay, is the objective to, it depends on the domain. If you have an e-commerce site, are people buying things successfully? If not, there's probably a problem on the site and people probably aren't happy for some reason. But you have to have a domain specific, just like business outcome. Same with agents.
51:29Jeff Meyerson:It's not an API. These are things that users are using to drive a certain business outcome. That's tricky. Some things that we see around agents are humans entering like a thumbs up, thumbs down, which is a useful signal. It helps, you know, but then like, is the user really going to provide a really detailed description of why they weren't happy? Like we don't want to overcorrect on that. I'm never a huge fan of the, I mean, it's important and necessary, but I don't want to just rely on the thumbs up, thumbs down. It's kind of the airport bathroom cleanliness button as you leave. is it clean or not?
52:01Jeff Meyerson:I don't like that. In fact, I liken it to that because it is sort of people wrinkle their nose at the idea of that. It's like, I kind of intentionally give a sort of gross comparison because I think what we need to be doing is figuring this out more naturally. What is the objective that people have with an agent and how close is it to achieving that objective? And where it didn't achieve it, why? So this evaluation and learning are both very related about how do you evaluate whether or not this is succeeding for customers? And then when it did, how do you make sure that we are good at that next time if there's a new discovery?
52:35Jeff Meyerson:And if we didn't, that we don't go down that route next time, unless it's just a situational thing where maybe like an operational investigation, and maybe we do need to still check to see if it was a bad deployment that triggered the event, even if it wasn't a bad deployment this time.
52:50David Yanacek:I love that. I recently set a challenge to my team. I said, one of our goals for this product is it should be delightful. You have to figure out how we measure that, right? But it's like, these things are so nuanced and subjective now.
53:03Jeff Meyerson:Yeah, it's a nice thing where we can borrow the page from website and mobile device, mobile application monitoring, like real user monitoring is sort of this rum, as industry called is like kind of, I'd say it was certainly appreciated by people who are doing website operations and development and mobile app website development. But I think the rest of the industry might not have really appreciated the need for that. Maybe because it didn't apply to what they were doing so much, but I mean, it could have. But now it brings that kind of technique into the forefront, but with new technology to be able to do real user monitoring better now.
53:38Jeff Meyerson:So yeah, it's interesting how it just, it's a nice shift toward understanding your customer's happy. It's one of always drawn me to observability because I used to work on observability on CloudWatch of observability as a lens with which to understand your customer's happiness. And so I think that's very true with agents.
53:57David Yanacek:So we're coming to the end of our time at this point. Is there anything we haven't talked about that you think would be important to leave folks with?
54:05Jeff Meyerson:I'd say when looking at coding tools that help you with, like, that's a really great starting point into kind of rethinking how you're building and operating software. But look at, I guess, one, when you start with a tool, especially if you're new to AI coding tools, you're not going to get what you want from it on the first try. It's like when search engines came out, like this is like a little bit of a story time thing. But when search engines came out, I remember some people, you know, you weren't just magically good at using search engines. You had to use the right syntax. Maybe search engines were a little primitive when they first came out, too.
54:43Jeff Meyerson:So you had to learn how to use Boolean expressions to be able to filter out things that were uninteresting to you. But the more you did searching and the more you really studied, like, okay, why didn't I, and introspect, why didn't I get what I was looking for from that search? The better you get at using the tool. And, of course, the tools get better over time that you don't need the expertise on, but they also, like, introspecting, well, why didn't I get what I want? And I see that the more that people think about that, the more successful they are with a coding tool, like a coding agent, like for any coding agent like Kiro.
55:15Jeff Meyerson:Because maybe like, well, why wasn't it able to use this new framework that I made five minutes ago? It's okay. Well, you need to remind it to. Like, okay, so that's what steering is for. That's what powers are for. Like, that's what MCP servers are for. So just think about, well, why didn't I get what I wanted? And so because other people are getting that. We found that a team did a 30 people for 18 month level of replatforming of a service, like 30 people for 18 months. We're able to get that done in like with six people in six weeks, something like that. I'm citing this statistic, but it's just a significantly shorter time when the tools are really used in this.
55:55Jeff Meyerson:So my advice for people is these tools are extremely powerful, but they take you kind of learning how to use it and incorporating it into your development practice and then changing your development practices. Because these tools can accelerate coding so much that they force the need to change the rest of the practices around the coding. It doesn't become the bottleneck anymore. And this is why then we've built these frontier agents to handle things beyond the coding because we don't want to just shift the bottleneck. When you're doing a bunch of production-grade coding, you also need production-grade security, pen testing, and review.
56:28Jeff Meyerson:You also need production-grade operations. And so you kind of need to accelerate all these things at once. And so look beyond. Basically, distilling this into two things is keep using tools and ask yourself, well, why didn't I get what I wanted? Because other people are. So why didn't you get what you wanted out? There's maybe some trick of using a tool in a better way. And then second is look beyond that one tool and look for the other bottlenecks that you can speed up and make your life easier with, like around DevOps and security.
56:58David Yanacek:Yeah, absolutely. Well, and I think the stuff we've talked about today, like all these things, Cure was baking in some of the practices that six months ago you had to learn, right? Like I have to learn, prompt the thing for a spec. There it goes. I love the property-based testing because it does fit into this idea that I've been playing with a lot of, like, forcing the LLM out of its groove. Instead of here, confirm your, you know, jump down this confirmation bias loop, it's go across all of those things. And as we talked about, like, that concept of what needs to be verified is just baked in there.
57:30David Yanacek:And we did check while we were talking, like, there are property-based testing libraries for all sorts of different environments, and we can just throw them. Sounds like Kiro should be able to just generate tests in whatever environment. So super cool. I know we've been seeing incredible speed ups. It sounds like, yeah, Amazon, you guys are seeing incredible speed ups. So excited to see where this goes.
57:50Jeff Meyerson:Yeah, same. It's a new frontier. It's the next big speed up in making developers lives easier. In this case, a whole lot.
58:07Yes, sir.
From the publisher
AI-assisted coding tools have made it easier than ever to spin up prototypes, but turning those prototypes into reliable, production-grade systems remains a major challenge. Large language models are non-deterministic, prone to drift, and often lose track of intent over long development sessions. Kiro is an AI-powered IDE that’s built around a spec-driven development workflow.
The post Amazon’s IDE for Spec-Driven Development with David Yanacek appeared first on Software Engineering Daily.
