963: Reinforcement Learning for Agents, with Amazon AGI Labs’ Antje Barth

3 Feb 2026 · 51 min · 21 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Amazon AGI Labs’ Antje Barth discusses Nova Act, a new AWS service for building reliable browser UI automation agents. She argues that “flashy” agents that work ~60% of the time are effectively useless for production, and claims Nova Act targets 90%+ reliability.

Guest backgrounds

Antje Barth is a technical staff member at Amazon AGI Labs in San Francisco, a developer relations lead there, and a multi-time O’Reilly best-selling author. She has taught GenAI to 400,000+ students.

Key claims

Nova Act provides a free playground (no AWS account) to prototype tasks via natural language, shows reasoning traces for debugging, generates Python scripts, and supports deployment on AWS with observability (step views, traces, CloudWatch). Reliability is improved via reinforcement learning “web gyms” (replicated UIs) where agents self-play and learn to generalize across UI changes.

Notable examples

PGA Tour uses Nova Act for website QA (widgets/buttons/weather). Hertz uses it for rental workflow testing, claiming 5x shipping velocity. One example of secure auth handling is OnePassword-style sign-in flows with user-controlled password entry. Benchmarks mentioned include Work Arena and RealBench.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Exploring Nova Act

0:46 to 2:42

Antje Barth explains Nova Act, its features, and how it aids developers.

“I'm calling in from the Amazon AGI Lab in San Francisco.”

User Journey with Nova Act

3:03 to 3:43

Detailed walkthrough of the user journey through Nova Act's playground experience.

“And then you iterate there, and then you can move into the next step, for example, if you want to customize more in IDE environments.”

User Journey with Nova Act

3:48 to 6:42

Detailed walkthrough of the user journey through Nova Act's playground experience.

“There's also links to all the dev tooling we're offering, IDE extensions, SDK downloads.”

Improving Agent Reliability

6:43 to 11:09

Discussion on how Nova Act improves the reliability of AI agents through training and testing methodologies.

“Watch it, do that task and then get the Python code, use it in whatever environment you're comfortable writing your code in and then that allows you to easily scale up whatever you're doing.”

Evaluating Agent Performance

11:10 to 14:01

Antje explains methods to evaluate the performance of agents and the importance of real-world testing.

“But how can we evaluate that to be sure?”

Debugging and Observability in AI Workflows

14:01 to 16:00

Learn about the importance of debugging and observability in AI workflows, especially when deployed on AWS.

“because we know as a developer, you need to be able to debug.”

Balancing Customization and Convenience in AI Development

16:30 to 18:55

Explore the balance between customization and convenience when using AI tools and platforms for development.

“And now I'm starting to understand the model here, I think, as well, the business model from Amazon's perspective, which is, you know, a lot of organizations will have open source initiatives.”

Building Trust with AI Coding Tools

18:55 to 20:04

Understand how trust develops in AI coding tools and the importance of oversight in their use.

“We're giving you this fully integrated way to hopefully take a lot of this heavy, undifferentiated building away and give you a faster time to value with this.”

Automating Tedious Tasks with AI

20:04 to 22:30

Learn how AI can help automate tedious tasks to improve developer productivity and workflow.

“Every one of us has like those really kind of boring, tedious tasks we have to do during the week, right?”

Use Cases for Nova Act in UI Automation

22:30 to 26:32

Discover various use cases for Nova Act, specifically in automating UI tasks and QA testing.

“Like my job shifted to less reviewing, but really kind of just the supervisor, right?”
Show all 21 chapters

Exploring Normcore Agents in Automation

26:32 to 28:00

Delve into the concept of normcore agents and their role in automating core but mundane tasks.

“maybe the UI example is a really good place to go because it seems like Nova Act specifically is designed for browser-based tasks.”

Introduction to Normcore Agents

28:00 to 30:00

Learn about normcore agents and their role in automating mundane tasks.

“They're doing automated testing for a lot of their core rental workflows, which help them to just increase the shipping velocity 5x.”

The Evolution of QA Engineering

30:00 to 32:00

Explore how QA engineering is changing with the rise of AI and automation.

“And when you were talking about norm core, you were talking about QA engineering and something that's interesting with this shift to agentic is that traditionally QA engineers were focused on finding bugs, really.”

Secure Interactions with Nova Act

32:00 to 34:00

Discover how Nova Act ensures secure interactions in authentication workflows.

“Reliability, security, to be able to build a trust.”

Evaluating AI Agents in QA

34:00 to 36:50

Understand the challenges of evaluating stochastic AI agents in QA roles.

“very important to make sure, you know, you're not sending any sensitive data.”

The Future of Digital Teammates

36:50 to 40:30

Discuss the evolving role of AI agents as digital teammates in the workplace.

“But right now, a lot of those tasks, there is an end state where you can check and then give the agent a little bit of like freedom, right?”

Autonomous Systems and AI Models

40:30 to 42:01

Learn about the future of AI model architectures and their applications.

“Is that kind of in the namesake of Amazon AGI Labs?”

Exploring Context in AI Models

42:01 to 44:26

Discusses the balance between context size in AI models and retrieval systems.

“context, or whether it even makes sense to have million token contexts, considering retrieval systems and tools are continually improving to fill in the gaps.”

Connecting with Antje Barth

44:26 to 45:14

The host engages with Antje, discussing her insights and where to follow her work.

“maybe even like, you know, a couple of years from now.”

Book Recommendation from Antje

45:14 to 46:04

Antje shares a book recommendation on AI systems performance engineering.

“There's definitely a lot of content I'm sharing.”

Wrapping Up the Conversation

46:04 to 46:51

The host expresses gratitude to Antje and discusses the future of AI collaboration.

“I haven't managed a thousand pages yet, but hopefully over the next couple of months, I will achieve that.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Jon Krohn:An AI agent that's reliable 60 % of the time for nearly all real-world use cases is 0 % useful. Welcome to the Super Data Science Podcast. I'm your host, Jon Krohn. In today's outstanding episode, I'm joined by Antje Barth, a multi-time, best-selling O 'Reilly author, an instructor on Gen AI with over 400 ,000 students, and a member of the technical staff in Amazon's prestigious AGI Labs, where they're focused on building reliable AI agents that feel like a digital coworker, not just a tool. Learn how in today's great episode. This episode of Super Data Science is made possible by Dell, Intel, Fabi, and Cisco.

0:40Jon Krohn:Angea, welcome to the Super Data Science podcast. It's a treat to have you on. Where are you calling in from today? Thanks so much for having me. I'm calling in from the Amazon AGI Lab in San Francisco. Very, very nice. Amazon AGI Labs is the focus of our episode and all the cool things that you guys are doing there. Our research is fascinating on what you're up to at the Amazon AGI Labs. I can't wait to dig into it over the course of this episode. Let's do it. Yeah. So you're a member of the technical staff there at Amazon AGI Labs, which is a very cool role. You're also a multi-time O 'Reilly book author.

1:14Jon Krohn:That's not really going to be the focus of this episode, but the point is you've got deep technical chops. You're a really well-known individual. I guess maybe from those kinds of things like making those O 'Reilly books, You're also a developer relations lead at Amazon AGI Labs, right? Right. Yes. Nice. And so from there, you drive the vision, strategy, execution for how developers engage with Amazon's next generation AI products. And the latest exciting AI product you're promoting is something called Nova Act. So tell us about Nova Act. Yes. So Nova Act is a service that we just recently launched.

1:50We had a research preview going on since March last year and are super excited that this past December at our AWS reInvent conference, we launched this as a GA service. Nova Act helps you to build UI automation tasks at scale very reliably and helps you to really kind of start prototyping and putting it in production really fast. So you can get started, which I love as a developer, really fast in a playground experience and then, you know, iterate on it, debug it. And once you're ready, push it onto the AWS site and run it there reliably and safely at scale.

2:28Jon Krohn:Really cool. And it's free to get started, right? Absolutely. So one of the key things, again, I'm excited as a developer is we want to make it really easy for you out there to get started, right? Like in this industry, we know the speed to delivery, speed to shipping is kind of the mode. So you don't want to like, you know, spend too much time spending up the right environments and infrastructure and integrations. You really just want to validate your ideas, right? A lot of startups, you just want to go build the ideas you have, validate them really quick, iterate on them. So you can get started really quick in our playground experience to especially do that.

3:06And then you iterate there, and then you can move into the next step, for example, if you want to customize more in IDE environments. So we really want to keep also the surface area really kind of, you know, where you're doing your day-to-day jobs as a developer.

3:20Jon Krohn:Really interesting. Would you be able to walk us through the typical user journey through a Nova Act experience? Obviously, I'll have a link to Nova Act in the show notes for this episode so people can go there and describe to them what it would be like to us to experience it as we go there for the first time and we're playing around in the playground all the way through to deploying? Absolutely. So you can go to nova.amazon.com slash act. And then there's the playground experience. There's also links to all the dev tooling we're offering, IDE extensions, SDK downloads. But really kind of the first step is you go into the playground free of charge.

3:58You don't need any AWS account or anything. So you just go in there, you provide a website. For example, you're going to, let's say, a booking website or a specific event, maybe sign up site. Conference season is starting soon here in the Bay Area and everywhere. So you just put in the website. And then in natural language, you can decide and put in the actions to take. Let's say I want to sign up to an upcoming meetup. So maybe I put in the Luma website and I'm going to say, hey, search for a specific meetup. Maybe I want to join, you know, an AI performance meetup. I look for that. And then you can also like, you know, put the actions in to fill out, like click the RSVP, click the join it and have it fill out the form for you.

4:46And on the same playground, you see an embedded UI, the browser environment. So you don't have to set that up any manual way. So you can see it right there. So you can observe how Nova Act, how the agent is going to that website. It's performing your actions. you see also the reasoning traces. So what it is doing, which is exciting, especially important for developers to be able, you know, to debug, to troubleshoot, to really see what's happening there. And you can tweak it if it's not getting at it the first time. So you can optimize your instructions, see it. And then once this is performing well to what you want to achieve, you can then download the scripts on the background.

5:26It's writing a Python script that captures all of those steps. So you have the ready script. You're using natural language. It converts it to the code for you. And then we have IDE extensions, for example, or just an SDK. You get that into your preferred IDE of choice. You import that script. And then you're basically back in your coding world as a developer, right? So you can customize it. You can tweak it in there. And also the IDE extension has this embedded live preview. So, you know, a lot of times when we're building those automation workflows, you have like a separate window popping up that shows you the browser.

6:03And we got a lot of like developer feedback, but that's just a little bit too much, right? We want to stay in the flow when we're building something. So we included that in your IDE. So you can really stay in there, have a unified window with all the troubleshooting, the debugging, the traces, so you can keep really close eye, develop, customize. and from the IDE, you can then, if you want to, also connect to your AWS account if you have one and for example, deploy it there to run it in production.

6:30Jon Krohn:Nice, so that sounds like a really easy way for somebody to be experimenting with developing an agent because you can go to nova.amazon.com slash act and then be able to use your natural language to describe what you'd like your agent to do. Watch it, do that task and then get the Python code, use it in whatever environment you're comfortable writing your code in and then that allows you to easily scale up whatever you're doing. It sounded like there was also, it sounded like you mentioned also being able to productionize on AWS with this solution. Right. So a core kind of motivation for us, you know, talking to a lot of customers, talking to developers out there, we've seen so many flashy demos, right?

7:08Like you go to any meetup, especially here in the Bay Area, but I assume now in other parts of the world in a similar way. And everyone has a flashy demo, right? And all the kind of fun stuff that agents can do. But what we observed and the feedback we received as well from customers is that on average, those agents work maybe 60 % of the time. And to be honest, an agent that is reliable 60 % is 0 % useful, right? Especially if you want to productionize this, you need to really have reliability that you can trust those workflows and that agent to complete that one task. So this was the core motivation for us, kind of really the P0 to make sure Nova Act can reliably 90 % and more deliver on those workflow executions.

7:55Jon Krohn:Awesome. So how do you do that? How do you get that kind of confidence? How do you go from a 60 % reliability to better than that? What kinds of tooling do you have in Nova Act to ensure that? Yeah. So I want to talk a little bit how we train over ACT, right, which is exciting. At least I find it super exciting. So in the past, when we trained, you know, AI models and things, it's a lot about imitation learning, right? You collect data and you show it that data. Now, as we're moving from kind of the chat-based conversational models that we all know and work really well into this space where we're building agents that need to take an action, that doesn't work anymore, right?

8:37Because the agent is not just predicting a next word, next token anymore. The agent needs to predict the next action to take. So what we're doing is we're building out those reinforcement learning based web gyms, as we're calling them. We have actually one in the playground. You can actually play around and observe one. And those web gyms are replicas of like, you know, very typical UIs. So like maybe a form filling UI, a shopping workflow, etc. And we're letting the agent train in those web gyms. So imagine hundreds of those gyms, like typical UI design elements, they're typical tasks to do, and then give the agent like 1000s of tasks, right?

9:19And they basically self play in there. So this is similar, like in past, you know, how AI learned to play chess, how it learned to play Go. It's really kind of a trial and error approach. So the agent goes in there, tries to do that form filling, it might fail a couple of times, right, but it self corrects, it does it again, and it learns a better way to achieve the task. So this way, the agent understands to, to reason, you know, through this UI and understands how to accomplish the task and helps it also to generalize well, right? Because UIs change. So that's super important. If you're building that agent, yes, you have a specific site, maybe you're, you know, in those gyms you're training it on.

10:01But then again, you wanted to

10:02Jon Krohn:generalize if the website changes, if the checkout button moves, if the sign in button moves somewhere else, if it's using different icons, because a lot of UIs are really designed for us humans, right to navigate. So and for us, this is a simple task. If you think about maybe you go and write an email. And depending on which email program you're using, which application, sometimes the button says draft or new or create, right? Or it's just a little pen icon to create a new email. For us humans, we learned that. We understand that it's not a hard task for us. But if you're sending an agent to that environment, to that UI.

10:41The agent needs to be able to generalize and understand and reason in a similar way, even if some buttons and some tool says draft, the other one says create. So this is really exciting. So you're training the models and Nova Act is trained like this on those web gems to then really kind of be able to generalize and navigate those web pages, even in real life as they're changing the structure and then they look. Cool.

11:04Jon Krohn:So a big problem that occurs to me that I think a lot of people have with using agents and maybe the kind of training that you've been describing, it helps increase the reliability in a lot of common scenarios. But how can we evaluate that to be sure? Is there anything within Nova Act that allows us to evaluate the performance of whatever agentic workflow we're automating? Absolutely. So obviously you can check the leaderboards we've done, the public benchmarks on Work Arena, on RealBench. That is important to just be able to understand the capabilities of the solutions. But even more important, and what we focus a lot on, is really working closely with our customers.

11:44Because we want to deliver agents that work on real use cases with real company tasks and for real users. So we're not focusing on this kind of the flashy demo side. We're really kind of working closely in collaboration with the companies, with the customers we have, and then making sure the evaluations are running on their specific tasks. And this is where we achieve those over 90 % really kind of on those early customer enterprise workflows where we're really kind of focusing down what needs to be done and evaluating and making sure we're really delivering this reliability there. Love it.

12:22Jon Krohn:Tell us a bit more about Work Arena and RealBench specifically. I have heard of those benchmarks, but I don't know too much about them myself. And I bet they're new to a lot of our listeners. Yeah, so you can check it out. The RealBench, for example, super exciting. They have in a similar way how those agents work, right? They have those web gyms, those replicas of websites. So you basically can submit your agent there. And then it's navigating a similar environment where there's a bunch of tasks across a different set of those gym environments to solve. And this is kind of this new benchmark, right?

12:53Especially for those UI automation agents to see how they perform on those unseen new environments.

13:00Jon Krohn:Yeah, we'll have a link to both of those in the show notes for sure. So a related question to the evals is there's lots of other things that you need to worry about when an agent goes from being in a playground to going to production, security, governance, access controls, change management. How do you handle those kinds of things? Right. And this is where we really leverage the AWS integration, right? Running as an AWS service. So we fully use, you know, we're building on the foundation security that AWS delivers. All the security, all the data is running in your account. So you have the full control and we're making sure, you know, that is really running in a very reliable, in a safe and controlled environment that you as a customer can control.

13:43So all the goodness of, you know, that AWS delivers in terms of reliability and security, that comes as well. And the important thing is really kind of, you know, have this seamless journey. So yes, there is a playground that helps you to get started and evaluate ideas quickly. But we also wanted to make sure, because we know as a developer, you need to be able to debug. You need to have traces, right? You need to see what is happening. If any of those steps fail, why did they fail, right? So across this journey, across those different surfaces, whether you're in the playground, whether you are in the IDE.

14:21And then obviously when you're running it on AWS, you have that observability, you have those traces. It's really exciting. So for example, once you deploy this workflow onto AWS, and even if you're running it in the IDE, we're creating very detailed step views. So you see for every run you're making, so basically you're doing an act call, what we're calling it. So you're triggering this workflow. and then within that workflow, there's a couple of steps, right? Look for the search button, insert this, and then grab maybe, I don't know, the cheapest coffee machine I can find somewhere. Whatever your task might be or fill out this form.

14:56And for every step, you see the reasoning traces so you can understand why the model took a specific action. You see also the next action prediction. Let's, for example, click on those coordinates, put this in. And this is super important because you can really debug and troubleshoot as you're developing how the agent behaves, where it might get tripped over something and then correct that. And then as you're running that in production on AWS, you have the similar views. So you can, for example, run the same workflow. If you're thinking like enterprise cases, form filling that needs to be done maybe a thousand times in parallel, you can schedule those runs.

15:35And for each run, you have the detailed view. So in any way you have the audibility, you can look into the locks there, integration with CloudWatch and all the other tooling to really kind of give you the full control and also the visibility into what's happening.

15:52Jon Krohn:Data scientists, it's time to talk about your tech. With Windows 10 support coming to an end, now is the perfect moment to rethink your setup. Enter Dell AI PCs powered by Intel Core Ultra processors. These devices are built for the demands of modern data science, delivering faster performance, smoother multitasking, and the power to handle even the most complex workflows. Whether you're training machine learning models or analyzing massive data sets, these PCs are designed to keep you ahead of the curve. Don't let outdated tech slow you down. Visit dell.com slash shoppcs to explore how you can upgrade your device and elevate your work.

16:30Jon Krohn:That's dell.com slash SHOPPCS. I like that. And now I'm starting to understand the model here, I think, as well, the business model from Amazon's perspective, which is, you know, a lot of organizations will have open source initiatives. But even when an enterprise has an open source initiative, they're looking for some kind of goodwill or some kind of benefit downstream. and maybe that's something is that when people want to scale up their agents on AWS, Nova Act is free, gets them going, but then when it's time to scale up, it's convenient that you have AWS available to do that if you so choose to use AWS to do that scaling up.

17:14And I would add to that, it's also like giving customers, giving developers choice. Like some developers really enjoy like building it themselves, right? And using the open source tooling and just do a very customized approach. But I think as you're scaling this out, you really kind of want to consider like how much of this work do I want to do myself? And there might be very well, you know, very important cases where you want to have that customization ability. And that has always like as well, Amazon's and AWS's approach in this, like given customers choice in the building blocks, whether they're using open source, whether they're using kind of, you know, kind of more individual services and capabilities, on the AWS side, there is this trans agents SDK that you can use to build agents and you can use any kind, you know, of tooling of models there gives you full flexibility.

18:05But then again, there might be customers that say, well, I don't want to go this route, right? I don't want to spend that much time customizing. I'm really looking for a solution that has the best parts integrated and I don't have to worry about, you know, sometimes we refer to this as kind of a Frankenstein AI agent, where you kind of, you know, you grab a model, you have to develop, look for the orchestrator, and then with the UI actions, right, an actuator that actually taking those actions in the browser, in the UI, and you have to bring those together. Our approach at the Amazon AGI lab is to train this all together.

18:42So basically, we're training the brain and the body in one to give really kind of, you know, optimize it and give you that reliability. So again, there's plenty of solutions. You can achieve this with Nova Act. We're giving you this fully integrated way to hopefully take a lot of this heavy, undifferentiated building away and give you a faster time to value with this.

19:06Jon Krohn:Yeah, there's a lot of thought to be put into what's the right balance to allow people to be doing things independently or using pre-made tools. Now there's trade-offs between customizability versus speed and reliability. And we dug up something, a project that you've done where you used, and you can correct me if I'm wrong on any of this, but our research indicates that you used Amazon Q and VS Code to develop an app called Summarize Me. And that Summarize.me app produces an Antje avatar, animated video summaries of your meetings. And so you argue that developer environment agents act like a personal tutor sitting next to you while you work, and that they help you.

19:46Jon Krohn:They help developers stay in flow, which is always a nice place to be. So drawing on your experience building that Summarize.me app, as well as all of your Nova Act experiences with Amazon customers, given the subjective nature of developer flow, what do you think is the ideal balance of agency and oversight? Yes. I want to also bring in another example here in a bit, but yes, correctly, I built this thing as kind of like, you know, I worked as a developer advocate for many years and with AI, like many of us probably were thinking like, what tasks can I actually automate? Right. Every one of us has like those really kind of boring, tedious tasks we have to do during the week, right?

20:27It might be filling out an expense report at the end of the week because we were traveling somewhere for business and we're really dreading this exercise, but we have to do it. Or something similar, right? Or researching the web for like, you know, AI news that's coming out. Like, I don't know, we cannot even keep track of what's happening throughout the day, right? So I think a lot of us have those little side projects where we're trying to just to automate some tedious parts of our daily jobs. And this was also kind of a motivation for me when I put this project together just out of fun, really.

20:58Can I actually, you know, using some avatar technology and like meeting summarization and stuff to kind of automate parts of my job where I'm like, you know, I'm in so many meetings throughout the day. But then again, you know, with AI summaries, you're getting like then now a million of those AI summaries. It's also sometimes too much. So I was trying to like think, how can I, you know, streamline all of this? And to your point about the agency and the control of things, right? I have another example. Like if you're looking at AI-assisted coding, for example, I think many of us are using those tools now.

21:36And if I just look back like a year ago when I kind of worked with the tools or even like, you know, one and a half years ago, we started out very early in the days with tap completion, right? So we were writing the code pretty much. Yeah, we did a tap complete. then in the next phase with those AI coding agents assistance becoming more powerful we then had them to maybe scan our repo and then also like you know co-develop that we're saying hey can you create this capability for me within this larger project within this larger application but for example I still kind of reviewed the code right like once I got it bad I was like okay what did you do there?

22:16So I went line by line through that. And to be honest, fast forward 2026, probably most of us are using some sort of AI tool. I cannot even keep track, right? There's like so many software agents that are helping you build your applications, build projects. Like my job shifted to less reviewing, but really kind of just the supervisor, right? So I cannot even keep track of reading every line. So to your point with the agency and the control I think it's really about building trust with a tool. And I think we went through that period over the last one and a half, two years where we started to work with the i-assistant coding.

22:57And now we have this trust, but with trust also comes, you know, we need to verify. So a lot of focus, what I'm doing is like when I'm coding like this, I want to make sure that I have really solid unit tests and verifications in place. but as we shifted from this vibe coding to kind of more spectrum development, we're still interacting with AI in the same way. We're telling it what to do, mostly a natural language. We're writing some specs now and then we have the testing framework. And with this all together, we have developed a decent amount of trust that I think we're giving more and more agency.

Read the full transcript

23:36So this is an example from the coding world but I think a similar pattern will arise with those UI automations. Right now, of course, we want to have, you know, for specific things, a human in the loop that maybe approves a checkout workflow or just takes care of some of the important decisions that are part of this task. But ultimately, as we're building up the trust and we have, you know, specific verifications in place, I think this agency will slide a little bit to, you know, to give it more.

24:07Jon Krohn:I think that makes perfect sense. I do think that is where we're headed. You mentioned at the beginning of your most recent response that you had another example other than summarize me. And now would be a good time to go into that. Yeah, that's kind of how I code, right? This is what I want to get in. It's shifting how we're still using natural language, talking to all those tools, but we're giving it a little bit more agency. Right, right. Gotcha, gotcha, gotcha. So we have a quote from you from a blog. We'll have a link to that in the show notes where you point out to where you point to over a thousand generative AI applications already built or in development at Amazon.

24:44Jon Krohn:So from those, do you have some kinds of takeaways for our listeners from those, from a thousand plus internal deployments of Gen AI, where agents can actually deliver value and where they still struggle? Yes. So it's really exciting, you know, being part of Amazon. Also, we're working and collaborating with a lot of the internal teams here and seeing the different use cases coming together. And one of the very exciting ones is from Amazon Leo. You might know Leo. Previously, they were known as Project Kuiper. They're working on those low orbit satellite networks to bring connectivity into the underserved regions.

25:25So very exciting work that they're doing. And they've been looking at NOVA Act as well, how they can integrate it. And they built some amazing automations that help them do testing QA frameworks across different, from the web to Android iOS apps as well. And then also build additional tooling on top of that, where, for example, they can now, product managers in the team can use Nova Act to, from a FICMA design, from a wireframe design, create test cases even before those get implemented. So there's a lot of like exciting work that's happening. And yeah, I'm also curious, you know, out there, like a lot of you, like startups, developers, data scientists that are listening, like how you think UI automation can come together and which processes you can automate.

26:20But yeah, certainly there's, I think, so much opportunity space. And we're really kind of really just at the very beginning of what's possible.

26:27Jon Krohn:For sure. Are there some kinds of general things you've seen that work or don't work? maybe the UI example is a really good place to go because it seems like Nova Act specifically is designed for browser-based tasks. We started with a browser because it's the easiest way to give us kind of the closest to this universal action space. If you look at like how we humans work, right, a lot of our work is actually happening in the browser where we're working across different applications. We have the different tabs open. So this gives us a very great first point to get to this universal space, but we're also looking beyond the browser, right?

27:05So again, we have great research teams here. I really enjoy talking to a lot of our cognitive scientists who are also looking into those problem space. And one of the main use cases really is, to your point, is the QA testing, right? We see a lot of things happening there for a couple of various reasons, I think, but really kind of, it helps you really quickly to automate some of those, you know, as you develop new applications, you're pushing out an update and you want to make sure that the experience delivers. For example, one of our customers, PGA Tour, they're using Nova Act to make sure the website correctly works, especially before, you know, any golf tournaments where there's a lot of excitement and people coming to the website.

27:50So they're making sure that, for example, the weather information, widget, everything works, all the buttons, everything in there. So a lot of this QA testing on websites, but also like, for example, another one, Hertz uses Nova Act. They're doing automated testing for a lot of their core rental workflows, which help them to just increase the shipping velocity 5x. So it's really kind of, and sometimes internally, we use this playful term that we're calling those norm core agents, which really automates some of the most more boring stuff, right? Like, it might not be those flashy cool things that we sometimes come up with at like hackathons or stuff.

28:33And really kind of what are kind of the core workflows that seem a little bit more boring, seem a little bit more tedious, but you can actually automate pretty well. And we see a lot of interest as a first use case in this automated QA testing, but we can imagine like a ton of more use cases out there.

28:51Jon Krohn:Yeah, I was going to eventually ask you about this normcore stuff because we did have some research on it. What does it mean? What does it mean, normcore? Where does that come from? Yeah, so this is kind of, I think it was coined in the fashion industry, kind of normcore looks, which is a little bit more, you know, like kind of a simple white T-shirt and jeans, something. and norm core for us is really kind of you know those agents focus on more kind of the the traditional kind of the boring use cases you would call them maybe sometimes like exactly that form that i need to fill out a thousand times that application that i need to do maybe you need to apply for something you need to navigate you know websites to get a new driver license renewal or whatever it might be.

29:38So you have those tasks that you're dreading, as I mentioned before, or that is just some work that you need to do. And if we can automate those use cases for you, those tasks in a very reliable way, I think that would be amazing to just give you all that time back to focus on the more higher level, a more creative task we're all excited about.

29:58Jon Krohn:All right, norm core. Yeah, I was thinking maybe it's kind of like the opposite of hardcore. And when you were talking about norm core, you were talking about QA engineering and something that's interesting with this shift to agentic is that traditionally QA engineers were focused on finding bugs, really. But now I feel like a lot of it is probably related to auditing reasoning traces that LLMs spit out. I don't know if you have any thoughts on that. This is kind of a tangential question. I think there's a lot of like things that will change how we're doing our jobs, right? Like we learned this in the software engineering world that we're like more talking to AI than we're actually writing the lines of code sometimes.

30:38And yeah, I think similar things for QA engineers that they don't have to write those brittle scripts anymore, right? If you're looking like how you achieve those tasks in the past, you had to write those very kind of rule-based scripts, like Selenium scripts or something. If you wanted to automate those processes, right, RPA, where you said, okay, here's the button, click that, do that here and here. And whenever some element changed, you had to rewrite those. So hopefully we can, again, like simplify this for those QA engineers, and you can use much more QA through natural language, creating those scripts for you.

31:18And then also with those agents, they're much more resilient to adapt to those changes, right? So you don't have to spend that much time in keeping those scripts up to date. And hopefully QA engineers can focus on the real things. And that might include, yes, looking like, you know, a different type of traces, for example, reasoning traces, etc. But hopefully this gives them also like, you know, a little bit more focus on what's exciting in their job, like figuring out what's going on rather than just spending hours and hours and hours in writing those scripts and updating them.

31:49Jon Krohn:That makes a lot of sense. So going back a little bit, going back a few minutes, at least, you're giving us an example, a real world example of how Amazon Nova Act can be useful for the PGA Tour. And we dug up another example here, which I thought of which i thought was kind of interesting which is with one password and so one password highlights nova act handling complex sign-in patterns which is a high-stakes surface area especially for an application that you're trusting with all of your passwords so how has nova act dealt with secure interaction with authentication flows so agents don't become the weakest link in identity and access right and security is really kind of you know one of them that if not the most critical part as well, right?

32:32Reliability, security, to be able to build a trust. So for Nova Act, we're really kind of building on the security foundations of AWS. I briefly mentioned this before. So you can be assured that all of those executions are running in your AWS account. They are protected. You have full control over them. And also like we give you back control, right? So for example, if you're encountering a password that you need to fill in, you can give back to the user. So basically the user would then be in control and you would, for example, with a playwright integration, you can then prompt like locally in the terminal to, hey, get the password from the user.

33:11It's never being sent over the network, but then the agent can fill it in. So a lot of thinking really goes around that. And then one password, obviously given like they are the password manager, one of the leading ones. So they build their framework around that as well to make sure, for example, the password manager can intelligently navigate the site and then they have their part, how they deal in a very safe and secure way with the passwords. But that is definitely an area where you want to make sure you're not just sending any things over network, et cetera, and really want to make sure you have a very trusted and secure way to deal with those.

33:48In a similar way, like with captures, less sensitive maybe, but also this is another case where we would give, you know, user the control back to solve those. And we will probably see a lot of development in that area as well, what's happening, but very critical and very important to make sure, you know, you're not sending any sensitive data.

34:08Jon Krohn:Definitely. Another critical area with agents, we did talk about this earlier when you were talking about the reinforcement learning gyms that teach agents how to behave. But related to that, something that, and it also ties into the QA engineering thing we were talking about more recently, which is that even though QA engineering is easier in many ways than before, thanks to agents and to LLMs, one of the things that's tough and potentially becoming tougher is evaluating agents because they're so stochastic. In a traditional QA role, you could say the result should be this as a result of this test.

34:47Jon Krohn:And when we run an automation, this thing should output as the results. But with agents backed by LLMs, you're getting maybe slightly different responses can make that part of the QA engineering job trickier, I'd think. So how do you choose the right success criteria so that the agent doesn't learn shortcuts that pass the test but aren't really doing what you intended? For that, it's really important, like on a couple of different layers, I think, to tackle this. So for one, you want to make sure that the agents can generalize well, right? So even if things are changing, they're not tripping over.

35:23So as you're building out those web gyms, you want to make sure they're really high fidelity gyms and the agent learns a lot of things. Like for example, it sounds simple, but like, you know, one webpage is light mode, either one is in dark mode that is already kind of a totally, even if it's the same UI environment, right? It's a totally different experience maybe for an agent. So there's a ton of variations that You can even with the same task in UI, just mix it up a little to make sure the agents can handle all of those. And then also to your point, like how do I evaluate the success? You really want to focus on the end result, right?

36:02Because there might be a couple of different ways to achieve it. And you can look into it like, you know, maybe one took a little bit longer than the other one. But ultimately what you're interested in is A, it's a successful task completion. and then there's a couple of ways to verify, right? Let's say you have an agent and you're asking it, hey, I want to schedule a one-on-one with my colleague and the agent needs to figure out which calendar is to open, where to look like for free spots. And then the end state would be, okay, you can confirm that the meeting is actually scheduled in that calendar.

36:36So it's going to be somewhere in a database entry. Like if you're doing shopping, like checkout workflow or something in that sort. So you have a very verifiable end state to do that. It gets a little bit trickier when there's not an easy way to verify that, right? But right now, a lot of those tasks, there is an end state where you can check and then give the agent a little bit of like freedom, right? For the generalization part, like which way they go to figure this out.

37:05Jon Krohn:I like that answer. It makes a lot of sense. Thank you for making something that seems so complex, easy to understand with a relatively simple solution, actually. I appreciate that. So let's look ahead a bit now to what's next for agents as we kind of reach near the end of the episode. Let's talk about where this is going. Obviously, it's a fast-moving space, very difficult to predict what's going to happen. But in a recent podcast appearance, you describe AI evolving from a helper into a kind of a peer in the coding experience. And we talked about that earlier in the episode together already.

37:37Jon Krohn:In presentations, you've expressed the idea that the atomic unit of all digital interactions will be an agent call. And virtually every customer experience we know will be reinvented with AI. So how will this changing environment not only change how software is built, which we've talked a lot about in this episode, but also how AI engineers, data scientists, software developers, how they see themselves and how businesses value their efforts? So to start like where I see the space evolving, right, which I'm really, really bullish about, is we want to build, especially here at the Amazon AGI Lab, we want to build useful AI, useful and practical AI.

38:15So this is really our core focus here. And what we think this might look like and will look like is that the agents become much more of a digital teammate, a digital coworker of yours. And we spend a lot of time here in our research teams to think about how do humans actually learn, right? So the future of agents, as we see it, is definitely like a multi-agent, a very collaborative environment, interacting with other humans, right? Like imagine you have this digital coworker of yours, how do they learn how to do the tasks, everyday tasks? And my colleague, cognitive scientist, Dr. Danielle Persik, she actually talks a lot about that in her podcast, Making a Mind.

39:03And she really looks into those aspects, right? How humans learn things, how humans collaborate on a task together. And then we're working together on how can we translate that to building those agents? And how can we teach those agents to learn like a human, to collaborate, to ask for help, to being humble? So this is a super exciting space. And if we think about what this can bring us, right? It's like we might be able to just, you know, select the agent and say, hey, can you help me with this task? and the agent maybe checks back a couple of times, but then also it learns from those interactions like humans would do and then can do those tasks more independent in the future.

39:45So I think a lot of things will go in this direction, working with agents together to help us multiply our own productivity, which is exciting. So I think for everyone, whether you're a developer out there, whether you're an AI engineer, I'm super excited about the space. I think how we interact with technology will change. And it has, you know, looking in the past over all the decades and the technology innovations, like we always had adapt. We will still have, continue to do that. But I think there's a lot of positive in there. Again, doing it obviously very responsible and safe and reliable, but getting to this stage where we can all use AI to then augment our own capabilities.

40:28So I think this is what I'm excited about, at least for the future.

40:31Jon Krohn:Wow, it's a really cool mission. Is that kind of in the namesake of Amazon AGI Labs? Obviously, artificial general intelligence is this idea of an algorithm or system that can replicate all the aspects of human intelligence. So it kind of sounds like that aligns with this idea that you have someone that feels more like a coworker working alongside you that could be helping you out on any kind of task. Yes. And again, the focus is really kind of what we call the useful AI, right? So we want to make sure everyone gets value out of working with AI, useful and practical for day-to-day tasks. And this will hopefully reinvent how we as knowledge workers work, how our work looks like, and hopefully loading off those really kind of tasks we dread throughout the week and we can focus on kind of more the creative and higher level tasks of our jobs.

41:27Cool.

41:28Jon Krohn:And one last technical question, I think here before I start getting to wrap up my questions, which is, this is tying together a few different things that you've said in different places into one kind of forward-looking question where on where you see the underlying models and the architectures going behind autonomous systems. For the last couple of years, There's two active areas of discussion have been about how much to rely on large language models, like really big ones, versus smaller specialized LLMs. And then another big area of discussion, too, is how much to push the boundaries of model context, or whether it even makes sense to have million token contexts, considering retrieval systems and tools are continually improving to fill in the gaps.

42:16Jon Krohn:So you recently mentioned in a presentation that Alexa Plus relies on hundreds of specialized expert systems orchestrated together, and that an internal agent in AWS manages over 6 ,000 tools using retrieval instead of stuffing everything into a context window. So yeah, tell us a bit about that. How much context is too much? Where do you see that going? And where do you see this idea of lots of specialized models versus, you know, big monolith LLMs with trillions of parameters that can do everything? I think there will be a diverse environment where all of the different kinds will have a place, right?

42:54If we're thinking like AI at the edge, that will most likely be kind of smaller models, right? Just giving the constraints there. I think for a lot of tasks, the models will be the larger models that have the reasoning capabilities and all kinds of different shapes and sizes. And going back to this thousand, and specialists and experts. So my thinking is like, you have your AI within maybe your company, but then your company does business with another company, right? So we saw this last year, like a lot of hype around the protocols, like how can I connect my models to tools? And then how can I connect one agent to the other agent, interagent communication?

43:32So I think it's pretty much still an open area for research. What's the best approach here to build this? And there are a lot of components, I think, still being built. But my understanding and my thinking of how the future looks there is like, you know, we're basically entering a stage where we'll have what I call like every single interaction will be an agentic AI call. And then if you think about this, you know, not just within your team or your company, how you do business with the other companies. So there will be a lot of, I think, level and layers up there where the agents communicate with each other.

44:10and figuring out, you know, the right technology, the right protocols in that space. And that will be super exciting. So I think, yeah, there will be a lot of place for all the different kinds of things we're developing, which is exciting. And I'm really curious to see how the world looks like maybe even like, you know, a couple of years from now.

44:28Jon Krohn:Great answer. It makes me feel silly about the question in the first place because it's kind of like, which way are we going? And you're like, obviously this is a complex ecosystem where it's going to involve lots of different kinds of solutions for different scenarios. Makes perfect sense. Well, Anje, this has been an absolute treat for me. I'm sure it has been for our audience as well. You're an outstanding speaker on all these topics. It's such a joy to listen to you. And for all our video viewers, you'll be able to see, of course, that Anje was smiling this entire interview that probably even comes through in the audio in the way it sounds.

45:02Jon Krohn:You just seem like such a happy person talking through all this stuff. So for people who want to be able to follow you after this episode and get more of your happy, insightful thoughts, where should they do that? Follow me on LinkedIn. There's definitely a lot of content I'm sharing. Follow me on X. And yeah, you will see me hopefully also in person if you're in the Bay Area. I'm planning on going to a lot of the conferences this year, popping into meetups. So yeah, hopefully we have a chance to meet in person. And I'm super interested in learning what you all out there are building. Perfect.

45:33Jon Krohn:Yeah. So we'll have links to all your social media in the show notes. And then my very, very last question for you, which I actually usually ask second last, but for some reason it just felt right to ask you the followers one first. And so, yes. So my final question today is, do you have a book recommendation for us? I do. So it's actually on my desk right now as well. So my former co-author from the O 'Reilly books, Chris Brackley, just released a huge, I think it's a thousand pages book on AI systems performance engineering. And this is definitely an area, you know, as we're developing systems, performance becomes so important and how to optimize performance for the whole stack, like really kind of from the GPU level to the application level.

46:19And that's on my reading list. I haven't managed a thousand pages yet, but hopefully over the next couple of months, I will achieve that.

46:27Jon Krohn:Well, maybe we can have Chris Fregley on the air to discuss this. He's someone that's been on my radar for about a decade, but he's never been a guest on the show. So maybe we'll have to get Chris to talk about that new tome that he's put together. Cool. Thanks, Antje. And hopefully we'll have you on air again soon. We'd love to check in and see how everything's going over there at HEI Labs. We'd love to. Yeah. Thanks so much for having me, John.

46:55Jon Krohn:Well, that was a terrific episode with Antje Bart in it. Antje covered Nova Act, Amazon's new service for building reliable AI agents, available free to prototype now. Check it out. I've got a link for you in the show notes there. She went into detail on how Nova Act achieves over 90 % reliability by training on reinforcement learning web gyms, where agents self-play through thousands of relevant tasks, making them norm core agents that focus on boring but high value tasks like QA testing and form filling rather than flashy demos. She talked about how the future of agents involves multi-agent collaboration, where AI becomes a digital coworker that learns from interactions the way humans do, and how Amazon has over a thousand Gen AI applications built or in development internally, providing real-world validation for all of these approaches.

47:42Jon Krohn:All right, as always, you can get all the show notes, including the transcript for this episode, the video recording, any materials mentioned on the show, the URLs for Anje's social media profiles, as well as my own, at superdatascience.com slash slash 963. And yeah, thanks to everyone on the Super Data Science Podcast team, podcast manager Sonja Breivich, media editor Mojir Pombo, partnerships manager Natalie Jaisky, researcher Serge Macisse, writer Dr. Zara Karchet, and our founder Kirill Aromenko. Thanks to all of them for producing another great episode for us today, for enabling them to create this free podcast for you.

48:16Jon Krohn:We're deeply grateful to you and to our sponsors. You can support the show by checking out our sponsors links, which are in the show notes. and if you ever want to sponsor the show yourself, head to johnkrone.com slash podcast to find out how. Otherwise, support us by sharing, reviewing, subscribing, but most importantly, just keep on tuning in. I'm so grateful to have you listening and I hope I can continue to make episodes you love for years and years to come. Till next time, keep on rocking it out there and I'm looking forward to enjoying another round of the Super Data Science Podcast with you very soon.

From the publisher

Bestselling author and Gen AI instructor Antje Barth talks to Jon Krohn about her work at Amazon’s AGI Labs and their newest product Nova Act, as well as where we will see the most success with AI agents and how AI developers can reap those rewards. 

This episode is brought to you by the ⁠⁠Dell⁠⁠, by ⁠⁠Intel⁠⁠, by Fabi and by Cisco.

Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/963⁠⁠⁠⁠⁠⁠⁠⁠⁠

Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.

In this episode you will learn:

(01:23) Amazon’s latest product, Nova Act

(11:05) How Nova Act tests reliability

(24:01) Where Amazon’s 1000s of gen AI deployments succeed

(31:32) How Nova Act maintains its security

(36:32) The increasing value of agentic AI developers

More from Super Data Science: ML & AI Podcast with Jon Krohn

All 130 episodes
963: Reinforcement Learning for Agents, with Amazon AGI Labs’ Antje BarthSuper Data Science: ML & AI Podcast with Jon Krohn · 51 min
Listen in VO