In short
Eye On A.I. - Episode #256: Stephen Schmidt: Inside Amazon’s AI-Powered Cybersecurity Strategy
Episode Overview In this episode, Craig S. Smith interviews Stephen Schmidt, Amazon's Chief Security Officer, discussing Amazon's approach to AI-powered cybersecurity. Schmidt outlines how Amazon uses advanced technologies to protect its infrastructure and customer data, especially in the context of generative AI.
Key Topics Discussed
- The Role of AI in Cybersecurity
- AI is utilized in various areas including threat detection, alarm triage, and code validation.
- Discusses the use of AI as a tool to address security challenges, emphasizing the importance of context and use case.
- Amazon’s Global Honeypot Network (MadPot)
- MadPot Overview: A vast network of tens of thousands of honeypots designed to detect and analyze cyber threats globally.
- Adversaries can discover a honeypot within 90 seconds of it going online.
- The speed from discovery to exploitation is often under 3 minutes.
- Data Collection: Logs from honeypots are used to generate actionable intelligence regarding adversary behavior.
- Threat Intelligence Sharing
- Amazon shares threat intelligence through its service, GuardDuty, allowing customers to leverage insights from MadPot.
- GuardDuty acts as a combination of firewall orchestration and alerting systems, helping to identify potential vulnerabilities in customer environments.
- Security Debate: Open Source vs. Closed Source
- Discusses the ongoing debate about the security implications of open-source models versus closed-source.
- Schmidt emphasizes that both can be secure if implemented and monitored properly.
- AI Agent Deployment and Future Trends
- Current limitations of AI agents in terms of decision-making reliability (around 65-80% accuracy).
- Future expectations for AI agents to take on more autonomous roles in security processes, particularly in initial alarm triage.
- Data Privacy and AI Use
- Critical importance of data privacy and secure usage of large language models (LLMs) in AI applications.
- Essential for organizations to understand how data is handled when using AI services.
- Nova Trusted AI Challenge
- Amazon’s initiative investing over $5 million to engage universities in AI security research, focusing on improving the security of code generation.
- The challenge format encourages collaboration and competition between teams, aimed at developing innovative solutions.
Key Takeaways
- Security is Both a Responsibility and a Process: Organizations must prioritize robust security practices, including maintaining minimal access privileges and monitoring user behavior.
- Generative AI Introduces New Threats: As AI capabilities evolve, so do the associated risks, requiring continuous adaptation in security strategies.
- Collaboration is Crucial: The Nova Trusted AI Challenge highlights the importance of interdisciplinary collaboration in addressing AI security challenges.
- Transparency and Monitoring: Organizations using AI must have clear visibility into how their data is used and shared to safeguard against potential threats.
Conclusion Stephen Schmidt provides valuable insights into how Amazon approaches cybersecurity in the age of AI. The conversation emphasizes the critical need for organizations to be proactive in their security measures, understand the implications of AI technologies, and continuously adapt to the evolving landscape of cyber threats.
---
*For further insights and updates on AI, security, and innovation, listeners are encouraged to follow Craig Smith and Eye on A.I. on social media.*
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Security teams, as we talked earlier, can use Generative AI as one of a group of tools to address challenges. But from the security perspective, it's the use case that defines the relative risk. If you're using LLMs to generate custom code, is that code well written? Is it free of vulnerabilities? And does it follow the best practices that you've implemented for your company? And associated with that is, do you have the proper security controls and threat models in place for your particular use case? There are two major classes of vulnerabilities that our sensor infrastructure emulates. The first is sort of the oldies and the goodies.
0:35The big ones out there that we know our adversaries have exploited in the past, maybe an old version of a database engine or a vulnerable web server component or things like that. And that's to look for the kind of people who are doing mass exploitation of infrastructure around the world. Hey, Craig, Steve Schmidt here. I'm the chief security officer for Amazon. I've been with Amazon about 18 years, which in Amazon terms is a long time. And I'm responsible for information security, the security of our customers' data, wherever it comes from, as well as the physical security of our facilities and our personnel.
1:10Prior to that, I was in the FBI where I concentrated on counterintelligence activities of primarily the Russian and Chinese intelligence services. That's fascinating. Was that obviously electronic intelligence? And were you running systems similar to what Amazon does for its cybersecurity? I mean, was that directly portable, that experience? So what we spent most of our time doing was understanding how our adversaries exploited systems, whether they are a national laboratory computer system or the defense department, et cetera, what are the tools that they use, the techniques? How could we figure out what they were after?
2:00And whether a particular activity was just something with straight network exploitation or might have been human assisted, because obviously there's a different set of work you have to do in those situations. Yeah. And actually, that leads me to my first question before you jump into what Amazon does. I heard you speak at a conference called HumanX in Las Vegas, and you talked about running this massive network of nodes that you described as MadPot. It was HoneyPot, as they say in the cybersecurity industry, that I presume exhibited vulnerabilities that would attract cyber criminals. and you were talking about how you had built AI tools that could track the initial type of exploit and all the way back to the probable actor according to some database.
3:08Can you talk about that first? That was really interesting. So Madpot is a threat intelligence sensor network that we operate. It's tens of thousands of nodes that operate across the world. And their job, the individual node's job, is to be a honeypot. A honeypot is a system made to look like a customer system so that our adversaries out there try and connect to it, try and exploit it. And we get to gather intelligence about who they are, where they are, the kind of tools that they're using, what they're after, et cetera. And the nice thing about Amazon's scale is it really does give us visibility into the entirety of the internet so we can see what's going on everywhere around the world.
3:52Now, some interesting facts about the MadPod infrastructure. It takes about 90 seconds, 9-0 seconds, for one of our honeypots between when it comes online to when our adversary first tries to discover it. Wow. It's incredibly quickly. That gives you an idea. When people are talking about security on the internet, they're like, well, I'll just put this thing up there. It won't be up for long. People won't find it. No, you've got 90 seconds before an adversary discovers something. And the pivot from when they discover it to when they try and exploit it is incredibly sure as well. It's less than three minutes.
4:29So discovery to exploitation, three minutes. what we needed to do was to build a set of software and services that allowed us to digest all of the signals from these honeypots and to turn it into actionable intelligence. What I mean by that is you have to take all the logs from all these honeypots, put them together, build a model of what our adversaries are doing. And then in our case, we use some of our foundation models, our AI models, to take that raw information and turn it into indications of who is doing what using what tools, where in the world, focusing on what customers. Now, AI is not perfect.
5:08AI is not the solution to everything. But what it can do is it can dramatically reduce the amount of sort of grunt work that our security engineers have to do and to give them much more polished intelligence out the other end so that they can make good decisions about how we're going to protect people. To close that circle all the way, you've got sensors, you understand what's going on, you produce actionable intelligence out the other end of it. The question is, what do you do with that intelligence? What we've chosen to do is we use it internally, for sure, to protect ourselves, but we also build that into the services that our customers consume.
5:44So if you're using, for example, Amazon GuardDuty and you're an AWS customer, you get to take advantage of the threat intelligence that that MadPod infrastructure produces and is refined using those AI tools. Yeah. Okay. Can we drill down a little bit on that for listeners? Because, you know, I've had many episodes on cybersecurity and it's usually at a 40 ,000 foot level and, or it's deep in the, you know, the data stock and whether you're, you know, guarding what goes in and out of a database and things. But in this case, can you describe what kind of a vulnerability, as much as you can, you would put out there?
6:32And what kind of an exploit? Just give sort of a use case or an example so that people can really visualize it. Certainly. So there are two major classes of vulnerabilities that our sensor infrastructure emulates. The first is sort of the oldies and the goodies. The big ones out there that we know are adversaries that exploited in the past, maybe an old version of a database engine or a vulnerable web server component or things like that. And that's to look for the kind of people who are doing mass exploitation of infrastructure around the world. Those, of course, we build into rules in the firewall services that we've done so that customers can automatically just say, throw that traffic away.
7:16We don't need it. It's garbage. The other piece that we do, the second major tranche, is emerging things, new things that are coming out. So if a very interesting new vulnerability was just discovered in a pipeline engine that takes raw text and turns it into something valuable out the other end, we'll build an emulator for that, install it in the sensors, so that we can look for people who are trying to exploit the newest stuff, the most interesting things out there. And then lastly, there are situations where we'll build sensors which have no vulnerabilities that we know of, and we will look for people trying to exploit them because the people who are building the zero days who know something about the backend engines are always interesting to anybody who's collecting threat intelligence.
8:05Yeah, and these vulnerabilities are on servers, not necessarily on public-facing websites that were tied to public-facing websites. So is that right? I mean, these cyber criminals are scanning servers. They're not looking at Amazon webpages for vulnerabilities. Is that right? Lots of people try and look at Amazon webpages for vulnerabilities. It would be foolish of us to think that didn't happen. We know that it happens periodically because we see people attempting things. But yes, for what we're talking about here with MadPod, these are servers that we put out on the internet with the express intent of having them attempted to be exploited by an adversary.
8:54Yeah. In that 90 seconds, the adversaries are using AI tools to scan servers. Is that right? Looking for opportunities. of these? What we've seen with AI tooling is that there isn't a great deal of tooling focused on the exploitation of servers yet. What we do see quite a bit in terms of the use of AI by adversaries tends to be focused on the things where it used to require a human to take an action. So for example, building a very convincing phishing site or building an email that you can use to target somebody for something. Quite often, that took a lot of skill to make the language sound right, to make the imaging layouts work well, etc.
9:40AI tools can help adversaries succeed more quickly in that space. Similarly, AI tooling can help less skilled adversaries take certain actions more readily or more quickly. What we don't see a lot of, however, is the really higher-end adversaries taking advantage of AI to do things differently than they have before. Because remember, AI is not necessarily inventing something new. It's repeating a pattern that it's learned previously, largely, in training material. So it's hard to completely invent something new in that circumstance. But it is much easier for AI to build things more rapidly for adversaries than they could have before right the uh this relates to an ongoing conversation in the community the debate uh about open source and closed source and and with regards to cyber security closed source closed source models uh one of the drawbacks is that uh their security profile is not transparent uh so you have to trust the vendor but then again the vendor has a lot invested and presumably they're spending a lot of time on protecting the system on on the open source side people say well you've given away uh in some cases the source code uh and and that's going to give an advantage to bad actors but there's this worldwide community that's focused on the source code that can quickly find and fix vulnerabilities how do you see that debate or is that too close to what amazon does to talk about we've seen both the use of closed source and open source software writ large for many decades it's something that we do internally we use closed source software we use open source software it's all a question of how you implement it, what you're using it for, and how you monitor it.
11:56When you look at Amazon Bedrock, which is our service that allows customers to gain access to foundation models, we offer closed source models and we offer open source models because customers want choice. They want the ability to choose something that's right for their job, for their particular activity. I mean, some of the models are better at generating code while other models may be better at generating images, for example. Some are better when I can tune it myself because it's open source, but some are better because it's just so well tuned from the start by the creator of the model. So my opinion is use RITs right for your job.
12:34Pick the one that's appropriate for you, then build guardrails around it so that you know that your use of that model is safe and appropriate. If you're in an area, for example, that has high regulatory oversight, you had better build in visibility so that you can ensure when a regulator comes to you and says, hey, prove you're using this thing appropriately and safely, that you've got the logs that allow you to do that. And those are the areas where customers tend to trip is in, oh gosh, I didn't think that once I build this into my workflow, I'm going to have to prove what it did out the other end.
13:08Yeah.
13:12All of this data that you're collecting through your Madpot is fed into GuardDuty. Can you talk about GuardDuty? That's a product that you use internally, but is also available to Amazon customers. Certainly. So Amazon GuardDuty is a service that allows customers to build protections around their use of AWS and the apps that run on top of AWS. Think of it as a combination of a firewall orchestration, an alerting system, a SIM, et cetera, all put together in one place that allows customers to say, hey, is there anything weird going on with my use of AWS infrastructures and the apps that run on it?
13:55It alarms, and more importantly, it offers recommendations on how to fix the potential problem. So it'll say, for example, that hey you've got this particular database running on aws did you intend to expose it to the internet was that what you meant to do or should you lock this down that sort of thing yeah does that pick up for example a few years ago i was working with a guy that writing an article about uh i can't remember the name of the system but it's a medical uh imaging uh system that stores and shares medical images. And a lot of these places had left their S3 buckets open, and this guy was showing how he could scan the Internet for this particular opportunity and then go in and look at the images stored in that system.
14:57And a lot of it was lung scans of COVID patients, and you could get their PII there. You could identify who they were. I'm not sure that anyone was actually exploiting it, but it was pretty fascinating. Is that the sort of thing that GuardDuty would immediately flag, hey, you've left the door open? It will. So it's important to note that every Amazon service is built secure from the start. So for example, when you use S3, there is no public access to the content by default. That is the way it comes out of the box, as it were. Customers can choose to open that bucket to the internet if they want, but it is an intentional choice that they make.
15:42More importantly, though, there is a switch, a configuration that customers can set, which says block public access. And that means no matter what I do, don't allow this stuff to go out on the public internet. And that gives customers a sort of a hard fail safe in there. Now, Now, GuardDuty does detect open S3 buckets, and it will alert customers to that if they make a configuration change. It'll say, do you mean to do this? And customers then get an alert to do that. Lastly, though, and this is something that we haven't talked about a lot, is we have an internal program that we call Active Defense.
16:19Say again, Active? Active Defense. And so, for example, we look for people who are like the person you're talking about scanning for open S3 buckets, and we will intentionally mislead them. So we will allow them to scan our infrastructure. We will provide them with answers which are not correct about which buckets are open and which buckets aren't. and that has removed it's about 99.96 99.97 depends at the the day of the week the month etc of the scanners on the internet access to s3 buckets so basically what we do is we lie back to the the people who are doing the bucket scanning saying nope not there sorry go away or oh yeah there's an open bucket then they try and access and it doesn't work So it's a way to mess with them.
17:10Wow. You know, I jumped right in on this. I want to talk about alarm dashboards and agents, which you talked about at the conference. But I want to hear, before I do that, can you go back and just talk about, on a broader level, what would you do to protect Amazon infrastructure and how that then flows to the protection of customers? Certainly. So our job on my team is the protection of customer data, wherever that data comes from. And to do that, we have to protect our infrastructure. And protecting our infrastructure is what a lot of people think in designing secure systems, deploying them appropriately, testing them regularly, ensuring we got the configurations right.
18:03But it's also the process of making sure the access that our authorized humans have, our engineers, is appropriate for their job and confined to the minimum set of access as possible. That's an area that a lot of people in business often fail to think about adequately. You start off saying, I'll just give everybody root access to this because it's an easy way to get things done. But then you have a large number of people who have access to data, which is very sensitive. Now, our adversaries are intelligent. If they can't break in the front door, they're going to try and compromise one of your staff who have authorized access to that information.
18:40If you look at the Chinese intelligence services as an example, that is one of their favorite things to do. If you look at the way the Russians attacked some of the other big service providers out there recently and succeeded in doing so, they were going after that authorized access of people to data. So we have some very aggressive programs that we've run for a long time on scoping that access down as much as we possibly can. And it's something that I recommend strongly to our customers out there that they do the same thing. You know, if Steve doesn't need access to all of this data, just remove it because he can't be exploited for the purposes of access at that point.
19:19And deciding, first of all, scanning who has access and deciding whether or not to close access to something. Do you use AI tools for that? Because, again, in a large organization, that's a massive undertaking for human actors. It is something where we use tooling to look at who has access. What we do, for example, is we'll build models of team or human behavior. So here is Steve. He is a software development engineer. He is working on this team on this particular project. What do his peers have access to? Is his access similar to his peers or is it different? It may be completely appropriate that it's different.
20:01But what I want to do is have a high-judgment individual look at that difference and say, yes, it is appropriate in this circumstance. Now, if you're a small organization, you can do that kind of work using human eyeballs. At our size, you can't. You have to do it using tools. Yeah. You talked about agents. I think at the point when you gave the talk, you were citing an accuracy or dependability figure of 65 % for agents, which was too low to deploy them on sensitive actions or workflows. But that is changing quickly. I think it's up to 80 % now. Now, how long do you think before agents are reliable enough to give them autonomy or increase their autonomy in, for example, closing loopholes in a security system?
21:12I think it's going to be very situationally specific, meaning the individual agent, the individual action, and the situation around that in terms of when the accuracy will be good enough. What I expect to happen is a progression. So you get agents that are getting better, more accurate, more complete, et cetera. And you have to look at both the accuracy of the action and the false negative rate. Is it doing what it needs to when it's appropriate to do so? Is that action correct for the circumstances? When both of those are at the correct level of completeness, we'll see us move to a okay double check situation, meaning a human looks at the actions, you know, the AI proposes, let's do all these things because of these problems with this outcome set expected, etc.
22:00And the human says, yeah, yeah, yeah, yeah, that looks good, fire. And we'll see that work for a while, probably a couple of years at least. Now, there will be basic things that we can peel off of that and make more frequently automated without human involvement over time because they're lower risk. So as you determine, you know, this particular action has a low failure probability and the outcome of a failure is not catastrophic. Therefore, allow the agent to take the action because we can always go and undo it later if we have to. You'll also see people build tooling, which allows them to check the efficacy of agents' actions and automatically roll back if there is a problem in the action.
22:47And as that relates to alarms, alarm dashboards, this has been a problem for a very long time that cybersecurity systems have all these alarms. Yeah. And the engineers are constantly, they have all these blinking lights or beeping. In the worst case, they're getting calls in the middle of the night. How do you guys reduce that alarm exhaustion right now? And do you see agents playing a role? I mean, I can imagine a day where an agent is monitoring the alarms and deciding which ones are false and which ones are easily resolvable and doing that autonomously. I think actually that'll be the first set of stuff that we see accomplished completely using agents is the initial triage process.
23:50Meaning, all right, this one's sort of the more important stuff that I got to go elevate to a human right now. This set of stuff can wait for a minute. it. This set of stuff I'll just take care of as an agent down at the bottom. That's going to be a very logical thing for agents to do as agents become more and more accurate and people can implement them more easily. Then they'll move further up the stack of just saying that that middle tier will let the agent handle it completely. I don't foresee a situation anywhere in the near future, meaning five years, 10 years, where we will let agents run completely alone in an infrastructure like ours the consequences are simply too high to do that if something goes wrong now there are a lot of different kinds of businesses out there there are a lot of different ways to do things it's quite possible that at some point for example um there may be situations where people are saying you know what this is my coffee shop this is my bookstore this is whatever i'm completely okay letting an ai agent run security for me why because it's better than anything else I can do myself because I don't have security staff.
24:57That's where we're going to get a lot of big wins is on empowering small and medium businesses to have more skilled security staff in the form of an AI agent than they could ever have otherwise. Yeah. And is that being, or do you expect that to be productized by Amazon? So a small business, instead of having to hire a cybersecurity engineer, can implement an agent to handle that. I mean, they're not particularly high-stake enterprises, I mean, in the broader scheme, but they are for the individual business owner. Is that something that Amazon is providing or will provide? Yeah, I would say that it's super high stakes, actually, for an individual business owner.
25:52You know, it is literally their livelihood in a lot of circumstances. And yes, we build into our services the ability to take actions so that we can help you protect yourself. We can help you scale your business as you need to. We can help you save money by de-scaling your business when you don't need to be spending money. And all of those can be driven by agents. We will see more and more of them be driven by agents as the agents get more accurate and more complete. Yeah. What tools are you guys using to build agents, presumably their Amazon tools? We use a lot of our own stuff internally. The Amazon Q stuff and the Nova models behind are pieces that we choose to use because it gives us a lot of control over how we build things, what we build.
26:39That being said, we'll use a lot of different kinds of models as appropriate, given the circumstances that we're working in. Yeah. More generally, generative AI has increased the attack surface, as they say. I mean, people don't yet quite understand even how generative AI works internally. internally uh can you talk about that attack surface and and how much you guys have covered that attack surface at least with within amazon sure and let's talk about what our customers can do whether it's our internal customers who are amazonians or it's somebody on the outside i think that the questions that they need to ask as part of the development process in using ai are the same Number one, where is our data?
27:39Now, business teams are sending data to an LLM for processing, whether it's for training to help build and customize the model or through queries of that model. So how is that data handled throughout that workflow? How is it secured? That kind of info is really critical to understand. Anybody who's a user, a company, ourselves, et cetera, need to be confident that their data is secured, that it remains confidential, and they understand if the model provider will be able to access or use that data for any other purposes. Number two, what happens with my query and any associated data? So training data isn't the only sensitive material you need to be concerned about.
28:19When users start to embrace generative AI, and I'm sure you've done this yourself, they quickly learn what makes an effective query. They start adding more details and specific requirements because that leads to better results. if your user queries an ai engine is the output of that query and the user's reaction to the results used to train the model further what about that file that the user submitted as part of the query you need to be able to question answer questions about how the lm provider is going to use that data and if you're comfortable with it the query itself can be sensitive and should be part of the data protection plan however you implement it there's a lot that you can infer from a question that a user asks.
29:02And then going back to something we touched on earlier, is the output of these models accurate enough? The quality of the output from models is steadily improving. And security teams, as we talked earlier, can use generative AI as one of a group of tools to address challenges. But from the security perspective, it's the use case that defines the relative risk. If you're using LLMs to generate custom code, is that code well-written? Is it free of vulnerability? and does it follow the best practices that you've implemented for your company? Do you have the proper double checks in place before that code goes into production?
29:41And associated with that is, do you have the proper security controls and threat models in place for your particular use case? This is crucial and often overlooked, frankly. You have to start with a strong foundational security posture for anything you want to use for actual business. While novel generative AI-specific security risks like prompt injection get highlights and headlines, most of the security issues that we're likely to see are going to be because more standard security best practices were either hurried, rushed, or not even followed at all. So, for example, implement robust authentication and comprehensive logging is so important, whether you're doing something using AI or not, and associated with that is effective vulnerability management.
30:32These basics are essential, as most of the issues companies face with AI are really similar to traditional security challenges, just with different tools and in often different names for things. After establishing those foundational elements, then you can focus on threat modeling specific to generative AI. This involves identifying some unique risks, like prompt injection, where attackers can manipulate AI responses, and model poisoning, where training data is tampered with. So consider advanced AI-specific threats that are relevant to your particular use case. That might include model inversion attacks that attempt to reconstruct training data, or adversarial examples designed to fool AI systems.
31:21People have to develop and implement targeted strategies to mitigate these risks, like using differential privacy techniques. Or, of course, my all-favorite is conduct adversarial testing. Yeah, I want to talk about your program on adversarial testing and see how that's going. But first, you talked about where is your data going. how does i mean you know i i'm not uh concerned about my data particularly but but i'm uploading documents all the time to chat gpt or claude or perplexity one of the models available through written asking for a summary and yeah where does that document that i uploaded go does anyone really know.
Read the full transcript
32:12And the model providers sure do know. Yeah, that's right. If you read the terms of service for whatever you're using, it's often illuminating and people are often surprised. What do you mean you can use the content that I shared with you to make your stuff better? And by the way, we've all seen the attacks where adversaries can get models to reveal the data they're trained on. So I'm going to give my private stuff to some third party who's going to train their model on it. And there's a chance an adversary will be able to go and get that back later on. If you're asking for a summary of something that you got at a conference or is otherwise public material, et cetera, hey, awesome.
32:59No problem at all using whatever the services of your choice. If you're doing something that involves actual business value or proprietary information, customer data, or heaven forbid, PII or other regulated stuff, you have got to make sure that the terms of service for the service you're using are consistent with the requirements of that particular data set. Yeah. The other question, I mean, unanswerable, but can you trust the terms of service? I mean, there've been plenty of examples of corporations who are found to be using customer data outside the stated what's stated as allowable. I mean, how do you know?
33:53Yeah, well, all I can talk about is what Amazon does. What other people do with information, I don't know. But we have always been super rigorous with enforcement of the terms of service and our respect of the terms of service. AWS was founded on the very notion that AWS cannot get to customers' information. We literally build our systems so that it's physically impossible to do that. If you look at our Nitro hypervisor, for example, there's no interactive login to that the hypervisor. You can't SSH to it, which means I can't go do arbitrary things on it. It is designed so that even if a government comes to us with a court order and says, give us the content of this person's virtual machine, the answer is, sorry, we can't.
34:40It is not physically possible to do. Now, other people may do different things. I don't know. But that has always been foundational to the way we operate. Yeah. I mean, the one that comes to mind, and I'm sure you're very familiar with, are the coding assistants. I mean, it seems to me that's incredibly invaluable, incredibly valuable data, the code that people are writing that the coding assistant is helping complete. presumably the coding assistant, the model behind that, is gathering that data. I mean, it has to in order to act on it. Do you have concerns about that? Absolutely. So when you look at coding assistants, there are a bunch of specific controls that should be in place when you use a coding assistant.
35:41I'll talk about Amazon Q Business, which is our Q developer, sorry, which is our coding assistant internally. We treat the AI-generated content as untrusted input until it's validated and goes through testing, whether that's a syntax check or a business rule validation or automated reasoning checks for things that need that kind of level of scrutiny. And that multi-layered approach helps ensure the reliability and the factual accuracy of the AI output, which is crucial when these systems are making impactful decisions on building code for you. We implemented, I think it was about 50-ish security verification rules for input and output validation in Alexa Plus, for example.
36:26The rules act as guardrails, validating things before they enter the system and responses before they're delivered to customers. And speaking of code itself, for AI systems that generate code, we implement isolation protocols internally. And I think customers should think about that themselves. So all of our AI-generated code is executed in secure sandbox environments before it's given back to the developer. So if you're using QDeveloper to build your own code, we execute it in a sandbox first. This approach allows us to leverage AI capabilities and software development while also mitigating risks associated with potential harmful or buggy code and preventing those problems before they get into production.
37:09I think it's a crucial safeguard, especially in complex development pipelines. If I'm creating a hello world program, that's not a big deal. But if I'm doing something which interacts with or manipulates important information for my business, I think that's a necessary requirement. Yeah. the guys at Anthropic just came out with a fascinating paper where they have built a tool that can graph kind of like a functional MRI, which nodes in a network are being activated and when to respond to a query. And presumably, I mean, what we've heard is that models do not update their parameters. I mean, they may do it at test time temporarily, but it doesn't change the underlying parameters as they've been trained.
38:16But is there any, we're always discovering that these models internally are working in ways that we hadn't expected. Right. Is it possible that models are learning from inference queries? I don't know whether they learn from inference queries or not. I'm not the right person to answer that particular question. But I will say that one of the reasons that we do have an interstitial layer built in between the model and the consumer in Amazon Bedrock is to give us a place to enforce the guardrails that allow us to prevent certain kinds of unexpected behavior. I want to touch on something else that you brought up, which is interesting.
39:03Many people don't realize that most AI models are not deterministic. meaning they don't often produce exactly the same answer for the same question and a lot of people just don't realize that but it's so important when you're making impactful decisions using ai is especially in an area like security where i need to have the correct answer every time if that answer changes it may not be appropriate for me to take an action anymore so building in the safeguards that allow us to detect that kind of drift and responsiveness are as important as building the AI engine itself. Yeah. And on, on, um, knowing where your data goes, the other thing that I, uh, I'm curious about, uh, your, your people are already, and, and soon it'll be prevalent building networks of agents that talk to each other outside of human purview.
40:06and that requires those agents to be sharing data back and forth. How do you know that an agent isn't sharing data that it shouldn't share, or how do you know that data that's shared to another agent isn't being passed along somehow? Is there a way to track data through an agent network? There are ways, and it's individually specific, so you have to really get down to the details for your implementation. But observability and audit logging are something that we spend a lot of time implementing at Amazon. So we've implemented exhaustive logging mechanisms that capture every aspect of AI agent operations.
40:51That includes the initial prompt, the reasoning steps, tool invocations along the way, and then final outputs out the other end. The logging serves multiple purposes. Number one, it's essential for fixing any security issues. because it allows us for a thorough analysis of the decision-making process. It also helps us identify potential biases and give us a way to build a fix for that. And lastly, it maintains a clear audit trail for accountability purposes. Just to use an example there, in Amazon Bedrock, which is our fully managed service for building generative AI apps, we've implemented comprehensive observability features that include detailed logins of all API calls, all model invocations and the responses from those models.
41:39So Amazon builders, if it's an internal person or our customers on the outside, can then use CloudWatch to monitor these logs in real time to set up alerts for unusual patterns or behavior, to conduct analyses of their particular AI model's behavior after the fact. And that observability helps customers ensure their AI apps are functioning as they intended and allows for quick identification and resolution of issues. I wanted to ask about you had a challenge or I'm not sure what you call it, a contest of sort of red teaming activity with some big money attached. That was underway when I heard you talk.
42:28I don't know when it's going to end. Can you talk about that and whether it's reached a conclusion? Yeah, that's the Amazon Nova Trusted AI Challenge. It's an initiative where we're investing over$5 million with a bunch of universities around the globe. It's a competition intended to push the boundaries of secure and responsible innovation and generative AI. So we recognized a while ago that advancing the field of AI security requires collaboration and a bunch of fresh perspectives on things. That's why we're bringing together some of the brightest minds in academia to tackle these particular challenges head on.
43:08The challenge is intended to address a critical gap in AI research. So many academic institutions simply don't have the resources to conduct meaningful research in this field. We're changing that dynamic. We're focusing on a crucial area, which is improving the security of code generation by large language models. We've all heard about the poisoning of code that's being generated by adversaries who are trying to break into supply chains or modify the training material that's used. This is an area where AI is making huge strides, but also introducing new security challenges along the way. The challenge has a tournament style format that challenges the teams to create tools to identify security issues and techniques to make the coding models less susceptible to attack in the first place.
43:55We have 10 university teams, five are developing models, five are trying to break them. This is sort of like a high stakes chess match for AI security. We created a customized foundation model for coding that the defending teams were then asked to develop protections for, and the offensive teams would work to break. So along the way, each team gets$250 ,000 in sponsorship and AWS credits, and the top prizes are another$250 ,000 for the best teams in each category. The teams will compete in, I think it's four tournaments over the next six months. That's designed to really drive some rapid innovation and adaptability.
44:37and the adaptability is one of the things that I think is one of the key skills in the world of AI security because it is so fast moving. Personally, I love this approach versus just sort of funding academic research because it gives the students and the researchers real hands-on experience and the findings they produce can be used to directly improve real world security out the other end. The teams that we've seen being most successful are multidisciplinary. They're bringing together expertise in responsible AI and security, conversational AI, and automated code generation. It's the mix of perspectives that we believe is going to lead to some real breakthrough solutions there.
45:19Now, most interesting here is that all of the teams will publish their findings. This means that we're going to be contributing to the broader field of AI security around the world. By encouraging these really bright minds to think like adversaries, I also think we're helping to proactively address potential risks before they become real world problems. This is about us finding ways to shape the future of AI security, but also to build the next generation of AI security leaders and setting new benchmarks in the industry. I hope that we're going to see some of these folks come to work at Amazon afterwards.
45:54And that the final results of that, I mean, you said there are four, is it four cohorts or four benchmarks along the way? It is four sets of work, and there are five teams on each side. Okay. And when will the results be published? At the conclusion of each set of work or at the conclusion of the competition? So it's going to depend on the individual set of work, but I believe we're going to do a roll-up publication at the end. Yeah. Another thing that is, I'm sure a lot of people wonder, I mean, you were talking about, you know, people poisoning code generators so they, you know, they generate bad code or whatever.
46:58Why would someone do that? Well, if you think about what an adversary wants to do, they have an aim that's usually to give themselves an entry point into an infrastructure, or it may be to weaken a cryptographic function, or to disable an authentication or authorization check at some point. Breaking in using an exploit is something that's relatively noisy. An alarm may go off. Someone may see it. If you're a really, really competent adversary, what you want to do is you want to get into that infrastructure behind the scenes where nobody notices. One of the most effective ways to do that is a supply chain attack.
47:38Way back in the beginning of building a system is to give yourself that entry point through a set of code that looks like it's normal, looks like it's appropriate, looks like it should be there, but takes an action for you. And there are some great examples of that in history where we've seen people who have changed cryptographic functions, where we've seen people who have changed authentication or authorization infrastructure so that when they send a magic packet, for example, to an interface, it automatically opens up and says, here, come on in. So a really skilled adversary will try and get into the supply chain of software long in advance of it being deployed to a system on the end.
48:19uh and these are state actors presumably because that's pretty pretty i think yeah and the the line is blurry between a state actor and a criminal in many areas especially in russia for example where the the intelligence services in russia don't particularly pay well and so a lot of those actors will use their skills and the tools that they've got from their day job to make money at night. But yes, traditionally these have been state actors because there's so much patience and time and investment that's required to do that. For criminals, they don't need to spend that. There's a lot of money to be had by just going after the lower level stuff that's easily available.
49:04And so two questions. One, there is some amazing open source models coming out of China and they're remarkably inexpensive compared certainly to proprietary models coming out of the U.S. But there is a trust issue. And do you see a reason for that trust issue, given what you just said? You know, I've spent most of my adult life in China and I I have no doubt that the leaders of some of these companies are good people, but there's a trust issue. And do you think that's valid, that trust issue? I think that companies and individuals have to make decisions about which models they use based on the risk factors that are appropriate for their particular circumstances.
50:06And we offer open source models from a variety of different places through Amazon Bedrock because there are customers who want that super inexpensive, very low overhead opportunity to use an AI model. And that may be appropriate for what they do. Internally, we're going to be very circumspect about the models that we choose to use to make sure that they meet our business requirements and the need for security guarantees that we can only get in certain ways. So I think this is a situation that it's dependent on the user. Some people may be fine with an open source model out of China. That's awesome.
50:44Some may choose to use something else instead. Yeah. And the other question that you may or may not be able to answer, but it's bothered me ever since, you know, the early days of viruses, computer viruses. is how many people are out there working as adversaries? I don't mean in red teams. I mean, you know, the people that you're trying to protect systems from. Are we talking about, you know, several hundred people spread across the world? Are you talking about tens of thousands of people spread across the world? do you have a sense of that at all? I think that you can stratify it. So there are probably very large numbers who are using AI to improve the way that they do things down at the bottom of the pyramid.
51:40Up at the top where they're really skilled actors, that's a very small number, simply because the skill set required for this particular work hasn't been around that long. And there are not many people who have it. There are not many people who are really good at generating these models around the world. So our adversaries are in the same position. And remember that they've had a little less time to use the models than the people who've been building them that being said the adversaries will always get better over time and the pyramid will probably change shape a little bit and flat out and some well that's interesting yeah so right now the the high level adversaries and people that you're uh talking out about that are in some way affiliated with a state are in the dozens or hundreds?
52:28I don't know. I can't give you a real number, but they're certainly not a large number. Yeah. And even if you look at all of the population of state actors, the population who are really good at AI is going to be a tiny percentage of that. Yeah. Well, you know, you've done a great job of answering my varied questions, but I didn't really give you a chance to give the talk that you gave, which was wonderful at the conference. Do you want to talk for a few minutes about what you want to talk about rather than? Actually, I think we've hit all the important points, really. You know, it's make sure when you're using models that you're doing so intentionally.
53:12It's understanding where your data goes, how it's handled, what happens with it. Is the output of the models accurate enough? And do you have the right security controls in place? Those are the really important things that I wanted to get across to people in the process of having a conversation with them. Okay. Okay. Well, Steve, this has been wonderful.
From the publisher
Can Generative AI Be Secured? Amazon's Chief Security Officer Weighs In
In this episode of Eye on AI, Amazon's Chief Security Officer Stephen Schmidt pulls back the curtain on how Amazon is using AI-powered cybersecurity to defend against real-world threats. From global honeypots to intelligent alarm systems and secure AI agent networks, Steve shares never-before-heard details on how Amazon is protecting both its infrastructure and your data in the age of generative AI.
We dive deep into:
-
Amazon's MadPot honeypot network and how it tracks adversaries in 90 seconds
-
The role of AI in threat detection, alarm triage, and code validation
-
Why open-source vs. closed-source models are a real security debate
-
The critical need for data privacy, secure LLM usage, and agent oversight
-
Amazon's $5M+ Nova Trusted AI Challenge to battle adversarial code generation
Whether you're building AI tools, deploying models at scale, or just want to understand how the future of cybersecurity is evolving—this episode is a must-listen.
Don’t forget to like, subscribe, and turn on notifications to stay updated on the latest in AI, security, and innovation.
Stay Updated:
Craig Smith on X:https://x.com/craigss
Eye on A.I. on X: https://x.com/EyeOn_AI
(00:00) Preview
(00:52) Stephen Schmidt’s Role and Background at Amazon
(02:11) Inside Amazon's Global Honeypot Network (MadPot)
(05:26) How Amazon Shares Threat Intel Through GuardDuty
(08:06) Are Cybercriminals Using AI?
(10:28) Open Source vs Closed Source AI Security Debate
(13:09) What Is Amazon GuardDuty
(17:44) How Amazon Protects Customer Data at Scale
(20:18) Can Autonomous AI Agents Handle Security?
(25:14) How Amazon Empowers SMBs with Agent-Driven Security
(26:18) What Tools Power Amazon’s Security Agents?
(29:25) AI Security Basics
(35:34) Securing AI-Generated Code
(37:26) Are Models Learning from Our Queries?
(39:44) Risks of Agent-to-Agent Data Sharing
(42:08) Inside the $5M Nova Trusted AI Security Challenge
(47:01) Supply Chain Attacks and State Actor Tactics
(51:32) How Many True Adversaries Are Out There?
(53:04) What Everyone Needs to Know About AI Security




