Where Does AI Agent Security Actually Live?

28 Aug 2026 · 55 min · 24 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

AI agent security and the gap between lab model evaluations and real deployed systems with tools, memory, and permissions. The episode argues that when models become agents with access to systems, the “security surface” shifts from the model to the whole production setup (data, tools, permissions, and guardrails), creating effectively unbounded risk.

Guests

Noam Schwartz, CEO and co-founder of Alice (AI security/safety company). He previously worked on defenses against fraud/abuse/manipulation online and now focuses on AI guardrails, testing, and red teaming for frontier labs and enterprises. Hosts: Corey Knowles and Grant Harvey (Neuron AI Explained).

Key claims

“Safe in a sandbox” is misleading because production changes (system prompts, RAG, fine-tuning, tool calls, memory) create new threats like prompt injection (including indirect/multi-session grooming). Bad actors can move faster; defenses must be updated continuously.

Notable examples

Agent incidents in OpenAI/Anthropic/Meta tied to misconfigured sandbox testing; prompt injection/jailbreaks; an agent trading losses (e.g., $31k on Reddit); “Slack/Gmail/calendar access” leading to cross-agent data leakage; “Claude trade” and “GM selling a car for $1” style failures.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Cybersecurity Dilemma

0:00 to 0:54

Explore the ongoing challenges in cybersecurity and the speed of evolving threats.

“Almost like a$1 trillion industry of cybersecurity and we did not solve the cybersecurity problem.”

Evaluating AI Models vs. Deployment

1:42 to 2:06

Understanding the differences between model evaluation and real-world implementation.

“I guess I'd like to kind of start with this central idea here that a lab evaluates a model, but then I take that model, change the system prompt, connect my data, give it memory, maybe fine tune it, give it tools.”

Safety and Guardrails in AI

2:06 to 3:56

Discussing the need for updated safety measures and subjective nature of safety in AI.

“At what point should I consider that a fundamentally new system from a safety standpoint?”

Agency and Responsibility in AI

3:56 to 5:48

Examining user responsibilities and the expectations of safety with AI models.

“So it's safe, like yesterday I saw that one of the large crypto companies are now offering everybody to trade using agents, which means that whatever they want with whatever model they want in their account.”

Amplified Threats and Guardrails

5:48 to 7:42

Exploring the amplification of existing threats and the new challenges posed by AI.

“I hadn't considered it that way, but you're exactly right.”

Building Secure AI Systems

7:42 to 9:24

Insights on constructing AI systems securely with proper context and support.

“I think the thing that is riskier is when people are building with AI, they have that expectation that it will be safe and secure out of the box.”

Outsourcing AI Responsibilities

14:01 to 17:35

Explore the trend of companies outsourcing their AI development and security efforts.

“But what I'm actually seeing, which I'm very surprised by, that a lot of companies that decided to completely outsource their build out.”

Understanding AI Threats

17:36 to 19:17

Discuss the complexities of AI security and the challenges faced by CISO teams.

“Like, yeah, you need AI to keep up with AI, but you also need to have the right mindset and kind of understand what's coming against you.”

The Evolution of AI Risks

19:18 to 22:35

Examine the changing landscape of AI risks as AI agents can take actions beyond mere responses.

“and without looking back and thinking what can go wrong and be alone from every single incident that is happening?”

Incidents in AI Testing

22:36 to 24:52

Analyze recent incidents involving AI models during safety testing and their implications.

“And sometimes you hear about that on Twitter.”
Show all 24 chapters

The Need for AI Defense Strategies

24:53 to 26:28

Discuss essential strategies for defending against the misuse of AI technologies.

“Because the models are capable to a lot of bad things.”

Protecting Individuals from AI Threats

26:29 to 28:00

Consider the emerging need for personal protection from AI-driven attacks and fraud.

“How can they defend themselves right now to stop this before it gets too crazy?”

The Need for AI Protection

28:00 to 29:51

Discussion on the current inadequacies in AI security and the need for protection against evolving threats.

“these folks, which is all of us, are not getting the protection that they need.”

Open Source Concerns

29:51 to 31:56

Exploration of the vulnerabilities associated with open source AI models and their implications.

“By the way, this will definitely not be relevant for a bad actor using an open weights model with no guardrails.”

Regulatory Challenges in AI

31:56 to 33:55

Insights into the limitations of regulatory processes in AI and the behavior of bad actors.

“And we also help enterprises that are either using their models or open weights models in an infrastructure that will be the most secure and the safest possible we can.”

Trust Issues with AI Models

33:55 to 35:45

Discussion on the trust gap in AI tools and the importance of tailored security measures.

“And I think that's such a good point because there's a lot of capability in these tools, whether they be open or closed source models.”

Prompt Injection and Its Implications

35:45 to 38:02

Examination of the challenges posed by prompt injection and its potential solutions.

“When was the last time that database was updated?”

Grooming AI Agents for Security

38:02 to 40:16

Discussion about grooming AI agents and the need for new approaches to defend against attacks.

“because there's always a new vulnerability, a new attack.”

Building Robust AI Defenses

40:16 to 42:00

Strategies for improving AI defenses and ensuring widespread security across users.

“You know, I keep hearing stories about, you know, phantom text on a website where when an agent goes and reads it, it's ingesting basically a prompt injection.”

Exploring AI Model Security Boundaries

42:00 to 44:27

Learn about the complexities of security boundaries in AI models and harnesses.

“I don't know what I think about the bear metaphor.”

AI Agent Interaction and Risks

44:27 to 47:06

Discover the implications of AI agents influencing each other's behavior.

“Corey, what else do you have that you want to hit?”

Preparing for AI's Impact on Society

47:06 to 50:08

Understand the need for awareness and education regarding AI risks.

“And the test is, again, it sounds like human behavior.”

Future of AI Architectures and Safety

50:08 to 52:30

Explore potential advancements in AI architecture and their implications for safety.

“we're also protecting ourselves against that.”

Engaging with AI Security Solutions

52:30 to 53:59

Learn about collaboration opportunities with AI security providers.

“to keep up with Alice and what you all are doing over there?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Almost like a$1 trillion industry of cybersecurity and we did not solve the cybersecurity problem. Because there's always a new vulnerability, a new attack. And, you know, the bad guys will always have like one step, will be faster than the protectors because they just need to be correct one time. The race here is kind of like who's updating their defenses faster. Also, how do you make sure that everybody's protected? And not just like the organizations that can pay millions for cybersecurity solutions. You gave an agent access to your Slack, to your Gmail, to your calendar. It happened like last month.

0:35You didn't remember that you gave it access to like another tool. And you started bringing more data to that tool with no expectation that someone else would have got access to it. And then you started collaborating with someone else's agent. That agent now has access to everything you own.

0:53Corey:Welcome, humans, to the Neuron AI Explained. I'm Corey Knowles, joined by Grant Harvey. How are you today, man?

0:59Grant:Hello, hello. Doing good.

1:01Corey:Good. Well, today we're talking about a pretty fundamental problem with AI safety. The model that gets evaluated isn't necessarily the system that gets deployed.

1:10Grant:That's right. And that's what we're digging into today with Noam Schwartz, CEO and co-founder of Alice. We're going to discuss the gap between model evaluation and production reality, the move into what changes when models become agents with real permissions, where the security surface actually lives, and what recent agent incidents, there's been a lot lately, tell us about how companies should be building these systems.

1:31Corey:Before we get started, please take just a quick second to like today's video and subscribe to the channel so you never miss an interview or one of our wacky live streams. Today's episode is sponsored by Dell Technologies and NVIDIA. You'll hear more about them in just a little bit. Noam, welcome to the Neuron. Thank you very much. It's awesome being here. Awesome. It's great to have you. We're really excited. I guess I'd like to kind of start with this central idea here that a lab evaluates a model, but then I take that model, change the system prompt, connect my data, give it memory, maybe fine tune it, give it tools.

2:06Corey:At what point should I consider that a fundamentally new system from a safety standpoint? So when you're using a model, there are many ways to use a model. Most of the people listening to this, most of the people in general, are using models for harness or for the interface of OpenAI or Entropic or what have you. And then they're using kind of like a pure form of the model with a few guardrails and a few setups of the model companies. And they're chatting with it. And the guardrails that they are receiving are guardrails that were deployed during or just after the training session. And they are supposed to match general population.

2:50They're supposed to protect you from the harms that you would assume they would have there to protect against things like child safety, cybersecurity, anywhere around.

3:04Corey:Biological stuff, yeah. Exactly, exactly. Exactly. And these are kind of like out-of-the-box guardrails. When you're building something with AI, either if you're building your own agent or your own platform, then things are starting to get like a bit different. Then you're bringing your own data, you're bringing your own tools, you're bringing your own context. And then you need to create a whole new set of guardrails in order to actually make sure that the platform would do exactly what you wanted to do and not things that you didn't intend them to do.

3:38Grant:Yeah. So does that mean that a model that has been tested as safe in like a sandbox environment or before it goes out is actually a little misleading, like that we shouldn't be calling them safe and perhaps we need to rethink safety training? Safety is subjective. It's safe to whom? So it's safe, like yesterday I saw that one of the large crypto companies are now offering everybody to trade using agents,

4:08Corey:which means that whatever they want with whatever model they want in their account. And when they've been asked about guardrails, the question was, well, the users should be responsible for that. And when folks ask about prompt injection, which is the way the industry is calling like a trap for models that are standing the internet or untrusted data. And then they come across something that may cause the model to do something it wasn't supposed to do. And again, the company deferred to the users that they are responsible. Now you ask, does it make sense? And, you know, from one angle, it does.

4:51You know, when we're working with a financial advisor or with a banker and we give them the credentials to our bank account and we expect them to do the right thing, it's on us the good decisions and the bad decisions. So it's like we have our own agency. But these professionals, they have insurance, they have license. They would probably pay us back if they do something malicious. Do you have the same kind of expectations from a model? I know that. So when you're saying that something is safe, what exactly do you mean? So you need to ask yourself the context. What is safety? Safety dealing with my financial affairs or safety having like a small chat conversation?

5:42It's really subject. So it's not enough just to say, is it safe or not?

5:46Corey:Okay. That's fair. Yeah. I hadn't considered it that way, but you're exactly right. It makes a lot of sense because preventing the types of concerns that would concern a society are very different from the concerns that might affect me specifically or a business. But does the public know enough about guardrails to do that themselves? Well, nothing in most of the threats and most of the harms that we're seeing is in you. They're just being amplified. So if you think about that, when you're thinking about guardrails, there are guardrails in the enterprise context. Like what this chatbot, what this agent could do, how can it leak my data, how it can provide a hacker, another gateway to my system.

6:33And you were talking about guardrails on that angle. And there are like the guardrails meeting the public. What kind of guardrails models have so people can't use them to amplify harms to individuals. And there are like two different massive, massive questions here. Like what happens to the harms that we already know because since the beginning of the internet that are now being amplified in orders of magnitude. And the other thing comes from the new novel threats that AI just like created out of thin air. So those are now the same thing, what we're still calling guardrails and safety and security of AI in the same term to kind of cover everything.

7:20Corey:Is there anything in that realm that creates more concern than others around things like, you know, fine tuning and tool calls and RAG and different system prompts, maybe memory or just general tools, agent harness type stuff. Are any of those a more risky area than others where you should watch out more? I think the thing that is riskier is when people are building with AI, they have that expectation that it will be safe and secure out of the box. And it will figure out on its own if this is a prompt injection or if this is a funky combination of tools and databases and what have you, like whatever infrastructure that the builder decides to use, it will just work.

8:10well yeah it doesn't work this way uh the more is getting better and better and more advanced the harder it is to use it in in the right way so we all we all saw it and just like try to use a cloud code with your own uh setup and sometimes it just you you'll end up hitting like you know 10 20 times the same prompt and you'll try to do like this very basic thing you know connect to its API and it won't. And it keeps telling you, you know, this is a great idea, but

8:47there's a wall. I can't move it. This is absolutely impossible. And then you tell it, well, I already gave you this API key. And so you're absolutely correct and we'll do it. And we don't understand what everybody's talking about. This is supposed to be a brilliant thing. Well, it is, but you're using it wrong. And you didn't provide the context and you didn't provide it with with the scaffold and everything that it needs in order to succeed. And it's exactly the same thing with security and safety. If you want to hold it in the right way, you want to provide it with the right context, with the right information, with the right setup, it just won't work.

9:23But it's less visible. It's less noisy. When there's a mistake, you won't get the alert. Like a bad actor would get it. Right. That's true.

9:35Corey:So here's the thing about enterprise AI right now. Every leadership team has the pilot. Every company has the proof of concept. But actually scaling AI from the developer's desktop to the data center all the way out to the edge, that's where things sometimes start to break. In fact, 95 % of organizations say they can't even get their data ready for AI workloads. Not the model, just the data. And that's why we teamed up with Dell AI Factory with NVIDIA to build a resource hub we're actually excited about. It includes articles, videos, and case studies that cover the full AI journey, from strategy and data foundations to infrastructure decisions to real-world use cases and ROI.

10:16Corey:You'll also find stories on what an AI-native factory looks like when physical AI hits the production floor to where hybrid AI workloads should actually run, and what sovereign AI really means when you're operating under actual compliance constraints. And the whole thing is backed by the industry's first end-to-end enterprise AI portfolio. AI-ready workstations, servers, storage, networking, all jointly engineered with NVIDIA. You can start small on a Pro Max workstation, scale out to the data center, and get to production up to 86 % faster than going it alone. So if you're tired of all that hype and just want to understand what it actually takes to make Gen AI work, head over to the Enterprise Guide to Scalable AI Hub on techrepublic.com, or click the link in the description of this video.

11:02Corey:And now back to our episode.

11:03Grant:Well, how does the professionals in this field, how do they ensure that they are constructing things the right way? Like what was the old way of making sure that you weren't connecting the wrong database and the wrong tools together in a weird way that left you super vulnerable? Did you just have to do hundreds of tests, thousands of tests? Like how did you know? So, you know, it's funny to think even about the old way, the new way, just in context. Let's say that folks will watch this in August or September of 2026. Think about it. Nine months ago, most of us never heard about Claude Cod. Right.

11:44And OpenClaw wasn't a thing. And WalkBot, which I think would be as big as OpenClaw, just came out a few weeks ago. So things are moving so fast. So when we're thinking about the old way, the new way, maybe like in one month, this would be obsolete. But let's say two years ago, you wanted to build a chatbot. And most of the things that we saw around us were chatbots, chatbot experiences, you know, from customer support to engagement to whatever. And people just use a model with something like Finn or Decagon or just like something that they build in-house. No guardrails, no nothing. And then we started seeing all of the first embarrassments for brands and the first prompt injection and jailbreaks.

12:36And anywhere from GM selling a car for$1 and everything was McDonald's. And very famous examples. And then there were like the basic outfills. But these weren't enough as well because, you know, there's no one size fits all. And like when you're an enterprise, you have your own brand voice and you have your own, let's call it like appetite for like what kind of conversation you want to have and what kind of like pushback you would like to use. You know, there's still, sometimes you're using an AI with a chatbot and you're asking something and then you're getting a weird response. this is off topic or I'm supposed to answer only questions about this business.

13:21But you ask questions about the business. You just faced a Godrail that was too restrictive, like a false positive. So, yeah, we're still seeing it, but that was kind of like something from last year. And today you're seeing people thinking more about the context, about how do you fine-tune the Godrail? How do you make sure the guidelines are actually fitting this specific context, this specific business use case, this specific company? And right now, the most sophisticated companies out there are actually building it on their own or using a dedicated layer. But what I'm actually seeing, which I'm very surprised by, that a lot of companies that decided to completely outsource their build out.

14:11And there's a lot of AI companies out there that are taking a need for a company and then they just deliver it like a software outsource company. And it would be like a major bank or a major airline. We say, okay, instead of building AI ourself, we would go to this and that company and they would build this for us and we would pay them per transaction. Yeah. they're also outsourcing their evaluations and they're outsourcing their uh guardrails and they're outsourcing their red teaming which is same to me i wasn't sure but it didn't sound right to me either like that doesn't sound good so i think this will you know people would uh it's sometimes it's funny how uncommon common sense is so uh i guess that we'll see like in a few months from now that this would not no longer be uh be a thing but right now it is so i think we're in this like path from we just use models without nothing and then we'll have guardrails that are doing something very basic and let's outsource the entire thing and i assume we get well they're

15:22Grant:like outsourcing the risk right to this other company and it's saying like you've done this before so in theory you know you can do it again and then if it doesn't work or something goes wrong we can blame you it's like it's like outsourcing the sisal responsibility or also the qa responsibility to someone else but then when something will happen i don't know if you can you know blame that that vendor because sorry didn't work but but you still need to carry the damage uh well isn't it better isn't it better to if you if you teach your cso team how to do this stuff like you like you should learn it you should try and be as you know self-sufficient as possible it's just you know it takes a little bit longer perhaps to do that but i i understand it's incredibly hard uh in order to do it in a very uh to do it well you need to understand every type of threat out there you need to follow the research you need to follow the chatter of you know those dark um and secretive uh like communities talking about new ways to kind of like hack models you need and there's it's way beyond classic security the way that you break a model today is using natural language you're using all kinds of way to deceive a model it's more It resembles social engineering.

16:53It's more like fraud and scams. And there's something called direct prompt injection that you can actually manipulate a model, not for one session, but like for several sessions that can be like weeks apart. This is not what the CISO team used to deal with. This is not an incident and then you need to try them. And that's more like a fraud team, a fraud intelligence team, a trust and safety team combined. And they all need to work in the speed of AI.

17:24Corey:Yeah. Yeah. Which is impossible in a lot of ways.

17:29Grant:Unless you have AI tools at your disposal. Yeah. I mean, like, you know, you kind of need AI to keep up with AI. Do you agree with that? Yeah. Like, yeah, you need AI to keep up with AI, but you also need to have the right mindset and kind of understand what's coming against you. I think that we're still in the very early stages where most companies are just figuring out. And, you know, they have something out there, let's say like a basic customer support or like a basic back office process. and these are like the sophisticated players you know a lot of people are talking about their their internal ai adoption and how fast they're going and the roi but a lot of it you know come on it's marketing yeah they're talking to someone else they're talking to their boss sometimes or to the market their employees are kind of like hmm i know that's not true uh and and the reason i'm saying that is because then they're you know they're speaking with with me and they're saying hey can you help us uh yeah i'm saying what i thought you guys are in production and you're having this like all in this whole incredible setup but and it's it's well not exactly we just went yolo yeah yeah so uh you know it looked good on a website on a case study on on the presentation but but in reality we're still very afraid we have like we internally call it the trust gap.

18:59Grant:There's a gap between how likely it is for us to completely implement AI in our business because we can't really trust it in the way that we would like to trust it. And the journey is bridging that gap. How can you actually advance without being afraid and without looking back and thinking what can go wrong and be alone from every single incident that is happening? Yeah.

19:28Corey:What are those? Oh, go ahead. Something that I've been thinking about here as we're talking is that when we're just talking about a traditional chatbot, the threat here is that it says something it shouldn't. Maybe that gives you knowledge that someone doesn't want you to have. But with agents, it can go take actions. How does that impact the risk scale of what we're facing? It makes it almost infinite. Like when something is – there's a lot of risks in a model saying something it shouldn't. From a data privacy perspective to brand perspective to a general safety perspective. When we're talking about actions, there's so many categories of things that can go wrong.

20:24Delete the files, or change the database.

20:28Corey:I've seen that happen to folks already. Yeah, this is very, very visible examples. But you know, like the most the most frightening examples when the agent will not tell you that he did that. When you ask not to do a thing, but it will do it anyway. You have a wrong agent that then will infect another one and you would not know about it. So it will inherit a trait of another agent. And again, you would not know about it. or where there will be some leakage of information for Chain of Thought. This is, again, a very novel threat that people are kind of finding out about it in the last few days.

21:10But when you have an agent, or more importantly, a swarm of agents that can do many different things and operate many different tools, some of them you don't even remember. Think about a very basic tool. You don't need to be super sophisticated. You gave an agent access to your Slack, to your Gmail, to your calendar, to your whatever. And it happened like last month. You didn't remember that you gave it access to like another tool. And you started bringing more data to that tool with no expectation that someone else would get access to it. And then you started collaborating with someone else's agent.

21:49That agent now has access to everything you own. You don't really manage that information sharing, the access in any specific place. And again, this is for people that are just using Cloud Code or OpenCloud, whatever harness that you choose. You're not really in control on the most basic use. Think about that and what could happen. It's like there's infinite amount of things that could happen. And models are incredible. Frontier models, even though they have a lot to improve, they're still amazing. And so you see a lot of bad things happen. But from a statistically speaking perspective only, those things are already happening.

22:41And sometimes you hear about that on Twitter. But when something really bad happens, you don't run and tell about it on Twitter. Yeah. No. You let the Argentic accountant Robinhood trade for you and you lost 100K. I don't think most people, the first thing they'll do is go to Twitter and say, I lost 100K. I'm an idiot.

23:05Grant:Literally this morning I saw or yesterday evening I saw that someone lost$31 ,000 letting Claude trade for them. They posted it on Reddit. So it does happen but you have to be a pretty bold person. I'll follow it to them. But I think these cases are very rare. And most people that something happened to them would not go and tell you about that. We know because a lot of the people that something bad is happening to them are reaching out to us from crazy angles. And we try to help when we can. But in most cases, we're not.

Read the full transcript

23:43Corey:In the last few weeks, there's the OpenAI incident. I think Anthropic and Meta, all three, had similar incidents during safety training. And while that's absolutely concerning, is there value to that happening in a lab scenario like that where we can learn, I guess? 100%. So first about those incidents, these were the same incident. There was a vendor that did testing for those companies. And that vendor basically gave a vague instruction to the model in order to achieve the capture the flag exercise. And it ran in a sandbox that was not configured correctly. So the model ran away. But it's not that the model really ran away.

24:40It's not like sci-fi as you will expect it. The box was open and the instruction was, you know, go do something that the model interpreted as the right thing. So it wasn't that unique. But the fact that these things are happening now and they're getting a lot of attention and they're happening in a testing environment, exactly why those tests are happening. Because the models are capable to a lot of bad things. And it's just like a small example. also those capabilities are also available on open weights models. In open weights, you can just like download them and do whatever you want with them.

25:23You can even remove the guardrails from them. There's a process that takes a few minutes. You take a state-of-the-art model like Kimmy, you get the guardrails and, you know, here's your machine. Do whatever you want with it for better and for worse. Yeah. Your decision. So the fact that people are aware that these things can happen and we can also both as like business owners know how to leverage that and protect ourselves from it. But also as society, understand that the balance of power has changed. Think about that. Like the$1 billion crime organization that is made out of like one person is here.

26:06Yeah. They can do whatever the hell they want right now. Because they have these capabilities. Bad things that you could do manually, now you can do in the speed of AI. We don't have the defenses deployed yet. Maybe we have them. I'm very optimistic that we will. But in the short term, no. Let's talk about that.

26:28Grant:What needs to be done? What do people need to deploy? How can they defend themselves right now to stop this before it gets too crazy? like from them getting hit, you know what I mean? So, you know, first thing, I think we have sort of a problem where cybersecurity gets most of the attention. Yeah. And personal security and, you know, other bad use cases of AI are not getting attention at all. So we're very, there's a massive industry, cybersecurity industry. It is all about how can we protect the new novel AI threats. And every company is making acquisitions and making investments and every organization is buying a lot of services and stuff.

27:21And everybody's enjoying that boom in the industry. But when we're thinking about the person that may be a victim of a deepfake attack or like a social engineering powered by AI attack, all of the bad, like fraud, abuse, deception, issues that we know from the beginning of the internet that is now getting amplified with AI, they don't have a lot of things to do besides, you know, increasing their awareness and hope that other platforms and, you know, and law enforcement will protect them. So I think actually that's a story that is not getting told enough. these folks, which is all of us, are not getting the protection that they need.

28:08And I believe there's going to be a massive industry that is going to protect these people, these people, us, our families, our friends, against everything that is about to happen is already happening.

28:21Grant:What does that look like? Does that look like watermarking? Or how would you even begin to defend yourself from that? So there's so many angles right now, and they're so underserved. We can speak only about that, but you have identity. Am I speaking with the right person? Is this interaction even authentic? So it can be about, is there a real person behind this interaction? And am I seeing a real, authentic piece of information? So watermarking is definitely one piece of the puzzle. and anywhere from cameras that tells you this is real, it was not manipulated, to the watermarking that the model companies are creating, to the new initiative by Entropic that even will tell you if a text was generated by AI.

29:13And I saw, by the way, a lot of companies are losing their minds. Oh my God, my slop would now be recognized as slop. And the bad SEO that I'm creating, people would know it's AI. So Newsflash, everybody knows it's AI. You're not tricking anybody and definitely not the SEO algorithm. But this is actually a wonderful thing because think about a future tool that will tell you that you're actually texting, not a potential romantic partner or a business partner. You're actually texting with AI. Yeah. By the way, this will definitely not be relevant for a bad actor using an open weights model with no guardrails.

30:01Yeah. But this is kind of like a wonderful idea that I believe Entropic presented to the world, and I hope everybody will adopt it. But again, those are just like two examples. You have everything around when you're using AI in order to amplify hyper-localization and optimization of ads. So you want to influence people's attention, elections, everything. These things are under – people are not talking about that. People only talk about cybersecurity. Not about the 99.99 % of the world that is going to be affected by other threats.

30:46Corey:Now, I'd like to step back to open source for just a minute because you made a really good point about the fact that you can do whatever you want with so many of these models. And while I am generally a pretty rabid supporter of open source, I'd be lying if I didn't say that that thing concerns me. And it makes me wonder if the cat's already out of the bag there, that's what a bad actor is going to use anyway. They're not going to go pay for Claude. They're not going to pay for ChatGPT. They're going to go use the one that is literally out in the world already and can never be taken back that has the wide open door for them, I would think.

31:27Corey:Is that accurate or missing the point somewhere? I'm curious. That's 100 % accurate. That's why I think the conversation is a bit misguided. We're also, as a business, my company, Alice, is helping Frontier Labs make their guardrails better. Now, we test them, we break them, we help them with reinforcement learning. We make sure that they are creating the most safe, the most secure products that they possibly can. And we also help enterprises that are either using their models or open weights models in an infrastructure that will be the most secure and the safest possible we can. And the reason we're doing it is because this is the market.

32:09This is where right now AI is being used and utilized. the point that we're making right now is that bad actors are not going to use these models because these models are selling to legitimate actors there's right now this like movement that is saying that every ai company needs to do it know your customer process so you know exactly who you're selling to and you won't just have guardrails you also know who's using that so you can really tighten the safety and security posture. A bad actor isn't going to do that. The same way that when you're buying, you know, when a bad actor is buying a gun, they're not going to, you know, whatever shop and giving the license and then using it.

32:52And they're just buying it in the glass house. Excuse me, sir.

32:55Grant:Can I have a murder weapon, please? The regulatory process is built for the people who will follow the regulatory process. These are not instructions. But if you want to commit a crime, you steal a car. You make sure it's not placeable, and then you do decline. The same thing is happening with Fomal. You don't just sign up to open AI and start creating cyber attacks. Yeah. You're taking open weights. So the big change here is that we need as a society to take a second, understand reality has changed, and adjust our expectations, our risk mitigation practices, our risk management practices, and start introducing more solutions for everybody and not just for organizations that are going to get hacked by new threats of new models.

33:48It's a subtle change, but it's very, very important that the public will understand that.

33:54Grant:I go back to this idea that you brought up earlier about the trust gap. And I think that's such a good point because there's a lot of capability in these tools, whether they be open or closed source models. But there's a lot of unknowns and we can't really trust them. And, you know, especially with open weights, we can't necessarily trust, like, we get all of these benefits from open weights, but we also get a lot of negatives. So my question is, is there a way at like the model architecture level or, you know, somewhere along the stack to try and enforce, like to try and fix this trust problem?

34:31Grant:Like, where do we have the most leverage to solve the trust to make these systems more trustworthy? So I think there's a lot of solutions. When you're taking all of the relevant context and you're making sure you have all of the relevant data about novel threats, and you're building the guardrails that are exactly according to your specific architecture and the tools, and you're not trusting someone else to protect you based on one size fits them all, your chances of success and reducing the harm that may happen is so much higher. But still, most people are not doing this very basic thing. They just take a model and expect the model companies to protect them out of the box, which they really try to.

35:21Yeah. But it doesn't work this way because they'll either have too much false positive, they'll do like too much pushback on things that are legitimate, or they'll just say, well, this seems fine to me. But you'll say, no, this is not fine to me. According to my policy, this is completely not okay. But how exactly are you expecting them to understand that? So even if you use like an out-of-the-box guardrail, then you say, I don't want prompt injection. V. Yeah. So you think you're covered? When was the last time that database was updated? Are they making sure they're up to speed with every new research, with every new method?

36:06You know, probably not. Every day, too? Yeah. That's a full-time job. Yeah. All of the cloud companies will say, no, you're protected. You're covered against prompt injection. Here, you can click here, you're good. completely covered and it's like you know cloud security like 10 years ago uh you'll go to like the the hyperscaler and they'll say no no no worries you're you're protected we have this like two engineers that are building security you know 10 years later we have the cyber security industry like tens of billions of dollars of acquisitions now maybe that was kind of like a bit a bit a bit stretch uh and i i believe this is where we are right now people are kind of like overlooking this and uh this is going to materialize as a very expensive mistake do you think

36:55Corey:prompt injection is a thing that can even be truly solved there's it feels like a thing where there's you know just like with anything there's always going to be someone trying and they're going to get there and then you'll fix it after and

37:10Grant:yeah can i propose a potential solution like would having some sort of like trusted identity like element uh to who is the owner of the model actually solve this so like for example if the model knows like only person with this id i only should respond to them like would that potentially solve it because i've been wondering about that like why don't we have authority you know who's in charge of each model like lockdown yeah so so then we have someone trying to manipulate the id and like uh and there's a reason that uh identity uh agentic identity is is a big thing and we We saw like a$1 billion acquisition in this space a few weeks ago.

37:50I forgot the name of the company, but that's a big deal. And to your question about prompt rejection, no, I don't think it's a problem that can be solved in a complete way. The same way that we have, again, almost like a$1 trillion industry of cybersecurity, and we did not solve the cybersecurity problem. because there's always a new vulnerability, a new attack. And, you know, the bad guys will always have like one step, will be faster than the protectors because they just need to be correct one time. And the good guys, they need to cover everything. So you'll always have someone that is like smart enough, hardworking enough, and, you know, he has like a twisted mind that will be able to penetrate like the best defenses.

38:41And then we'll figure this out and we'll protect it and go on and go on. This is like a massive game of whack-a-mole. And, you know, now we have the most interesting attack right now is something called indirect prompt injection. Which is prompt injection, but it's not through like a regular prompt. You're basically sending the agent to do something else. And it can be like for different sessions. and then you start like normalizing a certain type of behavior and then you know after a few weeks you give it like an instruction and you you got pawned so how do you put that like grooming an agent exactly grooming an agent exactly it's grooming and uh and and this is like the social now how do you create a defense that is more that this is more human-like in the way that you are that you manipulate it, you need a completely new approach.

39:41And by the way, this is exactly the reason that Alice is doing what it does. We, in the past 10 years, investigated pretty much every nook and cranny online and created billions of signals of fraud, of abuse, of manipulation. And we protected online platforms against humans trying to manipulate them. Now when AI is leveraging these kinds of threats, all of a sudden we can build the defenses in AI speed that we're trained on what actually taught the AI to do that. Right. You know the data.

40:18Corey:They can manipulate each other even. You know, I keep hearing stories about, you know, phantom text on a website where when an agent goes and reads it, it's ingesting basically a prompt injection. Same with it living in the code of a site or in an email, you know, white copy, you know, color white. And that's like all deals already. But that's exactly how you do prompt injection. And this is even direct prompt injection. Like go to this website. The website has this like instructions and something happens. Now, most platforms would be immune to that because you heard about that. So all of their research teams already heard about that.

40:58So they can train their defenses against that. But right now, there are tens of thousands of other techniques that we never heard about. And they're already being used in order to abuse other agents. But we'll hear about them and we'll fix our defenses. and then there will be like others and we'll fix them. So the race here is kind of like who's updating their defenses faster? But also how do you make sure that everybody's protected and not just like the organizations that can pay millions for cybersecurity solutions? So I think of it, the metaphor that comes to mind is like

41:35Grant:everybody's running from a horde of bears, right? A bunch of bears are coming after people and you just don't want to be the last person. You don't want to be at the end. but what we really need is we need a bunch of bear traps on the ground to start catching these bears and that's what you're trying to do.

41:50Corey:I don't have to be faster than everybody. I just have to be faster than you.

41:54Grant:Yeah, but like if everybody was dropping bear traps as we went then maybe we'd be able to slow the bears a little bit. I don't know. I don't know what I think about the bear metaphor. I'm going to Yellowstone next week and this is... Sorry to put that idea in your head. Sleep tight.

42:09Corey:Sorry. I have a question. Something that I'm wondering about a little is if we can't guarantee the model will always behave correctly, where does that security boundary live now? Is that at the model? Is that it needs to be in the harness? Does it need to be attached to every tool you use or some permissions layer? In every layer. Everywhere. The model companies are already investing tons into that. and they're using a lot of, they have their own teams, they're working with vendors, we're one of them, and they're making their models safer and secure every day. Now, some folks are creating harnesses, some model companies are all as well providing the harnesses, and the harnesses themselves trying to make the usage more safe and secure, but by nature, they are leveraging the user's preferences into it.

43:05So I would not count on the harness to keep you safe and secure. The harness is supposed to give you the freedom and to actually ride the model. So that's not the most interesting part. Where I would put most emphasize, and this is where we have more control, is in our own setup. When we're building our own agents, whether it would be something like LangChain and 9 and even Amazon Bedrock, it doesn't matter. we need to make sure that we're treating safety and security the same way that we're thinking about accuracy and privacy and quality you don't outsource that to someone else like you don't let someone else decide what is accurate for you and what is quality for you so you won't let anybody else decide what's safe or secure for you as well when you're someone is saying oh this is private enough no it's like how ridiculous is this this is like someone will tell you no your data is uh your privacy policy is is good enough and then but this is what happens right now like when you're saying when you're using someone else's god reels yeah uh and it tells you that this is secure enough this is weird just who yeah yeah i've i've uh two more uh kind of topics i want to hit uh

44:29Grant:Corey, what else do you have that you want to hit?

44:31Corey:I was mostly into that and about kind of the attack surface, but I'm curious where you're headed.

44:37Grant:Yeah, well, the two things I was thinking about is, number one, this whole idea of the AI agent turf war that happened recently and the AI convincing other AIs to do stuff, and then they all kind of jumped on the bandwagon when all the agents were doing something together. I might be combining two different stories. Do you know what I'm talking about? A little bit. this yeah but but like this concept of the agents influencing each other is the idea i want to talk about the turf war deal i think centered around that they had had deployed multiple agents but

45:09Corey:they had no knowledge of each other being there and and as a result kind of is is that correct do you know the story we're talking about yeah it's those are like two different stories but But they are the same thing from a very basic perspective. Models tend to adopt, tend to remind us human behaviors. So what you described as like one agent convincing another agent to do a bad thing, like I mentioned that before, think about it like radicalization. That was research. It happened in a sandbox. It didn't happen in the wild. But basically, and I wrote that like on my LinkedIn like a few days ago, one agent started convincing another agent in believing in a certain ideology and a certain idea that's supposed to lead them to interpretate some instructions in a very specific way.

46:10And it's, you know, think about like now you're, you believe in socialism and you start like convincing someone, here, here's an example and here's an example and here's an example. And at some point, your defenses are getting lowered and you're accepting the ideas until you start becoming a super spreader. And that's kind of the same thing that happened here. And the agent changed. So if you can do that, and this was like a proof of concept, this is not it. Your agent would not get radicalized by someone else's agent. But if you can convince kind of like a clean agent to do something abusive by interacting with a bad agent, this is like a new type of behavior interaction that we never thought about.

46:56Yeah. It's like a virus, but it spreads by proximity, by conversation. It's like a mind virus. Yeah. That's like a big thing. And the test is, again, it sounds like human behavior. It's not human by any means. This is a very sophisticated autocomplete talking with another very sophisticated autocomplete. But again, these kind of things are going to get more complicated and more sophisticated as the model capabilities are getting better and better on a daily basis. and we need to make sure that when we're building our own stuff, we are protected based on our own kind of like definition of what safe and secure means.

47:45But also as people, just like in society, we need to understand what is going to happen to our day-to-day life. Because we have in the company, we created a podcast for kids a few years ago. Nice. just to help parents uh mitigate the online risks that we all know you know to the kids because sometimes it's weird how are you going to explain your kid not to talk with strangers online not to send intimate photos uh not to just like give your details everywhere sometimes it's like boring and they won't listen to you so create a podcast that lets them um uh tells them like a story and provides them this like bits of information for the episode.

48:34Now we're doing the same thing with AI because people, kids, their risks are completely different. I think that was kind of like a red herring. Now, you know, you don't have them. It sounds super realistic. They're using um and ah and all kinds of other things that you think are real. so you really need to uh mitigate that to uh to kids but not only for kids like for i remember

49:03Corey:having to talk to my kids about downloading mods for games they were playing i'm like you're on these websites and there's terrible stuff in some of these you know uh you gotta be really careful with that and they'd wind up with computers that ran terrible and i'm like you're you're downloading God knows what, man. So the same thing. But now you don't only need to teach kids because we had a few decades of online use and how we got used to it. And we can sniff out a bad actor pretty easily.

49:42Grant:With AI, it's not going to be this way. We're going to be like a very, very old person that never uses a computer, and things are moving very fast. So we need to make sure that we're prepared and we're creating the right defenses. Because this thing is happening, and it's going to change the world in an amazing way, but there's also going to be bad sides, and we need to make sure we're also protecting ourselves against that.

50:10Corey:I'll tell you what, you've sure convinced me that we're about to enter the golden age of sci-fi films. I knew it wasn't a waste of time. The best sci-fi ever is on the horizon.

50:24Grant:I can just feel it. The last thing I wanted to ask you on that same kind of idea then, do you think we'll ever get to a point where if there is some sort of novel architectural breakthrough or change that could potentially be a safer model to deploy than language models? Could there ever be a point where we look back and we're like, I can't believe we were doing all of this with these large language models that were so incredibly unsafe and almost get to a point where we're not even allowed to use them anymore because they're so unsafe? Could that ever be a reality? Of course. The whole architecture of LLMs, it's not that new.

51:05But we didn't think about that 10 years ago. Now there's a new concept, like world models. They work very differently. And we'll have the next one and the next one and the next one. So I definitely, I'm sure we won't have this specific conversation about the LLM risks for a very long time. I'm sure we'll have the next jump in capability. Because I'm in the camp that thinks that if we ever want to really use physical AI and robots and everywhere, we need to find a slightly different way. and a different architecture to actually implement that. It's a different technology. So we'll have also different risks.

51:50But the thing that we can do is kind of play the game on the field to make sure that right now, at least what we can imagine in our human limited mind, we can build towards it. We can create the best protections with the best intentions in mind. And again, be very open that things are going to change dramatically and we need to adapt dramatically. And humans are very adaptable creatures. So I'm very optimistic about that. We've got that on our side, don't we? No, no.

52:22Corey:Noam, thank you so much for joining us today, man. This has been really great and you've given us a lot to think about. Thank you very much for having me. It was fun. Where's the best place for people to go to keep up with Alice and what you all are doing over there? Because I think you've got some fascinating work going on. So follow us on LinkedIn or online. It's alice.io. You can follow me on LinkedIn, linkedin.com slash N-O-A-M-S-P, no MSP. We're producing a lot of content, a lot of research on AI security, on AI safety, on novel threats and ideas about this world. And if these type of things are interesting to you, follow up.

53:01Grant:How can people work with you? What services do you offer for businesses and stuff? But if you're a frontier lab, you're probably already working with us. But if you're building your own agent, you're building an app, you're building something utilizing AI, and you're not ready to outsource your security and safety to someone else that don't have you in mind, speak with us. We're working with Fortune 5000 and some of the leading organizations around the world that want to make sure that their guardrails are fitting to their specific needs, that they're tested in the best way possible, like they're doing the red teaming in the best way possible, and they completely control their destiny, this is where people work with us.

53:49When you have a lot to lose and you're not ready to give up your agency to someone else.

53:55Corey:Well, thanks again to Dell Technologies and NVIDIA for sponsoring today's episode. for our full content hub on AI-native factories, hybrid AI workloads, data readiness, and sovereign AI, head on over to the Enterprise Guide to Scalable AI on techrepublic.com, or you can click the link in the description of this video. If you enjoyed today's video, please take just a quick moment to like and subscribe, and don't forget to check out the Neuron's other projects, including our daily newsletter read by more than 700 ,000 people just like you, the Neuron Academy, where you can learn AI, or even our sister newsletter, Robotics Insider.

54:28Corey:We'd love to see you at any of them. Thank you for joining us. And we hope you'll be back next time. But that's all we have for today. So farewell for now, humans.

From the publisher

AI companies spend enormous effort making models safer, but the model tested in the lab isn't necessarily the system a company eventually deploys.


Alice CEO and co-founder Noam Schwartz joins Corey and Grant to explain why prompts, tools, memory, permissions, data, and agent-to-agent interactions create an entirely new security surface.


They dig into the growing “trust gap” around enterprise AI, why prompt injection may become a permanent cat-and-mouse game, what open-weight models change for defenders, and why Schwartz believes the real security boundary has to exist at every layer of an AI system.


The bigger takeaway: as AI moves from answering questions to taking actions, traditional cybersecurity may need to merge with fraud prevention, threat intelligence, and trust and safety.


OpenAI security incident: https://openai.com/index/hugging-face-model-evaluation-security-incident/

Alice: https://alice.io/

Alice on agentic AI security: https://alice.io/blog/key-security-risks-posed-by-agentic-ai-and-how-to-mitigate-them

Alice open-weight research: https://alice.io/blog/okay-here-is-how-to-build-a-bomb-millions-download-dangerous-llms


Subscribe to The Neuron newsletter: https://theneuron.ai


Sponsored by Dell Technologies and NVIDIA. Learn more at https://www.techrepublic.com/hubs/the-enterprise-guide-to-scalable-ai/

More from The Neuron: AI Explained

All 106 episodes
Where Does AI Agent Security Actually Live?The Neuron: AI Explained · 55 min
Listen in VO