Controlling AI Models from the Inside

20 Jan 2026 · 44 min · 17 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Practical AI Podcast Episode Notes: Controlling AI Models from the Inside

Episode Overview

  • Podcast Title: Practical AI
  • Episode Title: Controlling AI Models from the Inside
  • Hosts: Daniel Whitenack and Chris Benson
  • Guest: Alizishaan Khatri (Founder of Wrynx)
  • Release Date: January 2024
  • Description: This episode explores innovative approaches to AI safety and interpretability, emphasizing the limitations of traditional guardrails and how model-native, runtime signals can enhance AI systems.

Key Participants

  • Alizishaan Khatri: Expert in AI safety with a background at Meta and Roblox.
  • Chris Benson: Principal AI Research Engineer at Lockheed Martin.
  • Daniel Whitenack: CEO at Prediction Guard.

Introduction

  • The episode emphasizes making AI practical and accessible, focusing on the intersection of technology and real-world applications.
  • Discussions revolve around novel safety measures for AI models, especially as they move into production.

Key Topics Discussed

  1. AI Safety and Security
  2. Distinction Between AI for Security and Security for AI:
  3. AI for Security: Utilizing AI to tackle existing security challenges.
  4. Security for AI: Ensuring that AI models themselves are secure and safe from misuse or harmful outputs.
  1. Challenges with Current Approaches
  2. Black-Box Models: Current AI models operate as "black boxes," making it difficult to address malfunctions or harmful outputs effectively.
  3. Traditional Guardrails:
  4. Primarily rely on input/output filters that are often too slow, costly, or limited.
  5. Example: Analyzing prompts and responses is ineffective as the damage may already occur.
  1. Innovative Safety Approaches
  2. Model-Native Safety:
  3. Focus on understanding and intervening within the model’s internals rather than solely at the input/output level.
  4. The goal is to build a framework that provides safety by integrating insights directly from the model's operation.
  1. Interpretability and Mechanistic Insights
  2. Interpretability:
  3. Understanding how models generate outputs is crucial for identifying and mitigating harmful behavior.
  4. Mechanistic Interpretability: A subset focused on identifying how specific components or regions of a model contribute to its outputs.
  1. Customizable Safety Solutions
  2. Context-Specific Needs:
  3. Different industries (e.g., healthcare, finance) require tailored safety measures that go beyond generic guardrails.
  4. Wrynx's Approach:
  5. Provides a "safety module" that integrates with existing models to enhance security without requiring users to create or train new models.
  1. Cost and Efficiency
  2. Economic Considerations:
  3. Traditional safety measures can lead to high latency and cost, making them impractical for many applications.
  4. Wrynx's framework aims to reduce computational costs significantly while maintaining or improving safety performance.
  1. Future Aspirations
  2. Alizishaan Khatri’s vision includes establishing a standard for model-native safety across diverse AI applications.
  3. Emphasizes the need for dynamic, real-time safety solutions that adapt to specific use cases.

Conclusion

  • The episode highlights the importance of evolving AI safety measures beyond traditional methods.
  • Emphasizes an integrated approach that leverages real-time insights into the model’s workings to enhance safety.
  • Encourages listeners to think about AI safety as a layer that can and should be customized to fit various industry needs.

Additional Resources

  • Listen to the Podcast: [Practical AI Podcast](https://practicalai.fm)
  • Connect with Guests:
  • [Alizishaan Khatri on LinkedIn](https://www.linkedin.com/in/alizishaan-khatri-32a20637/)
  • [Chris Benson's Website](https://chrisbenson.com/)
  • [Daniel Whitenack's Website](https://www.datadan.io/)

Upcoming Events

  • Register for upcoming webinars: [Practical AI Webinars](https://practicalai.fm/webinars)

---

This structured overview provides insights into the episode's discussions, key arguments, and takeaway messages relevant to AI safety and its applications.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Ali Khatri's Background

1:24 to 2:26

Ali shares his experience in AI safety and how he founded RINX.

“To be honest, I'm really excited about talking about the future because our guest today is really thinking very innovatively about how we can secure our AI models and have safety.”

Understanding AI Model Safety

2:26 to 4:32

Ali defines safety in AI models and discusses potential harmful scenarios.

“I spent about three years at Meta where I built infrastructure that serves about half of the world's population.”

Current AI Model Security Landscape

4:32 to 6:44

Exploration of existing capabilities and challenges in AI model security.

“As models have entered the tech stack, they also bring in a bunch of security challenges, and that is what security for AI solves.”

Risk Assessment and Mitigations

6:44 to 10:47

Ali discusses how to assess risks and develop mitigations for AI use cases.

“So imagine you are in a, like a story, like a built, like a giant apartment building with over, like say a thousand condos, right?”

Terminology in AI Safety

10:47 to 12:54

Clarification of terms related to AI safety and guardrails in the industry.

“Or the inputs or outputs, managing the prompts and the outputs.”

Understanding Mechanistic Interpretability in AI

15:39 to 18:06

Explore the facets of interpretability and its implications on AI safety.

“So Ali, you were just getting into these ideas of interpretability, mechanistic interpretability, I think you called it.”

Intervening in AI Models: The Butterfly Effect

18:07 to 21:18

Discover how small changes within AI models can lead to significant outcomes.

“So you know that, okay, well, this is what's happening in this hallway and we got to stop it.”

Instrumenting AI: Real-Time Monitoring and Control

21:19 to 24:08

Learn about the methodologies for monitoring AI models at runtime.

“Am I understanding that right in terms of like various ways of, I guess, either intervening or instrumenting?”

Building Safety Modules for AI without Model Modification

24:09 to 28:00

Understand how to enhance AI safety with minimal customer burden.

“needed, I guess, from the customer side to create this sort of instrumentation?”

Innovations in AI Model Safety

28:00 to 30:55

Learn about the advancements in AI model safety and how internal analysis can reduce resource needs.

“research breakthrough where we can build safety about thousand times cheaper.”
Show all 17 chapters

The Limitations of External Guardrails

30:55 to 35:09

Explore the challenges and limitations of traditional guardrail approaches in AI safety.

“why they can't use guardrails in certain cases.”

Combining Security Layers for Improved Safety

35:09 to 37:53

Discover how layered security approaches can enhance AI model safety in various applications.

“So without information, you're not going to be able to do anything, right?”

Future Aspirations for AI Model Safety

37:53 to 42:00

Understand the future direction of AI model safety and the vision for native safety solutions.

“And then in that, let's say we detect some sort of misrepresentation or some sort of lying.”

Challenges of Fine-Tuning LLMs on PII Data

42:00 to 42:18

Learn about the complications of fine-tuning large language models on PII data.

“And no one in their right minds today would fine-tune an LLM on PII data.”

Vision for a Model Safety Layer

42:18 to 42:30

Discover the vision for creating a universal model safety layer.

“I want to build model, become the go-to for model safety.”

Encouragement to Explore Rinks

42:30 to 42:48

A recommendation to explore Rinks and their contributions to model safety.

“Well, I definitely encourage people, check out the show notes.”

Thanking the Guest and Future Topics

42:48 to 43:06

Reflecting on the guest's contributions and future discussions.

“And from the community, Ali, just thank you for digging into this topic and bringing a fresh look at things.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:03Welcome to the Practical AI Podcast where we break down the real world applications konsine intelligence and how it's shaping the way we live, work, and create. Our goal is to help make AI technology practical, productive, and accessible to everyone. Whether you're a developer, business leader, or just curious about the tech behind the buzz, you're in the right place. Be sure to connect with us on LinkedIn, X, or Blue Sky to stay up to date with episode drops, behind-the-scenes content, and AI insights. You can learn more at practicalai.fm. Now, on to the show.

0:48Welcome to another episode of the Practical AI Podcast. This is Daniel Whitenack. I am CEO at Prediction Guard, and I'm joined as always by my co-host, Chris Benson, who is a principal AI research engineer at Lockheed Martin. How are you doing, Chris? Hey, doing great today, Daniel. How's it going? It's going well. Well, done a little bit of snow shoveling on the ground. As we speak, we're kind of headed into winter break or the holiday Christmas season here in the US. And I think this episode will be released in the new year. So if you're listening to this, you're listening in the future. To be honest, I'm really excited about talking about the future because our guest today is really thinking very innovatively about how we can secure our AI models and have safety.

1:36as we move into that future. Really excited to welcome to the show today, Ali Khatri, who is founder of RINX. Welcome, Ali. Thanks. Thanks for having me, Daniel. Yeah. Yeah. We met earlier this fall. I'm really fascinated by your kind of line of work and technological innovation at RINX. But I also know that you've been thinking about these topics around AI safety, guardrails, disallowed content, et cetera, for quite some time. Could you give us a little bit of a background of how you got into these topics and what you've done in the past? Yeah. So I have been in the machine learning for AI safety or anti-abuse use cases in general for the past eight or so years of my career.

2:26I spent about three years at Meta where I built infrastructure that serves about half of the world's population. Basically, any safety, anytime you type a message on Facebook, it goes through tens of safety checks, which are powered by thousands of models. I built the infra that these models run on. And then I moved on to Roblox, where I built systems to protect about$3 billion in payments against fraud. So I've been in this space. I've been using AI models to sort of protect against abuse. And during this time, I realized that the models that I'm using themselves are susceptible to abuse. So that's what led me to founding ranks.

3:13And I know that now you're thinking about those actual models. So I often tell people also on the AI security side, there's kind of AI for security and then there's security for AI. It sounds like there's something similar kind of in, I guess, model or safety anti-abuse space. Could you give us a little bit of an understanding of like when we're talking about the safety or security of AI models? Could you kind of define that for us? What do you have in mind as the kind of bad case scenarios or worst case scenarios of what a model could do, why it might not be secure or safe? Yeah. So thanks for making that distinction between AI for security and security for AI.

4:04And they're two very different things, right? Like AI for security basically means using AI to solve existing security challenges in a more effective or a better way, right? So that's a very different and a very linearly separable body of work from security for AI, which focuses on making the AI models themselves and like AI-based use cases secure. As models have entered the tech stack, they also bring in a bunch of security challenges, and that is what security for AI solves. Now, in terms of the safety aspects that we talked about, models today will generate anything. We'll literally tell, like, these generative models, like, if it's a tech model, it'll generate any form of vile content known to man.

4:54And it could tell you, like, openly I got into trouble for, like, where a teen was asked to commit suicide, right? Or was encouraged to commit suicide. So you have self-harm. You have different other categories of harm where you can, you know, generate pornographic content. You can generate other forms of inappropriate content like violence, gore. And sometimes it doesn't even have to be inappropriate because safety is very context-specific, right? As a law firm, safety looks very different for you than what it does for a medical shop versus, say, a customer service setting versus like a co-generation environment.

5:29So each one of these use cases come in with a set of permissible and non-permissible behaviors. And that's kind of what safety really is, that making sure that the technology works in ways that you intended to and minimize the unintended outcomes of it. And I guess people are addressing, I mean, these are known issues in the sense that at least a segment of the people that are working on these models know about these issues. Could you give us a little bit of a sense of like, as we speak, how these, at least in a production sense, like what's capable now? Like how does the landscape look in terms of capabilities to defend against or align with certain policies or allowed content, disallowed content?

6:28What are our choices right now? If I'm looking at the availability currently of both kind of open source projects and what's in maybe closed platforms, what's available to me to deal with this issue? Okay. So before I answer that, I'll start with an analogy, which will help make the rest of the response much more clearer and add some context. So imagine you are in a, like a story, like a built, like a giant apartment building with over, like say a thousand condos, right? Now, let's say your neighbor for some, for some reason or no reason at all, really decides to pull out a golf club and starts violently assaulting you with it.

7:13Now for a situation like this, will the good guys be able to protect you just by checking IDs at the gate? Obviously not. That's kind of where we are. You're past that point. So that's kind of where we are with AI safety today. That's kind of how jailbreaks work, right? So here, the giant thousand-story building is the model. So today, what we are able to do is just today's solutions, analyze what's going into the model, also known as the prompt, and analyze what's coming out of the model, which is the response. But by then, the damage has already been done. So now, if you're using video generative models, right, You can put in a text prompt, get a video output.

7:52Your video output is going to be, it's too expensive to analyze, number one. And the video has already been generated. So we already spent a large amount of compute generating that bad content. Again, if you're talking with audio models, you can trick audio models into generating bad content. Like a seemingly innocuous looking prompt can be tricked today using a multitude of techniques to generate really malicious output. So now, unless you have visibility into what's going on inside of the model, you're not going to be able to catch a lot of these things. That's where jailbreaks come from. That's where adversarial machine learning comes from, if you look at it through the context of predictive models versus generative models.

8:36It's essentially the same core phenomenon where operating these models as black boxes, and we have no idea of what's going on inside of them. So we're trying to change that. Looking at that, it seems like in the way that you just kind of phrased it with the context there, it seems like a very intractable problem. As a follow-up to Dan's question, like how should I as a new user potentially or someone coming into a use case in my company where a model is desirable, but I'm worried about whatever bad or abuse means in the context that I'm operating in, how should I start off thinking about that?

9:14Like, like, what's my starting point? Because I, I gotta say, you know, coming into it, I'm not even sure where to start. So like, can you level set that a little bit in terms of how, you know, what, what's, what's square one? Yeah. So that's a good question. So normally what you do is the way I like to think about it is there's a general category of bad stuff that no one really wants or the law doesn't allow and things like that. So there's like a general category of stuff like that, right? Like here, I'm including porn, hate speech, yada, yada, yada, which they're a general category of undesirables that no one wants on the platform, like child safety, for example.

9:50That's non-negotiable, no matter what context you're in. Now, there's also another aspect of categories, which is about context-specific safety. So now if you're in a banking use case, you've got to think about, okay, money laundering. You might not have to think about money laundering, say, in a code generation setting, for example. So you want to think about these very specific categories of risk or issues that come from your use case. And people tend to usually have a very good idea of that. If you understand your use case well, which most people do, which is why they're exploring models and they're trying to solve some problem within their use case.

10:28So within that use case, you also would understand the problems that you're facing. And that's another category of risk that you want to think about. Then once you've thought broadly about these two, then you've got to figure out about once you've identified both of these categories, then you want to think about mitigations, detections and things like that. And I guess there's like if we're just kind of defining some terms for people that they might have heard of this approach that you talked about related to kind of guarding the gate to the apartment complex, right? Or the inputs or outputs, managing the prompts and the outputs.

11:08Is that what people refer to as like guardrails, safeguarding? What is some of the terminology that's being used? And then I know that you all that you have a different way of thinking about this. I guess just setting the stage jargon wise. How do we define these kind of guarding the gate things? And then as we filter into actual safety within the apartment complex or within the model, what kind of terminology? And I guess like there's a body of research building up to what you're doing. what terms are used to describe that? And as people wanted to research that, what would they look for? Yeah, so guardrails essentially are a catch-all term.

11:52They can refer to prompt and response filters. Now today, one of the ways of doing, there are multiple guardrail type solutions out there. There are guard models out there. Meta has one, IBM has one, Google has one, OpenAI has one. And I'm talking about public releases, right? Internally, most people have their own. But these essentially are prompt and response filters. So they look at the data going in, they look at the data coming out. Now, so that is like one thing that guardrails are used to refer to. Less commonly, guardrails are also used to use, refer to like static checks, where you just look at the output of the model and say something like, okay, the word, like forbidden word, let's say the F word, right?

12:35The F word appeared in output, this is not permissible. So that's a simple regex filter that you can use. So that would also be called a guardrail. In terms of looking at the internal state of the models, there's a whole field of research, area of research that's developing. It's called interpretability. There's a subset of that called mechanistic interpretability, where they try to figure out what subcomponent of the model led to this particular output and try to change it at the source or try to alter or modify behavior while it's happening as opposed to after or before.

13:29Well, friends, here's your hot take of the day. Your team's AI tools, they might be making collaboration messier, not faster. You probably know this. You feel this. Think about it. You've got AI literally everywhere now summarizing, generating, suggesting. But if there's no structure, no shared context, you're just creating more noise, more outputs, It's more stuff to wade through. The gap between having a great idea and actually shipping that idea, that gap isn't a speed problem. It's a clarity problem. Well, this is where Miro comes in. And honestly, it's shifted how I think about team workspaces.

14:06Miro's innovation workspace is not about brain dumping everything into an infinite canvas and just hoping for the best. It's about giving your work context, intentional structure, So your team knows what to focus on and where to find what they need without playing detective across 12 different tools, clicking and moving and tabbing. Just too messy. And the AI piece, well, Miro AI actually gets this right. They've got these things called AI sidekicks that think like specific roles, like product leaders, agile coaches, product marketers, reviewing your materials and recommending where to double down or where to clarify.

14:46You can even build custom sidekicks that are tailored to your team's exact workflow if you desire. And then there's Miro Insights. It sorts through sticky notes, research docs, random ideas in different formats, and synthesizes them into structured summaries and product briefs. And Miro prototypes, they let you generate and iterate on concepts directly from your board. Test 20 variations before you ever touch your hi-fi design tools, saving you time, giving you ideas, and getting it right. This whole thing is built around the idea that teamwork that normally takes weeks can get done in days, not by going faster, but by eliminating the noise and the chaos.

15:29So help your teams get great done with Miro. Check it out at Miro.com to find out how. That's M-I-R-O dot com. Again, Miro dot com.

15:44So Ali, you were just getting into these ideas of interpretability, mechanistic interpretability, I think you called it. We've talked about interpretability on the show before, I think mostly in relation to trying to figure out why a model made a certain decision in relation maybe to certain concerns around bias or other things. So like if I have a risk model for approving insurance or something like that, then I might need to have some interpretability around that. Or maybe in the case of healthcare, there's a burden for interpretability of like how decisions were made. Here, it sounds like kind of the interpretability is being applied to, I guess, why the, is a good way to put it, why the model generated some problematic output?

16:41Or is there a better way to think about that? So there's multiple overlapping aspects here, right? Interpretability is like an umbrella research area. It's an umbrella term. So what you alluded to earlier is more of, it's also described as explainability. So why was this credit card denied. So you're trying to explain it in human concepts. That is a part of interpretability, no doubt. It's a subset of it. There's another subset, which is how was this generated? So for example, if your generative model outputs, say, if you say, how are you? Hi, and the model says, how are you? You want to care about how was this generated internally?

17:25How were these tokens produced? And the reason you care about that is you want to know if, instead of saying, like responding to, like responding with how are you, it could respond with howdy or with something else, right? So you want to know what caused those differences and you want to be able to control that. So that aspect is also interpretability. Now, where this interplays with safety is when you have these prompts, which look bad, which look good to a human, but bad, but result in bad outputs, which is how jailbreaks work. When you analyze how the data flows inside of this black box, you're able to control it and stop it at the source.

18:06So think of this as, continuing our analogy from earlier on, think of this as cameras at every gate or every path. So you know that, okay, well, this is what's happening in this hallway and we got to stop it. We got to put an end to it. So it's a very different class of defenses. As you're saying this, you're actually talking about manipulating the internals of the model and the flows that are there, kind of the cameras on the doors and stuff like that, and making it maybe a gray box rather than a black box to some degree, as opposed to kind of the more traditional guardrail approach where you have programmatic, you know, you use the word guardrails around the models, inputs and outputs to try to handle things that way.

18:49So it's kind of a whole different thing about instead of treating the model as a black box, you're saying you're diving into it and trying to affect an improvement there? Yeah. So intervening is one, like stopping generation or modifying that is one form of intervention, right? Intervention does not have to necessarily be in that form. It can take various other forms. So once you understand today, we have no idea of what's going on inside of the model. We provide an input to the model, get an output. We have no idea what's going on. So this, what we're trying to build, or what Interpretability tries to do is understand what's happening inside.

19:29Now you could control it, or you could use that to make a risk quantification and use that downstream. You don't have to do anything in the moment necessarily. It can be leveraged downstream. So now it's basically like you have a whole new set of information or a whole new set, like a whole new class of data points that you can leverage in creative ways downstream. This is something that's not available today. And this is what interpretability builds really at runtime. Am I correct? So part of my assumption in the past, and I'm fascinated by this whole subject, is like some small changes in the, for example, the weights or individual layers or individual pieces of the model can produce very large changes in the output behavior of the model, which it's kind of...

20:24I'm blanking on this. What is the thing? It's like a butterfly flaps, it swings and Oh, the butterfly effect. Yeah. The butterfly effect or, or, or whatever. So like these, cause you can make a change, whether that's quantization or other things to the model. And that, you know, may produce unclear and sometimes catastrophic changes in the, in the behavior of, of the output. And so if I'm understanding what you're saying, right, Ali, it's one way you could try to use the information about how the model is producing certain outputs is to intervene by actually making a modification or preventing something in the model.

21:05But that could produce other changes that you may not want, I'm assuming. But you could also instrument the model to understand potentially when it is kind of firing those certain neurons or lighting up in a certain way that is indicative of problematic behavior. Am I understanding that right in terms of like various ways of, I guess, either intervening or instrumenting? I don't know if I'm using the right terms. We are instrumented. Like the way we're approaching this is we are trying to understand what happens inside of a model at runtime. Models are a monolith, right? But we're breaking it down into different spaces or subspaces.

21:47And we look at the subspaces that get activated during bad generation. So now when you're, let's say, generating non-permitted content versus permitted content, different sub-regions of the model get triggered. So we're building visibility into that, and we're trying to identify them at runtime. So now there are some sub-regions that you wouldn't care about, right? Like, for example, if you take a general-purpose LLM, it's trained on everything ranging from Python code to 15th century Chinese poetry, right? Now, when you're using it in a customer service setting, you care about neither one of them.

22:26And if those sub-regions of the model are getting activated, then you want to be able to arrest it while it's happening. So this is similar to, if you go back to the analogy that I had made about the apartment building, you want to have visibility at all times into what's going on at each level, right? So you find out that a bad thing is going to happen way before it actually happens. Like, for example, if someone's going to conduct a bank robbery, people just don't get up and conduct a bank robbery, right? There's some searching going on. There's cycles of planning going on. There's purchases of firearms or whatever going on.

23:02So now if you stop them at these bad activities at different levels, the police don't have to deal with a shootout situation in a bank at the very end. So similarly, defense works in depth, and we're building a whole new layer of safety that hasn't been tapped into just yet. Yeah, that's fascinating. And I'm guessing certain folks are probably wondering, like I am out of curiosity, like this might be the first time that they're hearing about such an approach to this kind of problem. And they might be thinking, you know, how is this possible? Like how. So one thing I think the general concept makes sense, right?

23:44Like instrumenting the interior of the apartment complex, understanding what's happening, kind of retrieving that, you know, that intelligence to for you to make decisions or determine if you want to mitigate something. What are the, I guess, have there been multiple attempts to try this sort of thing? And from your perspective, in terms of how you all are approaching it, what is kind of needed, I guess, from the customer side to create this sort of instrumentation? So I could imagine, like in one scenario, like I could train, I could tell the customer, well, we're not going to train on Chinese poetry or code.

24:28we're just going to train a whole new model. And that burden on the customer is very, very heavy, right? And on another scenario, you could say, oh, take this model off the shelf and do this to it. And then, you know, that's a less burdensome. So there's probably a spectrum of here. Could you help us understand like the, I guess the burden and what might be required to get to this instrumentation? Yeah. So the way we're approaching this today is we do not modify the model. We do not even build models nor require the customer to build a model. We take an off-the-shelf model and we build a safety module that sits on top of it.

25:08So essentially for the customer, it's like a very low friction approach where they take a model that they use and love, like Lama or Granite or any Mistral or any of the open-weight models for image generation. You have WAN and you have a host, other host of Chinese models for the audio video setting. So any model that you love, we sort of make it more secure and tailor it for your context. So now, again, remember, if you're like a lower law firm, right? An off the shelf Lama or off the shelf Mistral is not going to have the protection that you have. Like if you're like, say, a shoe company, let's say you're Adidas or Nike, right?

25:49You as a user want to talk about Nike, but not talk about Adidas, right you can't expect the model maker to put that in for you because the model maker is trying to sell to everybody so we help build that customization and we do that without changing your primary model like your primary model will continue to be as it is if you make any modifications to it uh like fine tuning or anything like that that's on you you control that we don't require you to do it but even if you do that we can still support you could could you uh so totally recognizing that there's proprietary stuff that you're not going to dive into and respect that.

26:26Could you talk a little bit about just kind of clarifying as we were kind of talking earlier about kind of the buttressing with guardrails on the external side versus the going into the model. And as you're talking about adding a component, like, so in my conclusion, it seems a little bit like it's on the outside. Can you talk a little bit about what you mean by that without diving into places you can't go? Yeah. So I'll give you a very, I'll try to address that as much as possible without going into the specific. Fair enough. Specific details. So today what you have, right? You have these filters, which analyze the inputs and the outputs.

27:04Now, analyzing the economics here is all messed up. Like if you were to analyze video or audio, right? Those tend to be very expensive computationally. Those models are very expensive. So now if your inference itself costs X, and if you're expecting someone to pay another X to analyze it, number one, it's slow. Number two, it's like paying someone$1 ,000 to guard a$100 bill. You're just not going to do that, right? So what people end up doing is people end up like shipping unsafe models. So today, we have tested models from different audio companies. We've tested models from different video companies, different image generation models, each and every one of them with little to no trickery, which means that average user can just go there and ask them to generate bad stuff and they will do it.

27:51There are little to no guards there. And that's because the economics does not make sense the way thing is today. So what we've built is a scientific, like we've sort of had a research breakthrough where we can build safety about thousand times cheaper. So just to give you some concrete numbers. What we've done with LAMA, we've taken a LAMA model, which is like an 8 billion parameter off-the-shelf LAMA model. And today, if you had to protect it, you would have to use LAMA Guard 3, which is an 8 billion parameter model. And assuming it generates 10 tokens, that is, you're running about 80 billion parameters of inference at runtime.

Read the full transcript

28:27Now, if you do that on your prompt and response vote, that number balloons to 160 billion parameters of inference. That is two extra GPUs or one extra GPU, depending on how you've wired or deployed the models. So now what we are essentially doing is we're analyzing the internal states of the primary model as it makes the prediction. So in doing so, we don't need any of those two extra GPUs. And that 160 billion parameters of inference that I counted, we have succeeded in bringing it down to 20 mil with an M. So we're essentially a rounding error today because of this expensive safety profile. You cannot even deploy them on edge devices.

29:10Like on edge devices, guardrails are non-existent because you can barely, like edge, when people are working on the edge, they work really hard. A lot of people work really, really hard to squeeze that one device through techniques like quantization, that one model onto the limited memory of the device. So you have no room to deploy a safety model, right? So that's why we've like built tech, which literally is like a rounding error. 20 million parameters on 8 billion is nothing. And we can sort of deliver safety, like comprehensive safety. And our safety performance is comparable to a standalone guard model.

29:44And it is significantly faster because our latency just doesn't exist. It's parallelized and it becomes, in practice, the latency of the primary model is the latency that the user sees. With today, you have to like sort of account for the latency of the primary model. and you also have to account for the latency of the response filter and the prompt filter. Also, the response filter cannot kick in until the primary model has finished generating. So you're looking at like very high latencies. So from the perspective of the end user, you're looking at a lot of added friction in terms of slow speed.

30:18You're looking at increased costs because ultimately the cost will be passed on to the user, right? So you're paying for two extra GPUs that you don't have to, and your quality will be substandard. Again, remember, all these models are able to do is check IDs at the gate. So that is the protection you're getting. So Ali, it's very fascinating and encouraging, the results that you're seeing and what you're able to do with this sort of approach. I'm also wondering the size or latency or behavior that you just described. That is certainly a huge component of what people are thinking about and why they can't use guardrails in certain cases.

30:58Another question that might come up though, and I actually think that you can validate me, but I think you have a very good answer to this, is what about the kind of accuracy or reliability of kind of one approach or the other? So a person, I guess, could argue and say, well, if I have a guard at the gate and he's 100 percent accurate in making sure no gun ever gets into the building, then there'll never be a shooting. And that's a very robust guardrail or something like that. But I think what you're saying is you still wouldn't know what happens inside the building with 100 % certainty. So could you address that side of things like the accuracy side or the quality side, I guess, of the performance of guarding and safety with this kind of instrumented model approach versus kind of an exterior guardrail?

32:03Yeah. So an exterior guardrail, as you pointed out, is limited in visibility, right? So even if they do 100 % great job of checking someone's IDs, they only have limited information. There's only so much you can do, right? The kind of defenses you're expecting is out of the scope. You're limited by information there. If you don't know how to drive a car, you're just not going to be able to do it, right? No matter how fit you get, no matter how much you train, you're not going to become a race car driver if you cannot drive a car. Similarly, in this security setting, the example that I use where somebody decided to pick up a golf club for some reason or no reason at all, you weren't checking for...

32:52Golf clubs are permitted items to bring into a home. So the security have done their job right. They haven't done anything wrong. But there's a fundamental limitation that exists here. So you can only do so much looking at artifacts. So there's a whole layer of safety that is untapped or unaddressed or inaccessible because of scientific limitations. But that's changing fast. And we're sort of at the leading edge of it. And just to make sure that I have it right. So would it be a good way to describe it that let's say that I have or I want to prevent toxicity or something coming out of the model of a certain type?

33:37There are a variety of inputs to the model that could result in that type of output, but I'm never going to know all of them. Or there's always an edge case. And so by instrumenting and saying this part of the model lights up when toxicity is being produced, then I no longer have to worry that I have all of the possible inputs in the world put together that might trigger toxicity. I just know when there's toxicity. Is that an appropriate way to put it? Yeah. So when you're on the defense, when you're playing defense, right, you don't have, like when you're defending models against abuse in any scenario, whether you're defending models or just protecting a platform in a classical trust and safety sense, you're never going to have an exhaustive list of the million ways in which things go wrong.

34:35But you can gauge, you can develop a fair understanding through past examples or through data points that you have. That's the benefit of machine learning, right? But with the way jailbreaks work or with the way adversarial examples work is that they do certain things inside of a model that are not possible to predict in a different model, which is being used as a guard, right? So the guard is model type A. The primary model that you're predicting is model type B. So the guard is not going to be able to predict what's happening in model type B, simply because they don't have visibility into what's happening inside the model.

35:11So without information, you're not going to be able to do anything, right? Like, for example, if you're the SEC and somebody takes away your access to bank accounts, you're not going to be able to prevent money laundering, no matter how many books you've written on that subject. There's only so much you can do with guns and badges, right? You would need visibility. If you want to try preventing money laundering, you will need visibility into the financial system. So that's kind of what we're building here. In terms of accuracy, the numbers speak for themselves. We're able to match and beat the performance of standalone guard models over a thousand times our size.

35:48And that's because we are exploiting this unique insight. Yeah, it sounds a lot like what is an analogy that in neuroscience would be, if I can pronounce it right, an electrocellulogram, an EEGC where they monitor the synaptic connections and they can see it lighting up. Since I can't pronounce the word, I'll just describe it to the best of my ability. It's a long word. I actually had it written down. I was like, oh, crap, I still can't pronounce it here. So it sounds like that as you're thinking about how you can apply that, are there any ways that that can kind of tie back also into the hybrid approaches?

36:29If we talk about some of the other more traditional forms of guardrails that are out there, can you combine them into sort of a hybrid approach where you do have different types of guardrails that are in place, but people can add this kind of capability from you into it and thus enhance their overall security model? What would that world look like in your view if that's valid? Yeah. I mean, I do. I'm a firm believer in defense and depth. So one product does not miraculously solve everything, just like with our human society, right? Like if you think about law enforcement, it's a very good parallel where for national security, you need army to protect you.

37:11You need different forms of the military to protect you from external threats. You need the border police to make sure that entrance is regulated. But at the same time, you also need state and local law enforcement. You also need federal civilian law enforcement. So you need different levels of that to make sure that security on the whole is ensured. Similarly, in the context of AI models, yes, you have guardrails which look at prompts and responses. They are valuable, right? But then there's ways to do them efficiently. And then you also oftentimes would need to combine, say, system-level features with model-level features.

37:49So we build those model-level features. So you could do something complex, like let's say you're running some sort of customer service bot and a customer has a history of refunds. And then in that, let's say we detect some sort of misrepresentation or some sort of lying. So you can say, okay, block if lying score. You can compose rules like saying block if lying score greater than 0.8 and this customer has refunded more than$1 ,000 worth of merge. So there's potential to combine mix and match as well. And I think that's how you improve the overall safety profile of any system. It's always in depth.

38:26And those layers have to work together. I can just look at web applications in parallel, right? You need static code analyzers and you need AI firewall. They don't replace each other. If anything, they complement each other and build an overall robust system. And before we leave the subject and maybe look towards the future a little bit, But I do want to highlight, I think one of the fascinating things that comes out of the work that you're doing, Ali, is some of the, I guess, customization that's possible in the sense that we've talked a lot about kind of the, quote, traditional types of things that you would want to prevent, whether that's jailbreaking or toxicity.

39:12But there's also, like you were saying, in every industry and actually at every company, there might be custom types of policies that they want to enforce or certain things that they want to instrument. Does this, I guess, how extendable is this approach to those sorts of situations? Yeah, so this is very extendable to those. I mean, our approach is designed for these situations. We've identified this niche in the market. With models, you're shipping a one-size-fits-all solution. So to put it in a different way, the model needs of most companies are similar-ish. But the safety needs are dramatically different.

39:59So you cannot have a one-size-fits-all safety stack that works for everybody. Yeah, sure. There's a general category of undesirables that everybody would want to keep off their platform. But that's a very small subset in with generative models, which are capable of doing so much. You need to be able to customize safety as well. So that's that's the space we thrive in. And that's what we're building for. Well, that's very interesting to me. I've learned a lot today. As you're thinking, Dan kind of telegraphed, we like to kind of finish up by asking about the future. And as you're thinking about the future, both in terms of kind of your specific approach that you're doing, but also kind of like where security is going for models in general, in the large.

40:51um what like what kind of what kind of guideposts do you have that you're thinking about you know i like to say um at the end of the day when you know you're taking a shower or you're lying in bed at night about to go to sleep and your mind's just kind of going loosely you know where do you see things going and what are you really passionate about um pursuing or exploring uh going forward and in time you know what's that aspirational hey i'd like to go do that um that you have in mind So I think when you look at safety, when you look at like runtime safety, like when you look at safety, right, there's different aspects.

41:25So there's build time safety, which is a whole different class of safety products. But in terms of runtime safety today, runtime safety only exists at the data layer, which is at the prompt and response layer. The model layer is missing. My aspiration is to sort of build that out. That has been my vision. That is the guiding vision behind RINX, where we want to build model native safety. And I see that will have to exist for models to be adopted into different settings. If you're in healthcare, for instance, today, it's very hard to use a public LLM because of data concerns. And no one in their right minds today would fine-tune an LLM on PII data.

42:05That's because you just can't. It's just not possible. So it's locking a lot of people out of the ecosystem. And that's the problem we exist to solve. So I want to become, like, my vision is to sort of build that de facto model safety layer for no matter what model you're using. I want to build model, become the go-to for model safety. That's awesome. Well, I definitely encourage people, check out the show notes. And I should have said this at the beginning, but just to make sure people, Rinks, W-R-Y-X. N-X. N-X, sorry. I even missed that. W-R-Y-N-X. So check out Rinks. We'll have the link in the show notes.

42:50And yeah, just really fascinating work. And from the community, Ali, just thank you for digging into this topic and bringing a fresh look at things. It's awesome. and hope to see you back on the show to find out where things have advanced. Yeah. Thank you for having me on the show, guys. We'll talk to you soon. Thanks. All right.

43:17All right. That's our show for this week. If you haven't checked out our website, head to practicalai.fm and be sure to connect with us on LinkedIn, X or Blue Sky. You'll see us posting insights related to the latest AI developments and we would love for you to join the conversation. Thanks to our partner, Prediction Guard, for providing operational support for the show. Check them out at predictionguard.com. Also, thanks to Breakmaster Cylinder for the beats and to you for listening. That's all for now. But you'll hear from us again next week.

From the publisher

As generative AI moves into production, traditional guardrails and input/output filters can prove too slow, too expensive, and/or too limited. In this episode, Alizishaan Khatri of Wrynx joins Daniel and Chris to explore a fundamentally different approach to AI safety and interpretability. They unpack the limits of today’s black-box defenses, the role of interpretability, and how model-native, runtime signals can enable safer AI systems. 

Featuring:

Upcoming Events: 

More from Practical AI

All 157 episodes
Controlling AI Models from the InsidePractical AI · 44 min
Listen in VO