Ensuring Privacy for Any LLM with Patricia Thaine - #716

28 Jan 2025 · 52 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Notes: Ensuring Privacy for Any LLM with Patricia Thaine - #716

Podcast Overview Podcast Title: The TWIML AI Podcast Host: Sam Charrington Guest: Patricia Thaine, Co-founder and CEO of Private AI Episode Description: The episode focuses on privacy, data minimization, and compliance when using third-party Large Language Models (LLMs) and AI services. It discusses risks of data leakage, challenges of entity recognition, and evolving AI regulations.

Key Topics Discussed

Introduction to Patricia Thaine and Private AI

  • Background in privacy-preserving natural language processing.
  • Focus on building a privacy layer for developers to comply with data protection regulations.

Understanding Personal Information in Data

  • Core technology identifies sensitive personal information (e.g., health, payment info).
  • Emphasis on data minimization: keeping only necessary data to mitigate risks.

Data Flows and Privacy Risks

  • Data flows to consider: input, output, and storage.
  • Risks when sending data to third parties; importance of removing sensitive information before processing.
  • Concerns about internal models collecting unnecessary data, increasing breach risks.

Risks of Data Leakage in LLMs

  • Data memorization issues where models can leak sensitive information.
  • Examples of data leakage from embeddings, emphasizing the need for careful data handling.

Techniques for Data Cleansing

  • Private AI’s approach: uses an API to cleanse data by redacting sensitive information.
  • Support for batch processing and screening modes.
  • Users can specify the types of entities to redact based on compliance needs (e.g., HIPAA, GDPR).

Challenges in Entity Recognition

  • Complexity in recognizing entities across various formats (OCR errors, multilingual data).
  • Importance of accuracy and speed in developing models for named entity recognition.
  • Discussed the balance between efficiency and comprehensiveness in recognizing over 50 entity types.

Generalization and Multimodal Data Handling

  • Private AI’s ability to handle multiple data formats (text, audio, images) while maintaining high accuracy.
  • Challenges in processing audio data: disfluencies in speech can complicate the detection of sensitive information.

Synthetic Data Use

  • Synthetic data as a tool for training models without exposing real sensitive information.
  • Importance of balancing synthetic and real-world data to maintain model effectiveness.

The Relationship Between Privacy and Bias

  • Discussed how personal identifiers in datasets can lead to biased outputs in models.
  • Emphasized the importance of preventing bias by anonymizing sensitive data in inputs.

Evolving Regulations and Compliance

  • Overview of current AI regulations (GDPR, EU AI Act, CPRA) and their implications for data privacy.
  • Need for clear guidelines in regulation to help organizations manage data responsibly.
  • Concerns regarding the lack of regulations leading to increased misuse of AI.

Key Takeaways

  • Data First Approach: Prioritizing data management is crucial for successful AI implementation.
  • Privacy by Design: Companies should integrate privacy considerations from the outset of AI development.
  • Dynamic Regulatory Landscape: Organizations must stay informed about evolving regulations to ensure compliance.
  • Mitigating Bias: Removing identifiable information can reduce bias in AI systems, aligning with ethical practices.

Conclusion The episode emphasizes the critical role of privacy in the development and deployment of AI technologies, particularly with LLMs. Patricia Thaine and Sam Charrington discuss a variety of challenges, solutions, and the importance of maintaining data integrity while navigating the complex landscape of AI and privacy regulations.

For complete show notes, visit [TWIML AI Podcast Episode 716](https://twimlai.com/go/716).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00I think that a lot of the misconceptions around what is easy or hard in AI when it comes to thinking that it's easy or hard to think it's easy or hard to think it's easy or hard to understand it's hard. high-quality representative data. And then the misunderstanding how huge a lift it is to get a product in production that is scalable, that also has a lot of corner cases in mind.

0:45All right, everyone, welcome to another episode of the Twimble AI podcast. I am your host, Sam Charrington. Today, I'm joined by Patricia Thain. Patricia is co-founder and CEO of Private AI. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Patricia, welcome to the podcast. Thanks so much, Sam. Great to be here. I'm looking forward to digging into our conversation. We'll be talking about some really interesting things you're doing with regard to AI and privacy and more. To get us started, I'd love to have you share a little bit about your background and how you came to start Private AI.

1:24Yeah, sounds great. Personally, my background is in doing research in privacy-preserving natural language processing. Specifically, it was originally focused on homomorphic encryption. So that lets you compute on encrypted data. and co-founded the first version of private AI on that topic. I have an academic hammer. Let's go find all the nails. It didn't work out. It turns out that's a very bad way of problem solving. And then started the second version of private AI really with the focus of building a privacy layer that's built for developers that can actually help them comply with data protection regulations.

2:01And now there are, I think, a couple of ways for us to dig into this topic. The first, I think, is to dig into what private AI is doing for AI developers and how those capabilities are expressed. And then also how you're using ML and AI to do all of those things. But let's start with the use cases and what you're trying to enable. Absolutely. So the core technology is understanding what's in your data. What kind of personal information you have, protected health information, payment card industry information, confidential information, all of these can play a really big role in whether or not you can end up getting data for your AI projects, whether or not you're going to be able to use third party tools.

2:51What exactly happens with regards to the output of the training data when it comes to sensitive information being displayed to your users or employees or otherwise? And so really what the product does is identifies this personal information and is able to redact whatever it is that you don't need. And that's really the concept of data minimization, which is core to a lot of data protection regulations. Got it. Got it. Minimization from the perspective of not having data that is not necessary in the process. Exactly. And it sounded like we're talking about data that comes into contact with a system in three different places, maybe even more, but certainly input and output and maybe storage.

3:47And can you kind of talk about data flows the way you see them? Absolutely. So if you look at large language models in particular, for example, you're going to be sending data maybe to a third party. in which case whatever information you're sending you're losing control of when you're sending it external to your organization or to your own laptop and so it's really important to think about what kind of information you're sending through what can you remove and then once that information is removed it can very often be reintegrated into the resulting output of the model and so you can still have a fairly seamless interaction with the model without that risk of sending that data externally.

4:28And then other aspects of it are even if you have a model internally within your own private environment, there are a few things to be concerned about. One, are you collecting data that you shouldn't be? Because whatever data that you do collect puts you at higher and higher risk of a more significant data breach down the line. That's part of the interesting part about data minimization. You keep what you need. And you don't keep what you don't need, meaning that you strip out that risk from your data in the very likely chance that at some point there will be a data breach. And then there are other aspects to it with regards to training.

5:09When it comes to trading models, I think this community probably is aware that that data can be quote unquote memorized by the model. And leaked out. Yeah, and as a result leaked out. And there's lots of research showing this, initially for character language models, then for GPT-2. And then, you know, there are more and more coming out around exactly the risks around training or fine-tuning on personal and confidential data. And we've got some episodes we can link to in the show notes about that topic. Perfect. Oh, yeah. I saw that you interviewed Professor Carlini. That's, yeah, he, yeah, great, great person to talk to about that.

5:51then there's also the risk of having personal and confidential data in the contextual data that you're using to create embeddings for RAG, for example. And that's the one where a lot of people actually don't think about. The assumption is that if it's an embedding, you know, it's generalized. It's kind of opaque, right? Yeah, exactly. But that is completely untrue. And there's research showing how if you have, for example, word embeddings that are created from healthcare data, the names and first names end up closer in distance to last names and disease that somebody had ends up closer in distance to the relevant name.

6:38and so even just if you can't inverse the embedding does that mean that is subject to leakage or it's subject to attack like i can get attacking that but yeah leak you know kind of passive leakage seems more difficult to less likely i guess no subject to both it's subject to really but i like that you're bringing that up because i i was doing a deep dive around uh inversion attacks of embeddings, right? There are some that are showing that when you invert an embedding, even a dense embedding, you can get up to 92 % of the original context. So that's an attack, right? That's an attack. However, what is more likely is that you are going to be leaking that information as part of the output of the model.

7:32So if you, for example, have embeddings that contain information about people's salaries, for example, that can get leaked out. In one of the examples for a data leak for one of these OLMs was finding out the salary of a CTO at a company. This can happen with embeddings just as much as with fine tuning. Interesting. Interesting. And was the example, the CTO salary example, that was an embedding leakage? I don't know. Okay. I don't know if that's public info. Can you talk through an example of how an embedding leaking data would work? I guess I'm imagining a scenario with regard to a particular condition and somehow the name that it uses as an example is a real person's name.

8:27Is it like that kind of thing? Yeah, so these embeddings can be more than just word embeddings, right? They could be representative of paragraphs, of documents. They could be chunks that are basically easy to search. And then when you send in your query through your large language model, it's going to go through an embedding model, look through your database of embedding, see what's closer to the query, and then produce an output with that information. And so if that query is asking for something sensitive, that sensitive piece of information could be leaked. And it could even be leaked accidentally, even if you're not asking for something particularly sensitive, but it is relevant to the context of the question.

9:10But then we need to distinguish between sensitive information being in the corpus of chunks that the embeddings are referring to and leaking information from the embeddings themselves. So we need to distinguish between the leak, if you have access to the embeddings directly, and how you can pull out information from there, whether or not it's an inversion attack, whether you're doing distance measurements or similar. And you have to distinguish that from output or risk with output from models that are using a database of embeddings to create responses. Does the risk require that the sensitive information still be in the data or can just the fact that the sensitive information was used to produce the embeddings, even though the database has been scrubbed, cause problems?

10:21So if the sensitive information was used to produce the embeddings, the chances are that there's sensitive information left over in the embeddings themselves. that could be output um and so that's why the the recommendation there is de-identifying data before you create embeddings or understanding what data you have in the first place to not include that personal information uh period um and just stripping out that entire entire piece it depends if you need the context or not uh around that personal information and the interesting part is that a lot of the times you don't need the personal information to get value from a lot of NLP applications.

11:02If you think topic modeling, sentiment analysis, and many, many other ones, the sensitive information is generally superfluous. Unless you're doing something very explicit in HR or need specific customer information for things like your master data management system, you're not going to you often don't need that personal information and in addition to that the personal information can lead to more bias in the responses and so that's something else to watch out for more and more research is coming out about how things like physical attribute or the country that you're from or the political affiliation etc that can all lead to biased responses from the model.

11:49And so by removing that, that removes the or reduces the risk of having biased responses. Got it. Got it. Got it. So that kind of brings us full circle to how folks might use private AI. So one is you are interacting with third party LLM and you've got sensitive information that you don't want to be exposed to that LLM. And so private AI would sit between your system and the LLM and kind of cleanse that data as it's passing through. And then there's another part that is about if you're training, doing fine tuning or training embeddings or some aspect of using data to build a model, then you can also cleanse that data kind of in, is it like in place or is it as part of the training loop that you're working with that data?

12:44It's an API. So it runs in your environment. You send a post request through and then out comes the redacted data and the list of personal information and sensitive data that was found. And so you could put it anywhere in your software pipeline. That makes sense. Got it. And are you typically doing it request by request or like are there batch or bulk APIs? There's a batch mode and also a screening mode as well. It sounds like a foundational part of what you're doing has to be kind of entity recognition, identification in this data because the user is not expected to tell you specifically the entities that are, well, what information does the user need to provide?

13:35This may be a better place to start. Yeah, the user can provide the, so we provide a list of entities that they can identify. and they could either choose to by default remove all of them they could choose which ones to remove if they don't want all of them or they could choose based on regulation that they're trying to comply with so if it's HIPAA or if it's GDPR CPRA you could choose which one you want and then it'll auto select which entities make most sense for that regulation thinking about entity recognition in general, like that's been a problem for a long time. And on the one hand, because it's been around for a while and there are lots of different solutions, I think it's tempting to think that it's like a solved problem.

14:23On the other hand, you know, if you have any experience trying to get LLMs or, you know, even traditional tools to do this, it is very hard. Talk a little bit about your experience with the entity recognition part. Yeah, I think that a lot of the misconceptions around what is easy or hard in AI, when it comes to thinking that it's easy, it tends to be around misunderstanding the huge lift that it takes to get high quality representative data. and then the misunderstanding how huge a lift it is to get a product in production that is scalable that also has a lot of corner cases in mind and part of the part of the problems that you have to deal with when building a named entity recognition system are you're going to deal with optical character recognition errors you're going to deal with automatic speech recognition errors, depending on, of course, your use case.

15:25But if you're dealing with a lot of data, like many of these large organizations are, the different types of contacts is huge. And then the different languages that it needs to apply to as well. And then all of the different entity types that you need to be able to perform well enough on for it to be good enough to comply with data protection regulations. So it's a huge balance between speed, accuracy, and then multimodality in addition to that, because it needs to generally deal with results from files and of course, text, transcripts, chat logs, but audio as well. And if you put all of that together, that's a giant problem space to deal with.

16:14So it's named entity recognition itself is an important piece of the puzzle. Everything else is also a huge lift. For the named entity recognition lift, if we double click on that, you have to think about the 50 plus entity types that you often need to recognize for some of these regulations multiplied by the number of languages that you need to function on. And you need to do that in a way that's also speedy, right? So we've created data pipelines that are very efficient with regards to adding new entity types. We've created data pipelines that are efficient with regards to creating synthetic data to add to our current data.

17:01Because there's no real personal information that you could go out and capture in the wild. you have to basically either beg customers for data or you have to create your own. And then finding out what kind of data you need to add to the model in order to make it more efficient or more effective while limiting the decrease in efficiency in your data process with the more data that you add. Is the customer training the entity recognition model? Or it's supposed to be zero shot with regard to the customer data? Yeah, absolutely no training by the customer. And for PII in particular. And here's why.

17:50It takes a lot of training to know how to annotate properly. if you introduce errors in the annotation process that can have bad effects in the models then we can't help you debug the model because we don't know what kind of data you put in there the in addition to that the

18:14basically we won't be able to provide the same model of that's that's consistent for the kind of corner cases that you have that might be relevant to other customers as well. Meaning if a customer's fine tuning, you know, what you're doing, then that it impacts repeatability across customers. And then are you as part of the service, like, are you making some type of guarantees to the customer or anything like that? We do have a warranty that's optional that folks can buy, and that's provided by Armilla AI. And it has a certain accuracy threshold that is guaranteed that we reach, if not your contract money is returned to you.

19:07And we have a certain time period in which to fix anything that falls below that threshold. And so your ability to honor kind of warranties like that is going to be impacted if you're customizing the model for each customer. Yeah. However, we do deal with extremely sensitive information, often things like credit card numbers for PCI compliance or health care information where folks do not want to have to go into the health care data and look around to see whether anything was missed. And so our mistakes are pretty high when it comes to accuracy. Can you talk a little bit about like how about generalizability and like how you've built the system to be able to deal with generalized data?

19:55You know, without without training or tuning to specific customer use cases, I would imagine there's a fair amount of even though we're limiting ourselves to entities, I would imagine there's a fair amount of variability in the types of files and documents. And that could be problematic. Correct. um lots of data is the answer to that okay lots of data collected over the years uh where uh oftentimes it's our customers saying hey you missed the stuff uh fix it and they send us a chunk of data and then we'll add synthetic data to that uh and do a lot of rounds of testing um But yeah, it's a big lift to be able to have a generalizable model like that.

20:42You mentioned multilingual as well. Is that also solved with more data and the models deal with it on their own? Or are there specific challenges and things that you have to do in order to support multilingual? There are specific challenges. and oftentimes it could be things like when you have multilingual data conversations, for example, so code switching. That can be fairly tricky to deal with. Also, the ability to detect what kind of language is being used to be able to provide the best model. That's something that we put in place as well, so a language detection system. the ability to also based on the language detection system provide different labels in the appropriate languages is also something that we've worked on and so there are more data definitely but also how do specific models deal with logographic languages like Chinese characters for example, versus alphabets like English.

21:57And being able to find models that are suitable for each of these, in addition to finding optical character recognition models that are suitable for each of these, and training them with the right data is very tricky. And you've mentioned OCR a couple of times, is the presumption that your customers are not responsible for the OCR themselves. They're giving you whatever their raw data is and PDFs and scans and all that, and you have to deal with it for them? It depends on the customer. But oftentimes, we have to deal with it. And OCR is another one of those things that is like, oh, that's a solved problem.

22:38And it's not.

22:44mainly again because of data complexity because you've got documents within documents or images embedded in documents and being able to it's not about just surface level comprehension right it's not just a summary that you're providing it's this is really detailed work that's being done and that's meaning you need to be able to find the name in the chart in the section and the you know on a page in an image basically exactly yeah and and when folks say oh i could just do that with a large language model it's really not taking into consideration uh the amount of lift it actually takes for the more detailed work at a large scale uh in in actual production when you're handling so many different variables at once.

23:34So large language models are out of the box, very good for things like summarization, for giving you ideas, stuff like that. But when it comes to really detailed work, the data is still incredibly important. And so as opposed to large language models, you're using more traditional approaches to these various problems, entity recognition and OCR and A, that question, but also like, are you using off the shelf open source or commercial things? Or have you rolled your own for most of the underlying models? Like, how do you think about your model supply chain? So can we think about it in the sense of what our customers really need?

24:24And they're dealing with massive volumes of data. And so large language models tend to be too slow for them and the amount of additional accuracy that one might get with a large language model with where we're trained on the data that we have and adapted for their tasks is often not worth the speed and cost increase for the customers and And so we spend a lot of time focusing on making the models as accurate as possible in a smaller package. And so there are multiple ways you can do that. You can start from a traditional base and kind of build up from there. Or you can take a large language model and distill it down to smaller language models.

25:17Is there a general direction that you prefer for these kinds of tasks? It's a continuous experimentation process. We're always trying different things to see how much juice we can squeeze out of this. You mentioned synthetic data a bit. Can you dig into that a little bit more and some of the ways you're thinking about that? I can a little bit. A lot of synthetic data is helpful when you don't have enough data for a particular problem. I think one thing I want to point out there is some folks might think that synthetic data is the right way to go for training their models, which might be true, and it worked for us.

26:00And the reason that it worked for us is because we also have a good balance of real-world data in addition to the synthetic data, and we've informed that synthetic data with the real-world coordinate cases. when it comes to fully synthetic data, the hesitation there is that if you're able to create fully synthetic data for a problem, that problem might, in some cases, depending on the problem that you're trying to solve, that problem might already be solved because you were able to create that data. In other cases, you might lose context where you might need things like conversation flow information or did a particular customer service agent need training in a particular case and medical information that's related to one another, to different parts of a different time, a full timeline of a patient, for example.

27:02And so the type of synthetic data that we use is usable because we have access to high quality original data. And we also have a very specific task where this works out well for it. And are you, like at what level are you primarily using synthetic data in the process of training the entity recognizers? Or are there other parts of the system where synthetic data comes in handy for you? We also generate synthetic personal information as part of an option for the output. And that's one of the interesting pieces where if you want to keep things like conversational flow or topic modeling and things like that, you still can because the majority of the context is still there.

27:52Meaning as opposed to masking, you can generate synthetic data that takes the place of the original tokens. Yeah. Interesting. Okay. Yeah. And so very similar to masking, you can still keep the context of the data. We've talked a little bit about OCR and how are you using that as one example of multimodality. Can you talk about other ways that you're using and incorporating multimodal data? Yeah. So we also do provide the ability to redact information generally for PCI compliance from audio recordings. And so you can bleep out the personal information from the audio recordings themselves. Meaning I've got a call center and I'm recording for quality assurance and people are talking about personal, they're incorporating PII into those conversations and you want to redact it before you store it, for example.

28:46Correct. That's right. And that's really important for PCI-D. assess compliance, where you do need to remove things like credit card numbers and account numbers. And the way that we do that is very much by, you know, we bleeped out the PII. One thing I will point out is that there is some research on anonymizing the voice within a recording. And that's interesting because if you anonymize information, it falls outside of certain data protection regulations like the GDPR. And for those listeners who don't know, the GDPR is the General Data Protection Regulation in the EU, and it basically covers what you're allowed to do with personal identifiable information and consent and all that jazz.

29:32So if you anonymize data, which basically means there's a very, very low risk of re-identifying the individual, then you don't have to comply with that regulation because you're no longer dealing with personal data. But anonymizing the voice doesn't remove responsibility to anonymize PII that's mentioned in the... Yeah. That's correct. And so the voice is a biometric identifier. And you can, I'd say at your own risk, modify the voice and perhaps call it anonymous, but there's not enough research out there, in my opinion, to say that a voice modification cannot be reversed. and so there are certain assumptions you might have to make with regards to whether or not you can consider an audio recording actually anonymized but data minimization is not just about anonymization it's about limiting the risk of the data and in situations like removing credit card numbers you're reducing the risk of fraud for example and if there's a data leak reducing the risk of how much information is going to be part of that data leak.

30:44And are you doing this on stored audio or do you have the ability to do it real time? It could be both. We are not an ASR company. That sounds a lot harder. Yeah, yeah, it is hard. It is hard. But we are not an ASR company, to be clear. So we can do it on audio. But if anybody needs something that's more specialized, like, you know, things like diarization of audio or things that, yeah, if you need actual control of what's going to happen to your audio aside from de-identification, there are lots of good companies out there like Deepgram, which I saw that you interviewed, Assembly, and Twilio and so on, who can do a great job at that.

31:29You have an audio stream coming from your call center. People are mentioning PII and credit card numbers and the like on your audio stream. You can real-time redact certain types of PII as it's happening. Is text being used as an intermediary or is it happening real-time on voice? Yeah, text is being used as an intermediary. If it were to happen real-time on voice, you'd be talking about probably keyword detection. and then you have to train a model specifically on keywords. That's just a very large search space of keywords to be looking for. Got it, got it, got it. I mean, you could train it for digits, for example, and probably be okay, but that's a big problem.

Read the full transcript

32:19So now I'm thinking about the, you know, even with text as an intermediary, you're either doing like word level resolution where each word has a time associated with it, but then I'm imagining there's certain types of PO, well, even like a string of credit card numbers, like you have to zoom out beyond individual numbers to know what's happening there, you know, or you can look at like complete utterances, but then you still have to figure out like how to align the credit card number that's in the middle of a phrase with the with the audio like how do you deal with that is that a big deal for for doing this or and how do you deal with that no it definitely is and and the reason is because of of disfluencies in speech right so somebody might say my credit card number is five nine oops i dropped my card sorry it's 3, 2, 5, 9, 2.

33:17Oh, no, that was a 2, not a 4. Yeah, and I wasn't even thinking about that. That would make it even a lot harder. I was just thinking about having been looking at, you know, looking at some transcription ASR data recently and the different levels that you can use it. Yeah, no, it's definitely tricky. So the more context, the better, definitely. uh the um the cleaner the transcript the better as well um and if it's uh only one language it's also easier uh but the the world isn't perfect and none of these things really apply most of the time maybe the unilang unilingual aspect but uh even then like what i guess i'm trying to get at some kind of practical detail there in terms of either, you know, models or, you know, things that people should be thinking about in terms of the data, or if, you know, someone's encountering problems like this, what should they be thinking about?

34:20Or what have you learned about tackling it? Real world examples are extremely valuable. And you might have all sorts of, so that's the That's a limitation of fully synthetic data. What you expect to encounter is very different from what you might encounter in the real world. And oftentimes, it's very hard to get access to that real world data because of personal information. And so one thing you can do is use our synthetic PII model to generate just the synthetic personal information and get access to the rest of the data, which might be relevant for your more minute use cases. Meaning just inject synthetic PII into random places and conversations?

35:12instead of the actual PII. Because oftentimes you can't get access to data because of the privacy concerns. Meaning if you're building models and part of the reason why you can't build strong enough models is because you can't get the data because of the PII concerns, you can use a tool like yours to mask the PII or make a synthetic version of it with synthetic PII and then convince somebody to give you their data based on that. Correct. Yes. We have customers who use us for that purpose. Okay. Okay. Got it. And there will be some who have contractual agreements for that with their customers where they promise to anonymize or reduce the PIA and the data in order to be able to train their models with it.

36:04And that's becoming more and more important as people realize the actual harms of training models with sensitive information. Also on the multimodal front, I think when we spoke previously, you mentioned something about like logo detection in documents was a challenge that you ran up against at one point. Yeah, yeah. So we have more and more requests for confidential information detection because also of large language models. And so we've built out a solution that deals with that. In the confidential information detection piece, That one we do handle with customers to adapt to their particular confidential information because that changes even on a daily basis for some.

36:51But logos is actually something that comes up as confidential information. So if you have, for example, customer presentation and you want to train a model or send that through to a model and you don't want to send that customer logo through, That's something that we can blur out so that you don't have to deal with that piece of information. But that's another thing about building a privacy solution. It's so many things that come up that you don't expect that you need to collect knowledge about year after year to keep improving the product. So it's not something that you could just build out of a box with a named entity recognition model that recognizes five entities.

37:34It's a real drag to actually figure out what's what. How do you maintain like a coherent, how do you maintain your product as a coherent system, as opposed to a patchwork of, you know, kind of individual things to address individual issues? I'm imagining that becomes a challenge just in terms of, you know, manageability of the system and like, you know, code manageability. Like, is there a particular approach that you found useful for that well one having great architects and engineers uh but i would highly recommend if you can afford them um but to um the there can definitely be product scope uh product creep or feature creep and uh we'll have folks requesting entity types that don't have to do with privacy and if we can't fit it into the framework of whether or not it has to do with privacy or confidentiality, then it's not something we really consider doing.

38:43We do have many folks using us for named entity recognition because of how highly accurate our system is. And that's where that creep can come in, in those requests. So it's really about keeping it focused. Otherwise, I think people already sometimes have a hard time figuring out what we can help them with, thinking we just do de-identification when we can also do this detection and incorporation into their risk pipelines and their data catalogs and MDM systems and so on. So not having that go beyond the realm of our goal is privacy and confidentiality and making sure that it's as important and as good as cybersecurity processes.

39:30Can you talk a little bit more about the integration with data catalogs and MDM systems? Is it just that you have masked the data before it gets put into those systems, or are we talking about something different? Different, different. In this case, it's using that named anti-recognition aspect of it to create metadata that goes into those systems. So while you're de-identifying data to go into your models, for example, you're cataloging that data and putting it into your data catalog so that you could go back to it and see what the risks were or are of specific data sources. And you don't have to go and redo the work every time.

40:09So while you're de-identifying, what is the data that you're putting into your data catalog? The entities that you pulled that you identified or the de-identified snippets? The entities that are identified, along with the data source, as an example. There can be other aspects to it as well, depending on what aspects of data protection regulations you want to comply with and how you want to build out that internal privacy pipeline. But yeah, that's very much it. So if you think about it, there's a really robust system right now in place within organizations to deal with compliance for structured data.

40:54And 80 to 90 % of the data they have is unstructured. There is no real robust system for folks to deal with that at the moment. And with our technology, what people can do is build out that system with that single API to pull that data from the unstructured side of things into their pre-existing data catalog and MDM system. And that means that suddenly you've got that robust data management system covering all your data, not just 10 to 20 % of it. And one really good place to start is if you're already dealing with that data for building up your large language models, you're looking at the risk of the data, you might as well use that information to start updating your internal systems.

41:42You mentioned earlier a bit about how the connection between privacy and ethical, responsible AI and how the personal information leaks or can leak into models and cause downstream biases. Can you elaborate on that a little bit? Sounds like that comes from research that you've come across. Yeah. So I used to, when I had a little bit more time, co-organize this workshop at various computational linguistics conferences. The last one was at ACL. And in this last one, there were a couple of papers talking about how the input to a model can affect the output based on bias. and one of them was about physical attributes and how that can affect the, if those physical attributes are about disabilities, that can affect the output of the model.

42:45And when we first came out with private GPT, we were also looking at what kind of biases can the model have depending on the PII. One of the experiments we ran, which has been since corrected, was if you write out, my friend is British and he's in prison, what are some of the reasons that he might be in prison? You get very different results from my friend is from Somalia and is in prison. What are the reasons my friend might be in prison? And if you just say my friend is from country one, what are some of the reasons he might be in prison? You get much more unbiased results as an output. And so what this is showing is the inherent bias that is in the internet is being showcased in the output of these large language models.

43:37We can try to minimize that by minimizing at input the pieces that might lead to that bias. Got it. So as opposed to trying to mitigate it via prompting techniques or various other things, if the input data is kind of has these inherent identifiers, whether they're names or countries or genders or less obvious things even, you can identify those things and mask them and that will reduce bias in the LLM output. Correct. Yes. And some of these that do end up leading to bias are actually classified as special categories of sensitive data under the GDPR. And you are not allowed to process them under the GDPR, which is interesting because technically that means that we might not be allowed to process it even if our goal is removing them.

44:35Right. So that's a little bit of a... That's always been a bit of a conundrum. Yeah. Yeah, yeah. But we're helping the problem anyway. Right, right. But things like your religious affiliate, political affiliation, religion, sexual orientation, those are all things that are special category data. And you have to be extra careful with those. So it's not just, I don't really want to use the word just, but it's not really about only complying, only trying to do, this doesn't sound good. It's not about just ethics. It's also about compliance. And both are very important. But the ethics are being reflected in the law in this case, which is nice.

45:22And on the topics of law and regulation, there are a host of new laws coming online, EU AI Act, the, you know, things in California and the U .S. are evolving. Like, I'm imagining that that's something that you keep an eye on. Absolutely. uh what what can you what's your kind of take on the landscape and um you know it's probably a whole different hour-long conversation but like how are you seeing things evolve interesting moment in time to ask that question

46:04so the the euai act when it comes to uh ethical ai uh it covers things like bias what is or is not an acceptable use of AI. And it really applies to the higher risk uses of AI. And with regards to sensitive or personal information, that's all covered still under the GDPR, regardless of whether or not the UAI Act applies to you. And one thing that's really nice about regulation, when it's clear, is that it provides guidelines to organizations that are struggling in figuring out what the best practices are, because they could go look at the regulation or the standard and say, okay, this is what is said is okay.

46:45Once we're done this, we can innovate and there's no limbo. I can't say there's no limbo with the GDPR. There are a few things. I don't even know if it's possible to comply with the GDPR at the moment with the technology that's available except for more constrained cases. But where part of the world is headed is there's this issue of trust And this issue of trust can be grappled with with appropriate regulation, with transparency, with making sure the public and the users are aware that their concerns are being taken seriously. And then this other part of the world is currently putting a lot of investments in AI and just dismantled the guardrails, if you will, that were recommended around AI.

47:39and that is not going to be particularly helpful with dealing with the huge and growing concerns that the population has around misuse of AI and that's not good for organizations that are trying to integrate AI into their business either because now they have to deal with that growing concern without any guardrail to point to that's official that so leads to inefficiencies down the line. So then taking a step back with all of the things that we talked about in mind, like how do you see the space evolving? How I see the space evolving is there are certain industries that are ahead of the game when it comes to responsible data practices and have already embedded them within their data practices for over a decade now.

48:30And there are other industries that are just catching up when it comes to unstructured data. And I think there's going to be a lot of cross-pollination once folks start catching up. What I'm seeing with regards to larger enterprise is that the ones that have already been thinking about AI for a while are now starting to think about the guardrails. The ones who are just starting to integrate AI are not doing necessarily trust, ethics, privacy by design. They're waiting to find out what they're going to be doing with AI first. And then they're thinking about what they're going to be doing with the privacy aspect of things or the ethical aspect of things.

49:12And that can be quite problematic because you're not building the pipeline that you need out of the bat. And so that's going to lead to you prototype, prototype, prototype, and then, oh no, we have to go back and reconstruct everything to be able to get this approved by the board, the chief AI officer, the chief privacy officer, the CISO, whatever. and if you start thinking about it from a this is the problem that i want to solve what does the pipeline look like that i need to build in order to solve that problem and then start addressing it from the data aspect of things that can lead to a bigger lift on the up front but down the line a much more efficient organization because you don't have to constantly just do band-aid solutions to the problems.

50:05Yeah, so data first. Data first. That's it. Surprise. I feel like something at every AI wave, we keep forgetting. And we try to ignore it. We try to pretend that we don't need it anymore, but it's not the case.

50:26Yeah. Awesome. Well, Patricia, thanks so much for jumping on and sharing a bit about what you're working on. Very cool and interesting stuff. Thanks so much, Tim. Thank you.

51:03Thank you.

From the publisher

Today, we're joined by Patricia Thaine, co-founder and CEO of Private AI to discuss techniques for ensuring privacy, data minimization, and compliance when using 3rd-party large language models (LLMs) and other AI services. We explore the risks of data leakage from LLMs and embeddings, the complexities of identifying and redacting personal information across various data flows, and the approach Private AI has taken to mitigate these risks. We also dig into the challenges of entity recognition in multimodal systems including OCR files, documents, images, and audio, and the importance of data quality and model accuracy. Additionally, Patricia shares insights on the limitations of data anonymization, the benefits of balancing real-world and synthetic data in model training and development, and the relationship between privacy and bias in AI. Finally, we touch on the evolving landscape of AI regulations like GDPR, CPRA, and the EU AI Act, and the future of privacy in artificial intelligence.

The complete show notes for this episode can be found at https://twimlai.com/go/716.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
Ensuring Privacy for Any LLM with Patricia Thaine - #716The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 52 min
Listen in VO