Using edge models to find sensitive data

13 Jun 2024 · 38 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Practical AI Podcast Episode Summary

Episode Title

Using Edge Models to Find Sensitive Data

Episode Description In this episode, Daniel Whitenack welcomes Ramin Mohammadi from Tausight to discuss the challenges healthcare organizations face regarding sensitive data, particularly Personal Health Information (PHI). The episode covers how edge AI models are deployed to search through vast amounts of records for PHI, highlighting the importance of privacy in the healthcare sector.

---

Key Participants

  • Ramin Mohammadi: AI and ML lead at Tausight, adjunct professor at Northeastern University.
  • Daniel Whitenack: Founder and CEO at Prediction Guard.

---

Key Concepts Discussed

  1. Understanding PHI
  2. Definition: PHI refers to any health information that can identify an individual, particularly as outlined by HIPAA regulations.
  3. Significance: PHI includes sensitive personal details like medical history and social security numbers, making it a prime target for hackers.
  1. Healthcare Data Breaches
  2. Statistics: In 2023, 133 million healthcare records were breached, affecting one in three Americans.
  3. Causes of Breaches:
  4. 78% of breaches originated from hacking.
  5. Breaches also occur due to human errors like misplaced laptops or phishing attacks.
  1. Challenges in Data Security
  2. Healthcare Organizations' Accountability: Only 30% of breaches come from healthcare providers; 70% are through third-party partners.
  3. Compliance: Breaches lead to legal ramifications, public disclosures, and significant financial losses.
  1. Existing Tools and Limitations
  2. Traditional Security Tools: Current tools often rely on heuristic patterns (e.g., regex) to detect PHI, which can lead to false positives and negatives.
  3. Dark PHI: Much of sensitive data is unstructured and exists in various formats, making it difficult to detect.
  1. The Role of AI and Machine Learning
  2. Application:
  3. AI models can analyze unstructured data and identify sensitive information without relying solely on predefined patterns.
  4. Models can run on the edge, meaning they operate where data is generated, reducing the risk of data exposure.
  1. Deployment Challenges
  2. Model Training: Limited access to real patient data necessitates the creation of curated datasets while avoiding biases.
  3. Edge Environment: The models must operate within constrained environments (limited CPU, RAM) typical of healthcare organizations.

---

Success Stories and Insights

  • Tausight's Impact: Ramin shared instances where their software helped organizations identify sensitive data on stolen laptops and reduce false positives in detecting PHI.

Future Directions

  • Emerging Technologies: Ramin expressed excitement about advancements in smaller, efficient AI models and the potential of federated learning, which allows models to learn from decentralized data without transferring it.

---

Conclusion The episode underscores the critical need for innovative solutions to protect sensitive health information using AI technologies. Ramin's insights reveal both the challenges and advancements in the healthcare sector's approach to data privacy and security.

Additional Resources

  • [Tausight](https://www.tausight.com/)
  • [Graphstuff Podcast](https://graphstuff.fm/episodes/2023-finale-llms-and-knowledge-graphs-throughout-the-year)
  • [Backblaze](https://www.backblaze.com/cloud-backup/personal/landing/podcast/practicalai)

---

Call to Action For listeners interested in further discussions, join the [Practical AI community](https://practicalai.fm/community) to engage with other technology professionals and enthusiasts.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:05Welcome to Practical AI. If you work in artificial intelligence, aspire to, or are curious how AI-related tech is changing the world, this is the show for you. Thank you to our partners at Fly.io, the home of changelog.com. Fly transforms containers into micro VMs that run on their hardware in 30 plus regions on six continents. So you can launch your app near your users. Learn more at Fly.io.

0:42Welcome to another episode of Practical AI. This is Daniel Whitenack. I am founder and CEO at Prediction Guard, where we're safeguarding private AI models. And I'm really excited today because I'm joined by a friend who I've had the pleasure of getting to know a little bit over the past weeks in the startup community. joined by Ramin Mohammadi, who is the AI and ML lead at TauCite and also an adjunct professor at Northeastern. Welcome, Ramin. Hi, and thanks for having me. Yeah, yeah. It's great to have you here. It's been cool to visit the Boston startup community a couple of times and participate in a few events together.

1:25I've been really fascinated to hear about some of the things that you're doing at TauCite, so I'm excited to dig into those a little bit. I'm wondering if you could share a little bit with us about kind of your really thinking deeply about the intersection of AI and privacy, but specifically as related to privacy and personally identifying information, but also personal health information. So PHI, so TauCite is thinking deeply about how companies are handling very private, sensitive data and knowing about where that data is, which is actually a huge problem. I'm wondering if you could talk a little bit about, I was kind of shocked when I heard one of your presentations and you were talking about just the size of this problem and the scope of this problem related to PHI.

2:17So could you give us just a little bit of a sense of what PHI is, why companies handling PHI is kind of a problem, and some of the challenges related to that? Sure, I can do that. So first to give introduction, what's a PHI or personal health identifiable? So based on the HIPAA rule, there are 18 identifiable, which can lead to identify an entity or a person within a healthcare organization. And this information is valuable and being targeted by hackers. One of the reasons is that they have high value because they contain your sensitive personal information such as medical history, social security number, and insurance details, which makes it very valuable on black market.

3:11They also use this for monetary gains. So hackers can sell this stolen PHI to criminals who use it for identity theft, insurance fraud, or other legal activities. They also use it for exploitation and extortion. Basically, they use this stolen health information to use for blackmailing individual organizations. So 133 million healthcare data was breached in 2023, which means one out of three Americans' life was affected. This means about 160 % increase compared to 2022 and about 240 % increase since 2018. So like for our listeners, if at least if they're in the US, one out of every three of those listeners has had some portion of their health information exposed in some type of breach.

4:07Is that right? That is correct. That record came out in 2023. Yeah, yeah, that's insane. And you mentioned hacking. How is this data being breached? Mostly by hackers or by sort of mistakes? What is the combination of ways that this data is getting exposed? That's a great question. So actually, based on the report on 23, 78 % comes from hacking of the network storage, where the data resides in healthcare. And there are a small amount, which is about like a 2 % happens by when someone stole a laptop, for example, and the laptop contained PHI, or sending an email via email, or basically phishing emails, stuff like that.

4:56So there's like a breakdown, but majority comes from hacking, 78%. And are these companies sort of mostly healthcare companies or like who has this data and how is it given? So that's really interesting because technically healthcare organization like hospitals are only, I think, accountable for 30 % of this incidence and the remaining 70 % happens by hacking through their business partners, like third-party organizations that they have some sort of softwares or they're storage for keep tracking of the most medical data. But at the end, the cost comes to the healthcare organizations. So that cost, what happens when this data is breached?

5:43What is the bad case scenario, I guess? Sure. So let me first tell you how the overall costs of this. So healthcare cybersecurity has spent about 28 billion over five years spent, and we are still not able to protect the PHI. And the way that it is that when an organization getting hacked or that there's some data breached, depending on the states, there are some cap. if, for example, if you have been breached and you have lost more than 500 entities or live basically data, you need to go public about it. And also you will basically get sued and also you need to also get fined. So we have this wall, which we call it wall of shame, unfortunately.

6:35And government posts the names of the organization where they got hacked or they lost basically PHI data. And this wall has been constantly updated. And the last thing any CISO or CIO wants is that to see their name on that wall. Yeah, and their brand is hurt by that, but also there's fines, right, for this sort of breach.

7:17What's up, friends? Do you remember when ChatGPT launched? I do. It felt like the LLM was this magical tool out of the box. However, the more you use it, the more you realize that's just not the case. The technology is brilliant. Don't get me wrong, but it's prone to issues like hallucination on its own. But there's hope. There is still hope. Feed the LLM reliable current data. ground it in the right data and context, then and only then can it make the right connections and give the right answers. The team at Neo4j has been exploring how to get results by pairing LLMs with knowledge graphs and vector search.

7:55Check out their podcast episode about LLMs and knowledge graphs throughout 2023 at graphstuff.fm. They share tips on retrieval methods, prompt engineering, and so much more. Don't miss it. Find a link in our show notes. Yes, check it out. Graphstuff.fm, episode 23.

8:28Could you talk a little bit about like, let's say that I'm a healthcare company and I I want to not be on that wall of shame and I want to do the best practices and all of that. Like, what is the reality of, you know, I know that, of course, you have been thinking very deeply about solving these issues with AI and machine learning or some related issues. But let's just say that doesn't exist. What are the choices that a company has in terms of, and what are the challenges that they face in terms of securing this data? So currently, healthcare organizations have a series of tools, some of them like maybe four or five tools that are going to do the same task.

9:14And the way that these traditional tools work is that they have a series of patterns like a regex, for example. And when someone tries to download the data, if this matches, for example, with some regex that they have, it will say, hey, you transferred PHI or hey, you downloaded PHI. So those types of files, which are coming directly out of EMR, electronic medical records, are easy to detect. The problem is with what we call dark PHI, a PHI that resides on your network, on your machines, but you are not able to detect it because the patterns that you are using are basically not capable of detecting.

10:00We have, for example, an organization that has millions of patients in your network. And I don't know, probably you have written Regex before. Regex is always as good as the person who's writing it. Yeah, yeah. And also, what is that meme? It's you decide to solve a problem with Regex, and then you just end up having another problem, which is Regex. That's correct. So current from the prospects that we have talked with is that updating these rules is costly and requires a dedicated engineers or IT professionals. And no one likes the right regimen. Very true. At least I can say that. So that's a problem.

10:47They have tools in place, but it's incapable of solving the problem. And when you say dark PHI, I'm assuming there's like, okay, you might have a regex for a social security number or something like that. But if I just think of like a doctor recording a dictation of a patient visit or something like that, there's a lot of natural text in there about diseases and all of those sorts of things. so is that more natural text uh sort of health information is that what leads to kind of the dark phi or does it also have to do with you know oh it's easy to detect in this file format because i know the the pattern i'm getting but then someone scanned in this document and it's in like a pdf or something and my script doesn't know how to scrape these different data types what could you kind of go into i'm super fascinated by this idea of this dark phi sitting around so the first thing is that point out that is the 80 percent of healthcare data are unstructured unstructured means that from image to audio transcript like all sort of pdfs and what we have seen also on the healthcare it's a variety of data extensions so you will be surprised that that you will see file extensions that don't exist.

12:18But what the clinicians or researchers do is that when they did a file, they put dot, their last name. Oh, jeez. So I think in the last study we did, we found 8 ,000 extensions on our MSQ prospects environment. So literally the personal information is in the file extension. That's correct. That's crazy. Personal information could be in the extension, but also random file extension that they might use in order to basically bypass some rules. Oh, okay. I gotcha. So they're doing a workaround because there's like this annoying tool that prevents this data from being transferred around is blocking this file type.

13:07So if I just change the file type. That's correct. So one thing also here to point out is that healthcare lives on data. Clinicians need to access that data. And you should not stop them basically from doing the job. You just need to have a better way to detect and basically maybe, for example, encrypt the file. So if someone else stole that file, they cannot open the file. Our goal here is not to prevent clinicians from doing something, it's just to make it more secure. Yeah, and maybe that starts to get a little bit to kind of transitioning to how AI and machine learning fit into this puzzle.

13:52Before we exactly describe how it does fit into that puzzle, I think there might be a lot of listeners out there that are very intrigued by what the challenges might be of applying AI or machine learning in the context of healthcare. What are the unique challenges if you're a data scientist and you're building a model or wanting to use an LLM or building your own model to use in a healthcare context? What's unique about that context that makes it more challenging to maybe it's on the deployment side or the model building side? What are some of the challenges related to working in healthcare specifically with this technology?

14:35This is an interesting question. As someone who's keen on MLOps, I always say the main challenge is the whole project. But for us, these challenges are a bit more than some other AI technology due to the space which we are in. For example, we don't have access to real patient data to train our model. And no healthcare organization will agree to let you use the data. And probably if one of them agreed to it, I'm guessing you couldn't use that same model for a different organization because you trained on specific data that's sensitive for one organization. Is that right? That's also correct. We have a huge data heterogeneity problem.

15:22And that comes, for example, one organization, it's for cancer organization. The other organization is like a dental organization. These datas are different. Other challenges that are against is that you also cannot collect or transfer any data to the cloud. That means everything needs to happen at the edge. Data labeling is highly difficult. Even human level performance has about 8 % to 10 % labeling error for detecting PHIs. Again, for example, you have lots of types and extensions. Data normally contains bias. Certain demographics have higher amount of data than others. model development is confined by first model performance and then optimization metrics.

16:16And model deployment on the edge has its own difficulties, which we can talk about later. I think lastly, unsupervised model monitoring makes it more challenging to detect drifts. And just to kind of define a couple of those things, when you say the edge, what is your, because people might have different definitions in their mind of whether that's a, you know, some staff member's laptop or a desktop in a lab or something like that, or like a phone or a microcontroller somewhere. Like people have a range of that. So in the context of healthcare, what is the edge environment? Edge could be from laptop that the clinician is working with, from the desktops, could be the tablet, for example, that you're using, and could be the server storage, basically.

17:07Gotcha. But all sort of on-site with some healthcare data center or where staff are working on-site. That is correct. Gotcha. And when you say unstructured model monitoring, what do you mean around that? So what are you monitoring for? You mentioned drift. So I'm assuming that that has to do with like this PHI might also be changing itself in terms of what it's like. There's a new form type or there's a new thing that starts being collected. Is that what you mean by model monitoring and that drift element? One of the things that we will say is that the change in the distribution of the data, as you scan across different groups within the same whole organization, for example, there's a group for radiology versus there's a group for like a normal PCPs.

18:00And the datas have different distribution, and we need to be able to detect any drift in the data distribution as early as possible. Sometimes you also might find something like concept drifts, where it's more like contextual, maybe a file that under certain scenarios considered as PHI. In some scenarios, it's not actually PHI. There are some rules over here, which makes it more difficult. Yeah, contextualization, I guess, is a challenge. Interesting.

18:45What's up, friends? I love Backblaze. I'm happy to have them as a sponsor. Backblaze makes backing up and accessing your data astonishingly easy. This is a service I personally use. Go to backblaze.com slash practical AI. You get unlimited cloud backups for Macs, PCs, businesses for just$99 a year. You can easily protect business data through a centrally managed admin, protect all the data on your machines automatically, easily deploy across multiple workstations with various deployment options. You can add on enterprise control, including granular access permissions, advanced single sign on group management controls and compliance support.

19:26They even offer multiple restore options, including rapid recovery in the event of data loss or ransomware. That sucks. You can access your backed up data from anywhere in the world using their web app or their iOS or Android app. You can even restore by mail. They'll give you a hard drive with all your data shipped to your door. You buy a hard drive restore, send the hard drive back within 30 days and get a full refund. Get one year file retention and version history. Over 55 billion with a B files restored for customers so far. Visit backblaze.com slash practical AI so they know where you came from and continue to support the show.

20:04This is a service obviously recommended by me, but also by New York Times, Inc. Magazine, Macworld, PCworld, LifeWire, Wired, Tom's Guide, 9to5Mac, and just so many more. You receive a fully featured or no risk trial at backblaze.com slash practical AI. Again, they're supporting the show. Go there, play with it, start protecting yourself from potential bad times. Start today.

20:43Yeah, so maybe we could get a little bit now kind of into some of how you've been thinking about and approaching this problem and thinking about it from the TauCite perspective. So in the context of this edge environment, in the context of this unstructured data, in the context of the constraints that we just talked about, how did you and your team specifically think about applying AI and machine learning in the context of detecting PHI? And maybe also, what is your goal here? Is your goal to stop breaches? Is your goal to provide insights about this PHI? How did you decide on what the main problem is you wanted to solve and why AI or machine learning was relevant to solve that problem?

21:33I'd like to actually first tell a short story. That'd be great. Yeah. Six months ago, I was going to the AI Summit in Austin. I got to the airport and I was passing the TSA pre-wear. The TSA agent asked me to check my bag. It turned out that the machine picked up on this pre-workout container, which I had in my bag. The agent used this device on the box and the result was positive. He was like, oh man, I need to call for this special unit to come and check this. I was like, sure. Then this special unit with something like a hazmat suit, they came and used a kit with a bunch of different reagents to more specifically test my pre-workout.

22:21They started sampling from the pre-workout and added a bunch of these test tubes. Long story short, the agent was like, yeah, you're good. He was like, this happened to my cousin also, these pre-workouts are causing false positives. So we are like that special unit with a bunch of different model trains to find and protect the dark PHI. While the current tools in the market are like the first and the second machine, which leads to false positive or unknown false negative. At Tao's side, we do see this problem as a personal problem. It's our PHI that's being targeted. And clearly, the current tools in the market, they cannot protect it.

23:06So HIPAA security rule says that you must do a complete and true assessment of all your risk and vulnerabilities to ePHI. It is so fundamental to what we need to do. And AI is such a critical piece of taking advantage of the newer technology around. Solving what used to be a labor-intensive problem that are could be much easier if you can define the scope of the problem and you can have machine learning models which can run effectively and accurately in a calibrated manner. That's what we do. We take advantage of the AI to find sensitive data. I think we get to the point where risk and threats and vulnerabilities are going to be detected at the edge using the AI.

23:54As opposed to, gee, I have all these heuristic rules, which is how do we lots of stuff today when it comes to recognizing patterns. So at our site, we use AI, for example, to recognize when the sensitive data is in unstructured content. It doesn't require us to say, hey, there's a keyword here. There's another keyword here that would be how you do heuristic programming. Combination of these three words must mean this. Combination of this form must mean this. you never get to do all the rules. You will never get to the variability you need. So in our model, for example, one of our model with about 50 million parameters model that can be set to recognize this stuff.

24:40You will never in years of programming get that much logic into your rejects, right? So now the other main factor is for us its ability to run these models right at the edge where the data is being created, emailed, printed, copied, or faxed. We bring the AI to the data rather than taking the data to the AI, which most of the current AI solution do that. By doing so, we can ensure that our data is always protected, agnostic of hardware, spec, or network connections. Yeah, so all of what you said makes a lot of sense in terms of the approach and how you're applying AI and machine learning. But also in my practical sort of data scientist mind, I'm like, oh, man, that's really, really difficult to sort of have these deployments of models, especially against sort of heterogeneous types of data, run them on the edge, run them across a diverse set of hardware.

25:45From your perspective as a practitioner, where do you think was the most challenging of those issues? Was it having to do with the deployment targets and the diversity of those? Was it having to do with the types of models that you could or couldn't run in those edge environments? Did it have to do with the actual training and labeling of the data? I imagine all of those are really difficult problems to solve and you had to tackle all of them. But what were you maybe, what was some of the hardest problems to solve with respect to those things? I definitely will say the first but the most challenging problem is the data labeling and data creation.

26:29Because we don't have access to real patient data. So we need to create our own curated data set, which we need to ensure that we don't introduce bias, creation bias in that data. The other thing comes around the model training. So our solution needs to be able to live alongside other programs that are running on a given machine within a certain performance boundaries. One example I give you is that IT has set of rules. If there's an application surpasses certain memory or CPU, it will block that application. So you need to be sure that all this ML models that you have or this ML pipeline that you have always remains below this basically boundary.

27:22And is that because like these are essentially, I mean, I might say mission critical, but these are sort of life critical systems, right? Like they're using these to treat patients, right? So if they, if you pull the memory and the thing stops working, then it's potentially a life threatening type of situation, or at least a very concerning situation in the healthcare context, right? That is absolutely correct. Gotcha. Yeah. And in light of those constraints, of course, some people now might just say, oh, well, we've got all these LLMs now and they're great at doing all of these things. But I'm guessing a lot of those aren't sort of fitting for this sort of environment, these memory constraints.

28:07So where do you go with that? Is it looking back to sort of traditional NLP sorts of things? Is it model optimization? Is it a combination of those? How are you balancing the constraints, but also kind of looking forward to these new generations of models and that sort of thing? Regarding the LLMs, I was reading about this Phi 3 by Microsoft, the small model. And even that model requires a certain amount of core RAM or GPU. None of the healthcare organizations have computer with those specs. They all have like four gigabytes of RAM maximum and some legacy CPUs. The other problem with LLM is that it introduced some additional risk to the health care.

29:04Some clinicians or researchers, they're using tools like, for example, ChatGPT to copy-paste patient data to get some summary extraction, which is not how it should be used. I know, for example, I know actually your company, PredictionGuard, prediction, Guard. You are trying to solve a problem like that. Yeah, there's certainly a lot of people pasting things into chat interfaces. That's very concerning. I'll definitely say that. Yeah, for sure. That is great. Now, when it comes to model optimization, we take a series of approaches to be sure that our models are optimized for such an environment.

29:44This could be from knowledge distulsion or student-teacher networks, quantization, and model pruning. We do a technique combination of all of these to ensure that every model that we have lives within a certain boundary. Gotcha. Yeah, yeah. That makes sense. So it's very important. And I'm guessing the model architectures and the approaches that you can only go so far. It's not like you're going to take LAMA 370 billion and do these optimization techniques and fit it into four gigabytes of memory and run it on a CPU. So it's super interesting. And I think, I don't know, what is your view? As maybe you observe in the marketplace, people are exploring these open models, exploring bigger models.

30:30But at least in the space that you work in, the only way that you kind of move forward is with small models or customized models, optimized models. How do you view that kind of shifting into the future? Do you think there will always be this sort of diverse set of environments in the healthcare space that you need to optimize models for? Will they eventually get over their hurdles of using kind of a cloud or large models? How do you see that developing moving forward and into the future? A report by Schneider Electric indicates that currently 95 % of AI workload operates on data centers. They have forecast that this number to go to 50 % between edge and cloud by 2028.

31:24When you're monitoring the current developments in the market, you can see that most of the chip manufacturers are moving towards creating much stronger chips or machines where they allow to run the AI at the edge. For example, Intel's Meteor Lake or the AI PCC, right? But it will take quite a while for healthcare organization to have that change adapted because it requires budgets. And I think healthcare organization, they go through machine update once every five years. I don't think they do it over all their machine, only maybe certain machines. But definitely, I do see the future that you can bring much larger models right at the edge.

Read the full transcript

32:14But I don't think we are there yet. Yeah, I appreciate that perspective because some people, I think, in our listener base, they're constantly overwhelmed by this news about these new big models. But it's harder to get this sort of story of a practitioner on the ground working with specific companies in certain constraints. There's still quite a diversity of constraints that practicing data scientists or AI engineer has to work within. And so I think that viewpoint's very important. As you kind of look at what you've done with TauCite and these tools that you've built in detecting PHI, helping companies know where their PHI is, reducing false positives, figuring out how to run these models on edge devices and all of those things.

33:05Do you have anything that stands out in your mind in terms of, you don't have to mention specific customers or anything, but success stories or really things that you're proud of that you're glad that you've been able to be a part of in terms of helping protect this PHI? Any sort of case studies or use cases that pop into your mind? Yeah, I can give some example without naming anyone. But we were in this meeting and this CISO was in the call and it's like, I had these laptops that was stolen, and I don't know what's on that laptop. But because we have our software on that laptop, we give them an inventory, and it turns out the laptop contains lots of PHIs.

33:52So they never have that view on this type of scenarios, that where the PHI is or who has access to it, but we can basically give you that. I was on another customer call and it was like, we are happy with the tools that we have. They have less false positive, but unknown false negative. And we are quite unhappy for writing the rules. Yeah. It takes us a while. But when you use our tool, there's no rules that you need to write. is basically out of the box after you're installing our product, it will start scanning all the files and also monitoring what's happened on the machine. So if someone copying or pasting, for example, PHI into an email or into another file, we can read that, we can detect that.

34:52If someone faxing it, we can basically detect that. So these customer calls that they have been on, And they also got to be positive and sometimes scary for customers because, you know, we are able to find really detailed information around your network. You're pulling the curtain back. There's work to do once you understand it. Yeah. Yeah, well, maybe as we kind of draw pretty close to a close here, as you are kind of plugged in both on the academic side, you're plugged in on the startup side with TauCite, you're investing in this healthcare industry from the perspective of AI and ML. What gets you excited as you look to the coming year?

35:39Maybe it's things with TauCite, maybe it's things more generally in the AI community. What are you excited about and what do you think is some of the positive things that you're seeing develop over the coming year? I think there are two things that I'm really interested. One is the development around these large models and the fact that they're getting smaller and smaller. I am looking for the day that I could work with those SLM or small large models and deploy those on right edge. And I think the other thing that I'm quite interested in right now actively working on is the federated learning.

36:20I know federated learning is kind of in the background. Not many companies actively are doing it due to all the challenges that it has and also some security concerns. but for a domain like healthcare where you cannot transfer data and you cannot see the data I found that absolutely necessary that for your models to be able to train themselves and update themselves so I think those are the two main things that I'm looking for the upcoming guys that's awesome yeah I'm definitely excited by both of those things as well and I know Chris and I on this podcast have mentioned federated learning for years.

37:02I hope that it kind of comes more to the forefront as people figure out paths to do this. I think that will be interesting. Well, Ramin, it's been great to have you on the show. I think we'll see each other again in Boston before too long, I think. But yeah, it was great to have you on the show. And thanks for taking time out of your schedule to share some of these insights with us. Thanks for having me, Daniel. And see you soon in Boston. Sounds good. See ya.

37:57community. Sign up today at practicalai.fm slash community. Thanks again to our partners at fly.io, to our Beat Freaking Residence, Breakmaster Cylinder, and to you for listening. We appreciate you spending time with us. That's all for now. We'll talk to you again next time.

From the publisher

We’ve all heard about breaches of privacy and leaks of private health information (PHI). For healthcare providers and those storing this data, knowing where all the sensitive data is stored is non-trivial. Ramin, from Tausight, joins us to discuss how they have deploy edge AI models to help company search through billions of records for PHI.

Join the discussion

Changelog++ members save 5 minutes on this episode because they made the ads disappear. Join today!

Sponsors:

  • Neo4j – Is your code getting dragged down by JOINs and long query times? The problem might be your database…Try simplifying the complex with graphs. Stop asking relational databases to do more than they were made for. Graphs work well for use cases with lots of data connections like supply chain, fraud detection, real-time analytics, and genAI. With Neo4j, you can code in your favorite programming language and against any driver. Plus, it’s easy to integrate into your tech stack. 
  • Backblaze – Unlimited cloud backup for Macs, PCs, and businesses for just $99/year. Easily protect business data through a centrally managed admin. Protect all the data on your machines automatically. Easy to deploy across multiple workstations with various deployment options. 

Featuring:

Show Notes:

Something missing or broken? PRs welcome!

More from Practical AI

All 157 episodes
Using edge models to find sensitive dataPractical AI · 38 min
Listen in VO