Controlled and compliant AI applications

31 May 2023 · 50 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Notes: Practical AI - Controlled and Compliant AI Applications

Episode Overview

  • Title: Controlled and Compliant AI Applications
  • Hosts: Chris Benson and Daniel Whitenack
  • Guest: Daniel Whitenack, Founder of PredictionGuard
  • Focus: Addressing issues related to the integration of Large Language Models (LLMs) in corporate environments and how to produce compliant, structured, and reliable AI outputs.

---

Key Themes

  1. Challenges of LLM Integration
  2. Inconsistent Output: The episodic nature of text outputs from LLMs creates difficulties in forming robust systems.
  3. Legal Concerns: Corporate lawyers and security professionals are wary of LLM integrations due to:
  4. Hallucinations (inaccurate outputs)
  5. IP/PII leakage risk
  6. Compliance issues (e.g., HIPAA)
  7. Injection vulnerabilities
  1. Pressure on Organizations
  2. Many businesses feel compelled to adopt AI technologies to remain competitive.
  3. There’s a growing challenge surrounding understanding and navigating the compliance and operational implications of deploying AI.

---

Insights from Daniel Whitenack

Foundational Concepts of PredictionGuard

  • Mission: To deliver structured and compliant outputs from AI applications.
  • Methodology: Uses open access models and LLM wrappers to ensure outputs are reliable and structured, which helps businesses utilize LLMs effectively.

Importance of Controlled Outputs

  • Structured outputs are essential for making meaningful business decisions.
  • Examples include producing consistent sentiment analysis outputs and valid JSON or Python code, which can replace unstructured text blobs.

Proliferation of Open Access Models

  • Open access models are improving and becoming more competitive with commercial offerings.
  • The need for robust model management and output structuring tools is highlighted.

---

Practical Applications

  • Use Cases: Businesses often face the challenge of extracting and processing information from unstructured text. For instance:
  • Automating sentiment analysis from customer feedback.
  • Extracting numerical data from textual information.

Key Features of PredictionGuard

  • Structured Output: Allows companies to define specific output requirements (e.g., floats, integers, valid JSON).
  • Compliance and Security: Ensures that hosted models meet compliance standards, protecting sensitive data.
  • Validation Checks: Implements factuality and toxicity checks to ensure output reliability.

---

Challenges in the Fast-Moving AI Landscape

  • Model Selection: Organizations struggle to choose the right models due to the rapid advancement of AI technologies.
  • Integration Complexity: Engaging with multiple AI models introduces complexities around structuring their outputs.

---

Future Outlook

  • Ease of Use: A vision for making AI tools more accessible, allowing users with limited technical skills to leverage AI capabilities.
  • Automation: Ongoing efforts to automate the structuring of output, reducing the need for deep technical knowledge.
  • Encouragement for Open Models: The increasing capability of open-source models alongside proprietary models offers a promising future for businesses willing to innovate.

---

Conclusion

  • The episode highlighted the tension between the rapid advancement of AI technologies and the need for structured, reliable outputs in business contexts. PredictionGuard is positioned as a solution to bridge these challenges, ensuring compliance, security, and usability of AI systems in organizational settings.

---

Sponsors

  • Fastly: Bandwidth partner for fast and reliable digital experiences.
  • Fly.io: Platform for deploying apps and databases globally without the need for extensive ops knowledge.
  • Changelog News: A podcast and newsletter combo providing insights into developer news.

---

Relevant Links

  • [Prediction Guard Website](https://www.predictionguard.com/)
  • [LLMs in Production II Event](https://home.mlops.community/public/events/llm-in-prod-part-ii-2023-06-20)

---

Note For a deeper discussion, listeners are encouraged to join the ongoing conversation on the Practical AI [discussion forum](https://changelog.zulipchat.com/#narrow/stream/456003-practicalai).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:06Welcome to Practical AI. If you work in artificial intelligence, aspire to, or are curious how AI-related technologies are changing the world, this is the show for you. Thank you to our partners at Fastly for shipping all of our pods super fast to wherever you listen. Check them out at Fastly.com. And to our friends at Fly, deploy your app servers and database close to your users. No ops required. Learn more at fly.io.

0:42Welcome to another episode of Practical AI. This is Daniel Whitenack. I'm a data scientist and founder of PredictionGuard. And I'm joined as always by my co-host, Chris Benson, who is a tech strategist at Lockheed Martin. How are you doing, Chris? Doing well today, Daniel. It just continues to be super interesting in this space, in the world of AI. So much change. This has been a year, I think it's a year that's for the history books in terms of the advances and the fact that AI is really making deep impact into the general population. People who normally might not be listening to our podcast as hard as it is to believe that.

1:24I was just going to say people like in companies like my wife's company, which is not a large company, it's not a tech company, but they're having conversations about how do we as a company leverage AI or leverage large language models in our content generation or what have you. And it's really permeated all industries at this point, I think. And people are wrestling with the idea of what do we do, not just if we do something in relation to AI. Agreed. I think that's a huge issue right now. It is potentially more confusing about how to handle everything that's coming at companies these days from a large language model and generative AI than it ever has been.

2:13and the problem is getting harder. And so we want to talk a little bit about that today. And I want to acknowledge a couple of things with our audience. Many of you have been with us for quite a long time. We've been doing this show for about five years on a weekly basis. I know it's been forever and it's just getting more and more interesting. Over that course of time, something I wanted to share with our audience is I've gotten to know Daniel pretty well. And we didn't know each other super well beforehand. We had met in the Go software development community as kind of the two people looking at data concerns.

2:49But Daniel over time has demonstrated not only the fact that he is repeatedly an incredibly smart man with a lot of capability, but he's also an incredibly good human being. And for anyone who's followed the show for a long time, they know that the idea of just being a good person and AI for good and such things are a huge repeating topic in the show. And I've also learned to trust where he's going and to understand that if Daniel is doing or interested in something, it's something that I want to know about. And so today, we want to hit this large language model kind of in the world and how you manage that.

3:28But I also want to acknowledge that we're going to have bits of the show that could be considered conflict of interest. And the reason I say that is we're going to talk about some work Daniel has been doing. And so if that bothers anybody, this is the point where you might want to shut off this particular episode. But I'm hoping that most of you trust us and have been with us for long enough to know that I'm not going to take you down a path that you wouldn't want to go. And I've asked Daniel to talk not only about the space of large language models being brought into production and trying to juggle all the things coming out, but talk about the work he's doing.

4:01And so we are unabashedly going to go that direction. And if anyone has hate mail to send, please send it to me because I have demanded, I have demanded that Daniel talk about this. Uh, and so, uh, thank you for bearing with us on that. And so Daniel, if you, uh, you're kind of both the co-host today and the guest, if you will. And if you might lay out a little bit of the landscape for us about, what this looks like as an insider, as someone who spends all your time focusing on this problem, that might be a good way to start. Yeah, thanks, Chris. And thanks for the kind words. I've over time learned so much from doing this show, and it's shaped a lot of what I think about.

4:43And certainly the things that I've been thinking about really since, well, this whole year since kind of Christmas time have been focused around these ideas of controlling large language models, guiding them, guarding them, making compliant AI systems. And a lot of that's led into the thing that I'm building right now, which is called prediction guard. So that's what you're referring to in terms of what I'm building. So I'm coming at it from that perspective and been thinking about this a lot, been talking about this a lot publicly and excited to do things like, you know, Upcoming, there's an LLMs in production event put on by our friends at the ML Ops community.

5:25That's really cool. I'm giving a talk there on controlled and compliant AI applications. So that's part of what I'll share here today as well. One question maybe that I have for you as we start out here, Chris, is what have you experienced in terms of the people that you're talking to with regard to the pressure that they're feeling either internal to their own company or from like market pressures, like jump into the AI waters, like implement something, make AI part of our stack. Like, what are you seeing there? So I don't think you'll be surprised when I say this, and we've alluded to this on some previous episodes, but it is a difficult business concern to navigate.

6:09I know all of us who straddle into the AI technical realm are incredibly excited. We're trying to figure out how to do the models and put them out there and everything like that. But if you are not in our shoes, if you're walking on a slightly different path, and let's say you work for a legal department or a compliance department or other business concerns, and suddenly these technologies are coming at you hard and fast week by week in 2023, and you're trying to navigate that and look at things like licensing on how the data that goes into models is used, and you're looking at compliance concerns and you're looking at protecting your intellectual property.

6:54There's a whole host of challenging business problems with essentially no guidance. This is a brave new world that has to be pioneered through. And so I have talked to a lot of business people in various roles, including attorneys, and this stuff is scary stuff. It is problematic stuff. It is challenging to navigate. And I definitely want to take you down the path today of talking about the space and prediction guard relative to how you actually get these models out there in a productive way in a business environment so that people can take advantage of the technology and understand what the pitfalls are and such.

7:37So that's the big thing that I've been hearing. I've of getting an earful of it lately. Like Chris, settle down. Stop taking us down this AI thing. We got to figure some things out first. So I'm coming to you for answers, man. Yeah, it's so tempting actually to have really easy to use systems. Like let's say the OpenAI API, right? I can go to the playground or I can go to chat GPT or I can go wherever, put in my prompt and get some like magical output right it's magical and immediately it triggers in your mind i can solve real business problems and i can create like actual solutions with this type of technology like it's so quick to make that connection but what i've seen both in sort of advising and consulting and conversations that I've been having is on maybe like a less stringent case, like people are struggling to make that connection to how they can build robust systems out of these technologies.

8:43So it's one thing to get text output and look at it with your eyes as a human, right? And say like, extract this piece of data or give me a summary of this or something like that. But as soon as you make that programmatic, right, and automated, then how do you know you're getting the right output? And if you actually want to do something with that, like you're outputting a number, you know, a vomit of text blob out of a large language model doesn't really actually do you that much good if you're trying to implement like a robust system that's making actual business decisions on top of the output of large language models.

9:21On the harder side of this, I'm getting feedback from people that either I know or I'm advising or other things that companies are actually telling them no. There's a full stop on using, quote, GPT models in this organization because of one of a few different reasons. Maybe that's a risk thing around, hey, we're going to hallucinate some name out of this or something that this person doesn't exist and that's going to get us in trouble. People are going to stop trusting our product and that sort of thing. So there's the hallucination or consistency of output sort of thing. There's also, as you mentioned, the IP or PII type of leakage scenario.

10:07So it is actually a problem for people to sit in a company. I'm sure this would be true whether you're at your company or a variety of other companies that I've talked to where I'm sitting there and I'm like, oh, I could solve this problem with ChatGPT. Let me copy and paste this user data into ChatGPT and have it summarize something or extract something or whatever it might be. It's sort of unclear in murky waters how that data is actually going to be used by OpenAI. and you're kind of leaking IP or company information, PAI to external systems, right? Which is a big, big no-no. Regardless of how that's used in the end, this data, it seems like is going to exist outside of your own systems.

10:58And so on the harder side of this problem, people are being told like, no, you have a full stop. Can't use GPT, can't use large language models. So to summarize how I would kind of think about this problem space, people are feeling the pressure that they need to or really want to implement these systems, either because they feel like they're getting left behind or there's an actual market pressure for them to do something. but in practice they don't know how to deal with the outputs of large language models and they might not even be able to connect to the kind of most common large language models because of these privacy security leaked IP type of issues.

11:39I think that's really really widespread it's funny you've kind of enumerated a whole set of risks associated with that yesterday just as a thing. I have a particular employer and thinking about public information, well-known public information about the lines of business that we have that is publicly acknowledged in multiple sources out there. I went to chat GPT. I should have done the 4.0 model, but I forgot. And I just let it default to the 3.5. And I simply asked for our 19 lines of business, which is incredibly public knowledge. And it got it wrong. It got it wrong the first time. And so I tried to steer it a little bit and it got it wrong the second time.

12:20And had I not known better about the intellectual property concerns with licensing, had I tried to put something in that might have been out there in the public. So I run into what you just said all the time and there's so many risks and yet there's so much value to extract from the space. And so I think putting your finger on the fact that if you can find a way to mitigate these risks in various ways, that will unlock a huge amount of value for a lot of organizations and users to do that. But it certainly, from my standpoint, feels like the Wild West right now. Yeah, I would say that that's true.

13:05And yet, there's these concerns like the money that people are able to save operating costs with AI in your business are significant. So I saw the study from Accenture estimating like insurance companies saving$1.5 million per 100 full-time employees. So if you're insurance company A and you're not trying to implement AI systems in your business, then you're actually introducing a liability, Right. Because insurance company B might be doing that and they're going to slash their prices and undercut you and put you out of business. Right. So even regardless of new features that might be implemented in like people's products and that sort of thing, there's this real liability around not considering AI solutions as part of your business strategy.

14:04I think that's a huge point. And that's the other side of the coin that I was just talking about. There was the risk of using and there is the potentially larger risk of not using at all. So we're seeing that in all markets in terms of the need to stay on top of what is gradually evolving over these past months and to be able to use that to promote your business. And if you don't do that, the risk is substantial. So the idea of navigating the licensing and the compliance concerns and being able to productively use these outputs is really crucial to being successful in almost any industry going forward.

14:46So definitely looking forward to finding out how we might do that.

14:55I'm Jared, and this is a Changelog News Break. In what appears to be a particularly security unaware move, Google has added eight new top-level domains, two of which are quite concerning,.zip and.mov. Yikes. Ars Technica writes, quote, While Google marketers say the aim is to designate tying things together or moving really fast and moving pictures and whatever moves you, these suffixes are already widely used to designate something altogether different. Specifically,.zip is an extension used in archive files that use a compression format known as zip. The format.mov, meanwhile, appears at the end of video files, usually when they were created in Apple's QuickTime format.

15:43End quote. Fishers and scammers rejoice. The rest of us, beware and be ready to help protect your family and friends from this otherwise completely avoidable new threat vector. The linked Ars Technica article demonstrates a few URLs scammers could now craft, and they're darn near indistinguishable from the legit URL, even to someone like myself with trained eyes. one such URL in the example is a Kubernetes release which yes, is distributed as a zip file you just heard one of our five top stories from Monday's Changelog News subscribe to the podcast to get all of the week's top stories and pop your email address in at changelog.com slash news to also receive our free companion email with even more developer news worth your attention Once again, that's changelog.com slash news.

16:42If I could summarize some of what's been said, we kind of talked about these two large categories of problems. One was the structuring, consistency, and validation of the output of these models to make them useful in actual business use cases. And the second was maybe compliance concerns, privacy security concerns, which really have to do with like how a model is hosted or how you access that model. So on the one side, it's how do you process the output of a model? And then on the other side, how do you access or host a model? Both of those things can be pretty big blockers. To kind of dive into the latter of those, the hosting, privacy, security thing, I actually am quite encouraged by where things are headed recently because we've seen this kind of proliferation and explosion of open access models that continue to be released day after day.

17:47the most recent one at the time of recording this i i might be missing one they seem to come out every week but one for example that came out recently is the mpt uh family of models from mosaic ml which is just really extraordinary i think that they have up to like context links or like you can think about that as kind of your prompt size for the model of like 60 000 tokens and they do quite well in various scenarios. So there are these increasing number of open access models, but I would say there's two problems with using these as a business. Let's say I wanted to host one of these and use it internally.

18:31Well, maybe three problems. It's always good to have three points, right? Three problems. One is like you still have to figure out like the weird GPU hosting and like scaling of that model, which is a challenge. right. The second is in reality, these open access models, at least according to most people, I think it's generally accepted that these aren't quite up to the standards of the larger commercial systems that like OpenAI and others are putting out there, Cohere and Anthropic and others. Sure. So there's like a performance concern, there's the hosting concern. And then the third, which is the same as our other major topic here is you still have to figure out how to use the output of them.

19:13They're still just going to vomit up text on you and you have to figure out how to deal with that. This has led some people to strike up these kind of expensive deals to host open AI models in Azure infrastructure. That's becoming easier over time. I hope that becomes increasingly easier. It's still a little bit like limited to Azure mainly in my understanding. And it's definitely not cheap, I would say, if you kind of compare all the costs and add in the engineering time to do that and all that. So some people are solving this model hosting issue by either hosting an open access model, maybe with a hidden performance, or implementing a really expensive kind of private version of open AI, something like that.

20:05And if you don't have that budget, or if you don't know about GPUs or how to host models, you're kind of out of luck in a lot of ways. Not only I agree with you, but I think that that's going to proliferate in terms of the challenges across there. I know speaking for myself and another friend that I talk to about this stuff a lot, we are experiencing the fact that as model updates come out, new models come out, they have different strengths and weaknesses. There are some things that I might, for instance, go to GPT-4 on. There are other things I might go to BARD on now. And, you know, those are just two.

20:41There's a whole bunch of open source ones that we were starting to talk about that. And with the acknowledgement of, for instance, OpenAI has kind of acknowledged that there is a practical limit in terms of how much data you can feed a model and that we need to start looking at other dimensions on that. So with practical limits in sight, the commercial advantage, for instance, may hit that ceiling and open source ones will gradually catch up. And so you're seeing the relationships of utility for a user between different models changing on a regular basis and us users having to make adjustments to that.

21:18How does that play into the landscape? Because if you're an organization and you're trying to make investments, like we bet on open AI and Microsoft, do we bet on Google? Do we bet on open source options? What are the options there? What are the different capabilities that might be available to us for doing that? And acknowledging up front, PredictionGuard may be one of those. What does the rest of the landscape look like and how does PredictionGuard fit into that? And what are some of the pros and cons that you see? I'll decouple a couple of these things and talk about the general landscape and then PredictionGuard.

21:52So in terms of this problem of the hosting, compliance, privacy, IP leakage, that sort of thing, I think if you're a company of a certain size and you can afford kind of a private open AI setup in Azure, it's probably a pretty reasonable solution. It will definitely work very well, but it's going to be very, very costly. And again, it's not going to solve this like structuring and usage of the output of language models problem. So you're going to have to put additional engineering effort into helping build layers on top of that that work for your business use cases. You could bet on certain open access models right now, but like you said, things are advancing so quickly.

22:39It's hard to say, like, I'm going to put all of this effort into one and hosting of the one and build a system around it. I do think that there's advantages if you're going that route to center your infrastructure around kind of model agnostic workflows like those in Langchain or others where you actually abstract away the model interface and can connect to multiple large language models with a lower switching cost than if you kind of have a one-off solution centered around a certain model. So I think there's there's some things that people could be encouraged about there in terms of that though if you think about okay now i'm going to go all in on these open access models like you say these models have different characters so i'm going to want to host maybe multiple of them and generate these model agnostic workflows on top of lang chain and other things you start to really add up the engineering effort to make this happen.

23:40A parallel might be I could create a data visualization solution for my company by assembling a database and hosting that, making the connection into like a layer that would run like Plotly plots or something like that. And then maybe some UI for my users that those are embedded in. And all of a sudden I'm now talking about an absolute fortune and engineering costs and support costs over time, which is why products like Tableau or other, you know, I remember a long time ago, I don't know how much people are still using it. One of the companies I was at was using Domo. This is one of these solutions where you quickly suck in data and visualize it and all of that.

24:25There's a reason why those products exist. So Prediction Guard, you could kind of think of as taking the best of open source models and the best of this kind of control and structuring of output, which we haven't talked about yet and we can get into here in a second, and assembling those together in an easy to access and cost efficient manner so people can get quality output out of the latest large language models that's structured and ready to be used in business use cases. And also with a guarantee if you want it around using only specific models that are hosted in a compliant way, even compliant in a certain way, like a HIPAA compliant way or in a data private sort of way where your data isn't leaked if you're putting data into models.

25:14So that's kind of how the landscape works and how Prediction Guard works as this kind of system that assembles the best of large language models with structured and typed output that can be deployed compliant without this whole huge engineering effort to build your, you know, roll your own system. You mentioned structured and typed output. And can you go ahead and kind of talk a little bit about that? Because I think, you know, for many of us that are listening, we're used to using the models that are out there kind of in the default interfaces on the web, you know, using chat GPT, using BARD.

25:50And we're not really dealing with that. You know, we get an output and but we're not at the level of sophistication where we're doing APIs and such as that. Can you talk a little bit about what structured output looks like when you're dealing with it from an API standpoint and how you unify that landscape? There's a lot of use cases where this may come up, but let's take one for example. Let's say that you're doing data extraction. You have a database with a column in it, which is basically, so this scenario has happened at every company that I've been with, so I know that it's very common. There's some database with a table in it, and there's a column that's like a comments column or something, and it's just like text blobs in there that are like notes from people or technician messages or user messages or like whatever it is.

26:41It's not structured. And you want to run a large language model over that to extract, you know, maybe it's phone numbers or prices or certain classes of information out of this column. Well, you could run your large language model and set up a prompt that says, you know, give me the sentiment of each of these pieces of text in my database. Well, that prompt, each time you run it through a large language model, maybe once it generates an output that says space positive sentiment, and the next time it creates an output that says positive, and the next time it creates an output that says this is positive sentiment.

27:26And you can start to see there's a consistency problem here. like how do I parse all of these strange outputs from my large language model? You can do a little bit of prompt engineering to get around that, but ultimately it doesn't solve the problem that you could have all sorts of weird output out of your large language model. So ultimately what you would want in that scenario is a system that lets you constrain and control what types of output you're going to get out of your large language model. so in the case of sentiment maybe I want to restrict my output to only pos neg and neu tags for sentiment there's only three choices I always want one of those three right I don't want it to say this is positive sentiment right so I want to actually structure or control the output of my large language model to produce one of these outputs another example that's maybe a little bit more complicated would be to say, I actually want to output a valid JSON blob out of my large language model or valid Python code out of my large language model.

28:29And these are structures that are very well defined, but you could have all sorts of variability coming out of your large language model. And if you want a specific type coming out of your large language model, maybe it's a float that you can do like greater than or add it to another number. Like you need that as a typed output, or you need very specific structured output to actually make automated decisions in your business. And so with PredictionGuard, what we're doing is we're kind of assembling the best of the recent advances in this kind of control and structuring of output and layering it on top of these open source large language models to allow you to say, here's my prompt.

29:11I'm going to send it to these five open source or open and or closed. We support open AI as well. So open and or closed models. And for each output, I want you to give me a float number. And that's the sort of rich output that you can get from large language models very quickly with a prediction guard kind of prompt because you can control the models that you're using, either ones that are more privacy conserving or the closed source options and provide constraints around the output that allow you to actually make business decisions on that. Now there's additional checks that could go along with that, like factuality checks and toxicity checks, which we also implement.

Read the full transcript

29:54But I've vomited up a lot of information, so I'll pause here. No, no, that sounds fascinating. It's the way I'm interpreting what you're saying is sort of like you have these kind of software filters that are creating boundaries, if you will, on how you structure input and what that output can be. So it's usable, which kind of goes back to one of the points that we're often talking about on the show is that the AI is to some degree inseparable from the software that you're using it within. And so you have a best of breed software product that's kind of shaping and constraining what that can be so that it's actually usable on that.

30:33So as we look forward at kind of where things are going and you, what, what are some of the problems that you see going in the space that we have? And like, what are some of the things that you would like to see prediction guard starting to address, uh, in the time ahead? And I don't mean so much as the far distance, but kind of like you're, you're busy putting this solution together now, uh, works pretty darn well. Oh, what you already have. What are some of the challenges when you're in this kind of a fast moving space? Because you're having the world change out from under you on a week by week basis right now.

31:08Yeah, I think maybe one of the things that we're thinking about is really at the forefront of our mind is ease of use and accessibility to both data scientists and developers. So the reality is that I think we had Kirsten Lum on the podcast talking about this, like the majority of data scientists out there are super constrained in the time that they have to put into one of these solutions. Right. So it's really, really important that there is an ease of use to this sort of controlled, compliant LLM output and generative AI output. But now what we're seeing, and I want to acknowledge this as well, is there are an increasing number of open source projects that are doing an amazing job at digging into this problem of controlled and guarded LLM output.

32:01So these are things like guardrails and guidance from Microsoft and Matt Rickards, RE LLM or Regex LLM. These projects are doing amazing things at really flexible ways for you to control the output of large language models. But I see this as kind of like a double edged sword a bit. The more flexible you become, it's also possible to become less easy to use. And there's more engineering involved in it. Yeah. So I saw Matt Rickard tweet about this related to his Regex LLM project, which is that sort of famous quote about Regex, which is, I have a problem. And so I decided to use Regex and now I have two problems.

32:51That's been around for a long time, actually. Yeah, it's actually so true, right? And some of these solutions are coming up with their own query languages to kind of deal with this structured output, which I think is great and it's really important. but there's a need for this abstraction layer on top where I know kind of what I want my output to look like. So I should be able to plug that into something and have it constrain the output of my large language model in an appropriate way. So with PredictionGuard, what we've started with is the kind of presets of structuring your output. So I want integer and float and JSON and Python or YAML.

33:37I want categorical output. These are things that we support now. Also supporting kind of these hosted models and access in a guarded kind of controlled way to these models. But let's say that I have a really specialized format that I want to work with. I would rather set up a solution with PredictionGuard. And this is actually what we're actively working on, where they could give examples of the structure that they want and we actually generate the right constraints for them on the large language model output which I think is very possible and our initial work on this which is kind of in a beta form is really good so let's say that I want a specific JSON with these specific fields or a specific CSV output with these specific columns right I should be able to give a few examples of that and generate the right underlying constraints for my large language model without the user having to think about special languages or regex or context-free grammar or these things that are a little bit harder to grasp.

34:40We'll handle that bit for you and you just get the right structured output from your models. So that's part of where I see us headed is leveraging these rich systems under the hood that are being produced around using context-free grammars, special query languages, regex, all of these things to structure output and combining those in a more automated way for users where they can just say, here's my examples, here's my query, and they just start getting the right formatted output from their language model. So that's kind of thing. One is this automation of some of the problem and the constraints.

35:18I think the thing too would really be around the validation and checking of output in addition to the structuring. So right now we support factuality and toxicity checks on the output of large language models. So could you talk a little bit about what each of those are? Yeah. Yeah. So let's say that I, I take a big piece of text and I generate a summary or I do a question answer prompt and get an answer, right? It doesn't mean the answer is factual, right? And we all know about the hallucination problems of these models. So the things that we have implemented in Prediction Guard are two things with respect to that.

35:58The first is a factuality checking score, which is built on these trained models under the hood that look at a reference piece of text and your text output to determine a likelihood of the answer being factual. So this is an estimate on the factuality of your output. The other thing that we're doing around hallucinations and factuality is making it really easy for people to do consistency checks. I kind of alluded to this earlier, but we have all of these different language models accessible under the hood. So you could combine the outputs of CAMEL 5 billion, MPT 7 billion, DALI, and OpenAI, restrict the output to say, give me the answer, but only if all of these agree on what the output is.

36:47If all of them don't agree, then I'm going to flag that as not a reliable output. And so you can actually gain a lot by not just leveraging one model, but ensembling these models together to do a check. The toxicity thing is something that's been studied for a while, and there's models out there, state-of-the-art models, for detecting whether an output is toxic or not, or includes hate speech or not, that sort of thing. So this is another layer of check that you can have on the output. And so if you put the whole pipeline together of prediction guard, you've got models on the output, which can be deployed compliant with HIPAA or just data privacy.

37:28Those structured or typed output that you can define very easily. And then you can run additional checks on that output for factuality, toxicity, consistency as a final sort of layer in the pipeline towards the output that's used in a business application. I appreciate the explanation. It's a very robust sounding pipeline that you have on that. Let me ask you this, and this could be whether it's prediction guard or whether it's the larger field. One of the challenges is certainly something I've been playing with, but I don't have a good rhyme or reason to it yet. With the proliferation of these models coming out and ever more coming, we know this space is going to get larger and larger.

38:12How does a user or how would a system like prediction guard be able to determine which is the right way to go in terms of which model you want to choose or which group of models? And you talked about the comparisons a moment ago. Like, how do you structure the input and know that you're going to get what you need from an output by putting the right model or collection of models together and then knowing how to evaluate them against each other? Does that make sense? Yeah, yeah, that makes sense. Actually, early on when we were building the prediction guard backend, this was actually front of my mind and has since kind of evolved a little bit.

38:52The fact that there's all of these models and I want to choose the right one for my use case, you can very much automate that process. And we actually, it's actually still implemented in the prediction guard backend where you can give some examples and evaluate a whole bunch of models on the backend. I think where this is headed, though, and where the prediction guard system is headed is making it easier for people to get output from multiple models in a typed way because they know how to do the evaluation. they're familiar with this sort of thing whether you're a developer doing sort of integration tests or unit tests and you're checking and you're asserting certain values or you're a data scientist that's running a larger scale test against the test set people kind of know what they want to do with that sort of thing what they need is an easy way to get that typed output from multiple models so like if i have a test set and i'm comparing two scores on the output like float numbers.

39:55I need to get float numbers out of a whole bunch of different large language models to compare them to my baseline or to my test set. Right now, that's very difficult because all of these different kind of structuring guidance control systems work not for all models, and they don't work in the same way for all models, and you have to implement it for all of the models. And so it becomes this compounding problem to figure out how to do that. And so how we're approaching that with the prediction guard system is there's a standardized API to all of these different models, along with the typed and structure control on the output.

40:36So I can do a query that says, give me the float output for these hundred prompts using these five models. And then I'll just compare all the float outputs and figure out which is the best. That's not the hard problem. It's the getting that structured output from the variety of models in a robust and consistent way that's actually a more difficult problem. Gotcha. Is it fair, you know, as we're talking about this, it sounds a lot like you're also solving one of the bigger challenges we've talked about over time, which is that there's so much domain expertise in the AI space in terms of being able to manage models.

41:19But if I'm understanding you correctly, it sounds like minimally, with some basic software skills, knowing how to use APIs and stuff, you can probably, without deep expertise and deep learning, manage to get some fairly productive output through ProductionGuard by implementing it that way. In other words, it becomes just another part of your software workflow. Is that a fair characterization, what I'm saying? I would say it is in the sense that there's still some sort of like integration testing and integration that will have to happen regardless. Right. But going back to my example before of like the the data visualization stack, it's a lot harder to implement the database and the visualization layer and the front end than it is to like log in and do the.

42:11there's still configuration that's needed in like a domo type solution or tableau it's just a lot more accessible right so here we have the language models hosted on the back end we have the structured guarded way to query those models via something that all developers know how to use a rest api or a python client maybe there'll be other clients over time and have the ability to configure that in the way that you want. So I want output from these five models. I want to ensemble them together, or I want this structured output. And so there's still configuration. And I think developers and data scientists, they want that.

42:50It's just that it's really hard to get all the other pieces in place. And we're hopefully making that a lot easier. So, and let me ask one final, I think this is an aspirational question, but I'm kind of curious. One of the things that we've seen with large language models is the ability for people who aren't even developers, you know, I was saying like developers who aren't even deep learning experts, but to have a certain amount of capability producing code, you know, that's the kind of avenue into a no code world has at least been started on this. It has a lot of maturing to do, obviously. Do you envision a point where someone with very limited skills can also use prediction guard in this way and be able to kind of generate apps using large language models that then kind of feed into a more mature workflow like what you've described.

43:42Do you think that that's attainable at some point in the not so distant future? It's hard to say how far this kind of automation will go. I think a lot of the agents that we've seen produce good demos, right? But they have an additional layer of this sort of additional problems around automating these various steps of the process. I think that in terms of what we're looking at, this sort of automated structuring of output is a step in the right direction in terms of, I don't have to define a special query language or a special specification, but I can say what sort of structure I want output and that gets output.

44:25I think then if you layer that on top of the agent sort of infrastructure that's in Langchain and the data augmentation, we just had the episode with Jerry from Llama Index, which is super fascinating. So if you layer the kind of structured guarded output with the chaining and agent and automation of Langchain and maybe the data augmentation of Llama Index, I think a lot of things become possible. I hope that that becomes some of the things that you mentioned become possible. It's yet to be seen. But I am really encouraged that adding in this sort of type safety for outputs and structuring of outputs gives a lot more confidence maybe in some of the checks that you could do on AI agents over time.

45:12And that increases our confidence in sort of releasing AI agents on various parts of the workflows that we'd like them to work on, right? So you've sort of already covered some of the territory. But for our listeners, Daniel and I often when we're talking to a guest, we'll kind of finish with what we roughly call the future question, you know, kind of wax poetic about where things are going. And, uh, and so Daniel, since you're knowing that there's, I've kind of hit some of that already, but what would you be asking yourself? You know, so you've kind of had me throwing these questions at you, uh, from a point of somewhat ignorance compared to where you're coming from, um, as the expert on it, what right now would you ask yourself that you haven't covered that you think is worthy of getting in before the episode is over.

46:01I'm putting you on the spot. Yeah, I've mentioned open access models quite a bit. And I think hopefully a lot of us are encouraged by the direction that that's going, that these models are getting better and better. But one thing that maybe I would ask myself or that I think is important to highlight and encourage people with is these open access models might not quite be at the level of open AI, Anthropic, etc. yet. But I think not only will they get there, but already like in the space where we're at now with some of these kind of structured control elements around open access models, you can actually boost the performance of open access models to be more in line with, you know, open AI level output.

46:53Because what you can do is say, well, I'm going to force my output to this. If I'm not able to produce it, you know, I can re-ask the question or I can try a variant of my prompt. And these kind of wrapping layers around open access models actually provide a way for you to operate in a data private compliant way with open access models that boost their performance closer to what these kind of closed and maybe more suspect in terms of IP leakage and that sort of thing systems are doing. So I think that's an encouragement that I've found recently. And I hope that's encouraging to others is we are really seeing a proliferation of these models and they're all going to have a little bit different character.

47:41But the ways that we wrap them and the way that we present them is a lot of provides the majority of the value of those models. And I think we'll see not only PredictionGuard, but other systems as well coming out that wrap these models and use them in really intelligent manners that boost their performance in a way that isn't reliant on sort of a centralized API. I appreciate that. I think you're right. And I am deeply appreciative of you not only telling us about PredictionGuard, but actually kind of laying out the space. even if someone is not chomping at the bit the way I am to use prediction guard, they hopefully kind of understand what some of the problems are that need to be addressed, whether by you or others out there.

48:28So thank you for allowing me to twist your arm and do this episode today. I appreciate you letting me go there. So anyway, thank you very much to my good co-host and my guest today for coming on Practical AI. Thanks so much, Chris.

49:15People find podcasts they love. Thanks once again to our partners, Fastly, Fly, and TypeSense for helping us bring you awesome pods each and every week. And to Breakmaster Cylinder for producing all the beats on all Changedog podcasts. That's all for now. We'll talk to you again next week.

From the publisher

You can’t build robust systems with inconsistent, unstructured text output from LLMs. Moreover, LLM integrations scare corporate lawyers, finance departments, and security professionals due to hallucinations, cost, lack of compliance (e.g., HIPAA), leaked IP/PII, and “injection” vulnerabilities.

In this episode, Chris interviews Daniel about his new company called Prediction Guard, which addresses these issues. They discuss some practical methodologies for getting consistent, structured output from compliant AI systems. These systems, driven by open access models and various kinds of LLM wrappers, can help you delight customers AND navigate the increasing restrictions on “GPT” models.

Join the discussion

Changelog++ members save 3 minutes on this episode because they made the ads disappear. Join today!

Sponsors:

  • Fastly – Our bandwidth partner. Fastly powers fast, secure, and scalable digital experiences. Move beyond your content delivery network to their powerful edge cloud platform. Learn more at fastly.com
  • Fly.io – The home of Changelog.com — Deploy your apps and databases close to your users. In minutes you can run your Ruby, Go, Node, Deno, Python, or Elixir app (and databases!) all over the world. No ops required. Learn more at fly.io/changelog and check out the speedrun in their docs. 
  • Changelog News – A podcast+newsletter combo that’s brief, entertaining & always on-point. Subscribe today. 

Featuring:

Show Notes:

Something missing or broken? PRs welcome!

More from Practical AI

All 157 episodes
Controlled and compliant AI applicationsPractical AI · 50 min
Listen in VO