In short
Practical AI Podcast - Episode Summary: The New AI App Stack
Episode Overview In this episode of *Practical AI*, hosts Chris Benson and Daniel Whitenack delve into the emerging architectures for large language model (LLM) applications, expanding on a16z's diagram outlining the new AI app stack. They discuss various components, including model "middleware," app orchestration, and the ecosystem surrounding generative AI applications.
Episode Details
- Hosts: Chris Benson (Tech Strategist at Lockheed Martin) and Daniel Whitenack (Founder of Prediction Guard)
- Focus: Understanding the new AI app stack and related technologies.
Key Concepts
- Understanding the AI App Stack
- Misconception: The model is often perceived as the application itself.
- Reality: There is a broader ecosystem of tools that surround the model, which are essential for creating functional applications.
- Components of the AI App Stack
- Playgrounds: Interactive environments where users can experiment with models (e.g., ChatGPT, Hugging Face Spaces).
- Characteristics:
- Browser-based interfaces.
- Useful for demos and experimentation, but not for building applications.
- App Hosting: The infrastructure for deploying applications.
- Examples include Vercel and traditional cloud providers.
- Shift toward integrating AI models within app hosting environments.
- Orchestration Layer
- Defined as a "convenience layer" that connects various components of the app stack.
- Includes:
- Prompt templates.
- Chains and agents for orchestrating calls to models.
- APIs and plugins for integrating external resources.
- Middleware for Models
- A new layer between the orchestration and model hosting, encompassing:
- Caching: Techniques to store and quickly retrieve outputs, reducing costs associated with expensive model calls.
- Logging: Monitoring model performance, latency, and resource usage.
- Validation: Ensuring outputs meet quality and security standards, and that inputs do not include sensitive data.
- Data and Resource Integration
- Emphasizes the importance of data pipelines:
- Embedding Models: Models that convert data into embeddings for searching and retrieval.
- Vector Databases: Databases optimized for handling embeddings and supporting semantic searches.
Discussion Points
- Emerging Practices: The conversation highlights the evolving landscape of AI engineering, encouraging practitioners to consider the entire ecosystem around models rather than focusing solely on the models themselves.
- Learning Resources: Listeners are encouraged to explore example applications and documentation to deepen their understanding of each component.
Key Takeaways
- The AI app stack comprises multiple layers, each essential for building functional applications.
- Understanding orchestration, middleware, and data/resource integration is crucial for effective AI application development.
- Practitioners should adopt a holistic approach to AI engineering, recognizing the importance of tools that facilitate interaction with models.
Conclusion The episode concludes with a recognition of the rapid advancements in AI technologies and the importance of staying informed about the evolving landscape of AI applications. The hosts encourage listeners to explore the linked resources in the show notes to gain practical insights.
--- For further discussions and resources, tune in to the full episode and check out the provided links.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:06Welcome to Practical AI. If you work in artificial intelligence, aspire to, or are curious how AI-related technologies are changing the world, this is the show for you. Thank you to our partners at Fastly for shipping all of our pods super fast to wherever you listen. Check them out at Fastly.com. And to our friends at Fly, deploy your app servers and database close to your users. No ops required. Learn more at fly.io.
0:42Welcome to another fully connected episode of Practical AI. In these episodes, Chris and I keep you fully connected with everything that's happening in the AI community. We'll cover some of the latest news and we'll cover some learning resources that will help you level up your machine learning game. I'm Daniel Whitenack. I'm the founder of Prediction Guard. And I'm joined as always by my co-host, Chris Benson. who is a tech strategist at Lockheed Martin. How are you doing, Chris? Doing very well today, Daniel. How are you? I'm doing great. I am uncharacteristically joining this episode from the lobby of Hampton Inn in Nashville, Tennessee.
1:24So if our listeners hear any background noise, they know what that is. You have a built-in audience right there. Built-in audience, the people in this lobby are unexpectedly learning about AI today, which I'm happy to do. Yeah, out here visiting a customer on site. And yeah, it's nice to sit back and take a break from that and talk about all the cool stuff going on. Excellent. Well, I'll tell you what, you know, we have had so many questions and about kind of sorting out all the things that have happened the last few months and over the last year. And we've done a couple of episodes where it was trying to kind of clear out like generative AI, what's in it, what LLMs are, how they relate and stuff like that.
2:12What do you think about taking a little bit of a deep dive into large language models and kind of all the things that make them up? because there's a lot of lingo being hurled about these days. Yeah, yeah. I think maybe even outside of LLMs, there's this perception that the model, whether it be for image generation or video generation or language generation, that the model is the application. So when you are creating value, the sort of model, whether that be, you know, LLAMA 2 or Stable Diffusion XL or whatever, that somehow the model is the application. Like it's providing the functionality that your users want.
2:57And that's basically a falsehood, I would say. And there's this whole ecosystem of tooling that's developing around this. And one of the things that I sent you recently, which I think does a good job at illustrating some of the various things that are part of this new ecosystem or this new generative AI app stack was created by Andresen Horwitz. They created a figure that's like emerging LLM app stack. We'll link it in our show notes. I think it goes though, maybe more generally than LLMs. But that provides maybe a nice framework to talk through some of these things. Now, of course, they're providing their own look at this stack, especially because they're invested in many of the companies that they highlight on the stack.
3:50But I think regardless of that, they're trying to help people understand how some of these things fit together. Have you seen this picture? I have. And I appreciated when you pointed it out a while back there. It definitely is an interesting, I haven't seen anything quite like it in terms of putting it together. And some things they seem to dive into more than others in the chart. It will be interesting to see how we parse it going forward here. Yeah. And maybe we could just take some of these categories and talk them through in terms of the terminology that's used and how they fit into the overall ecosystem.
4:27So, you know, we can take an easy example here, which is one of the things that they call out, which is playground. Now, I think this is probably the place where many people start their generative AI journey, let's say. So they either go to, I think within the playground category, there would be like chat GPT might fit in that category where you're prompting a model. It's interactive. It's a UI. Like you can put stuff in and put in a prompt and get an output, right? Now, chat GPT is maybe a little bit more than that because there's a chat thread and all of that. But there's other playgrounds as well.
5:09So you could think of spaces on Hugging Face that allow you to use stable diffusion or allow you to use other types of models. There's other proprietary kind of playgrounds that are either part of a product or are their own products. So OpenAI has their own playground within their platform. You can log in and try out your prompts. There's NAT.dev, which is a cool one that kind of allows you to compare one model to the other. There's other products, like I would say something like QuipDrop, which is a tool that lets you use stable diffusion. And you can just go there. You can try out prompts for free.
5:51You can pay up if you need to use it more. So there's a limit to that. But there's a lot of these playgrounds floating around, and that's often where people start things. it's funny the playground itself as uh as a category has a lot of subcategories i think to it because you know you've already kind of called out kind of the diversity of what you've you know in the cloud providers for instance all the big cloud providers have their own playground areas yeah nvidia has a playground area i think it's almost becoming a ubiquitous notion and of course all those playground areas for the commercial entities are focused on their products and services definitely but trying to you know trying to bring some cool factor to it so yeah it's almost like a demo or experimentation interface so if we define this playground category it's usually but not always a browser-based playground or a browser-based interface where you can try to prompt a model and see what the output is like i think that would kind of generally be true maybe there's some caveats to certain ones like mid journey for example is a discord bot or there's still a discord bot that you could use maybe that fits into the playground but generally these are interactive and useful for experimentation but not necessarily useful to like build an application yeah i agree and another thing to note about it from a characteristic standpoint is not only is it's not made for you to go build your own thing.
7:25It's made for you to try the stuff of whatever organization is doing. But they do do it, they provide the resources. So by being in a browser, you don't have to have a GPU on your laptop. You don't have to have resources, yeah. Yeah, you don't have to have all the things. Through various means, they set up all that for you on the back end, whether it be just calling a service or whether it be creating a temporary environment through virtualization. but it is a good way to either to test out a new product line or to or to just get your toes wet a little bit if you want to try some stuff out maybe you've been listening to the practical podcast for a little while and you want to a particular topic grabs you that would be a good place to go yeah and i think within that same vein you could transition to talk about this other category, which is not unique to the generative AI app stack, let's call it, but it's still part of the stack, which they have called out app hosting.
8:23So that's like very generic, right? So in here would fit things like Vercel, or I would say, you know, generally like the cloud providers, right? And the various ways that you can host things, whether that be in Amazon with ECS or AppRunner or whatever that is, or in even your own infrastructure, your own on-prem infrastructure, if you host things. Now there are, I would say a number of hosting providers that are kind of cool and trendy and people that are building new AI apps, they seem to gravitate towards like, let's say Vercel and a lot of front-end developers that use Vercel, which I I think it's an amazing platform.
9:05So cool. That hasn't traditionally been like a data science-y hosting way of doing things, but it represents, I think, this new wave of application developers that are developing applications, integrating AI. And you see some of those now kind of coming into or being exposed in this kind of wider app stack. Which is a good thing because we've talked for a long time. even as we open this conversation up saying the model is not the app. You know, you have to wrap the model with some goodness to get the value out of it, to be productive with it. And so I personally like the fact that we're seeing the model hosting and the app hosting are starting to merge because I think that's more manageable over time.
9:52It's less being in its own special category and it's more about, okay, every app in the future is going to have models in it. And so, you know, we're accommodating that notion. So I like seeing it go there. I've been waiting for that for a while. Yeah, and to really clarify and define things, you could kind of think about like the playground that we talked about as an app that has been developed by these different people that illustrates some LLM functionality, but it's usually not the app that you're going to build. You're going to build another app that is exposed to your users that uses the functionality and you'll need to host that either in ways that people have been hosting things for a long time or new interesting patterns that are popping up like things that modal is doing or maybe things that front-end developers really like to use like for cell and and other things but there's still that app hosting side now where i think things get interesting is you have the playground you have the app hosting but regardless of both of those what happens under the hood and And this is, I think, where things get quite interesting and where there's a lot of differences in kind of emerging generative AI stack compared to the maybe more traditional non-AI stack.
11:09In the middle of the diagram that we're talking about, this emerging LLM app stack diagram, which I think also is, again, more general, is this layer of orchestration. So I don't know about you, Chris, but I am old enough, I guess, you don't have to be that old, I don't think, to when someone says orchestration, I think of like Kubernetes or like container orchestration. And maybe that's my own bias coming from working in a few microservices oriented startups and that sort of thing. But this is distinctly not the orchestration that's being called out here in the generative AI app stack. There's a level of orchestration, which in some of my workshops, I've been kind of referring to as almost like a convenience layer.
12:04Think about like when you're interacting with a model. Let's give a really concrete example. Let's say I want to do question and answer with an LLM. I need to somehow get a context for answering the question. I need to insert the question in that context into a prompt. And then I need to send that prompt to a model. I need to get the result back and maybe do some cleanup on it. Like I have some stop tokens or I want it, you know, to end at a certain punctuation mark or whatever that is. That's all convenience. What I would consider sort of this convenience and what they're calling orchestration around the call to the model.
12:49And so this orchestration layer, I think, has to do with prompt templates, generating prompts, chains of prompts, agents, plugging in data sources like plugins. These are all things that kind of circle around your AI calls, but aren't the AI model. Yeah, I mean, it's the software around it, you know, just to simplify a little bit. Yeah. And maybe tooling. Yeah. Orchestration tooling. Yeah. Yeah. It's the stuff you have to wrap the model with to make it usable in a productive sense. And from the moment that I saw that word, that was almost the very first thing that grabbed me. You know, those, you know, little psychological quirk where you kind of notice the thing that sticks out.
13:35Yeah. That's the thing that stuck out was they, it's a big bucket that they're calling orchestration, which is a loaded word that can mean a lot of different things depending on what it is you're trying to do. and the examples that they list in that category are all somewhat diverse as well. I think that was the first point where I thought, well, it's a chart with the creator has a bias there. What are some of the ways, I'm just curious, when we think about this kind of orchestration, as they say, wrapping around and providing the convenience, any ways that you would break that up, like how you think about it?
14:10You mentioned convenience and stuff, but they go from something like Python as a programming language to LangChain to ChatGPT, all three very distinct kinds of entities. Yeah, I think that you're kind of seeing a number of things happen here. The first one that they call out is Python slash DIY. So you're seeing a lot of roll your own kind of convenience functionality built up around LLMs. But I do think one of the big players here would be like LangChain and what they're doing because if you look again at those kind of layers of what's available there you have maybe categories that i would call out if we just take lang chain as an example categories that i would call out of this sort of orchestration functionality would be templating so this would be like prompt templates for example or templating uh in terms of chains so manually setting up a chain of things that can be called in one call there's also an automation component of it maybe the or this is a way that orchestration kind of fits with the older way the orchestration term is used in like devops and other things where some of it could be automation related to with things like agents or something like that where you have an agent that automates certain functionality it's not the LLM itself, but it's really automations around calling the LLMs or the other generative AI models to generate an image or what have you.
15:50They also kind of have some separate call outs, you know, for APIs and plugins. And then they have, which we can hit in a moment, they kind of have a collection of the maintenance items, you know, the things to keep the lights on, if you will, logging and caching and things like that. How do you look at that breakdown the way they have it? Yeah, so I think this is where they kind of have the orchestration piece in the middle there as connecting a couple different things. One of those would be what I would consider, I think, more on the data or resource side. And then one is more on the model side.
16:29So I think we could split it into those two major categories. So what are you orchestrating when you're orchestrating something with Langchain or similar? Well, you're orchestrating connections to resources. I'll use the term resources because it might not be data per se. It might be like you say, like an API or another platform like Zapier or Wolfram Alpha, something like that. The other side of that is the model side. both the model hosting and some really useful tooling around that. But let's start on the resource side. So as you mentioned, you might orchestrate things. One of the things that I've found both really fun to do and useful is to orchestrate calls into a Google search.
17:18So if I want to pull in some context on the fly, then I might want to do a Google search. That's a call to an API. So that's a resource or a plugin that might be conveniently integrated into your orchestration layer, either via something like Langchain or via your own DIY code. Another side of this would be the actual data and the data pipelines, which are your own data or data that you've gathered or is relevant to your problem. So again, if we're thinking about this sort of set of resources that could be orchestrated into your app, maybe you have a set of documentation that you want to generate answers to questions out of.
18:05Or maybe you have a bunch of images that you want to use to fine tune stable diffusion or something like that. Having data and integrating it into models isn't new. And so the things that are called out in this particular image, like data pipelines, those are also not new and are part of this app stack if you're integrating your own data. So things like Databricks or Airflow or Pachyderm or tools to parse data. So PDF parsers or unstructured data parsers or image parsers or image resizing or all of that sort of stuff still fits into the data pipelining piece. And so you've either got your data coming from APIs, which might be a resource that you're orchestrating, or you've got your data coming from your data sources, which might be traditional data sources of any type from databases to unstructured data.
19:12This is a changelog news break. It's official. advancements in computer vision have rendered captchas obsolete as new research shows ai bots are 15 percent more accurate than humans at picking which images have a bridge or sign or bicycle or whatever in them the researchers recruited 1400 participants to test websites that use captcha puzzles which account for 120 of the world's 200 most popular websites the bots accuracy ranges from 85 to 100 % with the majority above 96%. Meanwhile, we mere mortals check in at a pathetic 50 to 85 % accuracy and we answer slower than the robots to add insult injury.
19:58I've surmised this for months now as we've been unable to ward off spam account creations on changelaw.com no matter which shiny new CAPTCHA service we tried. There are other efforts in the works besides CAPTCHA in order to differentiate between robots and humans. But so far, the robots are winning. You just heard one of our five top stories from Monday's Changelog News. Subscribe to the podcast to get all of the week's top stories and pop your email address in at changelog.com slash news to also receive our free companion email with even more developer news worth your attention. Once again, that's changelog.com slash news.
20:42Well, Chris, part of the data piece or the resource piece that is kind of unique within this new generative AI app stack is the embedding and the vector database piece. And I have to say, I've just got to recommend that our listeners, if they haven't, listen to our very recent episode about vector databases, because that episode goes into way more depth in terms of what a vector database is and why people are using it. But just for a quick recap, part of what you might want to do with generative AI models is find relevant data that's relevant to a user query and somehow orchestrate that into your LLM calls, either for chat or question answering or maybe even into image generation or video generation.
21:41In order to find relevant data, what people have found is that they would like to do a vector or an embedding search on their own data to find relevant data. And again, you can find out much more about that in our previous episode, but that's called out in this app stack as probably something unique that's developing, which is not just having data pipelines and databases, but having data flow through an embedding model and into a vector database where you're performing semantic searches. I mean, at the end of the day, it's a database that's very, it works well for the kind of operation that we're doing here.
22:23Whereas some of the traditional things that we had been working on for years before, there's kind of a context shifting in terms of how you're handling data, what data is, how it's organized. So this makes a lot more sense. Yeah, and it should call out here too, part of the stack here, and I'm glad that they called it out in this way in the schematic that we're looking at, is the embedding model. So a lot of people are talking about these vector databases, but in order to store a vector in a vector database, there is a very relevant component to this stack, which is the actual model that you're using to create embeddings.
23:03and not all are created equal. So think about if you are working on an image problem, right? You may use a pre-trained feature extractor type model from hugging face to extract vectors that your images, so put an image and get a vector out. But if you're working with both image and text, for example, maybe you're going to use something like clip or one of those, a related model that's able to embed both images and text in a similar semantic space. But if you're only using text, there's a whole bunch of, of course, choices. And all of those don't perform equally for different types of tasks as well.
23:50If you search on Hugging Face or just do a Google search for a Hugging Face embeddings leaderboard, there's actually a separate leaderboard. So Hugging Face has a leaderboard for open models and how those score in various metrics. They also have a leaderboard for embeddings and you can click through the different tasks. So let's say you're doing retrieval tasks, like we're talking about here from a vector database, you can see which embeddings perform the best according to a variety of benchmarks in retrieval or in summarization or other things. Do you use that a lot when you're putting models and storing them into vector databases and figuring out the embeddings?
24:34Do you tend to go and see what is going on? Because right now there's so much happening in that space. Does that make for a good guidepost for you? Yeah, yeah. And I think what is also useful is looking at those performance metrics, but also at least on the Hugging Face leaderboard and some other leaderboards. So if you search for, if you're working with text, one of the major tools for creating these embeddings in a really useful way is called sentence transformers. And they have their own table where they have measured and benchmark various embeddings that can be integrated in sentence transformers.
25:13That's useful, but it's also useful to look at the columns, whether you're looking in the hugging face leaderboard or the sentence transformers or wherever you're looking at the size of the embedding and the speed of the embedding because it was called out when we had our vector database discussion but only in passing let's say you want to embed you know 200 000 pdfs so i just ran across this use case with some of the work that we're doing and it can take a really really really really long time depending on how you implement it to both parse and embed a significant number of pdfs the same would be true for documents or other text or other types of data even and so when you're looking at that there's two implications here one is how fast am i able to generate these embeddings do i have to use a gpu or can i use a cpu because there's going to be a different speed on GPUs versus CPUs and how big are the embeddings this is another kind of interesting piece which is if I have got embeddings that are thousand or more in dimension that's going to take up a lot more room in my database and on disk than embeddings that are 256 or something like that so there's also storage and moving around data implications to how you choose this embedding space.
26:42So there's a lot of, I think, practical things that maybe people skip over here when they're just doing a prototype with Langchain and some vector database. It's easy. But then as soon as you try to put all your data in, it gets much harder. You raised a question in my mind, and I'm going to throw it out. You may or may not be familiar with what the answer would be. But when you're looking at vector databases, and you're looking at all this, you know, the diversity in embedding possibilities here, and the fact that that has kind of physical layer consequences, you know, in terms of storage and stuff like that, are we seeing that in vector or other database arenas where they're trying to accommodate this new approach to capturing data in terms of having embeddings the way of vector, you know, with the rise of vector databases?
27:33It seems that there would be a whole lot of kind of vendor related research on how you do that. Because to your point a moment ago, you're talking about data. It's such a volume that poor architecture in terms of what's under the hood could have some pretty big consequences there. Yeah, I think that's definitely true. And there was a point that was made in one of our last episodes that the vendors for these things are having different priorities that don't always align. So some are optimizing for how much data, how quickly you can get a large amount of data in, but maybe they're not as optimized for the query speed.
28:14Some are optimizing for query speed, but it might be really slow to get data in. And so that's one piece of it. I think another piece of it is how large of an embedding do you need and how complicated is your retrieval problem, right? I would recommend that people do some testing around this because let's say you have 100 ,000 documents that are very, very similar one to another or 100 ,000 images that are very, very similar to one another. And the retrieval problem is actually semantically very difficult. You might need a larger embedding and more kind of power, even like optimization around the query, like re-ranking and other things to get the data that you need.
29:04Whereas if you have, you know, 100 ,000 images and they're all fairly different, well, maybe you don't need to go to some of those links. So, yeah, I think that that's also part of this problem. And people are still feeling out the best practices around some of this. partially because it's this kind of new part of the AI stack and partially because things are constantly updating as well. So if you use this embedding today, there's a better one tomorrow and vector databases are updating all the time. So it's also just a very dynamic time here. As we look at the chart here, and there's kind of the three that we referred to earlier that are kind of together, and those are LLM cache, logging, slash LLM ops, and validation.
29:54First of all, could you kind of describe what's encompassed in each of those? And also kind of why are they fit together? Why are we seeing those lumped in one category here, one super category? Yeah. So if you think about what we've talked about so far, there's this new generative AI stack, whether you're doing images or language or whatever, there's an application side, which might just be the playground or it might be your own application. there's a data and resources side which is what we've talked about with integrating apis and data sources and then there's like a third arm here which is the model side and all of those are kind of connected through the orchestration layer the automation layer the convenience layer whatever you want to end up calling that so on now we're kind of going to this third arm of the model side and we can come back to it here in a second but one side of this is just hosting the models and having an api around them which we can come back to but between the model and your orchestration layer almost as maybe we could call it like model middleware i'll just go ahead coin that i just coined it on maybe people are already referring to it that way and i didn't coin it, but model middleware sits kind of either wrapping around or in between your orchestration layer and your model hosting.
31:18And these are the things that you're referring to around caching, logging, validation. Probably the one that people are most familiar with, if they are familiar with one of these would be the logging layer, which is again, something that is kind of a DevOps-y infrastructure term but here we might think of very specific type of logging like model logging which might be more natively supported in things like weights and biases or clear ml or these other kind of ml ops type of solutions where you're logging requests that are coming in prompts that are being provided, response time, GPU usage, all the kind of model related things.
32:09And you want to put those into, you know, graphs and other things. So there may be specific kinds of logging. So how quickly on average is my model responding? What is the latency between making a prompt or request and getting a response. How much GPU usage is my model using? And do I need more replicas of that model? These sorts of things can be really helpful as you're putting things into production. So that's a first of these middleware layers.
33:00so chris um the other middleware layers i would say that have been called out at least in what we're looking at is our validation and caching so i'll talk about caching and we can talk about validation a little bit which is close to my heart but caching uh let's say that again this already happens in a lot of different applications so think about like a general api application If someone makes a request for data in your database and you retrieve that data, and then the next user asks for the same data in your database, the proper and smart thing to do would not be to do two retrievals, right? But to cache that data in the application layer in memory so that you can respond very quickly and reduce the number of times that you're reaching out to your database and things like that.
33:52I notice in this chart that some of the examples that they put for caching, such as Redis and SQLite and such, are very typical and long-term players in the app dev world. Yep. So does that beg the question, or at least for me, begging the question that when you're caching, like you're really talking about for the input, here's an output, whether it goes to a model or not. Is it really just application data that you're caching at that point? So it's caching in that sense, but I think there's maybe implications to it that go beyond kind of normal caching. So if you, you know, running AI models is expensive most of the time because you have to run them on some type of specialized hardware, right?
Read the full transcript
34:37If I've got a model running on two A100s, I would rather not have four replicas of that model. I would rather just have one if I can because I don't want to pay for all those GPUs. So part of it is really related to cost and performance. So it's also for a large model. This is mainly for large models, I would say. you've got a lot of cost either because you're running that model on really specialized hardware or because like if i'm calling out to gpt4 it's really expensive to do a lot of requests to gpt4 right so in order to deal with that if you have a prompt input you can cache that prompt and if users are asking the same question i would rather just send them back the same response from gpt4 or my large llama to 70 billion model, or whatever it is, I'm going to respond to them the same way, based on the same or a similar input.
35:44The other implication to this, which in my mind, it sort of fits into caching, but maybe not in the traditional sense. So I normally think of caching as like, oh, I'm going to cache things in memory or locally at the application layer. But if you're caching prompts and responses, there's a real opportunity to leverage that data to build your own sort of competitive moat with your specific generative AI application. So for example, like you've got a user base, they're prompting all of these sorts of things. All of a sudden, if you're saving all of that data and the responses that you're giving, you're essentially starting to form your own domain-specific data set that you could kind of leverage in a very competitive way in kind of two senses.
36:40One is right now, if you're using a really expensive model to make those responses, maybe you start saving those responses from the really expensive model and you can use that data to fine-tune a smaller model that might be more performant and cost effective in the long term. So it's an operational kind of play. The other way is if you're gathering that over time and you actually have the resources to human label that or give your own human preferences on that or certain annotations on that, that now is your own kind of advantage in fine tuning either one of these generative models or your own internal model for the domain that you're working in.
37:26So it's caching, but that's almost like a feedback or data curation side of things as well. So you mentioned earlier that validation was close to your heart. Yeah. So as our users know, I think part of the tooling that I'm building with Prediction Guard would fit into this category. It would actually span, I think, more categories. It kind of spanned between validation and orchestration and model hosting. So there's kind of a little bit of overlap there. But this validation layer really has to do with the fact that generative AI models across the board, I think people would say, there's a lot of concerns around reliability, privacy, security, compliance, what have you.
38:15And so there's a rising number of tools that are addressing some or all of those issues. So whether it be putting controls on the output of your LLM, again, think about this like a middleware layer. My LLM produces something harmful as output, or my generative AI model generates an image that is not fit for my users. I want to somehow catch that and correct it if I can, right? or I want to put certain things into my model, but I want to make sure that I'm not putting in either private or sensitive data or I want to structure the output of my model in a certain way into certain structures or types like JSON or integer or float.
38:59All of these sorts of things kind of, I personally would break this apart probably into maybe like validation type and structure and then like security related things. Cause there's a lot here. There's validation, which is like, is my output what I want it to be? There's security related things, which is, am I okay with putting the current request into my model and, or sending the output back to my users? And then there's type and structuring things. So with images, like is the image upscaled appropriately for my use case? or with text, if I'm putting in something and wanting JSON back, is it actually valid JSON?
39:44That's more of a structure type checking type of thing. So there's a lot in this category and I think you're probably getting the fact that I'm thinking a lot about this and there's a lot here. But yeah, other things fitting into this category would I think cool one called rebuff, which is doing kind of checking for prompt injections for example that's like part of that security side of things there's things like prediction guard and guardrails guidance outlines now that do type and structure type of things there is also i would say a layer of this which a lot of people are implementing in the kind of roll your own python diy way as well which in prediction guard we implement some of these but also people are implementing them in their own systems, like self-consistency sampling, like calling a model multiple times and either choosing between the output or merging the output in some interesting way or things like that, those sort of consistency stuff.
40:52I think a lot of people are rolling their own too. What do you think as we start winding up here, what do you think are some of the takeaways from this chart, you know, or what brings top of mind things that people as they look at it might benefit from? How would you see it in the large? Yeah, it's a good question. I think one major takeaway, one thing to keep in mind is the model is only a small part of the whole app stack here in a similar way to like used to when a thing existed called data science. We would say training a model is only a very small part of the kind of end-to-end data science life cycle of a project.
41:37There's a lot of other things involved. And I think here, you know, you can make a similar conclusion that the tendency is to think of the model as the application, but there's really a lot more involved. And there's our friends over at Latent Space would say this is really where AI engineering comes into play. This space of AI engineering seems to be developing into a real thing, whether you call it that word or not, it is part of what this is. So that's one takeaway. I think the other takeaway is maybe just kind of forming this mental model around these three spokes of the stack. So you've got your app and app hosting, you've got your data and your resources, and you've got your model and your model middleware.
42:26And all that kind of middle hub would be some sort of orchestration that you're performing either in a DIY way or with things like Linkchain to connect all of those pieces together. So you're probably hoarse by now because we've pulled so much information out of you. This was a really, really good dive. You know, it's one particular publisher's way of looking at it, but we've never really dived into all the components of the infrastructure of a stack with this kind of. And I think most people haven't had a chance to see it yet because so much of this has really arisen in recent months. Thanks for kind of wearing half of a guest hat along the way here and taking us through this on this fully connected episode.
43:12Yeah, and I think in terms of learning about these things, I think people can check out our show notes. We'll have a link to the diagram that we've been discussing here. I would say learning wise, this helps you organize your thought process. But to really get an intuition around these things, you can look at various examples in this diagram and go to their docs and try out some of that. There's a variety of kind of end-to-end examples as well that are pretty typical these days. Like in language, if you're doing kind of a chat over your docs thing, that involves a model and a data layer and an application layer.
43:52So just building one of these example apps, I think, could give people the kind of learning and that sort of thing that they need. But yeah, it's been fun. It's always helpful to talk these things out loud with you, Chris. I find it very useful. Well, I learn a lot every time we do this. So thanks a lot, man. Yeah. Yeah. We'll see you next week. See you next week.
44:22Thank you for listening to Practical AI. Your next step is to subscribe now, if you haven't already. And if you're a longtime listener of the show, help us reach more people by sharing Practical AI with your friends and colleagues. Thanks once again to Fastly and Fly for partnering with us to bring you all Change Talk podcasts. Check out what they're up to at Fastly.com and Fly.io. And to our Beat Freakin' residents, Breakmaster Cylinder, for continuously cranking out the best beats in the biz. that's all for now we'll talk to you again next time
From the publisher
Recently a16z released a diagram showing the “Emerging Architectures for LLM Applications.” In this episode, we expand on things covered in that diagram to a more general mental model for the new AI app stack. We cover a variety of things from model “middleware” for caching and control to app orchestration.
Changelog++ members save 2 minutes on this episode because they made the ads disappear. Join today!
Sponsors:
- Fastly – Our bandwidth partner. Fastly powers fast, secure, and scalable digital experiences. Move beyond your content delivery network to their powerful edge cloud platform. Learn more at fastly.com
- Fly.io – The home of Changelog.com — Deploy your apps and databases close to your users. In minutes you can run your Ruby, Go, Node, Deno, Python, or Elixir app (and databases!) all over the world. No ops required. Learn more at fly.io/changelog and check out the speedrun in their docs.
- Typesense – Lightning fast, globally distributed Search-as-a-Service that runs in memory. You literally can’t get any faster!
- Changelog News – A podcast+newsletter combo that’s brief, entertaining & always on-point. Subscribe today.
Featuring:
Show Notes:
Emerging Architectures for LLM Applications
Something missing or broken? PRs welcome!




