In short
TWIML AI Podcast Episode #659 - Patterns and Middleware for LLM Applications with Kyle Roche
Episode Overview In this episode, host Sam Charrington interviews Kyle Roche, founder and CEO of Griptape. They discuss innovative patterns and middleware solutions designed for Large Language Model (LLM) applications. The conversation highlights the concept of "off-prompt" data retrieval, secure middleware stacks, and emerging use cases that optimize human augmentation tasks.
---
Key Topics Discussed
Introduction to Griptape
- Kyle Roche's Background: Former AWS GM with expertise in IoT, visual effects products, and more.
- Griptape's Mission: To develop middleware that connects LLM applications with internal and external data systems securely.
Emerging Patterns for LLM Applications
- Off-Prompt Data:
- A method that allows data retrieval without injecting it back into the LLM's chain of thought.
- Benefits include handling larger datasets while alleviating privacy and sovereignty concerns.
- Pipelines vs. Chains:
- Griptape uses "pipelines" for sequential tasks, allowing different models and tools for each step without managing chains directly.
- This offers more flexibility in orchestrating complex workflows.
- DAG Workflows:
- Griptape supports Directed Acyclic Graphs (DAGs) to manage large workflows and maintain context between steps.
Privacy and Customer Concerns
- Discussion on addressing privacy, retraining, and data sovereignty concerns through the off-prompt approach.
- Examples of role-based retrieval methods as a way to optimize data access while maintaining compliance.
Real-time Retrieval and Use Cases
- Human-driven Problems: Focus on optimizing processes in industries like construction and visual effects.
- Customer Example: Automation of email parsing and PDF data extraction for construction parts sales, reducing manual labor from 75 employees to more efficient workflows.
Integration with AWS Services
- How Griptape integrates with AWS services like Bedrock and SageMaker, and offers abstractions through drivers to facilitate various LLM interactions.
- Ensuring contextual relevance and data integrity without overwhelming the LLM's context window.
The Future of Griptape
- Prioritizing product-market fit and exploring offerings on the AWS marketplace for enterprise customers.
- The goal is to remain agile while building a robust framework that addresses long-term challenges in LLM applications.
---
Conclusion Kyle Roche emphasizes the need for organizations to rethink their approach to LLMs, focusing on established patterns from enterprise data systems and balancing innovation with practical application. The episode showcases how Griptape aims to streamline and enhance the integration of LLMs into existing business processes, improving efficiency and decision-making.
For more information on this episode, visit [TWIML AI Podcast](https://twimlai.com/go/659).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:09All right, everyone. Welcome to another episode of the TwiML AI podcast. I am of course your host Sam Charrington. Today I'm joined by Kyle Roche. Kyle is CEO and co-founder of Griptape. We are coming to you live from the AWS reInvent conference. I'm trying to say the Future Frequency podcast studio. I've been here at the conference covering the AI announcements via X and LinkedIn, and be sure to take a look at those for the latest updates. Kyle, welcome to the show. Hey, thanks so much for having me. I'm excited to dig into what you're doing at Griptape. It sounds like you've got some interesting approaches to allowing folks to get more out of large language models, and we'll dig deep into that.
0:51But before we do, I'd love to have you share a little bit about your background and how you came to start the company. Sure. Yeah. So my name is Kyle Roach. I'm based out of Seattle. The last eight and a half years, I was a GM at AWS. So I built AWS IoT, moved into some of the visual effects products. I ran Thinkbox, Nimble Studio, geolocation services, simulation, and a bunch of other kind of things in the 3D space over time. Started Grip Tape in March of this year. So we left Amazon. There's a team of 16. I think 12 of us are from AWS. Okay. Spread out around the country. We can get into that later.
1:25But yeah, a huge fan of the company. I've been in the AWS ecosystem even before there. I had a startup called Telemetry before that, which Amazon acquired in 2015. That became AWS IoT. So just kind of been in and out of the community from both sides of the table, I think for the last couple of decades there. So awesome. So grip tape kind of connotes glue or binding together. And you describe the company as building middleware for generative AI. Tell us a little bit about what you're trying to achieve. Sure. Yeah. So we call ourselves a middleware stack for Gen AI. We're focused mainly on the enterprise from the perspective of the patterns that we think will emerge if enterprises are to adopt LLMs and attach them to large data systems, internally or externally.
2:10Some of the key things that GripTape says differently, we focus on a pattern which we call off-prompt, which allows for retrieval of data that doesn't get injected back into the chain of thought for the LLM. And in that way, you can work with larger data sets. We can not only ensure there's no retraining, but completely negate the concern about data coming over the wire to the LLM. That removes the need to do things like obfuscation or data sovereignty concerns. So that's one of the particular patterns that we work with. It's a little bit unique for grip tape. Also, we have abstractions for things called drivers, which map to models and then model, I guess, APIs.
2:44So things like SageMaker and Bedrock and Claude with Anthropic directly. We have abstractions for memory management. So everything from the conversational memory to what we retrieve on the fly and how we keep that in context for the LM. We have an abstraction called a prompt stack, which lets you kind of construct where in the context window, you want some of the data to go back to the LLM, which is interesting for larger context windows where it loses some of the data in the center. Okay. Maybe a point of reference that a lot of folks will have heard of is Langchain. It's fairly popular out there for building LLM applications and essentially allows you to chain invocations of an LLM inference.
3:22And you're approaching it slightly differently with this idea of off prompt. Let's dig into that a little bit more and what exactly that allows you to do and why you think it's important? Sure. Yeah. I mean, first, I think Lanchi has done a great job building this community and this pattern. And one of the things they're focused on is breadth of support. So they have thousands of contributors adding all kinds of different projects and drivers and things like that for those chains. Their chains are managed in the context window for the most part. So even data you're retrieving and things like that are put back in.
3:52They rely a lot on pre-processing of data, creating embeddings and kind of prepping things like that. So CryptiP is a little bit different in that we don't manage chains. So we can do things called pipelines, which are sequential tasks that we can give to LMs. Those models can be different per task, per step in the pipeline. Is that a semantic difference? It sounds like a chain. Yeah. So imagine I want to use Claude to drive the pipeline. So I want Claude to decide how I'm going to approach this problem. but within that pipeline i might use a different model for each step and each of those steps might have different tools attached and each of those might have different retrieval patterns so like you can kind of split up and so it would be like i think if you were to project it onto line chain it would be like a bunch of chains sort of with an abstraction higher than that which allowed you to defend orchestrate them at one step higher so we also support pattern we call workflows which are like basically dags like from the data pipeline space so you can have these large sprawling workflows that, again, have different tool sets, different models per step.
4:53We keep all that context in GripTape on the side for the LLMs. And so kind of circling back to this idea of off-prompt, the typical pattern for using chains or chain of thought is you make one invocation to the LLM, you get back some information, you maybe manipulate it, and you kind of stick that back into the prompt for your next request. And you're suggesting that maybe we should be thinking about doing something different we should keep some of that context off or out of the prompt yeah so by default grip tape actually so we keep anything we retrieve on behalf of the lm we keep it off prompt so we have a parameter that you would have to sue off prompt equals false to get it to go back into the chain of okay so a couple things that protects one it protects like running over the context window and then the way grip tape interacts with the lm is we give it enough context about the data it has and then tools in which it can interact with the data so So as an example, the AWS GitHub page is a samples project that has Bedrock, GripTape, Redshift, and then a couple other models.
5:53So that particular example uses Claude behind Bedrock to retrieve a pretty large amount of data out of a Redshift database. So it could be up to like 90, 100 gigs of data. That's managed by GripTape. So GripTape's interacting with Redshift on behalf of the LM. We keep anything that we pull back in memory vector database, and then the LM can interact with it through GripTape. So none of it ever gets injected back into the prompt. And then you can do things like have Titan summarize it off prompts. But yeah, you can do things with other models in kind of more of a local environment there. And so is the main idea that it's really an inversion of the default, the historical effect, we can say that about something that's so new, or the typical way of doing these things is you get back some response and you stick it all into the prompt.
6:38And in your case, you're suggesting that we be more selective and kind of craft the response manually from the things that come back. Yeah, actually, I think that the overarching point is actually very important there, which is this is all very new. So, like, I think everyone's approaching the problem from different perspectives. And, you know, a lot of the initial customer conversations we had, there was a lot of privacy concerns. There was retraining concerns. You know, there was, you know, are you going to use my data for this or for that? And then you also have like this sovereignty issue as well.
7:09Like, can I keep it localized and still do something that an LLM is good at? So Off Prompt came out of a lot of those conversations, like how do we just negate that need? So I think there's companies approaching it from different perspectives where they'll still send it into the chain of thought, but they'll obfuscate data or they'll swap it out with kind of anonymized substitutions or whatever. Yeah, we just chose to kind of handle it differently. And then we built tooling that allows the LM to still have contacts interact with it, but just kind of removes that need. So it was sort of born out of the customer conversations from privacy.
7:39Yeah, I think there's some side effects that are interesting, like, you know, Anthropics 200k window. There's research papers that talk about how the LM is very good at the beginning of the context window and the end of it. And it sort of gets a little bit fuzzy in the middle. And if you're selective about what goes in and out and, you know, how you sort of structure that prompt, I think these types of new retrieval patterns might be interesting. You mentioned that this pattern allows you to, you're still able to provide context to the LLM without putting the context or the responses back into the LLM.
8:11Can you elaborate on that? What does that mean? What's the context that you're able to provide that isn't coming from the prior step in the chain or in the process? So it's all available to the developer to kind of decide how you want to do that. So I guess it depends really on the use case. But in that example that I mentioned on the GitHub page from AWS, there's a customer database there in Redshift. So you might not have access to that as whoever's running this particular pipeline. So you can say, I need to go get the customer data from this particular subset. I need to do something with it, summarize it.
8:46But maybe I can't see those records. So protecting those types of access rights and things like that are available. if you don't inject everything back into the prompts, right? And I think we're already starting to see, imagine role-based retrieval. So a company might have an agent they publish internally that can be used by multiple people, and then it's going to attach to your HR database or your customer database or have some kind of PII, PCI compliance or concern. And different people have different rights to the subsets of data, and what Dell could go retrieve. So you need to manage that kind of like we've been managing every other distributed system.
9:23I don't think these are new patterns. We're just like, we're taking things from like old school, just, you know, enterprise data systems and projecting them onto like, what's this like with an LLM? Right. No, it's true. And it's important. I often will talk to people about RAG, which is kind of a popular pattern for using LLMs and making them kind of constraining them so that they don't hallucinate as much among other things as well as personalizing their output just on that point like rag is interesting to you like one of the things that we're experimenting with and grip tape is sort of built with in mind right is retrieval like at runtime or just in time retrieval i think amazon called it in their blog post if you look at kind of how most of these articles and tutorials are written today you go out you find the data you create embeddings you put it somewhere yep and then you do retrieval against that but in reality like you're going to have data systems that are moving and they're creating new records as you're asking for it.
10:19So you can't really have that kind of pre-processing and creating all the embeddings and going through that pipeline before you pull something that just happened. So yeah, Griptapes allows for this pattern of like real-time retrieval and we'll create the embeddings on the fly for the LLM and still keep it up prompt. So it looks the same to the LLM and like the application still built the same way, but like you could attach us to a customer ticketing database that's getting new information all the time. That's awesome. The point that I was going to make about RAG is that it is really compelling because it's very easy to get from like zero to one, from zero to a demo, but to make it really useful and kind of incorporate some of these enterprise-y things that we might want, like role-based context or access control and, you know, real-time and access to other systems, like not to mention even just tuning the responses so that it reliably produces kind of on target responses.
11:11It's a lot of work and a lot of what we need to do to, to fully take advantage of it, you know, isn't quite here yet. I agree. And I think those projects are super interesting for us. Like we've been trying to focus on customers that have like a human driven problem that could be optimized. So as an example, we have an industrial customer that they sell basically plumbing and piping parts for large construction sites, right? Okay. They have 75 humans that sit on a email box and look through CAD files, PDFs, just random emails. They try to decipher what parts need to go into a bill of materials into an order.
11:45You can't just take those and then throw them through like any of these embedding engines and then try to like, there's a lot of this, like the page break is in the wrong place and you have to kind of tweak things. So there's still a lot of that work, I think, that goes into just getting these things to be like low error rate and high reliability. So with that use case as context, can you walk through kind of these various patterns that you're seeing and how they materially change the way folks have approached this kind of problem? Sure. Yeah. So this one's interesting. We can stay on that particular use case.
12:19So now Griptape, this customer has an agent that runs on Griptape Cloud, which is basically a managed service that will run any of these structures that are written in Griptape. The agent sits on that email box, looks for new emails, parses out the PDF attachments, the bodies of the email, Excel sheets that are attached, kind of any of these varying formats. Gryptip has an abstraction called rule sets, which is how we kind of steer the behavior of the input and output of the LM. There's small textual based rules. So there's no prompting with Gryptip. It's all Python and like rules are just small sentences like RF means raised face.
12:51If someone's talking about pipe or something like that, you kind of feed a little bit of context to the agent. And then we basically create these build material parts. We pass it on to the coding system. So like all of that happens in real time. Every time a new email comes in, all the parsing of the PDF is kept basically off prompt in grip tape. And then the - So the rule sets are, that's happening outside of, that's not an LLM capability. It's more traditional pattern matching of some sort, or are you using the two some way? Yeah, we're fusing the two. So we have this, I mentioned earlier, we have an abstraction called prompt stack, which so the presets go into the prompt.
13:28So we have a very well-structured kind of prompt template that comes out of good tape. That's what the LLM gets. I mean, I think that's how all these things work, right? So got it. So it's like hints that you're giving the LLM. I saw PF, that means this. Incorporate that into the response as opposed to pulling in the document or something that refers to that and expecting it to figure it all out with slang like this particular use case is a lot of customers are using abbreviations and words that don't like have a deterministic mapping database or something so i mean so you could just continue to build this massive lookup table or we kind of sear it towards like here's some patterns that like those are these lms are very good at right like write some patterns go find things that look like that too so it's interesting i joke about how you know there's a strong regex system behind every powerful LLM app or, you know, any production like NLP system has some kind of rules that somewhere it's interesting that you've made that kind of a first order component of the platform.
14:31Yeah. I mean, like back to kind of what we were hinting around earlier too, I think this is like an enormously transformational phase for everybody. Everyone wants to get into it. And a lot of the patterns that we know and trust in other systems are going to just emerge here too. So like, I think you can do things that we've trusted in the past and have worked well and just put them in context of LLM based applications. And most likely they're, they're going to be needed in short order. So. Got it. So you've got these rule sets that are applied to the email and the PDF that's helped the LLM kind of map between, you know, what it already knows and things that are specific to this particular use case.
15:11That's may be an interesting way to avoid having to do something like fine-tuning? Yeah. So when possible, we're pushing on RAG or something like RAG. This isn't exactly RAG, but like instead of fine-tuning. And I think if you project out a little bit longer, one of the, I think, most compelling use cases that we've heard from customers is like back to, again, like role-based retrieval or something like that, where if you push it into fine-tuning, the model has access to all this data. And what you really want is like a model that is kind of orchestrating and behaving like the brain of this sort of problem solver, but then the retrieval is still role-based.
15:48It looks a lot like it does today, right? We have permissions, I have access rights. Like you can't just dump it all into the LLM. So I think in those types of cases, like these other retrieval patterns are much more important than fine tuning. Are you managing kind of sessions, session state on behalf of the, you know, on behalf of the application, like the thousand people that are talking to the system, they're all coming through you. You're managing those sessions and applying ACLs or access controls to each of those sessions independently. Yeah. So the open source framework supports that. If you want to do that on your own, GripTape Cloud has a bunch of other kind of abstractions on top of that and APIs.
16:23So you can manage sessions could be a collection of runs of an agent or pipeline or workflow. And then we have a bunch of other data structures behind that to keep context history. We can switch between modalities on the fly too. So if you're going from one of the demos in our booth here, we're basically using Bedrock to summarize something that we retrieved on the fly from webpages. And then we're creating images with stable diffusion and Leonardo behind Bedrock. So switching from text to image, like it'll, all that will happen with Grip Tape on the fly. So we can move between those modalities and you can basically get an image inside your chat.
16:58You don't have to do a bunch of other kind of things around that. So those are all available when the agents or the structures run on Grip Tape Cloud or on your own. Yeah. You mentioned there's a DAG abstraction. Can you talk a little bit about how that comes into play and how that is distinct from, again, kind of this chaining that we've been using as a touchstone? Yeah, I think there's similar approaches with a blank chain as you can spawn out and do three things in parallel and then come back. So those look a lot like DAGs. I think we've spent a lot more time on how do you keep the data cached between steps passed appropriately between them?
17:35Like, can I isolate permissions, which model I use at which step? Like how these things are kind of constructed. And it's just a little bit more of, it's not deterministic obviously, because it's an LM, but like as close to like, I can map out what I want the LM to do, like step by step. So, you know, I think again, one of our sweet spots is the larger amounts of data, the more data that we're retrieving, the better we'll do in those types of scenarios. So because of context limits and keeping kind of watch on step by step how this thing is working. So does that make sense? Yeah. And is the DAG aspect of it, is it like, is it your own DAG kind of engine and language or are you using some existing thing like an airflow or something else?
18:21No, it's all in grip tape. Yeah. So, okay. Yeah. I think I've seen some posts about like, it's kind of like an airflow for LLMs. So we've been proxied as that before. Yeah. So I think that's a good pattern just to keep in mind. But yeah, it's all in the open source project and available on Coffee Tubs. All right, cool. So you mentioned something about real time and the context was if you're doing, again, this role-based access control, you're pulling information out based on the user's identity at a given time. Can you talk a little bit more about what you've seen in terms of usage patterns for real-time, integrating real-time information into LLMs?
19:00Well, I mean, yeah, I think we went through some of those. But just imagine any ticketing system that's a super easy example. Tickets are coming in. You have to be able to go get those, retrieve them, create. You still create embeddings. That's how the LLM best understands what's going on. So GripTip creates those. It keeps them in memory. or you can put them in PyCon or something else if you want. But I think those patterns, anywhere where we're going to hit an external API, we're hitting some system that might not be static or was pre-processed, they're all good fits for us. So we tend to kind of focus on finding customer examples that are proving out that RAG in a more real-time fashion is important or I want to interact with large amounts of data outside the context size.
19:45Yeah. And you had mentioned like pulling in information from a database and sticking it into some kind of embedding thing. Like, yeah, that seems counterintuitive, at least for structured data. I know. I was actually just going to say like LMs are not awesome at structured data. Right. So I think this is an interesting pattern that like we've been working on trying to get pretty stable and which the LM will create the query for us based on the schema of what it's looking at. Like as a redshift example earlier I mentioned. So it has enough context. Here's the schema. I can create a query. I can tell you, you know, it can tell Griptape what it wants to go get.
20:22And then we can go get it. And then after that, it'll have embeddings in some kind of structure that it can understand, do things like summarization or Q &A or generate something out of that. But, you know, I think we can go from structure to maybe a semi-structured in between. And that's a pattern I think is interesting. So are the embeddings useful on like traditional structured information or is it when there's also unstructured along with the structured? I guess I envision that for like, if you're pulling financial information or something like that, that's very structured, you'd want to just give the LLM that information as context and have it generate its summary or it's, you know, whatever text is going to generate as opposed to, I'm trying to envision like how sticking that structured financial information into some kind of embedding or factor database.
21:09Yeah, you might not have said it all, right? Okay. You might be able to just handle it with like LLM as a query. We go get data and we do something else with it. Yeah. So yeah, I think that flexibility is depending on the use case. But the developer has all that control. So yeah, so we have like Bedrock has what they call activities. Like LinkedIn has tools. We started off with tools and then we started moving them into what maps to something that looks more like an API. So like I think Bedrock is going in a similar way. Their agents, they have, I forgot what the exact word for it is, but it's an open API spec in a Lambda function.
21:39And if you give the LM context about an API, they seem to understand that better than like an ambiguous description of what a tool. and what functions it could map to. So I think they do last well with those. So GripTip, we've started to move it towards that same pattern that Bedrock introduced, which OpenAI was also using on their plugins last year. So I think it's a stable pattern to kind of project what an API looks like inside the context of the prompt. Okay. Along those lines, Bedrock has come up a few times in this conversation, certainly came up this morning in the keynote. Talk a little bit about, like are these competitive solutions i'm sure you're going to say no you're here as an aws customer slash partner but like it sounds like there's a lot of overlap i think that there's overlapping you know this whole ecosystem like we talked about link chain you know haystack grip tape bedrock agents i think we're all approaching how you take an lm inject chain of thought and then try to solve problems with it yeah i think we're all approaching it from maybe different angles.
22:40And like, there's ways that a lot of these tools and projects can work together. So like a lot of our examples, we have a more complicated workflow that uses SageMaker or Bedrock underneath or a combination of both of these. So yeah, we're focused on, like I said, just a little bit higher up the stack maybe than Bedrock Agent. And I think also just the ease of use from the developer's perspective, like that's always kind of an opening for, I think, third parties to help accelerate some of these patterns. So it's great if you're really familiar you're with AWS, but like usability is still, it's still an open, open opportunity.
23:13I think. Does Q fix that for us? I mean, Q is cool. Like that was really awesome. Yeah. I mean, I'll be referring to Amazon Q, which was also announced in the keynote this morning, which is a dialogue based chatbot that sits on top of AWS services and third party services as well. And And one of the potential benefits is helping AWS users better navigate the multitude of AWS services and capabilities, which can often be daunting. Yeah. No, I mean, that's always been a key focus. All those years I spent at Amazon, there's a lot of services, a lot of people working on them, a lot of separate teams.
23:53And how you can go from that to they have to keep velocity up, but also try to keep that user experience for developers somewhat coherent. I mean, I think Q, like it plugs a much needed, you know, much needed gap. So, yeah, I think that's a great project. I think for frameworks like ours, like not everything is conversational in nature. So, like, I think when you get to these like task based or event driven sort of workflows that still leverage LLMs or have like a very complicated pipeline or workflow after that, like those are still outside the scope of what Q is today anyway. So, yeah, I'm curious to maybe talk about you already talked about some use cases.
24:34What's the most either interesting or complex or one of each use cases that you've come across? Yeah. So, I mean, I think just as a company and as a team, we're not very attracted to like the demo sort of magic stuff. So we kind of focus on like, is there a core problem that like has human augmentation or like optimization or can I remove a redundant task? You know, you're saying all the real work and money's in boring stuff. I think we're very, we're very attracted to boring stuff. We've been doing this a lot. Yeah. So like, we're not, I'm sure you've seen some interesting things. I do. Yeah.
25:08So I think like those examples are meaningful. Like if you can take 75 sales reps and optimize that job. So they're working on quotes instead of reading PDFs and emails and CAD files. Like, yeah, that's a meaningful change. um we've been working a lot a lot of us have history in vfx also so so one of the other demos we brought here which we've been working on with some vfx animation studios is automating the production work so like basically did a bunch of integrations around autodesk shot grid which manages you know shots and sequences and props and assets that go on to like a visual effects production um you can imagine that like a producer might want to know like hey i just changed this character's outfit or prop or whatever like where else was it used summarize all the things that need to be reshot or re-rendered yeah so we're doing a lot of work in that space too and you know and i think for us it's a pipeline it's event driven it's not necessarily conversational and it has a lot of integrations and moving parts different modalities of data so that's a use case that we're very interested in but also does it lean on some of these same patterns that you've referred to that did those patterns kind of uniquely give you the ability to solve that particular problem and if so which like which ones and in what ways yeah i think not so much like the off prompt piece there i can see that being interesting when you get to like as an example i ran rendering for aws for a number of years and we have customers that they want to render something but like let's say like uh there's a new lightsaber or secret weapon or something in a production i'm trying to not say production names but that uh like can't be leaked out or whatever so like i can see how off prompts might be interesting in those types of cases where you want to query what's going on in your job but you might not want that to run back over the wire but in these particular cases i think the strength that we bring to the customer cases around the pipeline like event-driven workflows and then keeping context over long running and sprawled out jobs because you might have a thousand assets that need to get like refreshed or like have some kind of analysis done on them and that you know it's a very long running flow to to put into a or just some other kind of abstraction.
27:10So again, it's just a, it's a very like boring human problem that's been there for the last 20 years. So like people make movies the same way. There's ways to optimize some of how some of those jobs get done. And that's where we try to focus. Right. Talked a little bit about Bedrock and are you primarily used in conjunction with Bedrock and or AWS or is that one of, you know, some number of environments or usage patterns? We have an abstraction called Drivers, which maps to a whole bunch of LN providers. I think this is another area where we're maybe different from some of those other community projects in which we're kind of more curated.
27:48So we're not getting random PRs from 10 ,000 people. We've selected the ones that we think are appropriate for the types of businesses that we're going after. But we do a lot of work with Andropic directly and through Bedrock. So I think the off-prompt nature of the conversation leads to safe AI response. So like, I think a lot of the narrative that Anthropic and Bedrock together are speaking about, you know, seems to align with what we're talking to customers about too. Okay. But yeah, we also, I think, yeah, even just Bedrock and SageMaker, we kind of moved between those in the same workflows as well, which is interesting.
Read the full transcript
28:23In the context of embeddings and vector databases and all that, are you doing anything kind of interesting to help folks either construct those embeddings, load them, orchestrate them together to provide effective context for LLMs? Again, I alluded to this earlier. It's one of these things that is easier said than done in a lot of cases. Yeah. I mean, I think everyone's kind of scratching away at that problem in different ways. So yeah, we do have loaders and we will create embeddings for certain types of artifacts and things like that. We're not in that business. So like, you know, we still have drivers for Pinecone and Mongo and OpenSearch and whatever makes sense, right?
29:07So yeah, I think we're trying to be very disciplined about staying in that middleware space. So we'll all facilitate that. And just even back to that same example we talked about with the PDF, I mean, the way those PDFs are different per customer that's sending them in and then per month they're asking for, it wouldn't work if you dragged and dropped that into any other tool that's creating embeddings right now. So there's still a lot of manual kind of tweaking and we're trying to build tooling around, you know, how do you make that more accessible to developers? So like it's been, you know, a couple of days, like, like the page break is different and this guy formatted it different or put it inside way.
29:39Like there's, there's things like that, that you just have to kind of still look at and try to figure out if there's a way to optimize those. And do you think of that as like, I'm trying, I'm hearing you describe this use case and I'm trying to think about like what's different about it than the typical rag use case. Is it that there's like some inherent structure that you want to exploit that's not just like blob of text or is it like, how do you think about what's different about it? Well, I mean, like a simple example would be just like, here's a huge PDF that might probably not with like cloud two or something, but like could overrun the context window of what you're doing.
30:14Right. There's also just cost on like how much needs to go back to the LM versus like you said earlier, some of these problems are not LLM problems. Like, you know, I think you could probably get a lot of it done with just a normal rag kind of example. But yeah, there's just tweaking and there's interesting areas to kind of make sure that you're using LM for what it's good at. and you're not like just using it by default for every single thing that you're trying to get done. So it's not the answer to like every problem, right? It's an answer to like a nice set of new problems that we have now.
30:44Yeah, I think what I was like drilling in on was there's, I think, an important class of applications that aren't just like text manipulation and like rag style, but it's using the LLM potentially in addition to some context to generate queries, like you described and manipulate the results of those queries to present them back to the user in some way. So it's like I'm gluing together like the unstructured and the like the structured interactions. I think that's right. Yeah. API calls like running functions, creating new things. If you wrap all this together, like it looks like a normal middleware problem just happens to be in there.
31:26Yeah. I think I was like, if you personified or whatever, it's like one of the guys on our team had an example where he was like, it's a 10 year old who's can memorize any encyclopedia. He's got a set of them, but like has never left the house. So he still have to kind of like teach him basic. Like he's super smart in the context of things he, you know, he's, he or she isn't trained about, but like you have to teach him the next task and this task, like all the time. Right. So, yeah. So I think just some of those primitives just have to be, I mean, not reinvented because we already know them. We've, we've done them in other spaces.
31:56We just need to like reapply them to this particular problem space. Right. It's a lot of old boring stuff that we've already done right that's that's what we think anyway so yeah this is outside of the science the science is obviously like magic and amazing like but like other smarter people are doing that and then we're trying to build that the boring stuff on top so got it what do you think about the keynote this morning i thought it was great i think it was one of my favorite ones of the i mean i think there's my eighth to reinvent and like my only my second one as a as a customer but yeah i don't know i think adam did agree what jumped out at you i was so it was just cool to see like you know i think having like jensen there and like that partnership so important to like amazon customers and like just to see that there's like some synergy publicly with nvidia going forward i think is is great it wasn't all ai i think he didn't come out of the gate with like ai and they're like he sort of went to the core business and then you brought it in gracefully so i thought it was well done but ai was like main stage much earlier in the week than we usually see, right?
32:53Like I tweeted to Swami this morning, like what's left for you to talk about? I'm sure you'll come up with something. Yeah. Yeah. I'm really looking forward to his talk. He's an amazing speaker. But it was interesting. Someone on the expo floor was just kind of talking about like 10 years ago was you look at the booth since cloud, cloud, cloud, cloud, cloud. And then it was like data, data, data, and then data leak, data leak. Now it's like everything is AI. Like it doesn't matter if you're a staffing company or education, it's like education with that. And yeah, so I think it's, it's just that it's on everyone's mind right now.
33:22And I think that's what people are trying to figure out. How do they get through that more quickly? So where do you see grip tape going? What are some of the big priorities, directions, features, like what's top of mind for you looking forward? Yeah. So this week at the show is the first time we've shown grip tape cloud. So the managed service that runs anything built with the open source project. So today we've, we've only had the open source project out in public and we'd be doing kind of one-off projects with customers sort of quietly. So we're excited to launch that, like show developers what we could do.
33:55There's a couple of demos on the show for, like I said. We have some time in the AWS open source booth, so looking forward to like getting that out there. But yeah, the next phase for us is really just focus on product market fit, make sure we're building the right thing for the right customers. We're going to be working on AWS marketplace offerings so that you can run the backend stack in your own account context, which I think is important, especially if you get into like address and bees or enterprise customers you know they want to control that they don't want to like run that over a small like size provider so so yeah those are the next key steps for us i think we'll stay small through this next phase we're we're 16 people now we have a great group of investors you know we close the seed rounds like i said we've we're all startup veterans so responsible stewards of capital so like we will we'll take the next phase very tactically so i think there's a lot of noise in this space that you got to be careful not to like latch onto the short-term trends and keep your eye on the longer price.
34:48So awesome. Well, Kyle, thanks so much for joining us and sharing a little bit about what you're working on and the way you're thinking about helping folks get value out of LLMs. Yeah, thanks so much for having me. It was a great experience. Thank you. All right, everyone, that's our show for today. To learn more about today's guest or the topics mentioned in this interview, visit twimla.com. Of course, if you like what you hear on the podcast, please subscribe, rate, and review the show on your favorite podcatcher. Thanks so much for listening and catch you next time.
From the publisher
Today we’re joined by Kyle Roche, founder and CEO of Griptape to discuss patterns and middleware for LLM applications. We dive into the emerging patterns for developing LLM applications, such as off prompt data—which allows data retrieval without compromising the chain of thought within language models—and pipelines, which are sequential tasks that are given to LLMs that can involve different models for each task or step in the pipeline. We also explore Griptape, an open-source, Python-based middleware stack that aims to securely connect LLM applications to an organization’s internal and external data systems. We discuss the abstractions it offers, including drivers, memory management, rule sets, DAG-based workflows, and a prompt stack. Additionally, we touch on common customer concerns such as privacy, retraining, and sovereignty issues, and several use cases that leverage role-based retrieval methods to optimize human augmentation tasks.
The complete show notes for this episode can be found at twimlai.com/go/659.




