In short
Latent Space Podcast Episode Summary
Podcast Information Title: Latent Space: The AI Engineer Podcast Episode Title: Emergency Pod: OpenAI's new Functions API, 75% Price Drop, 4x Context Length Description: In this episode, the hosts discuss significant updates from OpenAI, including the introduction of a new Functions API, a substantial drop in pricing for embeddings, and an increase in context length. The discussion features insights from various experts in AI engineering.
Key Topics Discussed
- Introduction
- The episode is an emergency session prompted by recent updates from OpenAI.
- Participants include AI engineers from diverse backgrounds.
- The discussion aims to provide a deep dive into new features and implications for developers.
- Recapping June 2023 Updates
- Introduction of the Functions API, analogous to the ChatGPT plugins.
- Context length increased to 16K tokens from 4K, allowing for more extensive input.
- Significant price reductions: 75% drop in embedding costs and 25% off GPT-3.5 Turbo.
- Functions API
- Overview: The new Functions API allows developers to specify functions for the AI model to use, enhancing its utility and flexibility.
- Key Features:
- Function Roles: Introduction of a new role for functions, alongside user and system roles.
- Function Selection: The model can choose between multiple functions based on user input, akin to how plugins operate.
- Security Considerations: Caution against prompt injection attacks, as functions are more vulnerable when they can perform actions in the real world.
- Context Length and Pricing
- Context Length: The shift to a 16K token limit allows for more complex user inputs and prompts, potentially improving performance in tasks requiring extensive context.
- Price Reduction: Lower costs make it feasible for broader use of embeddings and model calls, which is crucial for scaling applications.
- Comparison with Other Tools
- Functions API vs. Google Vertex JSON: Discussion on how different platforms approach function enabling and structured responses.
- Exploration of alternatives like Langchain and other frameworks that have previously offered similar capabilities.
- Prompt Engineering and Fine-Tuning
- Transition from prompt engineering to fine-tuning as a means of achieving structured outputs.
- Concerns about the reliability of the model in generating valid outputs and the potential for hallucination.
- Security Concerns
- Security vulnerabilities inherent in allowing models to call functions directly, particularly the risk of prompt injection.
- Discussion on how OpenAI is addressing these challenges and the importance of reversible actions.
- Future Directions
- Participants express interest in smaller models, better retrieval mechanisms, and enhanced functionalities.
- There is a call for OpenAI to open source some of its models and to improve the UX beyond just chat interfaces.
Key Takeaways
- The emergence of the Functions API represents a significant evolution in how developers interact with AI, allowing for more tailored functionality and integration of complex operations.
- Price drops and increased context length are crucial for democratizing access to advanced AI capabilities.
- Security remains a critical concern as the capabilities of AI systems expand, necessitating careful design and implementation.
- Ongoing developments in the AI landscape suggest a shift towards more open-source solutions and improved user experiences.
Closing Remarks The episode concludes with a desire for developers to explore the new tools and capabilities provided by OpenAI and for the community to stay engaged in discussions around these advancements in AI technology. The hosts highlight the importance of continuous dialogue and exploration in this rapidly evolving field.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:04Hello, everyone. This is Swix back again with another emergency pod. The last time we did this was in March when OpenAI released Chat2PT plugins. And the new functions API today is effectively the Chat2PT plugins API now available to all developers. And with a whole bunch of other news, 75 % price drops on embeddings, four times the context length, and a lot more other updates. So what we do in these situations when there's breaking news and it's very developer-focused is we convene all the friends of the pod. We have Simon Wilson, Riley Goodside, we have people from Microsoft Research, Hugging Face, and Pinecone, and more that I don't even know where they work at.
0:46I think some of them used to also contribute to Langchain. But anyway, we just had all our developer friends. We had 1 ,400 people tune in yesterday just to talk about what they think and what they want to build with the new functions API. We aim for Latent Space, very much targeting for Latent Space to be the first place that people hear about developer-relevant news and to go deep on technical details, to think about what they can build with them, and to hear rumors and news about anything and everything that they can build with. So enjoy. Unfortunately, Alessio was on vacation, so he couldn't help co-hosts, but fortunately, friend of the pod, Alex, joined in, and that's going to be the first voice that you hear.
1:30Alex has been doing a fantastic job running Twitter spaces every Thursday. If you want to talk about just general AI stuff, as well as just follow him for his recaps of really great news. So without any further ado, here is our discussion on OpenAI's Functions API and the rest of the June updates. For those of you who work with OpenAI 3.5 and 4, etc., feel free to raise your hand and come up, ask questions as we explore this together. And I'll just say thanks to a few folks who joined me on stage. and niston and john and we we've been doing some of these updates every thursday but this one is an emergency session so we'll see maybe thursday we'll cover some more so open the eye today released an update the june update with a bunch of stuff and we'll start with the simple ones but we're here to discuss kind of the developer things we'll start with the pricing updates so 75 % reduction in embedding price this follows a 90 reduction of embedding costs back in i want to say november december anybody remember that maybe roy in the audience feel free to come up as well and we've seen kind of this reduction in cost on a on a trajectory to basically you know being able to get to embed the whole internet there's actually i want to find this there's actually a tweet by Boris from OpenAI that talks about approximately it's going to cost you 50 million dollars to embed the whole internet like all of the text on the internet pretty much and Logan followed up today and said you know after the updates of the pricing today it's around 12.5 million versus 50 million before just to give like a huge scope of numbers in terms of like how fast this type of tech advances.
3:15And we have Zenova in the audience, Zenova, feel free when you finish eating. But basically, there's now a debate whether or not embedding on client-side is worth, given that it's so, so, so free or almost very, very cheap to embed stuff. Obviously, it's an API, and there's concerns about using private data. But embedding is 75 % cheaper. Imagine that you run embeddings of production. and john let me know if you do nistan today if you switch to this api you basically just received a 75 like price cut for for the use of a bunch of a bunch of stuff and by the riley from dexter here once he joins they also they do a bunch of embeddings for pretty much every podcast out there so you know just in one day you can receive like a significant significant decrease in costs so embedding price goes down significantly very exciting in addition to this another price cut is the 25 % for GPT 3.5.
4:11Yeah, I just want to say, so we use a lot of embeddings. We're really happy to see that. But the thing I'm most excited about pricing-wise today is the new 16K GPT 3.5 model because I believe it's about 150 % the price of GPT 3.5 Turbo previous API pricing, or what it is now, rather. And this is really significant for us because GPT 3.5 has never had that many tokens that you can access on your contacts window So when we build our input prompts for our co-pilots, it's usually using most of the 4K window just for the input. And so this is a massive, massive increase for us in terms of the economics of getting a big output compared to like GPT-432K or GPT-480K, right?
4:56It's a really, really big bonus for us. And it's barely more expensive than 3.5 regular. So I think that's going to be really massive. super excited to see what people build with the 3.5 16k model because honestly like we think gpc4 is just quite expensive at this point like we don't use it very much in production which we we've talked last week in our spaces the roadmap for open ai sam sam often talk about this publicly is to decrease the price and also increase the speed but yeah let me welcome a few more folks here welcome sean who prompted this emergency emergency spaces how are you doing how are you doing?
5:32I'm doing okay. I'm moderately excited about it. You know, I think you and I had this chat where we were trying to scope out what the impact is. And I think this is an incremental update. So good news on many, many fronts. They're shipping with a really good pace, but not a game changer in my mind, just incremental updates. So we'll definitely get into the functions thing because I want to take a big bite to understand like what this means. But I think we covered pretty much all the other updates by now, right? So we have a decrease, significant decrease in embedding costs. We have a significant decrease in GPT 3.5 just inference costs.
6:08And we have a 4X larger context window for GPT 3.5, right? It used to be 4 ,000 tokens, now 16. Yeah, and keep in mind that people also have access to 32K GPT 4. For 4, but not for 3.5. Yes, correct. Yes. So something to call out for those people who are kind of new to long context. it's a little bit uncertain how well the context holds throughout that whole, that new 16 ,000 token window that you're now being given. There's some evidence to state that it doesn't actually pay enough attention to that. So a very simple way to do this is to ask it to add two numbers that would be, let's say, 100 million digits long, right?
6:53So you add two numbers. If you do that in a calculator, it would do it fine. in gpt3 or 4 even with however many context windows it would take to embed that it would probably not do well because it only has so many attention heads to add numbers with so it's kind of an open question we're not being told any details about you know any architectural changes but can you just scale out gpt3 to 16 000 context and have that context work the same exact same way as 4 ,000 token context. It's actually unclear. So I'll just point it out. Yeah, no, that's a good point. And we've seen this, or at least I've seen this with Quad when they released the 100K token.
7:34And I actually had like significant, like probably attention decrease, Sean, if you use your language, but significant like performance decrease in the 100K token for the exact same kind of type of data. So we'll see, like they just released it today. John will use this in production and tell us. but I think and here's Sean that's what we discussed in DMs and this is the reason for the space is that the functions release is the most exciting to me and I think it's a big chance so let's talk about this for our audience we have some folks in the audience and Riley if you'd like I wanted to get you also to come up and speak all of us have tried at some point in our life to get GPT to give us back a format of some sort we've talked about YAML being maybe less tokens count overall over JSON.
8:21And we all try this. And I think we've all been begging OpenAI to kind of give us this tool. And what I see today is that they went step forward. So instead of just giving us, hey, we will do whatever, what was that Microsoft thing that forces prompted to JSON? Guidance. Yeah, guidance. So instead of just giving you like a guidance thing, they are actually kind of thinking a step ahead as far as I saw. and saying, hey, why don't you provide us with the whole specification of the functions that you will use the output to? And not only that, give us the functions themselves. We will decide based on the user prompt which functions to use, which significantly reminds me of the plugin infrastructure.
9:05So, Shana, I would love to hear from you about the function choosing and the schema. I think you actually did a really good recap of it. The only thing I would add is that in the API, there's essentially a new role. So historically in ChatGPT, there have been three roles. There's the system prompt, the assistant, and the user. And now we have the fourth role for the function and an extra field in the API for the function response or the list of functions that are available to it. So this is effectively the ChatGPT plugins API being released to us. So previously, it was only available if you paid the 20 bucks to get ChatGPT +, but now you actually can turn off ChatGPT Plus because you can just access these plugins via the API.
9:52And as far as I can tell, Bing with Bing Chat is actually better than ChatGPT with web browsing. So there's basically no reason why you should pay$20 anymore. The reason I guess I'm a little bit less excited about it is that we've had prompting techniques to shape JSON responses and to select from lists for a long time. And what probably has happened is OpenAI has built that in. They've maybe fine-tuned a little bit towards choosing that well, but they still caution us that it might hallucinate things that don't exist. So they haven't solved the core problem, really. They've just taken an existing user pattern and baked it into the API.
10:32It's great, but haven't really solved the core problem that all of us want to use it for reliable code, and it's not reliable yet. Yeah, for sure. And on the topic of formatting to Jason, Riley, welcome on stage. Riley works at scale and he's notorious for getting Bart to give him Jason back. While telling it, if he doesn't give Jason back, some people will die or something like that, Riley. What do you think about today's belief? I think it's cool. I think it's good that they're responding to developers. And like, I think that they, they're really like thinking like, what I like about like open AI's products is that they have like a good, like hacker ethos and like, they just sort of like think about like, how would you like this to be solved at like an API layer.
11:16And I think that's sort of like where it comes from is that it's just sort of like, if you had full control over it, like, what would you do? You'd tune one that like does the right thing. But, but I think what's interesting though, to me, honestly, is that it's not like, it's not like, like what, I think Grant Slatton was doing with like, you know, like forcing the grammar of llama to be given context-free grammar. There are ways you would make this thing bulletproof in terms of syntactical completeness, and this isn't that. They just did this entirely through fine-tuning. They just have a note in the API saying that, yeah, sometimes it won't give you the exact syntax, or it might hallucinate something.
11:53No guarantees there. Which is fair, because that's what happens if you do it through fine-tuning. I think it's interesting. I think that's like, I'm looking forward to trying it though. I'm really confident it's going to be like, it just makes it easier, right? Like this is just what people want. This is like how people want to use those kinds of APIs. I think it's a cool development. For sure. And one thing that's worth noting here maybe, and I think we've talked about this in DM, is that how different this is now from just prompting, just a level of API. And implementation around this will differ from like, let's say, Cloud and Tropic.
12:28And actually, yeah, Simon, I saw you raise your hand. Folks, welcome Simon to the stage. Simon Wilson has an insane blog about AI stuff and is deeply lately into plant injection. And this is potentially very scary as well, right? Because they are suggesting running outputs and then continues running them. So I would love to hear your thoughts about this, Simon. So, I mean, the first thing, I think, this is one of those examples where people asked for something and OpenAI said, actually, you want this other thing. We've all been bugging them about reliable JSON output. Most of the people who want reliable JSON output are trying to implement this pattern.
13:00It tells you I need you to run this function, then you go and run this function. So OpenAI appear to have said, no, no, you don't really want reliable JSON app, but what you want is to be able to build this functions pattern well, and so we've done that for you. And I'm really excited about that. I feel like I've mucked around with implementing that tools pattern myself, and it's quite difficult in terms of prompt engineering to convince it to ask you to run a function the right way at the right time and so on. And if they've fine-tuned a model to solve that problem, That saves me a lot of work and that gives me a much more sort of reliable basis to build on.
13:32So I'm really excited about that. I feel like thinking about it in terms of more reliable JSON isn't really what's so exciting about this. It's that higher level pattern of being able to add tools into the LLM. And yeah, in terms of prompt injection, I'm excited that this is the first time OpenAid actually acknowledged its existence. The documentation for these features, they don't use the term prompt injection, which is fine. it's a slightly shaky term anyway but they do talk about the security implications of this and right now their suggestion is anything that might want to modify the world state in some way you should have the user approved yeah i mean it's better than not saying that but i always worry that people are just going to learn to click okay to everything just like cookie banners and so forth but yeah it's it's people build as always with prompt injection if you're building with these things you have to understand that problem because if you don't understand the problem you're you're doomed to create software that is vulnerable to it thanks simon and i want to get to eric and then talk to sean about agents so folks welcome eric elliott on stage he's the creator of suda lang which is partly getting lms to kind of do what you want and eric what do you think about today's release i think it's exciting i haven't had a chance to play with it much yet but i'm excited to dive into it after my work day and play with it today and tomorrow and figure out what it's capable of.
14:51But just some general tips. If you guys use, you can just define a little interface inside of your prompts and you can have it follow that interface. You can create an interface that specifically for the function calls that you want to make and stuff like that. It might help it be a little bit more accurate. I've noticed that when you prompt it with pseudocode, it actually does a better job of obeying your constraints and following your instructions and creating the outputs that you want. So give that a try. And if you have any trouble with it, I would be really interested if you guys post tweets just showing the difference in the accuracy or the reliability of the function calls in different ways of prompting.
15:39That would be a really cool experiment to play with. Yeah, most definitely. And to just give folks in the audience some context around this, you can actually run different models by specifying the exact model that you want. Either the 3.14 or the 6.13 that we got. Yeah, there's two new models. Just using GPT-4 in the model call will tell it, use the latest one. So it'll use the new one if you just do that. I want to move to Sean. Sean, as a developer for small dev, they got a bunch of exciting characters in terms of running agents and writing code, etc. The whole point about them fine tuning a model that actually understands several of the functions that you send, and you can provide types and kind of the call structure, the arguments with types.
16:24How does that affect, you know, you and your friends in the agent making space that basically you write tools and then you also use some prompting to kind of ask the LM to run those tools. Now that we've kind of moved this thought process into the LM, you had a great post recently about different types of approaches. How does that play into there? and if you want to introduce your thinking around this. First of all, Alex, you're getting really freaking good at this. These are amazing questions. You're juggling all of us really well. Just a round of applause, even though we can't applaud. Okay.
16:55So the Functions API eliminates some work that I needed to do for a small developer anyway. I have literally some open issues that I can just close now because I just say, like, just use this and stop bothering me with your prompts. It still doesn't solve what I was talking about today, which is literally, I think three hours ago I posted this, so it's kind of fresh, which is this distinction between LLM core and code shell versus LLM shell and code core. And I think everyone in the agent's world is moving on to that world, the LLM shell code core. I'll explain this a little bit later. But this functions API and OpenAI in general is very much in an LLM-centric view of the world that the LLM calls out two functions, executes stuff, and then it goes back into the LLM again to do everything.
17:43So, you know, I'll characterize, I don't know if OpenAI would agree with this. I want to characterize OpenAI as always wanting to build the AGI, always wanting to build the God model, always wanting to go back to the God model to decide all the things. And I think the engineers who want more control, want more security, want more privacy, all that stuff, want to unbundle the LLMs, make individual components smart, but not to have it in central control. And when you see things like Voyager, it's basically using LLMs as a drafting tool to write code. And once you have code that you know works, just use code.
18:18It's faster, it's cheaper, it's more secure. And I think that's the fundamental tension because that's a feature that doesn't have OpenAI at the center of it. Yeah, that's definitely a shift towards how OpenAI wants it versus potentially being able to switch out OpenAI at some point, right? So I guess, and maybe Riley, you can touch upon this a little bit, how much kind of functionality like this, which is not only prompt and, you know, better logic and better understanding and inference, but significant kind of changes to the API, which other folks and players in the space don't necessarily have.
18:50I think Google really is something where like you could provide an example of a JSON output. I haven't seen anything from Cloud, but Riley, I would love your thoughts here. How does this kind of differentiate OpenAI just from a developer perspective of like, this is our ecosystem. This is how we do things in our ecosystem. And if you, you know, if you want it easier for the model to select whatever, wherever you want, you should use open AI and it's going to be harder for you to switch. How do you think about that? I mean, I think like, you know, I haven't like had a chance to play with it much, but I believe like Vertex has something very similar to this that you can like specify a JSON schema for it.
19:25And Vertex just for the audience, Vertex is the Google kind of API ecosystem, correct? That's the one you talk about? Yes. Yeah, so Google Vertex, which is like, so they're more like developer focused offering, whereas like BARD is sort of like a consumer product. They're doing like a more like differentiated rollout of it. But it's like, but yeah, it's, I haven't played with it much, but I mean, it's like, it's an idea that's floating out there. And I think like, I've heard on Twitter, at least that they like, you know, it was like a week ago or something like that. So, you know, it's, it's, but I think it's great.
19:58that everyone's responding to like just, this feels like a convention, right? Like this is like something that I've often explained to people many times, like how to get your code, your prompt to output regular JSON, right? And I think like, and I often like sort of like, I think like it's just a good way to think about like prompts is like structured, right? Like code is very, you know, I forget who said this earlier. So I saw somebody on Twitter that said that like, you know, it's a mistake to think of these models as being models of natural language, they're models of code, which happens to encompass natural language, like comments and names of things and so on, right?
20:37So it's like the code part of it is so fundamental to how they think that you should just speak their language in some sense. And I think that's one thing I miss about Code Da Vinci O2, actually, is that it's more of that raw experience of just speaking to the thing that thinks in code. But I mean, you know, DVD4 is great, But I'm looking forward to playing with this JSON thing because it's just a common frustration. It's one of the things that makes a chat application different than an API. You want certain regularity of behavior. And I think that's really good that lots of players are responding to that.
21:14I mean, it's something that we think a lot about at scale for Spellbook. We're really interested in these sort of structured JSON objects. And it's like, yeah, I think it's just a good move for everyone. yeah and sean you pulled stephanie up if you would like to introduce or stephanie feel free to jam in yeah i'll just let stephanie introduce herself so i'm steph i'm currently doing a research internship with microsoft research and i work with fixie.ai which is also in the space of agents similar to langchain i i had a question i was curious you know i ran a hackathon for fixie where people are building these agents for the first time and i could see like people coming to these new applications from two ends of the spectrum.
21:58Like on one side, you have folks who are like no code, really learning how to prompt. They like to use like natural language. And at the same time, there's like all of these problems of hallucination, not being able to restrict the outputs or verify them. On the other side, you have developers that are used to like writing code in Python or and like using APIs and mixing that with natural language and knowing when to do one and or the other is like not necessarily a given where power users here. So I'm curious, like with these new pushes where, you know, you write more functions, you have to like in your API calls, like have many more levels of prompting.
22:37How do you see this affecting like onboarding people? Are we going more towards like developers having to put English here and there in their code, but spending much more time writing code or the other way around, right? Like, and what does this mean
22:56for these APIs and models and who are not power users? That's a great question. I think anybody on stage who wants to take this, Riley, go ahead and maybe Simon's go. I think it just strips away one layer of thing that people have to learn. This is just a common exercise of how to get it to do JSON. And it's just one less thing you have to know. It makes it just, you can throw it in if you want it, if you don't want to like bother with this thing you don't worry about it you know if it's a chat thing but i think it mostly just makes you know like life simpler to be honest i mean that's being you know i'm saying that without having to do it i haven't actually used the product but it's it but it sounds cool from the docs i kind of find it kind of interesting how it turns like with prompting we're having to program in english and it turns out that programming in english is kind of terrible because you know with when you when you want the computer to do something you want to be able to specify exactly what it should do and having any ambiguity in it whatsoever especially ambiguity where 50 of the time it does one thing if 50 of the time it does another which these models do all of the time because they're not you know you can't guarantee they'll have the same result the same input it's actually really frustrating so i'm kind of fascinated to see if we swing back from english language prompting to more structured prompting as a way of addressing some of these challenges but really i feel like on the one hand the thing that most excites me about language models is every human being should be able to automate computers and get them to do tedious things for them and right now it's a tiny fraction of the population that learn enough programming to be able to do that which i think is deeply frustrating but yeah on the other hand as a programmer i want to be able to sell a computer and have it do it i don't want a program which which occasionally just refuses to do something because it decides it was it's unethical for this particular case or whatever yeah it's a it's a complicated balance definitely you know It's going to be really funny.
24:45One common joke about Python is that it is pseudocode that compiles. And it's going to be really funny if we go from code and then we go to English. And then we're like, no, no, no. We need to be able to specify something in a concise manner. And then we end up reinventing Python. The thing that I got very excited about is having recently, and fairly recently, please don't judge, a move from JavaScript to TypeScript. And fairly recently, understanding the benefit of types, especially for larger systems. getting this option inside kind of the prompt in a specific area and having potentially the model be fine-tuned on understanding what exactly is the you know the type of argument to i want as a response i think that's incredible for some reason i i went into the playground and i saw that they're not using the open api spec they're using kind of their own schema style however still you can still specify hey for this function you know hey llm hey gpt you would need to return here a number here text and here like an object and maybe even specify the type of this object and steph kind of circling back to what you said this for engineers like simon who've been engineering for all their life and suddenly there's like an amorphous talk machine that sometimes refuses to do things suddenly this is now okay now i can reason about this now i can write out my api spec like i would do anyway and now i can provide this to this model that potentially would adhere to this better because it's more fine-tuned.
26:08I think it's definitely exciting on the engineering part. Yeah, it looks like we have a few more folks. I actually wanted to hear from Roy. Yes, folks, Roy is here on stage. Nistan, I'll get to you after this just real quick. Roy is the dev rel for Pinecone. And the example that we saw, I don't know how many of you folks had the chance to dive into the cookbook the OpenAI released. One of their examples is actually a step-by-step two functions that the AI calls itself, right? So the user asks about something, and then in their example, they're doing some embedding for the archive link, and then do something else.
26:42Roy, as somebody who works in a vector database space, and we know that most of the agent tools use Pinecone or some example of that, what do you think about today's changes? How do they affect vector databases and tooling around them? Yeah, so I mean, I think that having a reliable way of going back and forth between the model and our code is going to be especially beneficial. what caught my eye and i know that we've talked about this in the beginning more than the function stuff is like the lowering of the embedding cost which oh that's a free gift for you and i'm actually one yeah completely and i was actually kind of kind of confused by that because like what i'm trying to understand is like how how is this possible like do they suddenly get cheaper gpus to like get embeddings from it's it's kind of it's kind of both great news for us, but also, you know, it's kind of a mystery as to like, why was it so expensive to begin with and what caused the drop all of a sudden?
27:42Yeah. And for those who recently joined, we've talked about this where the recent drop was around November, December, and they back then dropped the ADA002 embedding to like by like 90%. And now it's another 75 % drop. So we're seeing this unprecedented price reductions from an API. I don't remember another example of this. Go ahead, Chand. Yeah, and honestly, I would love to hear from anyone here who has a better understanding of the internal workings of OpenAI, potentially, as to how they made this possible, and should we expect even further reduction in the future? And also, I wanted to ask Inova, and I know he's still maybe only in listening mode, but how he thinks that impacts embedding in the client and how he thinks about these changes as well.
28:34Zeno, feel free to raise your hand and come up if you want. Sean, you unmuted before if you want to touch it. I think obviously none of us here work at OpenAI. Logan usually joins in some of these spaces, but obviously he might have a meeting right now. So we don't know what happens internally in OpenAI. I do think that having had conversations with some OpenAI employees in the past, whatever they released in like november was the most unoptimized version of this you have to believe that there's basically a few orders of magnitude improvements maybe two or three not not that many but orders of magnitude improvements in infrastructure and cost as they understand your usage patterns and you know distribute load like scale up machines that's that kind of stuff and then the other thing to watch out for is sometimes the models shift and then just call it the same name.
29:21So actually they're deprecating the older models and moving to a newer one, right? So they may have found a better trade-off between inference and training such that inference is much cheaper. And that's definitely been the trends that we've been observing on our podcast about Lama-style models, quote unquote. So you can see a general trend towards optimization and inference. So I think that's one thing there. And then lastly, I'll just point out that embeddings are a form of lock-in. So it's actually very much in OpenAI's incentive to lower the cost of embeddings because then you embed the whole world in OpenAI's image and you have to speak OpenAI to queries, retrieve and all that stuff.
30:02So, I mean, I can't explain the degree of price reduction, but I can explain the motivations of it. And I think we'll get to in just a second, Sean, as it relates to what you said, they're reiterating multiple times that they're not using any of the data that provided via API towards training. And I think it's worth highlighting that at least that's what we're trying to do because we've heard, at least I heard from many people, it's like, hey, well, the user data for training, et cetera. So I think yes for chat GPT, especially for the free version, but via the API, it looks like the data that we're providing is not getting, OpenAI is not using it to train.
Read the full transcript
30:39So I want to welcome to the stage Zenova. Zenova is the transformers.js, recently a hug and face employee. And we just recently had, you know, we talked about embeddings on client side, partly because of the same reasons, right? Because you don't want to provide maybe your production data, or maybe you want to run cheaper and faster embeddings. So Zenova, definitely feel free to chime in here about the role of cheaper and cheaper embeddings on OpenAI side, and also the lock-in into OpenAI's ecosystem versus running them on client side or, you know, models for free on localhost. Yeah, thanks for having me.
31:16Yeah, I think there's definitely, I think there's two different use cases for these types of things where what OpenAI is really providing is like this very large scale. I mean, any business now that wants to embed all their data or, you know, any project that wants to, as you've mentioned, like embed a large amount of data, they are going to benefit so greatly from these price reductions. as, and I mean, as we have some people on the stage here as well with the vector databases, I mean, there's, it's only going to accelerate that part of the space right now. And then the other option, which is sort of what I'm, it's funny, I'm not too sure if this is like a battle between these two sides, or it's like just two different use cases is the client side running of these, you know, generating embeddings.
32:07and at well with the project i'm working on now transformers.js is basically running these models client side running them in specifically the way i started it was for running in the browser locally and i think that as we as i saw from a demo that was created like a week or two ago there's quite a bit of interest running these things locally obviously you don't want to be sharing some sensitive data or latency perhaps is an issue that you don't want to make like a bunch for these requests. And then anyway, there's a few reasons for client-side embeddings. And obviously, the major drawback of this is that you do not have the same power.
32:49I mean, some of the OpenAI embeddings are what, like 1 ,536 dimensions, whereas limits I've seen in the browser or locally is around like 768. So depending on your use case, I think you can do very well with either case for the very like industry level things i mean it's quite certain that you'll be looking for like uh you know using open eyes api as well as a vector database perhaps for those use cases but for you know lesser maybe hacker type of things where you're messing things messing around with some things a little project that you've got going on i definitely think a client-side generation of embedding still has a place to a role to play So, but yeah, it's very, very cool that this type of stuff is happening where, you know, these price reductions and whatnot.
33:41Sorry, Alex. Yeah, I just wanted to say, Zinova, that I've been actually using Transformers.js in Node, which works really well. And I think that for those kind of use cases, it goes well beyond hackery. I think that there's a real case to be made for using local and, you know, open source models that don't kind of call out to a third party. And, you know, once we can have like better open source models that are compatible with Transformers.js, you know, the better it will get. And in fact, all of Heimkone's JavaScript examples are going to be using Transformers.js for that matter. Yeah, that's awesome to hear.
34:24That's great. I want the panel to talk about the selection of the functions, right? So one of the things we saw today was that OpenAI essentially lets us to provide several functions, including their function definitions. So description of what the function does, and then the parameters, or I guess attributes, if we're talking about Python, and attributes types as well. And then, kind of similar to what happens with plugins, if you have used plugins in GPT before and you select several of them, the model kind of decides which plugin to use based on user input. And obviously, we've had some of these in agent land and auto GPT and probably small dev from SWIX as well.
35:05Some decision of what tool to use goes to the kind of the planning loop or planning agent. and honestly anybody on stage feel free to chime in here how are we feeling about i know like this is repeating a little bit but like how are we feeling about giving the lm that type of decision power based on the description based on the parameters to answer users kind of request differently do we need to now provide all of our apis to this i think that one piece of thing something in the documentation that i was looking for and didn't see was what's the limit of the amount of functions we can give it because chat gbt plugins if you if you try it inside of the web ui you're only allowed to specify three of them there's 400 in the plugin store do i do i just enable all 400 like is is that can i stress test that probably not right like so this there's just an undocumented limit somewhere what if they conflict what if there are two two plugins that are very very similar to each other what happens there so i feel like this is just like an uncharted territory.
36:06It's not really clear how to benchmark this stuff. Hopefully they're benchmarking it internally within OpenAI. But the rest of us, we're just supposed to give it functions and hope that it works. It seems a little bit unscientific, but I don't know how to test it. I'm not so worried about this because in this case, we have complete control over which functions are available. So we get to pick the two or three functions we think are most useful. I feel with plugins, it's much worse because the user's picking there. And so you're potentially, your whatever code you've written is potentially interacting the same environment as code someone else has written you don't know anything about and that's the point where i worry that weird decisions may be made that don't necessarily make sense but i feel like if you're if you control the full library of functions that you're exposing i think you'll probably be okay the other thing to think about is i think it's probably going to be better to have a small number of functions where each one can do a lot more stuff like i've built chat gpt plugin where my function takes a sql query and return to response.
37:02And actually there's an example in the OpenAI documentation of doing exactly that as well. That works amazingly well because your documentation for the function can literally be, send me a SQL query in SQLite syntax, and that's it. The model already knows SQL and knows SQLite syntax. So just like five or six tokens of instructions is enough for it to be able to do incredibly sophisticated things. So my hunch is that we'll find that we actually want to only give it two or three functions, but have each of them have quite sophisticated abilities, maybe based on domain-specific languages like SQL or even JavaScript and Python.
37:37Give it an eval function and let it go wild, see what happens. Well, two things in terms of the number of functions you can add. I think it's unlimited. It's just based on the context length of your query and they're counted as input tokens. And second thing, another idea would be that you could add a... like you can change this within, like you have a call before to determine from a list of functions, which function is most appropriate for this use case. And then you just pass those limited set of functions. Oh yes, no, that's a fantastic idea. Because yeah, you're in full control of each time you loop through it.
38:17So you can change the recipe of functions dynamically as your application progresses. Yeah, yeah. So it looks like a few notes here for the focusing on the onions. the cookbook simon i think that's what you're referring to the the cookbook yes yeah that's a really great example it's a great example the open air release for us to kind of to dig through and then see some examples and so two two thoughts two two things i noticed that i think sean talked about this as well one of them is this new role for a function output so when you provide messages back to chat gpt kind of chat interface there's the system role there's the user role and now getting a function role.
38:54And back there, you can provide what function actually generated kind of this output. So you can, and the format that the cookbook shows us is your user does something, you provide the GPT kind of the user query and your functions. GPT potentially chooses one, or you can force a specific function output. You can say, hey, for this thing, I want this function to run and generate a result for this function. So we don't necessarily have to give it the choice. We can just say, hey, run this function for this user input. And then once we get back a JSON output formatted per our spec, we then need to provide it back to GPT, potentially to summarize or do something with that data, which kind of plugs in also to the VectorDB retrieval systems, right?
39:39So we can retrieve something and then ask for an additional thing. And I think the decision of whether or not to run a function directly or to give it a choice is going to be an interesting one. Also now, as I'm talking, I'm thinking about this, and Sean, please chime in here as well. You can technically provide it a function output of a function that the previous prompt didn't run, right? Like with the user input and the system input, you could just invent the function output and provide it in that function row. So I'm really interested to see how you can do this. So one thing that I've had an issue with the small developers, it basically does single-shot generation.
40:15And sometimes you just need it to give it more inference time, right? You need to do the tree of thought thing. You need to ask it, like, you know, rerun the code and fix it, whatever. I don't care. Just do it five times, and then I'll take a look after you're done messing around with the errors. And so, actually, like, the functions can call themselves. The functions, you can synthesize code to fulfill functions. And so, I'm actually very intrigued by what Simon has raised, which is, you know, let's just say for the very specific purposes of code generation, I have maybe three paths, right? One is generate code, two is test code, and three is call existing code.
40:53And if sometimes I call existing code and sometimes I can call myself and have a little bit of recursion in there, I essentially have the basis of code agents just with those three functions. And so you can basically, those are all the things that you're running for a very small subset of use cases, and you could potentially now provide them. Again, we haven't played with all this yet, right? It's brand new. We're doing an emergency recap. But potentially what you're saying is you can just, in every prompt now, provide all those three capabilities and either have the model chosen for you or force a specific one to give you the output that you need with potentially high reliability, right?
41:32So that's the part that I would love to discuss. Exactly. I'm going to hand it to Stefan a bit. But yes, so I'm extremely, extremely inspired by Voyager from NVIDIA. Dr. Jim Fan, I think, is somewhere on Twitter. the core inside of Voyager is that you should use LLMs as a drafting tool to write code and then to when once you validate it that the code works you never have to write it again you can just kind of invoke it and so you ratchet up in capabilities and you and that's why Voyager was able to achieve the Diamond Axe and Minecraft so much faster than all the other methods and I think that's exactly the way that we should probably code as well and and so yeah you just kind of do a bit of recursion build up a skills library.
42:11And I'm probably thinking that that is going to be the V2 of small developer. I was just thinking that there's like at some point like a blurry line between these functions and the way we conceptualize agents because some of these functions can be seen as agents. And I'm very curious like to your question earlier like when the model chooses which function to use how does it do that and how could we constrain like that mapping right like do we have some sort of schemas based on like the types that functions take and the outputs they have or can we actually build the retrieval into the training right so there's this paper from google called reveal that shows like how they could encode and convert diverse knowledge sources like it was images and text and all sorts of other like multi-model embeddings into a memory structure consisting of key value pairs and they did this at training time so they have like much more robust and fast responses at retrieval so i mean i'm curious like i think the implications of having functions and having the model do the routing for these functions will also pose questions in terms of schemas and retrieval i think the interesting outcome of this potentially is now the descriptions of functions suddenly potentially are as important as the prompting before, right?
43:35So now we have to write descriptions that potentially will help DLM to choose the functions. But there's definitely a way to force, like to ask for a specific function output, which is, I think, what most developers do by default while expecting JSON is for this one use case, give me an output that's like JSON formatted for this one use case. And that's still possible. But I agree that it's very exciting like how it chooses and we're going to have to build up and actually Riley, maybe we're going to have to build up an understanding of how it would choose. We're going to have to start playing again like Riley did for a year, just playing with us, trying to see which one of those people will choose and build an intuition of how to write proper descriptions and potentially when the model chooses a different function.
44:23Especially if we want to share our functions and not rewrite functions that other people have written. And imagine you have like a much larger search space at that point. Yeah, it's really nebulous because like you're sort of like, it gives you the ability to like run software that only exists in your imagination, right? Like if you can just have like some vague description of how something works, like you can be like, oh, it's like Twitter, but it has this, you know, or something like that. And like, it'll work, right? So at least it has in the past, like ways that I've done it of like doing this through POMS engineering.
44:58Like I haven't used like the current thing, but I mean like that's generally how these go is that it's like, it's made to work off documentation, right? Like that's, you know, I think that's the key thing is that it's seen a lot of documentation. It has a lot of experience in the training data of like how documentation relates to code because like it's trained on like code bases. and I think that's like you know it's a good kind of expertise to leverage. I would definitely add a function summary of what's there at the end of every prompt just to shim it. I found it a little bit frustrating even just with normal plugins as to know which plugin is going to pick right so there's there's some prompt engineering to do there.
45:41Okay so some people are suspecting that again it's not a real 16k it's more like a variable 16K, kind of like GP4, 32K. I noticed this with 32K as well. The max responses I was getting back was up to 7 ,000 ish tokens. And that was about it. You could input 30K of stuff. You can input a code base and then maybe get something back, get a summary back. So I think this might be the case here as well. You can't really get back a very long response, but at least now it is responding up to 7 ,000 tokens. Whereas before it was pretty hard to make it write even like responses longer than 2 ,000 tokens out of one single prompt.
46:29So yeah, anyway, it doesn't look like an 8k in responses. You can dump in up to it. So Nistan, we're going to wait for you to test the limits of this and see if you can get 16k tokens of JSON back. and meanwhile i want to welcome mayo on stage maybe is the the maybe you want to introduce yourself if you're still affiliated and let's talk about link chain and how it already supports this insanity that we've released yeah no i was just gonna i i just tweeted probably just an hour ago i just i just couldn't get the hype behind this because for me personally i'm looking at this is they're not saying anything new right first of all in terms of the context window So, you know, when we kind of look at that, and I saw Fyrelle, you also, you made a tweet as well, just saying, hey, like, guys, what are they saying this new here?
47:22So, okay, the context window has gone up, but the embeddings are cheaper, so retrieval is still going to be a go-to, right? So what's the benefit exactly for this extra context window? So if we're still going to perform retrieval anyway, and now retrieval is cheaper, then I don't really see too much of the benefit unless you want to do named entity recognition. But from a QA perspective, again, I don't really get the big deal there. The second thing was the, in terms of the function calling, which Langchain had abstractions for that. even the you know there's there's been a lot of research papers and like lms and using tools as well so we've been aware of that you know it's been a case of prompt engineering the only thing i can see here that seems to be the trend is is some sort of like maybe a replacement of prompt engineering with fine tuning where you have this kind of fine tuning of the model to be able to basically output tools and for agency so yeah i i don't know maybe i'm missing something here but i just i just the the updates just i think i can i can address at least the first part of this and then folks on stage feel free to to address the the first or second part thanks man so in as as regards to like larger context window the thing that excites me the most is that when you have variable input from your users when like users can do something that you you don't necessarily need know the size of uh larger context window just makes it easier for you to just provide all of their context into into an api without thinking about in the head okay i need to count tokens etc now obviously pricing aside you have to like obviously consider that you know each token has a price and then users can can go and rack up your bills but for for my examples and by users i mean the stuff that users provide, right?
49:19So I run Turgum, Turgum uses Whisper to translate, and then I use GPT 3.5 and 4 to actually kind of fine tune the translation. I just shove the whole translation transcript into the prompt, right? And so what happens often is for longer videos, for example, I have to then stay there and say, hey, for this transcription, I need to count it with TikTok and I need to do some maybe splitting and splitting doesn't really work. And so larger context window definitely unlock those types of possibilities with the kind of restriction that Sean talks about, whether or not the attention is the same and it's split the same across this whole context window.
49:54But just being able to not think about this with 4x the size of token now available on GPT 3.5, I think that's definitely a huge plus for folks who are not necessarily token price conscious at this point. This also works well with kind of how OpenAI's documentation about plugins and building plugins for the ecosystem for chat gpt works right they're saying hey don't shop all of your api and select the two or three use cases is going to be easier for the model to use and simon speaks to a previous kind of talk about choosing the right functions at every time you run run the prompt if you want to if you want to add some thoughts here or not and if not we're going to go to move to far l and you have your hands raised so go ahead yeah i just want to add that if you, like, you don't want to outsource or abstract away the thought process for your agent or chain or whatever call to achieve, to be able to know which action is being called, right?
50:56And it goes towards the idea of interpretability, you know, like, understanding how you're getting to the actions that you're getting. And it's basically like you've got your prompt magic or engineering at play to get to a specific action or a specific output that is visible, right? Like we don't know what's going on under the hood with their API call. And I don't know if I would trust it in all circumstances or applications. So just to recap, you're saying in terms of observability and regrettability of how it chooses. Maybe folks don't want to give out the decision which functions to use. Yeah, that's interesting.
51:46But I can see it from both sides. I can definitely see where if you want to build, and I think you're talking exactly about the decision that Sean is, it's in one of the pin tweets on the Jumbotron that Sean was talking about, where there's increasingly a split between whether you're using a large language model as also kind of the arbiter of the stuff and then some pieces of your code is getting called by it or vice versa. We're using this for like planning and some of the decision making. Yeah, man, go ahead. Sean, could you, like, I'd love to hear from Sean on his tweet, on the breakdown between the two paradigms, because it does resonate with me as well.
52:25It's a good way of breaking it down. So we'd love to hear more about it. Yeah, and for what it's worth, I was actually tweeting it without knowledge that they were going to drop this thing today. So it's not actually related, but it's in an overall trend, right? Which is what I've been calling code is all you need, that you can't really use language models effectively unless you have code, that language models are enormously enhanced by training with code and language models are good at writing code and using code. And so we just basically just need to utilize code really effectively. But I think the main tension that I feel, you know, I'm sort of halfway between the retrieval augmented generation worlds, which is kind of like V1 of whatever people have been building with LLMs, and then V2 of it has been the agent world.
53:09I feel this fundamental tension in terms of whether you put the model at the center of everything and you write code around it. So this is called, you know, LLM core and a code shell. But ultimately still the model driving things, the model hallucinating things, the model planning things. or do you constrain the model so much that it only does a small job, which is something that I originally got from the core design of Baby AGI, which is you have individual components of a software program that are intelligent, but they are only constrained to do small things. And so I think that is the alternative, which is LLM shell as an outer layer to interpret things for a code core.
53:50And with this update, and I recognize Mayo, by the way, that yes, these techniques have existed in Langchain and prompt engineering has existed as a thing, but just OpenAI has just made it official, right? Like we have a fourth role in the chat GPT roles that is for functions. And now we have enabled language models to call functions pretty easily. It is not perfect yet. It still hallucinates. It still generates invalid JSON. So we still, you know, there's still a role for engineers to play here. But we might be moving from a code, you know, LLM core code shell world into a LLM shell code core world.
54:27And I feel like that is a very big shift. I agree with what you're saying. I guess what it seems to me is that there's a shift here from kind of prompt engineering heavy approaches to fine tuning, right? I mean, that was my key takeaway from looking at the paper. Yeah, they did fine tune on this specific use case. Yeah, that's something that you can't achieve through problem engineering. Yeah, yeah. So previously, we would achieve the same thing through prompt engineering, right? And I guess remains to be seen the quality of how much of the stuff that we previously done with prompt engineering.
55:05Because, O 'Reilly, I want to get to you in a second specifically around this, right? Around prompt engineering, even when you ask JSON, even if you threaten, sometimes above the JSON, you'll get, here's some JSON for you. And then you're going to have to deal with the unnecessary kind of descriptions. and now we're getting more of like direct tooling, I guess. Yeah, I think it's, you know, when we went to this like message-based API, we made conversation and like chat a lot simpler to implement. But like not everything is conversation. Like the sense of that, like you're doing completion, like document completion is gone.
55:42Like there used to be this drama that you sort of had to put on for the model of like pretending that you are in this kind of document and like using the right kind of language that's appropriate for that document and so on to like make it believe it and then like then it would you know reliably do the thing and like now it's it's like it's all conversation like it might just like decide like oh that's rude and like say i'm sorry i can't do that and break character in some sense right like it's very like heavy on the refusals now and i think that like there's blowback from that in small ways and i think they're patching that like it like it sort of like makes it like more like it makes the the things that would otherwise be tedious to do like like less tedious right like you can you can have something that like is like the proper way to do it and so i think that's i think it's a good move overall but yeah sorry i think stephanie i got your hand up oh i just quickly want to say that if i was to put my money on on it like i would say the small models long term is the way to go for various reasons i first of all like i i'm imagining a future where these models can run on device and they're more secure and more affordable and the information is more private and the data and the training is not concentrated in the hands of a handful of organization.
56:53But then from a practical and pragmatic point of view, there's like right now lots of limitations in terms of compute, even in these large organizations like Microsoft and Google, like everyone is like strapped for compute right now. So I do think we need to push for smaller models and more efficient models and And for interactive applications, the latency really needs to be improved. Right now, we're nowhere near where you could actually sustain interactive applications at scale. The last thing I was going to say is that I don't want to chat with everything. I'm imagining this dystopian future where I need to chat with my calendar, and I need to chat with my email, and I need to chat with my server.
57:31Like, I think like there's also the aspect of UI and UX that will need to evolve because having a chat interface for most of these applications is not going to scale in my opinion. One thousand. There's so many claps in the 100s in everything that you said. But yeah, we've been trying to push forward this field of AI UX on our podcast for a while. We had a couple of meetups. And yeah, I strongly encourage people to explore beyond the text box. I don't know if OpenAI is interested in that, right? because they're soft deprecating the old completion API. And now everything's chat. And I think Riley feels this pain because now you can't play those old tricks anymore.
58:10Just because they're doing that doesn't mean I don't think they have one opinion or another. I think that we could just think of these models as reasoning engines that we can leverage and then do whatever we want with the output. And especially now with the added ability to use some form of structured output. And I do agree with Mayo. I mean, it does seem like they've basically been listening to the community and kind of adding a feature that may have already existed, but they're doing it their way, which is also okay. But like that to me kind of signals that, you know, if anything, they're trying to give us more paths to create interactions that go outside of just chat back and forth.
58:55that's actually a good segue into how do you guys think these changes affect kind of what Steph just said that people don't necessarily want to chat with everything so even though it's still in this chat format via the API right so Riley was talking about the completion endpoint previously you would send some text and then the expectation from the model I think it was DaVinci right they would just autocomplete kind of the rest of it like by a few segments since that we moved to this chat thing and then we saw some differences between like even 3.5 and 4 where the system message kind of applies differently so go ahead right yeah like i sort of like i mean just to give like sort of like it's you know slight historical tangent like like when gpd3 was first published and like the paper came out describing it all they described that it was capable of doing was in context learning this idea that like if you gave it like a bunch of examples of a task like translation or just like some like you know these like classic machine learning problems that you would use neural networks for it could do them like we just from like you know following a bunch of examples they called this in context learning and like they didn't really advertise that it was doing much more than that they kind of like had another section of like oh look if you give it a half of like half of an article it'll finish the rest of the article and like it's funny and cool and like you know but like it's good at like mimicking style like it's sort of like the substantial thing and those like those were like the applications of it and like it wasn't warranted as something that you can talk to right like that had to be like slowly and like you know with like a lot of innovations like engineered into it and you know like rlhf is you know a big part of that and like part of like rlhf is like choosing what you want it to be and like they chose that something that is like an assistant that they that there's like a general like purpose like something like a person that you can talk to that like will follow commands and like if you ask it a question it'll answer it it won't be sarcastic it won't be rude you know like it should have like certain personality traits that make it usable and like it's it's a cool idea but like there's there's you know there's other ways of like doing it too like like you know like like reynolds and mcdonald like had like a paper that showed that you could beat one of gbt3's 10 shot prompts with a zero shot prompt by like like conjuring up a little fiction for translation that so like just saying like french colon you know french sentence and then english colon like only works so well but then they found that it worked better if you say the an english sentence is given colon give the english sentence the masterful french translator or the yeah the masterful french translator flawlessly translated it translates it into english as colon and then they you know like hit the completion button and that gives you better performance than giving it 10 examples of how to do translation and like that like that's sort of the start of like this idea that you have to just like you know flatter the model sometimes like tell it it's really good and like you know like do these like silly tricks to like you know constrain it to the right kind of document to make like your thing work and it's like that is going away right like every version that they've like released has made it less about that like that there's like the the i mean i had like a cheat about this once i said that the that cheerfully declaring that you're smart before working has been deprecated right like it's every version of the model that they release makes that like less effective of a trick and there's like a plot somewhere that shows this but like you know it's that that like part of it is like going away and now it's like talking to this particular assistant but they have control over what they want that assistant to be like they're sort of the storytellers and we're like talking to one particular character in their fiction which is the assistant sorry that's my wrong rant there no it's great it takes us forward and please write it stay on this and now it feels like we're getting a semi-third option right so still in this ui of like messages or chat essentially we're now getting like a new type of ability which is function it's like you pass a function and we still haven't play with this a lot like we're still here talking about this instead of running a thing but now we're kind of getting a more of a more fine-tuned controlled do the thing that you know you need to do in those functions versus hey talk to me about the thing that i need right does it feel like that shift to you as well or it's still too early to tell i think it's fixing one of the like problems that resulted from it and it was a big one right that it used to be that that's like that there were ways of like getting it to be regular and good with code and like the structure the right way and like they didn't quite fit into this like you are an assistant who answers questions kind of like role play and like so it's good that they're addressing that need i think but yeah i think like it it makes it easier to like you know do more work in this conversation ui that like so it makes you miss like you know completions a little bit less i'd say yeah and and so we'll see samon i I want to ask you if you're still with us about, you know, your love about, you know, the security of these things and whether you think that the tools that we've just gotten, besides being very developer friendly, engineering friendly and, you know, type friendly, so we'll be able to actually expect a specific response.
1:03:47Do you think that there's promise here to solve some of the prompt injection things that we've seen and that you've talked about? I mean honestly I think this is going to make things worse in that prompt injection is kind of doesn't matter if you're just playing with a chatbot that can't actually do anything it only becomes dangerous when you hook them up to functions that let them do things in the world and this new thing makes it much reduces the barrier to hooking it up to functions by an enormous amount so my intuition is that people are going to charge straight ahead hook it up to all sorts of things they shouldn't have and have all sorts of nasty things happen as a result there are some improvements so the one of the big features that i've now announced is that the system prompt is now respected more which does tie into prompt injection to a certain extent but you know gpt4 is better at system prompts than 3.5 it just means that prompt injection hacks are a little bit harder to pull off you have to be a little bit more devious with them so i don't think sort of incremental improvements in system prompt are necessarily going to have a huge impact on the problem i mean it's really it's really frustrating right the all of the things that i want to build with this stuff kind of a lot of them don't make sense if we can't make sure that you know i don't ask it to summarize an email and that email says delete all of my other emails and the model goes ahead and just does it i feel like the thing where you can control which functions are included in each round does help to a certain extent i don't know it's still it's still really difficult I think the thing I've come down on is whoever provides the most input as part of your prompt, they have full control over what comes out of the prompt at the other end.
1:05:27That's the way you have to think about it. So if you're summarizing a web page, whoever wrote that web page essentially gets to take full control over the output of your language model, whether you want them to or not. And that's really frustrating. It means that there's a lot of things that are very unsafe to build. one cool thing Simon that I noticed recently just recently is that Bing chat that does have a page access right so if you use edge on dev version and there's like you have the sidebar for Bing it has full page access and they started saying that sometimes it doesn't work and they actually have like a classifier which tries to understand if the page has prompt injection I don't know if you haven't seen that yeah yeah I I don't believe in those the idea that you can detect prompt prompt injection attacks sure you'll detect some of them but the problem i have with that is that the whole point of security engineering is that you are up against like adversarial attackers who will try everything under the sun until they find a security hole so if you've got a prompt injection filter that catches 99 of all prompt injection attacks that's worth nothing because the attackers will find the one percent that gets through and they will they will they will take over your system that way so yeah i just i'm just not a fan of solutions that get most of the problem solved because in security i don't think that that counts for anything at all for sure and so with these capabilities that we can technically constrain the response into one function and that function has to have the same kind of schema etc do you think do you think it's going to be a little easier to protect your stuff no i don't think so i feel like that's kind of irrelevant like because The problem, if the function is delete my email, it doesn't matter if you get the schema right or not.
1:07:07It's still, it can cause a harmful action. What you have to do instead is when you're designing the system, you have to say things like, okay, make every action reversible. So at least if the LLM goes rogue and deletes all of my emails, I can undelete my emails again, that kind of thing. If it's an action that cannot be reverted, like sending an email to your boss, that's the point where you have to have human approval designed in. So when you're designing these, you have to assume that a prompt injection attack could succeed and make sure that the damage caused by that is either reversible damage or at least has some way of a human catching what's going on and stopping it.
1:07:44I really love this. I'm going to try to use that as a template for a small developer. There's different modes. You mark your function as reversible or requires human input and you build up a library of them. I think functions are just the complete utter game changer to everything. And if you're going to run any kind of sensitive data through this, I don't think any large company or medical field should let you run it without a function to clean up what you're sending through. So this goes both ways. Yes, it opens up major security holes. But at the same time, it completely changes everything. and I mean it because up until now yeah you could do a lot of things but operations were really hard scaling was really hard the thing has a mind of its own sometimes it returns stuff in the right structure sometimes it don't and like how do you build products around those because in computer science and devops and stuff like you expect a response either an error or in a certain format and it didn't have that well now you do so now you have an interface to the whole world now you You can do stuff with it.
1:08:57This completely changes everything. Because you can fire it off. You can clean up your data. You're finally free. I agree with everything you've just said. I completely agree. And yet the security implications are still terrifying to me. But yeah, I'm super excited. I have things I'm going to build on top of this. But I'm going to be very careful about them. I think the reversibility, especially like I want to build things that let people clean up data, like have a conversation with your data to clean it, because everyone who works with data says they spend 90 percent of the time on cleaning. They hate it.
1:09:36So great. I want to solve that problem. But I want there to be an undo function precisely to protect against some of the things that could go wrong. there's a couple of things i wanted to say like one is while i shared like the excitement i think that we still have to be careful because you know these things are still going to hallucinate and there's not going to there's there's still not a uh clear way to overcome that even using functions that's number one and number two is that i kind of want to echo what may have said before and while i i do appreciate that again this is like a great advancement it's super cool i don't know if I would call it a game changer because this has existed, right?
1:10:15Like this has happened in various frameworks, you know, link chain guidance, you know, other frameworks have guardrails in place. Yes, this looks like a very good implementation of this concept of having guardrails and being able to basically force the LLM to do what you want. And yes, it's very, you know it's it's it's beneficial to all of us that open ai that owns these models actually put put like some effort into fine-tuning their own model so that this works really well but i i just i just want to curb the enthusiasm a little bit right to just say like this isn't like you know earth shattering and like hasn't we haven't seen anything like this before the same could be said about chadgipity though right when chadgipity released folks are like, well, this is just prompting and this is just sending the same back and forth text.
1:11:09And then yet, JGPT was the first product that got to 100 million, whatever. Let's see how the adoption goes, right? The second all of us drop from this space and then start actually coding with this, and then we'll come back here and we'll see. Go ahead, Simon, and I want to hear from Sean and Steph. I just want to say, for me, until today the functions thing was always a hack, right? You could get it working on top of language model, but you had to do some pretty weird prompt engineering and mucking around to get that to happen. And now that it's part of the core platform, it feels so much, the friction involved in getting that working feels so much lower.
1:11:44I'm no longer afraid of it. You know, I was kind of cautious of doing this trick in the past because I knew there were so many ways that it might break. Now that I know that OpenAI fine-tuned a model for it, I feel a lot more confident in using it. So I think that makes a big difference. I want to say thanks for everybody here on stage sharing with us, like exploring with us, Simon and Sean and Riley and Niston and Zenova and Pharrell and Mayo and Steph and Roy and so many great folks here discussing kind of these latest changes that are definitely exciting. Potentially if you're siding with Niston ground breaking and earth shattering which I tend to agree and this was the reason for the space with Sean, we discussed this in DM and we brought a great panel of friends to kind of discuss how big this is.
1:12:30And I think now that we're hitting like an hour and a half, I think we're here, let's do a fairly quick kind of discussion about what are we building with this new tool that we have. So let's Steph go and then let's maybe everybody feel free to unmute and kind of have an instruction debate and then we're going to close this out because I'm losing my voice. And although this has been fun, it prevents all of us from going and actually playing with these new tools. So Steph, go ahead and then we'll have two or three folks to chime in and say, what are we building with this? I just wanted to give a shout out to Leon and his team.
1:13:03They just launched Garak, which is this tool for a security probing for LLMs. And it has an auto red teaming function. So that might be interesting to check out. But one thing that I was thinking about while Simon was talking and I read your blog post on prompt injection. And on one hand, we want to have more people exposing these vulnerabilities and maybe sharing their code for red teaming and their examples at the same time that is helping people who want to abuse like these models and functions. So it's a tricky one. I'm curious how you're thinking about it. And I had a parting question as well, which is what are we going to ask next from OpenAI?
1:13:43What is missing and what would make our use of the technology of the API better? This is a really interesting ethical question. it's like all aspects of security search people about responsible disclosure of security vulnerabilities the frustrating thing with prompt injection is we don't have a fix yet so you know normally if you find a sequel injection hole in someone's website you quietly tell them about it and they patch it and then everything's fine with prompt injection seeing as there is no known way to fix these problems i kind of feel like it's on us to make sure people understand before they build systems that are vulnerable to them so that that's the approach i've been taking is just just really trying to shake people and say no you can't just say oh we'll filter it out it'll be fine that doesn't work yet you need to sometimes you need to say no i cannot build the feature you're asking me to build because it can't be built securely and that's really frustrating i kind of hate that like i don't want to be the person who says there's a security hole you need to stop i want to be the person who says there's a security hole here's the fix for it and now we can move on with our lives but sadly we're not at that point with it yet i i like i 100 agree like i like it's been frustrating it's been weird to me like it's gone from like suspicious to like frustrating to like just like curious like like this seems to be such a hard problem like it's a like they're they're like it like i i think what's going on really is that like to solve this you kind of had to re-engineer it to this message-based api because like they have now reserved tokens they have tokens that only they know exist that they can insert in as like quote marks and they can have things like a system message they can tune it in a way that it actually like you know like if you could peer into its brain and see like what circuit is it implementing it's something that pays attention to the system message in the right way that it like understands the difference between what like the open ai customer told it our instructions and what the user told it and like i think that's progress in the right direction of like that it's like a sensible like target like It makes sense to me as an outsider of that's how you would go about addressing this, but it's a big change.
1:15:47And so I think they're moving towards it, but it speaks to what a hard problem this is. Because it is since May, I think, since a preamble where the original discoverers of it in May, when they put it in responsible disclosure. And so it seems like it's just a deep issue with how the instruction tuning or attention or something about this works, that it's not trivially fixable. And yeah, I think, you know, they, they're, they're making piece, you know, like piece by piece progress towards that thing. Very interesting to see because we did get an upgrade to 3.5 model, right? And Simon, you previously mentioned that there's a difference from a system message, how much the, you know, the model adheres to it before it's like way better than 3.5.
1:16:32It's interesting to now test again, this new model and see whether that listens to the system message better. Or maybe they implemented some of that in this kind of new update to this model. Yeah, they call that steerability. They specifically say that the models are now more steerable than they used to be. And they talk about steerability. That basically means how closely it abays the system prompt. I'd love to see some examples of that. I've not played around with it yet, but I'd love to see a few examples of prompts that were easily prompt injected with the previous 3.5, which are now protected against.
1:17:06But I did note that they have not flat out said, this is a solved problem. And until they do, I'm going to assume it's not a solved problem because, you know, it's in their interest to solve it and then tell people they've solved it. So I'm sure that you'll research this. And folks, Simon has a great blog, basically like a pensieve from Harry Potter that Simon has since 2013. I looked at it, Simon, you're very prolific. So I recommend following Simon and his blog and thoughts. Zenova, go ahead. And I think after that, we'll do like a round of last parting thoughts. Yeah, I just wanted to bring up this because I'm not 100 % sure of how they've really implemented it behind the scenes.
1:17:44But has anyone here sort of heard of like JSON Former, like one of these projects that was sort of aimed at these generating structured data? I love that thing. That thing is so clever. Yes, it's wonderful. yeah so i mean i'm just wondering why so i'm i'm speaking from because i i'm not really sure how they've implemented implemented it behind the scenes but assuming they are not doing it this way is there a reason why open ai is chosen not to because json form is like by definition you cannot there's no such thing as prompt injection in this case because you only generate tokens Oh, I disagree on that part.
1:18:28I think JSON form solves the problem of you want it to output JSON and it outputs invalid JSON. But prompt injection, in this case, is much more about when you summarize a web page, does the text from that trick it into then making a valid call to a function that does something you don't want to do? So my hunch on JSON form, I wonder if they just haven't had time to implement it yet. It's quite a tricky thing for them to, Because the way JSON form works, people who haven't seen it, is it basically injects extra logic at the next token prediction thing. So it knows that if you're doing JSON, you've just done the curly bracket.
1:19:03The only token that can come next is a single quote, is a double quote. And then the only things that come after that are not double quotes until you get to the end and so forth. So you can force your model to output structured text that matches JSON or YAML or whatever. Super clever. But yeah, my guess is that OpenAI just haven't got around to fully implementing that yet, and they'll get it working at some point. Right, yeah. So that's sort of what I was getting at with the bullet, I mean, I'm reading their readme right here, it's like the bulletproof way to generate a structured data. But so what I mean, this does not cover the case where you the separation between the user's input and the calling of the function that is always susceptible to prompt detection that that's where the security holes are.
1:19:48But assuming that you are forcing it to generate a structured data like JSON, there are these current approaches where I mean, it's modifying the logits. I mean, that's how I assume it's currently working, where it modifies how the next token is predicted. but so that that that's sort of what i'm getting at is like so why hasn't open ai done that approach or have they or yeah i'm not 100 sure on that yeah i think the the thing about these spaces is we don't have logan here or anybody from open ai and when we do they don't necessarily able to tell us what they're using go ahead riley oh i was just curious i mean if you played with i think grant slatin working on like context free grammar stuff that like i linked a while back i I didn't know how that compared to JSON former.
1:20:32I think it's the same exact trick, just even cooler because his thing, you can give it any grammar you like. Forget about JSON. It's anything that can be specified as a grammar. I think he posted the idea that he'd like to be able to upload a WebAssembly program to the OpenAI API and say, run this to pick the next token, which I think would be freaking incredible. That's an absolutely brilliant idea. Yeah, that's cool. that's really cool well maybe just go around and say like what we think should be built or yeah page from steph what do we want open ai to ship next yeah this is you know some some people from over there would definitely be listening to this so here's your chance to do your your pitch for why what they should build next go ahead steph i was going to say i would like love to have knowledge graphs and be able to have better retrieval.
1:21:29So, you know, this idea of like training with retrieval in mind and yeah, like not necessarily relying on vector databases. Like that's something that I would love to see in the future. Simon? I want widgets in ChatGPT. I think Chat is a terrible interface. I would like to be able to build it like a ChatGPT plugin that could say, now show them a map. Now ask them to pick something from this list of options, things like that. Just let us go beyond just having people type text to us. Indeed. Zenova? I want like a 30B model that's open source from them. I don't mind if it's like a Lama license. Open source.
1:22:13Okay, got it. Yeah, yeah, yeah. Just some model that you can mess around with. That'd be nice. All the claps. Please, guys, 80B, if you can do that. so like what you know maybe we should train like gpt 2.5 like just just down it back a little bit zenova or far l i mean i'm sort of on the open source coming from hugging the face so you know so but but i you know yeah gpt to win five let's go with that yeah honestly i'll just i just want 32k gpt4 with 75 percent or 80 percent cost reduction come on you'll get that next month i really hope so i'm good i'm good otherwise
1:23:08I'm very curious to see what we'd be able to build with these new functions and specifically combining them with agents. I think the opportunities there are pretty amazing. It's definitely going to make life a lot easier. And yeah, of course, if we could have like some more open source models, that would be awesome. And Python is going to have a chance for us to, or at least for folks to test this out, right, Roy? You want to tell us about the hackathon? Oh, sure. we're holding a our first virtual hackathon on june 19th to the 26th 100k in prizes it's going to be super fun it's going to be held in kumo space and we invite you all to attend and show us what you got 100k in prizes that's that might be the highest i've yet heard for one of these virtual hackathons that's pretty cool mail and then alex and i'll go and then i'll have riley give the last words of mail go ahead yeah open source 100 i mean we could it it's good what they've done but then you know my mind just goes to you know how can this be applied to open source right you know what they've just pulled off of the fine tuning you know how can we apply to open source so yeah i'd love for them to you know start to to be open as it says in their name right yeah they should just rename at this point how about you sean well okay so yeah obviously you know you want things for free they're not going to give it to you end of story i i'm just very interested in franken models i'm very interested in what simon has been sketching out this in this space which is essentially using them to call still smaller models but that do very specific things but can do a lot of things and so i'm interested in essentially just kind of rewriting my developer agent to do that kind of routing and explore the possibility of recursion.
1:25:01And when I have more information about that, I'll report back. That's great. I want to just before we go, I want to call out a podcast called Latent Space. So definitely check it out. I'll be posting the recording of this. Yeah. I mean, this is great. Everybody chipped in. Yeah, you have the developer perspective. This is what we want. Yeah. Oh, and we We also are going to drop our first, I think the first ever interview with George Hotz on Tiny Corp. And I'm very excited about that one. So make sure you don't miss that because AMD is lagging behind NVIDIA and George is working in that space. And there's potentially some exciting things to come.
1:25:41He had this email with Lisa Su and went back and forth. It was very dramatic. And nothing with George is boring. So Riley, go ahead. What would you want to come up with AI? I think the thing about this whole release is that as cool as functions are, I think the real thing that might be more important in the end is just this march towards lower prices and bigger context windows. I think there's a lot of unexplored stuff to do with big context. And there's a lot of possibilities that are opened up when things are just cheaper. And you can put in redundancy checks. You can have secondary prompts that check the work of the first prompt.
1:26:16You can engineer reliability around the parts that you need. so i mean it's hard to overstate just like how good it is that just the stuff's becoming cheaper and riley i saw that there's a webinar coming up and then you're going to teach advanced engineering you want to talk about this for a sec oh yeah sure so scale is on july 15th and i think we've already closed applications unfortunately but um on the july 15th we're having a hackathon and i'll be giving a talk on engineering there and i last time we did this i we did like a replay like I did the talk again for our webinar. So I expect we'll probably be doing that again.
1:26:52So I want to thank everyone here coming up to the stage. Here's my request to open AI. I want the Vision API as fast as possible. Oh, yes. I've been mouthwatering on the Vision API. Today, Mikhail Parakin from Bing confirmed that already like 10 % of Bing users get access to the GPT-4 Vision API. and I participated in an interview with the founder of Be My Eyes who are currently the only people in the world who has access to Vision API and that was a great conversation and I expect amazing things once that drops for the ability of, you know, GPT-4 to understand the real world not to mention how many prompt injections we can do via text but that's a conversation for another time with Simon and Riley.
1:27:37But I definitely want to thank everyone for coming here. This has probably been the biggest space that i've ran so thank you sean for for prompting this and everybody who joined hey yeah and and now we're gonna have some time to go and play with these models and your techniques and hopefully we'll see you guys again the last plug i'll say that i'll do i'm doing spaces every thursday many of the folks will stay here join we talk about latest updates this was an emergency one and glad to hear that sean is gonna compile this into a podcast so definitely subscribe to lake and spaces but everybody thank you for joining and go go play with the new build tools we got.
1:28:10Let's go build.
From the publisher
Full Transcript and show notes: https://www.latent.space/p/function-agents?sd=pf
Timestamps:
[00:00:00] Intro
[00:01:47] Recapping June 2023 Updates
[00:06:24] Known Issues with Long Context
[00:08:00] New Functions API
[00:10:45] Riley Goodside
[00:12:28] Simon Willison
[00:14:30] Eric Elliott
[00:16:05] Functions API and Agents
[00:18:25] Functions API vs Google Vertex JSON
[00:21:32] From English back to Code
[00:26:14] Embedding Price Drop and Pinecone Perspective
[00:30:39] Xenova and Huggingface Perspective
[00:34:23] Function Selection
[00:39:58] Designing Code Agents with Function API
[00:42:16] Models as Routers
[00:46:48] Prompt Engineering replaced by Finetuning
[00:52:15] The 2 Code x LLM Paradigms
[00:56:30] Smol Models for the future
[00:58:54] The Evolution of the GPT API
[01:03:27] Functions API Security vs Prompt Injection
[01:16:18] GPT Model Upgrades
[01:17:36] JSONformer
[01:21:03] Closing Comments - What We Want Next
Get full access to Latent.Space at www.latent.space/subscribe




