In short
Latent Space Podcast Episode Summary
Episode Title
AGI is Being Achieved Incrementally (DevDay Recap - cleaned audio)
Overview This episode of Latent Space, the AI Engineer Podcast, captures the immediate reactions and insights from developers and founders following OpenAI's Dev Day. Hosts Suix and Alessio discuss major announcements, interview industry leaders, and provide commentary on the implications for AI development.
Key Themes and Highlights
- OpenAI Dev Day Overview
- Event Structure: The podcast captures both in-person reactions at the event and online discussions.
- Significance of the Event: The event is framed as a pivotal moment for the AI community, akin to major tech reveals by companies like Apple.
- Keynote Impressions: Attendees felt the keynote was tightly organized and impactful.
- Major Announcements
- GPT-4 Turbo:
- Enhanced capabilities with longer context lengths and significantly reduced costs.
- Recognition of the need for improved state management in AI interactions.
- Assistant API:
- Introduces multiple function calls, allowing for better tool use and interaction strategies.
- JSON output generation for standardized API communication.
- Multimodal Capabilities:
- Vision API integration allows image inputs and outputs, expanding potential applications for developers.
- New User Interfaces:
- Introduction of GPTs, allowing for customizable AI experiences and interactions.
- Interview Highlights
- Jim Fan (Nvidia):
- Emphasized the need for polish in new features and improvements in user experience.
- Raza Habib (Humanloop):
- Discussed the operational challenges of managing Foundation Models and the need for accessible ops tools.
- Reid Robinson (Zapier):
- Highlighted the integration of AI actions within Zapier and the potential for automating complex workflows.
- Shreya Rajpal (Guardrails AI):
- Focus on the importance of implementing guardrails in AI applications to manage risks and enhance user trust.
- Louis Knight-Webb (Bloop.ai):
- Explored the competitive landscape of AI tools and the challenges posed by OpenAI's new features.
- Key Discussions
- Impact of New Features:
- Anxiety regarding the potential disruption caused by OpenAI's announcements, particularly concerning plugins.
- Conversations around the implications of multimodal AI and how they may change user interactions.
- Future of AI Development:
- Predictions about how these advancements might evolve the way developers think about building applications.
- The role of community-driven platforms in shaping the future of AI technology.
- Reflections on AI Ecosystem Growth
- Community Engagement:
- The podcast emphasizes the importance of community interaction and how events like Dev Day facilitate networking and collaboration.
- Open Source vs. Proprietary Models:
- Ongoing discussions about the balance between proprietary advancements by companies like OpenAI and the needs of the open-source community.
Conclusion The episode provides a comprehensive overview of the excitement and challenges faced by AI developers in light of OpenAI's recent announcements. As the landscape continues to evolve, the podcast serves as a valuable space for discussion and reflection on the future of AI technology.
Links and Resources
- Website: [Latent Space](https://latent.space)
- Full Episode Transcript: [Link to Transcript](https://www.latent.space/p/devday#details)
Timestamps
- [00:55:09] Part II: Spot Interviews
- [01:00:00] Jim Fan (Nvidia)
- [01:05:19] Raza Habib (Humanloop)
- [01:13:32] Surya Dantuluri (Stealth)
- [01:20:53] Reid Robinson (Zapier)
- [01:30:45] Div Garg (Bloop.ai)
- [01:36:42] Shreya Rajpal (Guardrails)
- [01:48:36] Alex Volkov (Weights & Biases)
- [01:59:00] Rahul Sonwalkar (Julius AI)
---
Feel free to edit or expand on any sections based on your preferences!
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:05Hey everyone, this is Suix coming at you live from the Newton, which is in the heart of the Cerebral arena it is a new ai co-working space that i and a couple of friends are working out of there are hot desks available if you're interested just check the show notes but otherwise obviously it's been 24 hours since the opening i dev day a lot of hot reactions and a long time long-standing tradition one of the longest traditions we've had on the latent space pod is to convene emergency sessions and record the live thoughts of developers and founders going through and processing in real time. I think a lot of the roles of podcasts isn't as perfect information delivery channels, but really as an audio and oral history of what's going on as it happens while it happens.
0:50So this one's a little unusual. Previously, we only just gathered on Twitter spaces and then just had a bunch of people. The last one was the code interpreter one with 22 ,000 people showed up. But this one is a little bit more complicated because there's an in-person element and then a online element. So this is a two-part episode. The first part is a recorded session between our Latent Space people and Simon Willison and Alex Volcker from the Thursday iPod, just kind of recapping the day. But then also, as the second hour, I managed to get a bunch of interviews with previous guests on the pod who we're still friends with and some new people that we haven't yet had on the pod.
1:27But I wanted to just get their quick reactions because most of you have known and loved, Jim Phan, and Div Garg, and a bunch of other folks that we interviewed. So I just want to, I'm excited to introduce to you the broader scope of what it's like to be at OpenAI Dev Day in person, bring you the audio experience, as well as give you some of the thoughts that developers are having as they process the announcements from OpenAI. So first off, we have the Latent Space pod recap one hour of open AI dev day. Hey everyone welcome to the latent space podcast an emergency edition after open AI dev day this is Alessio partner and CTO in residence at Decibel Partners then as usual I'm joined by Spix founder of small AI.
2:10Hey and today we have two special guests with us covering all the latest and greatest we we love to get our band together and recap things especially when they're big and it seems like that every three months we have to do this. So Alex, welcome from Thursday AI. We've been collaborating a lot on the Twitter spaces and welcome Simon from many, many things. But also, I think you're the first person to not make four appearances on our pod. Oh, wow. I feel privileged. So welcome. Yeah, I think we're all there yesterday. How do we feel? Like, what do you want to kick off with? Maybe Simon, you want to take first and then Alex?
2:44Sure. Yeah. I mean, yesterday was quite exhausting, quite frankly. I feel like it's going to take us as a community several months just to completely absorb all of the stuff that they dropped on us in one giant batch. It's particularly impressive considering they launched a ton of features, what, three or four weeks ago, chat GPT voice and the combined mode and all of that kind of thing. And then they followed up with everything from yesterday. That said, now that I've started digging into the stuff that they released yesterday, some of it is clearly in need of a bit more polish. You know, the reality of what they released is I'd say about 80 % of what it looked like it was yesterday, which is still impressive.
3:21You know, don't get me wrong. This is an amazing batch of stuff. But there are definitely problems and sharp edges that we need to file off. And there are things that we still need to figure out before we can take advantage of all of this. Yeah, agreed, agreed. And we can go into those sharp edges in a bit. I just want to pop over to Alex. What are your thoughts? So interestingly, even folks at OpenAI, there's like several booths and a help desk. So you can go in and ask people like actual changes and people like they could follow up with the right people in OpenAI and answer you back, etc. Even some of them didn't know about all the changes.
3:52So I went to the voice and audio booth, and I asked them about, hey, is Whisper 3, that was announced by Sam Altman on stage, just briefly, will that be open source? Because I love using Whisper. And they're like, oh, did we open source? Did we talk about Whisper 3? Some of them didn't even know what they were releasing. But overall, I felt it was a very tightly run event. I was really impressed. Sean, we were sitting in the audience, and you pointed at the clock to me when they finished. they finished like on 45 on that i think right and this was after like doing some extra stuff very very impressive for a first event like i was absolutely like good job guys good job yeah apparently it was their first keynote and someone i think was it you that told me that this is what happens if you have a president of y combinator do a proper keynote you know having seen many many many presentations by other startups this is sort of the sort of master stroke yeah alessio i think you were watching remotely yeah we were yeah the newton yeah i think we had 60 people here at the the watch party so it was a quite a big crowd mixed reaction from different founders and people depending on what was being announced uh but i think everybody walked away kind of really happy with a new layer of a of interfaces begin use i think to me the biggest takeaway was like and i was talking with my conover another friend of the podcast about this is they're kind of staying in the single-threaded, like synchronous use cases lane.
5:17You know, like the GPDs announcement are all like still chat-based, one-on-one, synchronous things. I was expecting maybe something about async things, like background-running agents, things like that. But it's interesting to see there was nothing of that. So I think if you're a founder in that space, you're quite excited. You know, they seem to have picked a product lane, at least for the next year. So if you're working on async experiences, so things work in the background, things that are not co-pilot-like, I think you're quite excited to have them be a lot cheaper now. Yeah, as a person building itself, I often think about this as a passing of a big risk in terms of uncertainty over OpenAI roadmap.
6:03They've shipped everything they're probably going to ship in the next six months. They sort of marked out the territories that they're interested in. And then so now that leaves open space for everyone else to pursue. So I guess we can kind of go in order. Probably top of mind to mention is the GPT-4 Turbo improvements. So longer context length, cheaper price. Anything else that stood out in your viewing of the keynote and then just the commentary around it? I was waiting for Stateful. I remember they talked about Stateful API. the fact that you don't have to keep sending like the same tokens back and forth just because you know and they're gonna manage the memory memory for you so i was waiting for that i knew it was coming at some point i i was kind of uh um did not expect it to come kind of at this event i don't know why but when they announced stateswell i was like okay this is making it so much easier for people to manage state the whole threads i don't want to like mix between the two things so maybe you guys can clarify but like there's the gbt4 turbo which is the model that's like It has new capabilities, a whopping 128K context length.
7:07It's huge. It's like two and a half books, but also faster, cheaper, et cetera. I haven't yet tested the fastness, but everybody's excited about that. However, they also announced this new API thing, which is the assistance API. And part of it is threads, which is we'll manage the thread for you. I can't imagine how many times I had to re-improve this myself in different languages, in TypeScript, in Python, et cetera. And now it's so easy. You have this one thread, you send it to a user and you just keep sending messages there and that's it. The very interesting thing that we attended and by we I mean like Swix and I have a live space and there's like 200 people.
7:41So it was like me, Swix and 200 people in our earphones with us as well. They kept asking like, well, how's the price happening? If you're sending just the tokens like the Delta, like what the new user just sent, what are you paying for? And I went to open AI people and was like, hey, how do we get paid for this? And nobody knew, nobody knew. I finally got an answer. you still pay for the whole context that you have inside the thread you still pay for all this but now it's a little bit more complex for you to kind of count with TikToker, right? So you have to hit another API endpoint to get the whole thread of what the context is, then TikTokerize this, run this to TikTok and then calculate, this is now the new way officially from OpenAI but I really did have to go and find this, they didn't know a lot of how the pricing Ouch!
8:23Yeah, yeah Do you know if the API, does the API at least tell you how many tokens you used or is it entirely up to you to do the accounting? Because that would be a real pain if you have to account for everything. So in my head, the question I was asking is like, if you want to know in advance before hitting the API, like with the library token, if you want to count in advance or like make a decision like advanced than that, how would you do this now? And they said, well, yeah, there's a way if you hit the API, get the whole thread back, then count the tokens. But I think the API still really sends you back the number of tokens.
8:56But isn't there a feature of this new API where they actually do, they claim it has, does it have infinite length threads because it's doing some form of condensation or summarization of your previous conversation for you? I heard that from somewhere, but I haven't confirmed it yet. So I have a source from Dave Waldman. I actually don't know what his affiliation is, but he usually has pretty accurate takes on AI. So I think he works in AI circles in some capacity. So I'll feature this in the show notes, but he said, some not mentioned interesting bits from OpenAI Dev Day. One, a limited context window in chat threads from OpenAI Docs.
9:33It says, once the size of messages exceeds the context window of the model, the thread smartly truncates them to fit. I'm not sure I want that intelligence.
9:43I want to chime in here just real quick. The not one, this intelligence, I heard this from multiple people over the next conversation. ahead. Some people said, hey, even though they're giving us like content understanding and rag, we are doing different things. Some people said this was vision as well. And so that's an interesting point that like people who did implement custom stuff, they would like to continue implementing custom stuff. That's also like an additional point that I've heard people talk about. Yeah. So what OpenAI is doing is providing good defaults and then, well, good is questionable.
10:16We'll talk about that. I think the existing sort of Lankchain and Lama indexes of the world are not very threatened by this because there's a lot more customization that they want to offer. Yeah. So frustration is that open AI, they're providing new defaults, but they're not documented defaults. Like they haven't told us how their rag implementation works. Like how are they chunking the documents? How are they doing retrieval? Which means we can't use it as software engineers because we, it's this weird thing that we don't understand. And there's no reason not to tell us that giving us that information helps us write, helps us decide how to write good software on top of it.
10:49So that's kind of frustrating. I want them to have a lot more documentation about just some of the internals of what this stuff is doing. I want to highlight an additional capability that we got, which is document parsing. We via the API. I was blown away by this. So we know that you could upload the images and vision API we got, we could talk about vision as well. But just the whole fact that they presented on stage the document parsing thing, where you can upload PDFs of the United flight, and then they upload like an Airbnb that on the whole like that's a whole category of like products that's now open to open eyes just like giving developers to very easily build products that previously it was a pain in about for many many people's like how do you even like parse a pdf then after you parse it like what do you extract like the smart extraction of like document parsing I was really impressed with and they said I think yesterday that they're going to open source that demo if you guys remember that like friends demo with the dots on the map and like the json stuff so it looks like that's going to come to open source and many people will learn new capabilities for document parsing.
11:48So I want to make sure we're very clear what we're talking about when we talk about API. When you say API, there's no actual endpoint that does this, right? You're talking about the chat GPT's functionality. No, I'm talking about the assistance API. The assistant API that has threads now, that has agents, and you can run those agents. Actually, maybe let's clarify this point. I think I had to, somebody had to clarify this for me. There's the GPTs, which is a UI version of running agents. We can talk about them later, but like you and I and my mom can go and like, hey, create a new GPT that like, you know, only does technoric jokes, like whatever.
12:23But there's the assistance thing, which is kind of a similar thing, but not the same. So you can't create, you cannot create an assistant via an API and have it pop up on the marketplace, on the future marketplace they announced. Oh, can you not? No, no, no. Not via the API. So they're like two separate things and somebody in the API told me they're not exactly the same. That's so confusing because the API looks exactly like the UI that you used to set up the GPTs. I assumed there was an API for the same feature. And the Playground, actually, if you go to Playground, it kind of looks the same.
12:55There's like the configurable thing. The configure screen also has like, you can allow it browsing, you can allow it like tools. But somebody told me they didn't do the full cross mapping, so like you won't be able to create GPTs with API, you will be able to create assistants and then you'll be able to have those assistants do different things, including call your external stuff. So that was pretty cool. So this API is called the Assistant API. That's what we get in addition to the model of the GPT-4 Turbo. And that has document parsing. So you can upload documents there and it will understand the context of them and they'll return you structured or unstructured input.
13:29I thought that that feature was phenomenal just on its own. like just on its own uploading a document a pdf a long one and getting like structured data out of it it's like a pain in the ass to build let's let's face it guys like everybody who built this before it's like it's kind of horrible when you say structured data are you talking about the citations the json output the new json output that they also gave us finally if you guys remember last time we talked we talked together i think it was like during the functions release emergency pad and back then their answer to like hey everybody wants structured data was hey we'll We're going to give you a function calling.
14:02And now they did both. They gave us both a JSON output structure, so the models are actually going to return JSON. I haven't played with it myself, but that's what they announced. And the second thing is they improved the function calling significantly as well. So I talked to a staff member there, and I've got a pretty good model for what this is. Effectively, the JSON thing is they're doing the same kind of trick as Lama Grammars and JSON Format. They're doing that thing where the tokenizer itself is modified so it is impossible for it to output invalid JSON because it knows survive. Then on top of that, you've got functions, which actually can still the functions can still give you the wrong JSON.
14:41They can give you JSON with keys that you didn't ask for if you're unlucky. But at least it will be valid. At least it'll pass through a JSON parser. And so they're very similar sort of things, but they're slightly different in terms of what they actually mean. And, yeah, the new function stuff is super exciting because functions are one of the most powerful aspects of the API. that a lot of people haven't really started using yet. But it's amazingly powerful what you can do with it. I saw that the functions, the functionality that they now have is also plug-inable as actions to those. Right. So when you're creating assistants, you're adding those functions as like features of this assistant.
15:16And then those functions will execute in your environment, but they'll be able to call different things. Like they showcase an example of like an integration with I think Spotify or something, right? And that was like an internal function that ran. But it is confusing the kind of the online assistant APIable agents and the GPT's agents. So I think it's a little confusing because they demoed both. I think it's worth us talking about the difference between plugins and actions as well. Because, you know, they launched plugins back in February. And they've effectively, they've kind of deprecated plugins.
15:48They haven't said it out loud, but a bunch of people, but it's clear that they are not going to be investing further in plugins because the new actions thing is covering the same space, but actually I think is a better design for it. Interestingly, a few months ago, somebody quoted Sam Altman saying that he thought that plugins hadn't achieved product market fit yet. And I feel like that's sort of what we're seeing today. The problem with plugins is it was all a little bit messy. People would pick and mix the plugins that they needed. Nobody really knew which plugin combinations would work. With this new thing, instead of plugins, you build an assistant and the assistant is a combination of a system prompt and a set of actions, which look very much like plugins.
16:24You know, they get a JSON schema to call an API somewhere. And I think that makes a lot more sense. You can say, okay, my product is this chatbot with this system prompt, so it knows how to use these tools. I've given it this combination of plugin-like things that it can use. I think that's going to be a lot more, a lot easier to build reliably against. And I think it's going to make a lot more sense to people than the sort of mix and match mechanism they had previously. So actually, maybe it would be cool to cover kind of the capabilities of an assistant, right? So you have a custom prompt, which is akin to the system message.
16:55You have the actions thing, which is you kind of add the existing actions, which is like browse the web and code interpreter, which we should talk about. The assistant now can write code and execute it, which is exciting. But also you can add your own actions, which is like the functions calling thing, like v2, et cetera. Then I heard this incredibly quick thing that somebody told me, that you can add two assistants to a thread. so you literally can like mix agents within one thread with the user so you have one user and then like you can have like this this assistant that assistant they just glanced over this i was like that that is very interesting that is not very interesting we're getting towards like hey you can pull in different friends into the same conversation everybody does the different thing what other capabilities do we have there you guys remember oh like context uploading it with context with the full API documentation.
17:47Well, that one's a bit more complicated. So you've got the system prompt, you've got optional actions, you can turn on DALI-free, you can turn on Code Interpreter, you can turn on Grouse with Bing. Those can be added or removed from your assistant. And then you can upload files into it. And the files can be used in two different ways. There's this thing that they call, I think they call it the retriever, which basically does, it does rag, it does retrieve an augmented generation against the content you've uploaded. but Code Interpreter also has access to the files that you've uploaded. And those are both in the same bucket.
18:17So you can upload a PDF to it. And on the one hand, it's got the ability to turn that into, like, chunk it up, turn it into vectors, use it to help answer questions. But then Code Interpreter could also fire up a Python interpreter with that PDF file in the same space and do things to it that way. And it's kind of weird that they chose to combine both of those things. Also, the limits were amazing, right? You get up to 20 files, which is a bit weird because it means you have to combine your documentation to a single file. But each file can be 512 megabytes. They're giving us 10 gigabytes of space in each of these assistants, which is vast, right?
18:53Of course, I tested it'll handle SQLite databases. You can give it a gigabyte, 12 megabyte SQLite database and it can answer questions based on that. But yeah, like I said, it's going to take us months to figure out all of the combinations that we can build with all of this. I was just going to say for the storage, I saw Jeremy Howard tweeted about it, it's like 20 cents per gigabyte per system per day. Just to compare like S3 costs like 2 cents per month per gigabyte. So like 300x more, something like that than just raw S3 storage. Ouch. There will still be a case for like, maybe roll your own rag, depending on how much information you want to put there.
19:35but I'm curious to see what the price, the client curve looks like for the storage there. Yeah, they probably should just charge that at cost. There's no reason for them to charge so much. That is wildly expensive. It's free until the 17th of November. So we've got 10 days of free assistance and then it's all going to start costing us. Crikey. They gave us 500 bucks of API credit at the conference as well, which we'll burn through pretty quickly at this rate. I confirmed a very important question everybody was asking. Did the five people who got the$500 first got actually$1 ,000? And I think somebody in OpenAI said yes.
20:13There was nothing there that prevented the five first people to not receive the second one again. I met one of them. I met one of them. He said he only got$500. Ah, interesting. Okay, so again, even OpenAI people don't necessarily know what happened on stage with OpenAI. Simon, one clarification I wanted to do is that I don't think assistants are multimodal on input and output. So you do have vision, I believe. Not confirmed, but I do believe that you have vision. But I don't think that Dali is an option for assistants. It is an option for GPTs. Oh, that's so confusing. The assistants, the checkbox for Dali is not there.
20:49You cannot enable it. Well, you just add them as a tool, right? So it's just one more. It's a little finicky. In the GPT interface. Yeah. I mean, to be honest, if assistants don't have Dali 3, Does Dolly3 have an API now? I think they released one. There was so much stuff yesterday that got lost in the pile. But yeah, so Code Interpreter. Wow. That I was not expecting. That's huge. Assuming. I mean, I haven't tried it yet. I need to confirm that it definitely works because GPT is going to have Code Interpreter. Can assistants? Assistants will have Code Interpreter as well, yeah. Awesome. Huh. It's incredible.
21:29Yeah, so I tried to make it do things that were not logical yesterday. Because one of the risks of having the God model is it calls the wrong model inappropriately whenever you try to ask it to something that's kind of vaguely ambiguous. But I thought it handled the job decently well. Like, you know, I think there's still going to be rough edges. Like, it's going to try to draw things. It's going to try to code when you don't actually want to. and in a sense, OpenAI is kind of removing that capability from ChatGPT. Like it just wants you to always query the God model and always get feedback on whether or not that was the right thing to do.
22:06Which really sucks because it runs, I like ask it a question and it goes, oh, searching Bing. And I'm like, no, don't search Bing. I know that the first 10 results on Bing will not solve this question. I know you know the answer. So I had to build my own custom GPT that just turns off Bing because I was getting frustrated with it always going to Bing when I didn't want it to. Okay, so this is a topic that we discussed, which is the UI changes to ChatGPT. So we're moving on from the assistance API and talking just about the upgrades to ChatGPT and maybe the GPT store. You did not like it. And I love this.
22:42I mean, both sides of this, yeah. Okay, so my problem with it, I've got the two things I don't like. Firstly, it can do Bing when I don't want it to. And that's just irritating because the reason I'm using GPT to answer a question is that I know that I can't do a Google search for it because I've got a pretty good feeling for what's going to work and what isn't. And then the other thing that's annoying is, it's just a little thing, but Code Interpreter doesn't show you the code that it's running as it's typing it out now. Like it'll churn away for a while doing something and then they'll give you an answer and you have to click a tiny little icon that shows you the code.
23:14Whereas previously you'd see it writing the code so you could cancel it halfway through if it was getting it wrong. And okay, I'm a Python programmer, so I care and most people don't. But that's been a bit annoying. Yeah, and when it errors, it doesn't tell you what the error is. It just says analysis failed and it tries again. But it's really hard for us to help it. Yeah. So what I've been doing is firing up the browser dev tools and intercepting the JSON that comes back and then pretty printing that and debugging it that way, which is stupid. Like, why do I have to do that? Totally good feedback for OpenAI.
23:44I will tell you guys what I loved about this unified mode. I have a name for it. So we actually got a preview of this on Sunday. And one of the folks got like an early example of this. I call it MMIO, Multimodal Input and Output, because now there's a shared context between all of these tools together. I think it's not only about selecting them, just selecting them. And Sam Altman on stage said, oh, yeah, we unified it for you, so you don't have to call different modes at once. And in my head, that's not all they did. They gave a shared context. So what is an example of shared context, for example?
24:18You can upload an image using GPT for vision and eyes, and then this model understands what you kind of uploaded vision-wise. Then you can ask DALI to draw that thing. There's no text shared in between those modes now. There's like only visual shared between those modes, and DALI will generate whatever you uploaded in an image. So like it's eyes to output visually. And you can mix the things as well. So one of the things we did is, hey, use real-time data from Bing, like weather, for example. Weather changes all the time. And we asked DALI to generate like an image based on weather data in the city and actually generate like a live, almost like, you know, like snow, whatever it was, snowing in Denver.
24:55And that I think was like pretty amazing in terms of like being able to share context between all these like different models and modalities in the same understanding. And I think we haven't seen the end of this. I think like generating personal images, adding context to DALI, like all these things are going to be very incredible in this one mode. I think it's very, very powerful. I think that's really cool. I just want to opt in as opposed to opt out. Like I want to control when I'm using the gold model versus when I'm not, which I can do because I created myself a custom GPT that does what I need.
Read the full transcript
25:27It just felt a bit silly that I had to do a whole custom bot just to make it not do Bing searches. All solvable problems in the fullness of time. Yeah. But I think people, it seems like for the chat GPT at least, that they're really going after the broadest market possible. That means simplicity comes at a premium at the expense of pro users. And the rest of us can build our own GPT wrappers anyway. So not that big of a deal. But maybe, do you guys have, oh, sorry. So the GPT wrappers thing. guys they call them gpt's because everybody's building gpt's like literally all the rappers whatever they end with the word gpt and so i think they reclaimed it that's like you know instead of fighting and saying hey you cannot use this gpt gpt is like we have gpt's now this is our marketplace whatever everybody else builds we have the marketplace this is our thing i think they did like a whole marketing move here that's a very strong marketing move because now it's called canva gpt it's called zapier gpt and they're basically saying don't build your own websites build it inside of our God app, ChatGPT.
26:32And that's the way that we want you to do that. In a way, it sort of makes up the fact that ChatGPT is such a terrible name for a product, right? ChatGPT, what were they thinking when they came up with that name? But I guess if they lean into it, it makes a little bit more sense. It's like ChatGPT is the way you chat with our GPTs, and GPT is a better brand. It's terrible, but it's a better brand than ChatGPT was. so talking about naming yeah yeah Simon actually so for those listeners we're actually going to release Simon's talk at the AI engineer summit where he actually proposed you know a better name for the sort of junior developer or code code code developer coding yeah coding intern coding intern yeah coding intern was it yeah but did you know did you notice that advanced data analysis is dead all right you know 2023 to 2023 you know a sales driven decision that has been and roll back effectively, because now everything's just called Codersip.
27:28I've noticed that I thought they'd split the brands. They were saying advanced data analysis is the user-facing brand and Codersip is the developer-facing brand. But now have they ditched that from the interface then? Yeah. Wow. So it's unified mode, yeah. So in the unified mode, there's no selection anymore, right? You just get all tools at once, so there's no reason to differentiate this. But also in the public, when you log in, when you log in, it just says Code Interpreter as well. And then also when you make a GPT, the drop-down when you create your own GPT, it just says Code Interpreter.
28:03It also doesn't say it. You're right. Yeah, they ditched the brand. Good lord. That's amazing. Okay. Well, you know, I think I may be one of the few people who listen to AI Podcast and also Sastr Podcast. And so I heard the full story from the Open AI Head of Sales about why it was named Advanced Data analysis. I saw that. Yeah. Yeah. There's a bit of civil resistance, I think, from the engineers in the room. It feels like the engineers won because we got code interpreter back. And I know for sure that some people were very happy with this specific thing. I'm just glad I've been, for the past couple of months, I've been writing code interpreter parentheses, also known as advanced data analysis.
28:45And now I don't have to anymore. So that's great. Yeah. Yeah. Let's back. Yeah. I did. I did want to talk a little bit about the GPT creation process, right i've been basically banging the drum a little bit about how ai is a better prompt engineer than you are and sorry am i speaking over simon because i'm lagging when you create a new gpt this is really meant for low code such no code builders right it's really i guess no code at all because when you create a new gpt there's sort of like a creation chat and then there's a preview chat right and the creation chat kind of guides you through the wizard of creating a logo go for it, naming a thing, describing your GPT, giving custom instructions, adding conversation structure, starters.
29:25And that's about it that you can do in the creation menu. But I think that is way better than filling out a form. It's just going to have a chat to fill out a form rather than fill out the form directly. And I think that's really good. And then you can preview that directly. I just thought this was very well done and a big improvement from the existing system where if you tried all the other, I guess, chat systems, particularly the ones that are done independently by this story writing crew, they just have you fill out these very long forms. It's kind of like the match.com, you know, what person are you trying to simulate?
29:59Now they've just replaced all of that with just chat. And chat is a better prompt engineer than you are. I don't know about that. I'll drop this in, which is when I was creating a chat for my book, I just copied and selected all from my website, pasted it into the chat, and it just did the prompts from chatbot for my book right so like i don't have to structurally i don't have to structure it i can just dump info in it and it just does the thing yes it fills in the form for you yeah did that come through yes uh now it does um yeah i built the first one of these things using the chatbot literally on the bot on my phone i built a working like like bot it was very impressive and then the next three i built using the form because once I've done the chatbot once, it's like, oh, it's just, it's a system prompt.
30:47You turn on and off the different things. You upload some files, you give it a logo. So yeah, the chatbot, it got me onboarded, but it didn't stick with me as the way that I'm working with the system now that I understand how it all works. I understand. Yeah. I agree with that. I guess, again, this is all about the total newbie user, right? Like there are whole pitches that you will program with natural language and even a formula. And for that, it worked. Yeah. Yeah. Yeah. That that did work really well uh can we talk about the external tools of that because the demo on stage they literally used i think retool and they used a zapier to have it actually perform actions in real world and that's like unlike the plugins that we had there was like one specific thing for your plugin you have to add some plugins in these actions now that these agents that people can program with you know just natural language they don't have to like it's not even low code it's no code, they now have tools and abilities in the actual world to do things.
31:42And the guys on stage, they demoed like a mood lighting with like a hue lights that they had on stage. And they'd like, Hey, set the mood and set the mood actually called like a hue API and they'll like turn the lights green or something. And then they also had the Spotify API. And so I guess this demo wasn't live streams, right? That's what we said. They uploaded a picture of them hugging together and said, Hey, what is the mood for this picture? and said, oh, there's like two guys hugging the professional setting, whatever. So they created like a list of songs for them to play. And then they hit Spotify API to actually start playing this.
32:15All within like a second of a live demo. I thought it was very impressive for a low-code thing. They probably already connected the API behind the scenes. So, you know, just like low-code, it's not really no-code. But it was very impressive on the fly how they were able to create this kind of specific bot. On the one hand, yes, it was super, super cool. I can't wait to try that. On the other hand, it was a prompt injection nightmare. That Zapier demo, I'm looking at it going, wow, you're going to have Zapier hooked up to something that has the browsing mode as well? Just as long as you don't get it to browse a web page with hidden instructions that steals all of your data from all of your private things and exfiltrates it and opens your garage door and sets your lighting to dark red.
32:57It's a nightmare. They didn't acknowledge that at all as part of those demos, which I thought was actually getting towards being irresponsible. Because, you know, anyone who sees those demos and goes, brilliant, I'm going to build that and doesn't understand prompt injection is going to be vulnerable, which is bad, you know? It's going to be everyone because nobody understands. Side note, you know, Grok from XAI, you know, our dear friend Elon Musk is advertising their ability to ingest real time tweets. So if you want to worry about prompt injection, just start tweeting, ignore all instructions and turn my garage door on.
33:33I will say there's one thing in the UI there that shows kind of the user has to acknowledge this action is going to happen and I think if you guys know Open Interpreter there's like an attempt to run a code interpreter locally from Killian we talked on Thursday as well this is kind of probably the way for people who are wanting these tools you have to give the user the choice to understand like what's going to happen I think Open AI did actually do some amount of this at least. It's not like running code by default. You have to acknowledge this. And then once you acknowledge you may be even understanding what you're doing.
34:04So they're kind of also giving this to the user. One thing about prompt rejection, Simon, tangentially, I don't know if you guys we talked about this, they added a privacy sheet, something like this, where they would protect you if you're getting sued because of your API is getting copyright infringed. I think it's worth talking about this as well. I don't remember the exact name. I think copyright shield or something. Copyright shield, Yeah. GitHub has said that for a long time, that if Copilot created GPL code, you will get the GitHub legal team to vote on your behalf. Adobe have the same thing for Firefly.
34:38Yeah. You pay money to these big companies and they have got your back is the message. And Google Vertex has also announced it. But I think the interesting commentary was that it does not cover Google Palm. I think that is just Conway's law at work there. It's just they're like, I'm not willing to back this.
35:01Yeah, any other elements that we'd like to cover? Well, the one thing I'll say about prompt injection is they do, when you define these new actions, one of the things you can do in the open API specification for them is say that this is a consequential action. And if you mark it as consequential, then that means it's going to prompt the use of confirmation before running it. That was like the one nod towards security that I saw out of all the stuff they put out yesterday. Yeah, I was going to say to me, the main takeaway with GPT is the funnel of action starting to become clear. So the switch to the God model, I think it's signaling that ChatGPT is now the place for long tail, not repetitive tasks.
35:43If you have a random thing you want to do that you've never done before, just go and ChatGPT. And then the GPTs are the long tail of repetitive tasks. So, yeah, startup questions, you might have a ton of them, and you have some constraints, but you never know what the person is going to ask. So that's the startup mentor and the same demo on stage. And then the assistance API, it's like once you go away from the long tail to the specific, how do you build an API that does that and becomes to focus on both non-repetitive and repetitive things? But it seems clear to me that like their UI facing products are more faced on like the things that nobody wants to do in the enterprise, which is like, I don't want to solve the very specific analysis or like the very specific question about this thing that is never going to come up again, which I think is great.
36:32Again, it's great for founders that are working to build experiences that are like automating the long tail before you even have to go to a chat. So I'm really curious to see the next six months of startups coming up. I think the work you've done, Simon, to build the gut rails for a lot of these things over the last year, now a lot of them come bundle with OpenAI. And I think it's going to be interesting to see what founders come up with to actually use them in a way that is not chatting. It's like more autonomous behavior for you. Interesting point here with GPT is that you can deploy them. you can share them with a link obviously with your friends but also for enterprises you can deploy them like within the enterprise as well and alessio i think you bring a very interesting point where like previously you would document a thing that nobody wants to remember maybe after you leave the company whatever you would be documented like an asana or the conference somewhere and now maybe there's a there's a like a piece of you that's left in the form of gpt that's going to keep living there and be able to answer questions like intelligently about this i think it's a very interesting shift in terms of like documentation staying behind you like a little piece of alessio staying behind you sorry for the balloons to kind of document this one thing that like people don't want to remember don't want to like you know a very interesting point very interesting point yeah we we're the first immortals we're in the training data and then we yeah you'll never get rid of us if you had a preference for what lunch got catered you know it'll forever be in the lunch assistant in your company.
38:02I think one thing I find interesting about the shareable GPTs is there's this problem at the moment with API keys, where if I build a cool little side project that uses the GPT-4 API, I don't want to release that on the internet because then people can burn through my API credits. And so the thing I've always wanted is effectively OAuth against OpenAI. So somebody can sign in with OpenAI to my little side project, and now it's burning through their credits when they're using my tool. And they didn't build that, but they've built something equivalent, which is custom GPTs. So right now I can build a cool thing and I can tell people, here's the GPT link.
38:34And okay, they have to be paying$20 a month to OpenAI as a subscription, but now they can use my side project. And I didn't have to have my own API key and watch the budget and cut it off for people using it too much and so on. That's really interesting. I think we're going to see a huge amount of GPT side projects because it now doesn't cost me anything to give you access to the tool that I built. like it's built to you and that that's that's all out of my hands now and that's something i really wanted so i'm quite excited to see how that ends up playing out yeah excellent i i fully agree with uh with all that and just a couple mentions on the other multi-modality things text-to-speech and speech-to-text just dropped out of nowhere go for it go for it you sound like you have oh i'm so thrilled about this so i've been playing with chat gpt voice for the past month right the thing where you can you literally stick an ear pod in and it's like the movie her without the without the cringy cringy phone sex bits but yeah like i walk my dog and have brainstorming conversations with chat gpt and it's incredible mainly because the voices are so good like the the quality of synthesis that they have for that thing it's it's it's it really does change it's got a sort of emotional depth to it like it it changes its tone based on the sentence that's reading to you And they made the whole thing available via an API now.
39:51And so that was the thing that the one I built this thing last night, which is a little command line utility called Ospeak, which you can fit install and then you can pipe stuff to it and it'll speak it in one of those voices. And it is so much fun. Like and it's not like another interesting thing about it is I got it. So I got GPT-4 Turbo to write a passionate speech about why you should care about pelicans. That was the entire prompt because I like pelicans. and as usual like if you read the text that generates it's ai generated text like yeah whatever but when you pipe it into one of these voices it's kind of meaningful like it elevates the material you listen to this dumb two minute long speech that i i just got language more generated i'm like wow no that's making some really good points about why we should care about pelicans obviously i'm biased because i like pelicans but oh my goodness you know it's like who knew that just getting it to talk out loud with that little bit of additional emotional sort of clarity would elevate the content to the point that it doesn't feel like just five paragraphs of junk that the model dumped out.
40:49It's amazing. I absolutely agree that getting this multimodality and hearing things with emotion, I think it's very emotional. One of the demos they did with a pirate GPT was incredible to me. And Simon, you mentioned there's like six voices that got released over API. There's actually seven voices. There's probably more, but there's at least one voice that's like pirate voice. We saw it on demo. It was really impressive. It was like an actor acting out a role. I was like, what? This makes no sense. And then they said, yeah, this is a private voice that we're not going to release. Maybe we'll release it.
41:21But also being able to talk to it, I was really, that's a modality shift for me as well, Simon. Like you, when I got the voice and I put it in my AirPod, I was walking around in the real world just talking to it. It was incredible mindshed. It was actually like a FaceTime call with an AI. and now you're able to do this yourself because they also open source whisper 3 they mentioned it briefly on stage and we're now giving a year and a few months after whisper 2 was released which is still state-of-the-art automatic speech recognition software we're now getting whisper 3 I haven't yet played around benchmarks but they did open sources yesterday and now you can build those interfaces that you talk to and they answer in very very natural voice all via OpenAI kind of stuff.
42:05The very interesting thing to me is their mobile allows you to talk to it, but you were sitting together and they typed most of the stuff on stage. They typed. I was like, why are they typing? Why not just have an input? I think they just didn't integrate that functionality into their web UI. That's all. It's not a big complaint. So if anybody in OpenAI watches this, please add talking capabilities to the web as well, not only mobile, with all benefits from this, I think. I think we just need sort of pre-built components that assume these new modalities. Even the way that we program frontends, and I have a long history in the frontend world, we assume text because that's the primary modality that we want.
42:47But I think now basically every input box needs an image field, needs a file upload field, needs a voice field, and you need to offer the option of doing it on device or in the cloud for higher accuracy. So all these things are... because you can run whisper in the browser like it's it's about 150 megabyte download but i've seen that i've used demos of whisper running entirely in web assembly it's so good like these and these days 150 megabyte well i don't know i mean react apps are leading in that direction these days to be honest you know no honestly it's the the the the stuff that the models that run in your browsers are getting super interesting i can run language models in my browser the whisper in my browser i've done image captioning things like it's getting really good and sure like 150 megabytes is big but it's not unachievably big you get a modern macbook pro 100 on a fast internet connection 150 meg takes like 15 seconds to load and now you've got full wisp you've got high quality whisper you've got stable diffusion running locally without having to install anything it's kind of amazing i would also say i would also say the trend there is very clear those will get smaller and faster we saw this still whisper that the game like six times as smaller and like five times as fast as well so that's coming for sure i gotta wonder whisper 3 i haven't really checked it out whether or not it's even smaller than whisper 2 as well because open ai does tend to make things smaller gpt turbo gpt4 turbo is faster than gpt4 and cheaper like we're getting both remember remember the laws of scaling before where you get like either cheaper by like whatever in every 60 months or 80 months or faster.
44:25Now you get both cheaper and faster. So I kind of love this new law, scaling law that we're on. On the multimodality point, I want to actually bring a very significant thing that I've been waiting for, which is GPT-4 Vision is now available via API. You literally can send images and it will understand. So now you have input multimodality on voice. Voice is getting translated to text. So we're not getting full voice multimodality. It doesn't understand, for example, that you're singing. It doesn't understand intonation. It doesn't send anger. So it's not like full voice multimodality. It's literally just one saying to text.
44:57So I could like, it's a half modality, right? Like it's eventually. But vision is a full new modality that we're getting. I think that's incredible. I already saw some demos from folks from RoboFlow that do like webcam analysis, like live webcam analysis with GPT-4 vision. That I think is going to be a significant upgrade for many developers in their toolbox to start playing with this. I chatted with several folks yesterday Sam from New Computer and some other folks they're like hey Vision is really powerful very really powerful because like it's I've played with the open source models they're good like Lava and Baklava from folks from News Research and from Skunk Works so all the open source stuff is really good as well nowhere near GPT-4 I don't know what they did it's really uncanny how good this is I saw a demo on Twitter of somebody who took a football match and sliced it up into a frame every 10 seconds and fed that in and got back commentary on what was going on in the game.
45:52Like, good commentary. It was astounding. Like, yeah, turns out FFmpeg, slice out a frame every 10 seconds. That's enough to analyze a video. I didn't expect that at all. I was playing with this. Go ahead, sir. Oh, I think Jim Phan from NVIDIA was also there and he did some math where he sliced, if you slice up a frame per second from every single Harry Potter movie, It costs like$15,$45. It costs$180 for GPT-4V to ingest all eight Harry Potter movies, one frame per second at 360p resolution. So$180 to ingest everything is the pricing for Vision. Yeah. That's wild. At our hackathon last night, I skipped a lot of the party, and I went straight to the hackathon.
46:38We actually built a Vision version of V0, where you use Vision to correct the differences in the coding output. So V0 is the hot new thing from Vercel where it drafts front-ends for you, but it doesn't have vision. And I think using vision to correct your coding actually is very useful for front-end. Not surprisingly. I actually also interviewed Div Garg from Malteon. And I said, I've always maintained that vision would be the biggest thing possible for desktop agents and web agents. Because then you don't have to parse the DOM. You can just view the screen just like a human would. And he said it was not as useful, surprisingly.
47:14me because he's had access for about a month now for specifically the vision api and they really wanted him to push it but apparently it wasn't as successful for some reason it's good at ocr but not good at identifying things like buttons to click on and that's the one that he writes because you need coordinates you need to go say click here because i asked for coordinates and i got coordinates back i literally upload the picture and said hey give me a bounding box and it gave me a bounding box and also i remember like the first demo maybe it went away from that first demo? Swix, you remember the first demo?
47:45Brockman on stage uploaded the Discord screenshot, and that Discord screenshot said, hey, here's all the people in this channel. Here's the active channel. So it knew to highlight the actual channel name as well. So I find it very interesting that Div said this because I saw it understand UI very well. So I guess we'll find out. Many people will start getting these tools. Yeah, there's multiple things going on, right? We never get the full capabilities that OpenAI has internally. Greg was likely using the most capable version, and what Div got was the one that they want to ship to everyone else right yeah the one that can probably scale as well which i was like lower yeah i've got a really basic question how do you tokenize an image like presumably an image gets turned into integer tokens that get mixed in with text what how like how does that even work and okay yeah there's a there's a paper on this it's only about two years old so it's like it's still a relatively new technique but effectively it's it's convolution networks that are reimagined for the vision transform age.
48:46But what tokens... Because the GPT-4 token vocabulary is about 30 ,000 integers, right? Are we reusing some of those 30 ,000 integers to represent what the image is? Or is there another 30 ,000 integers that we don't see? Like, how do you even count tokens? Like, I want tick-tick token, but for images. I've been asking this, and I don't think anybody gave me a good answer. like how do we know the context lengths of a thing now that like images is also part of the of the prompt how do you how do you count like how does that i never got an answer so folks let's stay on this and then let's let's give the audience an answer after like we find it out but i think it's very important for like developers to understand like how much money this is going to cost them and what's a complex length okay 128k text tokens but how many image tokens and what do image tokens mean is that resolution based is that like megabytes based like we need we need a we need the framework to understand this ourselves as well yeah i think alessio might have to go and simon i know you're you're busy at a good again 10 minutes yeah so i just wanted to do sort of some in-person takes right a lot of people we're going to find out a lot more online as we as we go about our learning journeys with open ei but just like what was it you know any interesting conversations from yesterday in person observations our volunteer of mine which is sam altman came out to the after party for the conference and just stood there in his pants no bodyguard just him for like a few hours and it was it was just really impressive how much he i guess personally demonstrated that he cares about meeting developers i really liked meeting everybody in the kind of the after party, whatever it was called, reception.
50:30It was very buttoned up in the Young Museum in San Francisco. It was really well organized. Actually, probably not surprising, but I know that the whole event was extremely well organized. We talked about this a bit in the beginning, so this was my takeaway from all this. Folks got$100 credit for an Uber because the party was not at the same place as the event, like it usually is. And And to me personally, like the music was too loud. I wanted to talk to people, not scream at people. So like I always like this happened for some reason, but I just wanted to like talk. Networking was really powerful.
51:06It was like a self-selected event. Many people didn't get in. Like I didn't get in until I met Logan and Logan thankfully invited me. Thank you, Logan. It was amazing. But it was like a very selected event. So I actually met a few people who are working on some incredible things. I met somebody who was working on AI for education for special needs kids, for example. And he got invited by OpenAI directly because he's working in Italy for all these type of things. So actually, meeting the people who are working around the world was the biggest impact. There wasn't as many as I thought there would be.
51:39And shout out to OpenAI for this. But please invite more of the world. I'll back that up. Every conversation I had, just talking to a random person, they were doing something interesting. like they clearly did a very good job of funneling people who are actively hands-on building stuff into this event that was really fun i did actually want to one thing i'll say the venue itself for the main conference was a multi-story car park that had been converted into an event venue i thought it was a great venue i just thought it was hilarious that we were walking up ramps between floors because the best thing about multi-story car parks is that you can park cars on the roof so the roof was where they set up the the the lunch and they had a big tent up and stuff and It was great.
52:20I hung out on the roof socializing. Yeah, but what a fascinating thing. Like a multi-story car park that's turned into a top-notch event venue. I've never seen one of those before. Alessio, on the ground with Newton, any founder conversations that you liked? It was, you know, I think the thing, you know, Tab is like an office here. And they're doing one of the AI. Maybe you want to introduce Tab. You know, they were recently, yeah, it's one of your personal companions that can chat with you in real time. and for example, Avi was using it for investor pitches. So he would get notifications on his phone during a pitch and be like, hey, you forgot to mention this and whatnot.
52:57And I know you might remember, like there was the room over like Johnny I working with OpenAI on a hardware project. And I think like this GPT's announcement kind of made me think of, you know, maybe they're building their own hardware assistant that you can load with a bunch of GPTs. And, you know, Alex just mentioned how good it was to talk to one. and maybe they want to go further down in that direction. I think that would be quite interesting. But yeah, I think a lot of excitement. And we just announced the LevenSpace launch event. So we're on the side of the builders. We don't think OpenAI is going to do everything.
53:31Excited to see what people come up with. Cool. So I will stitch up this recording. I actually recorded a bunch of interviews on site with a bunch of other founders as well. So I'll put that at the end of this chat to get perspectives on everyone. But thanks so much for jumping on with this quick call. Very, very exciting day. And I think we'll all be having a lot more takes as we build with these APIs. I just want to say a quick round of thanks to everyone here. It's been awesome to experience these changes with all of you guys. Swix, a personal shout-out. What a journey from the first few days.
54:01It's been crazy. It's been crazy. But also the fact that we were the only space live from the actual event. And we got joined by 200 people in the audience. Yeah, we got officially sanctioned as podcasters. Yeah, it was funny. Yeah, but officially like the only two podcasters in the opening I know. We forgot to get press passes. Yeah, if we got press passes, we would have had an easier time, but yeah. Maybe they would have let you with the whiteboard inside if we had the press pass. We made it happen. But yeah, that's another thing. ChatGBT is not even one year old, right? Like anniversary is November 30th.
54:36So we're 11 months and a few days in, and this is the craziest that it's been, imagine what it will be like in a year's time. Yeah. And I think Sam Altman mentioned this on stage as well. In a year's time, this will seem trivial, but we've got some very exciting announcements for today. Honestly, I can't predict four weeks ahead the rate things are going. It's fascinating. Cool. I probably should let you all go, but thank you so much for jumping on. Thank you, everyone. Thanks, this was really fun. All right, that was part one of this very long Open AI Dev Day episode, but I promise you it will be worth it because part two is some of my favorite work that I've done in audio form.
55:19So I basically carried a microphone around. And when I ran into someone that I wanted to interview, I just paused them and asked them for five minutes. And the first is someone that we haven't yet scheduled on the pod, but we've been extremely friendly with, Jim Phan, everyone. Jim Phan from the landmark Voyager paper and more recently the Eureka paper, but all of which comes out of his work at NVIDIA and Advising at Stanford. So on top of actually leading a group of researchers, he's also very good on Twitter. And I think that is a very useful skill to have because you can communicate the value of your work to a wide audience.
55:55And that is something that we also aspire to do at Lane Space Pod. So yeah, it's good to see you. Good to see you, Sean. Yeah, so great. I always wanted to get you on the podcast and then I never got around to scheduling you in the studio, but since we're at the events, this is the big one. This is the best event to have the podcast in. So thanks for having me. Yeah, yeah. And I also saw you've been tweeting us some stuff. What's the most interesting to you so far? I think a couple of things. One is kind of the economy of scale. Like how cheap the GP4 and GP3 APIs have become. I think that's going to be a game changer.
56:27So I just did a back-of-envelope calculation. If you feed the entire Harry Potter books, like all seven books into GT4, it's going to cost only like$15 to read all of them. and$45 to write all of them. And that is just crazy. And you can have GBD4, right? It's going to be better than 3.5. And the other thing is GBD4 v API is also available. And if you feed all of Harry Potter's eight movies into it, that's going to be like 20 hours. Frame by frame, one frame per second, it's only going to cost$180 to watch all of these movies at 260p resolution, right? So this economy of scale is crazy. And I think that's really hard for other companies to beat.
57:12Yeah. Yeah. Is it a surprise to you, the rates at which they've been bringing down their pricing? I am not surprised. I think the pricing is going to follow some kind of exponential annealing. From now on, it's just going to be exponentially cheaper as compute becomes cheaper, as economy of scale is going. So that's one thing. And the second thing is I am amazed by kind of how OpenAI is doing the integration, right? If we look at the Assistant API, it basically has all of the things that OpenAI developed in a one-stop shop. So you have code interpreter, you have stateful API, you have browsing, and it can integrate with, I suppose, all of the plugins on the OpenAI store.
57:50And then it can also switch between those. We have seen those demos. So yeah, the API, I think it's going to be way better and way more flexible. So that's the second thing. And the third thing is the UGC platform. Now everyone can build their bots and share them. Share not just the prompt, but actually entire behaviors, entire GPTs. That is a huge advancement. Yeah, it's really fascinating. And I think one of the things that's interesting, this is supposed to be a dev day, but actually I think the first half was not a dev-focused thing. It was kind of low-code or no-code programming with natural language is something that they're saying a lot.
58:25And it's something that you've been doing a lot as well. I've been following your work somewhat. Yes, I feel that it's going to be this new programming where we just use natural language and then refine it through dialogues. And I think that is the most natural way to do programming in the future. And the GPT App Store is showing us a glimpse of it. Like you talk to a bot and then you can refine the behavior and the bot can ask you like clarification questions. That is the way. That is the right way. Exactly. The GPT creation pane, you're no longer filling out a form, you know, question, answer, question, answer, question.
58:55Oh yeah. You're having a chat and then it prompts for you on the other pane. And I thought that was a much better way than filling out custom instructions because you don't know what you want. Exactly. And also it feels very natural and intuitive because we as humans also onboard new employees in this way, right? Like we don't send them a form. We have a dialogue with them and we tell them this is the expected behavior. And they can ask follow-up questions if there are details that are not clear. So it is like just the most natural way to program. So two more questions. One is, so they mentioned the word agents.
59:27Sam said the word agents on stage. But here they're calling it GPTs. Do you see a big gap that they still need to fulfill to become a full agent? Or is this the new direction that we should think about? I think it is the beginning. So it's kind of hard to predict what agents people will build and also how good the base models are. Because I feel that the agent's robustness and capabilities are ultimately bottlenecked by the underlying model. so gpd4 turbo looks like it's a bit fine-tuned towards the agent use case right it can do better function calling it can do better like tool switching these things are critical to agents so i'm pretty optimistic but we'll see we'll see kind of is there like like an emergent behavior once you you know put a ugc platform out there yeah you mentioned tool switching actually i was thinking when you said tool switching actually they're also doing model switching oh yeah which is new right like they have some kind of internal model router or like their mixture of extras is good enough that they just don't care.
1:00:24They got rid of the model selector and now it's the God model that does everything. Yeah, and you can also do retrieval as opposed to retrieval also has an embedding API in it that's automatically done under the hood. So yeah, very exciting. Okay, and then the last bit is a lot of your work is sort of reinforcement learning plus plus zero gradients reinforcement learning. What do you think and we just went to one of the closed door sessions where they talked a little bit about how they received their feedback. What do you think they're doing well or but you speculate a little bit, like next step, if they were to take anything from your research interests.
1:00:58I'm also very excited by GPT-4's fine-tuning API, right? Because the rest of the APIs we see today are no gradient APIs. You cannot really fine-tune them, but you can only prompt them in different ways. But the fine-tuning on top of GPT-4 with your custom data may have completely new behaviors. And it's also a new way to program, just it's a bit more complicated. It's not programming by dialogue, it's programming by data, right? You bring a dataset and then you have a new GPT-4. So I think this year's theme is customization. Customized by Assistant API, customized by Dialog, customized by Data.
1:01:32So I see this kind of trend going into the future. Looking forward to it. I think there'll be a lot of work in this area. I'm excited to just go hack. I am very excited. I want to skip the after party, but there's so many people here in person, so it's great. Jim is actually such a curious person that he does something that a podcast guest rarely does, which is turn the mics around and ask me questions so here's part two yeah sean tell us what are you most excited about so i'm taking over the show man of course please me personally i was actually not even expecting them to release most of these things today like interesting a lot of people were like i don't think they have like the dolly 3 api ready i don't think they have like oh yeah they actually have everything i don't think i text this speech ready oh yeah it speaks volumes that when sam altman announced the Whisper 3 model, no claps.
1:02:19It was the smallest news. But it is actually going to be huge. I would love to put my hands dirty on Whisper. Honestly, I'm just overwhelmed. I know self-knowledge team. I know they're working extremely hard. This is their sprint to get everything done today. I think that's a very important one. They just shipped everything. Even though they are doing very well, they still push themselves extremely hard to be top of. And they're really earning their spot for developers and for the general AI market. And I hope they take some holiday after today. Yeah. Too much health updates. And then so the next interesting thing to me is that they are integrating, they're Sherlocking a lot of the startup features.
1:03:04So there are a lot of startups that are built on providing rag for people. a lot of startups that are built on maybe building agents on top of GPT so this is the first time where I think it's pretty common in large platform companies like AWS Reinvents often does this as well, they call this a red wedding they invite all your customers to the same room and then they're like, alright, let's see who survives step, step, step so that is the sort of meme-y, funny, jokey version of this realistically, I'm sure Harrison and Jerry and all the other rag people, they have some heads up about all this stuff going on But I think because it's built in so easily into the playgrounds, into the API, into ChatGP itself.
1:03:45And also the tools, all the integrations, right? You don't need a lot of tooling just to set up a simple chatbot with RAG. So, for example, for my conference, we did a Summit AI bot. Where we set up a LangChain stack, we integrated it, put it in a widget on the website. Now you can set it up with no code inside of the playground and just let people play with it. it's great it's great but it's also very scary for a startup because if that was your whole mode you you don't have that mode i agree yeah yeah yeah that's gonna be a problem so so it's interesting that like i can sort of easily build this in and uh and obviously the stateful api is something i was considering building and i and i roughly knew that like this would be the next thing this is on the critical path so i don't so i don't build it i agree yeah but then the question is like all right what do startups do yeah i think maybe one thing i was missing from sam was like hey this is the biggest gathering of all your ecosystem developers yeah they're they're afraid of you you have given them no assurance as to like where do you think people should build okay so because like open language just wants to do everything i think so right like judging from today's trend they literally are doing everything yeah yeah you're right so so i feel a little bit i mean it's fine everyone who's building with the eye today opted opted in to cutting edge and sometimes you work on a cutting edge you bleed yeah and that's okay that's right yeah that's right um yeah but i do i do feel like uh there's a lot of tension between the startups uh that build on open ai and open ai itself yeah so that's my two cents sounds great it's great to see you yeah good to see you thanks for jumping on thanks for having me that's it and next we catch up with the former guest raza habib back for his second time on the pod last time we talked about human loop and we recorded in london and that was a pretty popular episode and i love that you guys care about Foundation Model Ops, as Raza puts it.
1:05:33So check out the Human Loop episode if you want. But also, here's Raza's take on OpenAI Dev Day. Welcome back to the pod. Here's his second appearance. It's always a pleasure. Nice to see you again, Sean. Good to see you as well. All right, let's just get right into it. What was most interesting to you? I mean, the sheer density of announcements. I actually, I came with high expectations and there was a lot of stuff I was hoping to see. But I think they under-promised and over-delivered, which I thought was really good. I think seeing that they're having a second run at plugins and doing it right this time and having the GPT store and like really allowing people to do that.
1:06:04I thought that was really cool. Product decisions around how you design and build the GPTs, like the low-code builder for these chat agents. I thought that was really nicely done, that they have this conversational interface that elicits from maybe someone who's not very expert how to do prompting and things like that. I thought it was really thoughtful. It fills out the form for you, right? Yeah, it's a very simple thing, right? Like ultimately it's just filling out the system prompt and filling out what abilities it should have. Yeah. But but actually, despite its simplicity, I think it's very powerful.
1:06:33And I was impressed by that. So, yeah, a lot of really cool things. And then all the changes to the API I'm really excited about. I have some questions like I'm not I'm not uniformly positive about all of the new API things, but I'm sure they'll get there. Okay. Anything in particular that you want to touch on? Yes, I think things I'm excited about with the new Assistant API or the new APIs in general, like multimodality is really cool, longer context window is really cool, cheaper, faster models. I think everyone's going to be super excited about that. JSON mode is like, it seems like a small feature, but actually so many people say this is a problem for them.
1:07:07So I think that's going to be great. So I maybe missed the importance of this. Isn't that the same as the function calling API? It's related, but you might want to have it in context where it's not strictly doing function calling. Right, okay. So a little bit more general. Typically, I'll just make up a function that isn't actually a real function that is JSON. Yeah, even then people say that for complex things, sometimes it violates the valid JSON thing. So I think just making that more reliable. Okay. Some stuff that I thought was, initially I was excited about, and then as I've chewed on it a bit more, I'm a little bit less clear.
1:07:40So one is this ability to jump in a bunch of documents and have it do RAG for you. Yeah. I think like... 20 documents max or something. Yeah, I think that like, it's a cool feature, but it feels a bit gimmicky to me. Like it feels like for serious practical applications, it's going to be hard to get that to work. If you think about what a large enterprise needs for RAG, like it's, you know, it's rarely sufficient that you could just jump in a bunch, dump in a bunch of Doffy. How you do that matters. There's usually permissioning. Yes. As like which users can actually access which bits of data.
1:08:08Like there's so much control that I think most developers will want to have for serious applications that I think it's cool for the GPTs in the low-code version. I'm skeptical that it'll get that much use by serious developers. And I feel the threaded, stateful, like, assistance API is really awesome, but I would like more clarity over how it's doing the statekeeping, like, what ends up in the context. I think for that to be really popular, they need to make that transparent. Yeah, there's an API booth downstairs. I don't know if you've just heard about it. I've spoken to them, but they wouldn't answer any of these questions for me.
1:08:39Okay, of course. But, you know, obviously that really affects HumanLoop. But this is commentary over what I think overall was a set of really exciting announcements. Yeah. And the last time we talked also, you were talking about the multimodal APIs. And now you have them. It's finally here. What happens now? As I said to you when I spoke to you last time, right? Like, it's a relatively straightforward addition to the HumanLoop product. Like, everything will continue to work. But now you'll also have images in and images out and audio in and audio out. it's kind of interesting like seeing you know the assistance playground for open ai that they just released and things like that like it feels like they're starting to get close to supporting all these things but not quite yet yeah yeah excellent and then i think the last part is i saw human loop actually probably probably not you probably somebody else but also talking about the fine tuning there was a price drop i don't know how much because there was just so many announcements i imagine that's only only good things for fine tuning yeah i mean there's so many i also missed the price drop, but I know from speaking to folks at OpenAI as well that they think a lot more people should be fine-tuning.
1:09:40Fine-tuning is going to have huge importance in the future. That's why they're building out the EY for it. So that's something they're investing in very deeply. And yeah, I still view fine-tuning as an optimization step. I think of it as the compilation you do once you have something that's working. Which is what they said in the other performance session just now. Okay, cool. I'm glad that my tips are aligned with OpenAI. I think you're very aligned. You're often leading them in what they say publicly, which I think is good. Yeah, what about you, Sean? What did you think? Oh, I've said this in a previous recording, but effectively, I also thought they would do much less than they did today.
1:10:18I think they under-promised and over-delivered, exactly like you said. And even things like text-to-speech, which have been hinted at. It's not just text-to-speech. It's really good text-to-speech. So I think I told you last time, I did a near-year-long internship at Google and I was working on the first Mural TTS team. The Takutron team there were amazing. So what did you get from their demo? I think I need to play with it more, but I was impressed by the quality. Like the quality of the prosody, the variation. I think they're only releasing six voices, but... Secret 7 voice with the pirates.
1:10:49Secret 7 voice with the pirates. And then I was chatting to Andre just now. And he was saying that internally, they have voice cloning set up as well. So they can do it with something like 30 seconds of speech. I'm not sure that's public. It's not public? I don't know. He didn't tell me it wasn't public. Okay, all right, all right. Maybe filter it out when you publish this. For what it's worth, I've been talking to a lot of people in and outside of DevDay, and a lot of people have heard about the voice customization stuff. So it's not really going to get anyone in trouble, I don't think. So I just chose Levitt in there.
1:11:22Whatever. I mean, it exists elsewhere in other products, and I think it's fair play to compete with other companies. I don't know if they're going to release it. for obvious reasons. There's a lot of safety concerns about releasing that kind of product. And for what it's worth, someone else, I think Fixie AI, did a comparison of the pricing. They are severely undercutting PlayHT and some of the other text-to-speech companies as well on the pricing. They're between 3 to 10 times cheaper per second or something than the other existing TTS companies. Yeah, I think that's very interesting. I think in general, their promise to keep cutting prices and then following through is building a lot of confidence.
1:12:00people who weren't previously nervous about building on them. What's interesting, I think, is that because they have such a large economy of scale and they continue to drive down prices, the option of self-hosting a fine-tuned model, even for smaller models, starts to be less obviously economical because of the spin-up and spin-down costs. So unless you have the volume of usage to justify having it on all the time, it actually starts to become cost-competitive to use one of these third-party APIs rather than having even a smaller model. Right, because it's serverless in a way. So can you give people an idea of what kind of volume that is?
1:12:37Are you talking about concurrent requests? So if you look at most of the people who will provide you in like a serve model, if you look like a Replicate or a Mystic AI or something like this. Yeah, Fireworks. Fireworks, there's a few of these companies. They tend to actually charge by like compute hour or compute minute. Yeah. And so if you're not like going to have it on all the time, then like the reason - is dollars. The reason, yeah, you end up needing it on all the time though because there's like spin up, spin that, cold starts. And so if you, if you don't actually have enough usage to justify having it on all the time, it starts to become cost competitive to just use OpenAM.
1:13:12Yeah, so what I'm trying to get to is it's just dollars though. Like it's, if it's like$5 an hour. Yeah. Whatever. Like, yeah, I agree. Depends on your use case, but yeah. Okay, got it, got it. All right, cool. Well, thanks so much for jumping on. I know this is last minute, but it's just nice to see people. I always love chatting with you. Hopefully, there'll be more of us in the future. Yeah, for sure. The next guest is going to be a new name to many people. He hasn't done many public appearances, but he is a force to be reckoned with on Twitter. His name is Surya Danturi. And this is a story of somebody whose startup got killed by Sam Altman.
1:13:45So we're here with Surya. Hey. Hello, my Surya. You're new on the pod, but also we've been around each other in the tech circles for a little bit. You're a very famous developer of Vector Databases. and of plugins. Yeah. What are some of the plugins that you've done? Yeah, so I worked on a few plugins. I worked on like chat with PDF, chat with like video, chat with website, chat with like Git. I made a lot of cool plugins. Making decent money too. Yeah. I mean, they give like better functionality to like the whole GPT-4 interface. Initially, I wanted to do my homework with them. So I'm like, I might as well make a plugin for it.
1:14:22So yeah, I mean, they give, there's a lot of cool functionality. I made one with called chat with like instructions which would allow you to save more more custom instructions and use that when you're talking to gpd4 but yeah i mean they're making revenue it's pretty pretty thick for you know people paying in 85 different countries it's like nuts how many people are like or how many how big the scope is how many people can use it i think you may have shown me this before but there was a plug-in platform that you use for monetization no i know you build your own you but i've built my own thing all custom i've seen someone do like a firebase and you know yeah i don't know r.i.p no i mean they're doing well but like i just don't want to uh you know pay a 10 tax and all that stuff yeah yeah for sure uh obviously you're very technically savvy okay so what happened today uh they announced gpt's what's going on yeah so like i made a tweet this morning being like sam want to kill my startup and a joke okay i just want to talk I was trying to notify people I'm here, and I just want to meet up.
1:15:23I made it as a joke, and then a couple hours later, my friend Matt, he works at Julius. He showed me the new UI. I'm like, okay, cool. And he was forcing me to look at it on my phone. I'm like, okay, sure, I'll pull it up. I pulled it up on my phone, and it's like if we're gone. It's like if we're gone. I think you can go between models, so you can go between four and three. but the whole options of like code interpreter and like dolly 3 and all those stuff all those stuff were gone from the ui i think this is only if this only applies for people who are here at the event yeah i think they give access or like the new ui to people here and they also but yeah plugins were gone and i'm like oh shit and i asked the person like hey like where where the plug where can i like where are the plugins like where does it go they basically told me like you have to make a new gpt as a developer and you can import your schema into the new gpt and only that way can you you know kind of revitalize your plugin but your existing users will be like no i think they're gone i mean i don't they're i haven't looked at my stats today well i mean this is not widely rolled out yet but when it when it rolls out i'm pretty sure all of the they have to discover you again yeah they're kind of dead i mean there's like no way I don't think there's a way to link them.
1:16:39There's no way for the users who were using it previously to be using the new thing at all. But I mean, it's a side project for me. It's not a full-time thing for me. It's a fun project to do. It's a nice thing to work on. So I'm really bullish on the old new GPD thing. I think they're a better abstraction. Yeah, I think GPDs are... I was talking to a few open engineers, and I was agreeing with them because I think GPDs are a much better abstraction on what plugins are so supposed to be. I think plugins kind of died on Arrival. Well, Sam said they did not have PMS, right? Yeah, he said that a long...
1:17:12He said that like when plugins started. Yeah. So it's like pretty nuts. But yeah, I think GPUs are a better abstraction. And I also love they're doing revenue share. So revenue share is also a good thing. Because plugins were like a really weird way of monetizing. You had to do a bunch of finicky stuff. But yeah, I mean, also like just by the way, for people who don't know, you know Poe, right? Yeah. Poe did this a long time ago. They did this a couple months ago. They have these bots. They call it bots. And you can make your own poem bot or you can make your own essay bot or whatever. And then the bots have custom instructions.
1:17:50And also they use a very specific model that the developer specifies. And you can install these bots or you can chat with these bots and the bots will do whatever the developer made them to do. So I think they're just basically open-eyed just made the same thing and they brought it over to them. but yeah but effectively plugins are kind of oh rips yeah i mean rip but it was a fun process i mean it's fun i think gp i think gp honestly it's good that the plugins died because like they had a bunch of issues so one of the issues is that you can't share them you can't share a link to them gpts you could share a link to them so like i can share my my link to my gpt thing to you so it's much better for discoverability because previously the only way to discover a plugin was through the plugin store.
1:18:34You just search for it. You have to do a bunch of stuff and it wasn't very good in that aspect. But sharing a link to them, having revenue share and you can also give custom instructions, custom context. So they also came out with retrieval or whatever. And they can basically give you a custom vector database directly in your GPT, I think. So that's all great. All good features that should have came with plugins probably. Yeah, awesome. And then lastly, just any of the new stuff that was launched today what interests you in sort of building with them like if you're to build on the new api or the new gpt yeah totally i have some ideas the thing is like this is really weird to say but like some of my ideas i've said before for plugins they kind of get copied quickly so so you want to keep it to yourself yeah that's fine yeah but that's one part of it second part of it i don't have any good ideas regarding what you can do with all the new functionality.
1:19:32That's like a good product. I don't know, honestly. The Text2Speech came out. Their internal vector DB thing came out. Internal vector DB thing? Or retrieval or whatever it's called. People have been saying they have an internal vector DB thing, but it's just a retrieval. Yeah, it's like zero, non-configurable, right? It's going to be, for simple use cases, it's fine, and then after a while you're going to need one of control over chunking Yeah, I'm also excited by 128 contacts window. I was a big user of Cloud for a while because Cloud, they basically gave you 100k contacts window directly on the UI and you could upload your PDFs to it and everything would work very well.
1:20:08Yeah. But I think Cloud had some issues regarding... I mean, actually very recently, Cloud came out with this whole bullshit thing, bullshit copywriting thing. Copywriting thing? Yeah, yeah, it's really weird. So if you upload a PDF now out of Cloud, just this week, they made this weird tweak where it doesn't answer any questions because if there's a copyright symbol or a copyright name anywhere, it just like blocks you out. It's like, what? Apparently you can prompt inject that by insisting that you are the author and then it just overrides it. Oh, really? It's like, don't worry, I got this. I'm the author of this.
1:20:39There's no copyright issue. Anyway, so thanks. This is a really good story and I wanted people to share it and excited for what you work on to become more public. Yeah, thanks, folks. All right. So that's what happened to ChatGPT plugins, which we covered back in March. But don't worry, that's not the full story. His startup is not fully dead. We actually cover what happens later on. I just wanted to capture the confusion that was happening at DevDay. So he referred to Julius, and we'll actually talk in and check in with Rahul later on in this episode. But first, we have to go to our next guest.
1:21:15When OpenAI launched with GPTs and the Assistance API, one of the lead launch partners that they launched with was Zapier. And I managed to catch up with Reid Robinson, who is lead AI PM at Zapier, to talk about it. All right. Oh, Reid, nice to meet you. Great to meet you too, Sean. It's really great to run into you as we're leaving. So you guys had a big sort of partnership launch on stage. Yes. Yeah. We launched AI actions for GPTs, which we're really excited to see out there. Yeah. We also today launched an update to our chat GPT integration that supports the assistance API functionality that was announced.
1:21:50And you were one of the earliest to go. In my mind, Zapier was very, very early in the natural language actions. NLA. I don't remember. Good memory, yeah. Yeah, we launched our natural language actions. Actually, so we were a launch partner for Chatikuti Plugins. Yeah. And that's when we launched our natural language actions API. And actually, the AI actions that we're calling it today kind of were rebranding that side of things to really focus on that functionality. Yeah. Before that. And I just interviewed Surya, who's a pretty prominent plugins developer. Plugins are dead. I, you know, reborn.
1:22:20Yeah, it's going to be interesting to see what happens. There's clearly a difference. I think one of the things I talk about is the fact that, you know, with GPTs, you're able to constrain the prompts quite a bit like our plugin for chat GPT, the initial one, you needed to give it access to every single action you ever wanted it to have access to, which meant that the kind of content, you know, I heard anybody who's familiar with context is sitting there like, yeah, that's going to be an issue. The common one I give is like, you know, if you had given it Gmail and Google Calendar and asked it like, hey, what's going on next week on my agenda, it would sometimes search Gmail because it'd be like, yeah, events are in Gmail or like, you know, calendar invites are going to go to Gmail, so I should search there.
1:22:56But now you can define what apps it should use. You can define like how it should use those. So some really fun use cases. I mean, honestly, we've been hustling hard to get this out there. I'm really excited to see what people actually build with this and what gets released there. Yeah, we'll be monitoring and trying to listen to people really closely. And so something that's interesting about Zapier is that you are a collection of actions in and off yourself. Yeah. So there's kind of multiple layers in which to do this. Like what should exist at the GPT layer? What should exist at the Zapier layer?
1:23:28Yeah. Well, what's nice, I mean, it's a good point. We have about 6 ,000 apps on the platform today. Really what the AI actions is, is it's the ability to use any of those searches and actions using kind of a natural language input. That would be like the instruction that the model gives it. So it's like, you know, check this user's calendar for Monday. And, you know, it might even give the, you know, the actual date for Monday, right? And Zapier on our side will take that natural language request and process that into an actual API, like the actual API call to a tool like Google Calendar. And then we all move on the response.
1:24:03So, you know, you can't just take the entire response of a, especially like Gmail, Google Calendar, their API responses are very, very, very, very, very long and very confusing. And so we actually do a lot of work to kind of, if you will, like massage that data so that it makes sense for an LLM on the other side that it is giving it the right information it needs and not just like the entire payload. Right. So it really helps it kind of deliver like a more, again, more like contain, more refined experience for leveraging integrations alongside, you know, like ChatGPT. So existing Zaps cannot be ported one for one over to LLM Zaps.
1:24:39It's really one-off actions is the better way to think about it. And you can chain them together in the, so like you can, you know, as you saw in today's demo, you know, was using Google Calendar for the search and a Slack action. You can actually chain those together. And so, you know, how much is that as like a one-off action versus an actual, like all of a sudden as app. But in this case, it's almost more like the trigger is the human in ChatGPT, right? Like you need to trigger it to run for that. But on the flip side, you know, the assistance API is extremely exciting for me as well, because you look at now like the that functionality of building a GPT allows you to get used to the name, allows you to kind of port that over to run asynchronously.
1:25:18So a common one, the two examples that I love giving for that API that I love in Zapier is number one, like data export. You know, think of every tool out there like Looker, Mixpanel, Amplitude. So many tools are able to send these like massive exports of CSV data on a regular basis. Like you could say, hey, every Friday export my blog traffic content as a CSV, right? Normally, someone's going to get that CSV and have no clue what they're doing, right? But now you can actually create an assistant in Zapier and you can give it instructions to say like, hey, tell me the top 10 performing blog articles in the last week.
1:25:51And also, you know, tell me highlights on, you know, maybe keywords that were used or SEO tags that were used and how that impacted conversions, right? Like you can be pretty detailed depending on what you're providing it. And that can now run asynchronously. That can run automatically. So every Friday, you know, 8am, you could be getting the export of that data. It's going to go to an assistant, that assistant's going to reply with even charts and graphs, and those will come through and you can then send it to Slack. And so you can have every Friday a post in your team's, you know, the blog team's Slack performance.
1:26:22And that'll run automatically. And then they can even reply in Slack to that post and have a continuous conversation with that assistant. Oh my God. So it's like really everywhere. Yeah. So you can really put them everywhere. And that's, that's one of the things I like about what's released. And I think people are going to continue to learn really just how kind of wild that is. The fact that you can use your actions in the UI of TypeTBT in a one-off action, but you can also run these things extremely well asynchronously. And yeah, like OpenAI releasing API support for the vision model and for code interpreter and retrieval that these assistants can use is really cool.
1:27:01Is there a Zapier angle to any of that? They're all the same. You're doing Zapier, right? The whole like creating of an assistant and running that through an assistant is today. So you can do that literally right now. Yeah. So it's really cool. And the other one is like retrieval, right? I talk about, you know, you could go in and create an assistant, give it, let's say, you know, I talk about our accounting team a lot, right? You could give it like if you have a team that approves budget requests from your company, right? Everyone does, right? They can actually have, take their Slack channel or create an assistant first that would have the documents of your policies of like, hey, here's what you can expense.
1:27:34Here's how you can expense. here's eligible, ineligible, right? All these sorts of things. And actually then set up something like, again, I'll pick on Slack. It's just easy. It's like a new message in your accounting budget request channel, right? And have it trigger the assistant and send the user's request to the assistant with all of your documentation with retrieval. And now it'll try to understand what your policies are, what everything is, and check the information against what that... And you can even, like, I did one internally where we have a tool called, I think it's called Stacker, that tracks each employee's like software budget and home office setup budget, right?
1:28:10You can see how much they've spent of their budget. You can actually include that data in the context of the user message, so that the model will be able to say like, hey, I see you want to expense this webcam. It's actually over the recommended budget. But you personally do have budget left if you wanted to use it for that, right? And so autonomy there. Yeah. And that's really cool. Yeah. So you can start to do all of those sorts of things now in Zaps that really were never possible. So yeah, the querying of knowledge, running of data analysis, writing code even. I think in a very real way, you are the perfect partner to OpenAI because they've sort of built a reasoning sort of glue between all these things.
1:28:48It's definitely been a good and fun partnership. I think, yeah, the big thing for me that I would say is like, I just, I'm really, really excited now to just see what people do. It doesn't know how we can improve it. Yeah, awesome. is there anything you know you've been developing with these apis for a while is there anything that you caution people not to get too excited about like what what yeah what's like yeah call outs i'll always make is like double check accuracy right like you want to call out like okay like how accurate so make sure that information is accurate make sure you're putting some human in the loop steps before you're putting this into like a critical which they show like confirm deny yeah that sort of thing but even yeah all sorts of things you really want to make sure that You're comfortable with like what can go wrong, what is likely to go right, right?
1:29:29Like all those sorts of constraints. The other side that I often talk about is just like keep an eye on, you know, if you have free form human input somewhere in your application that is triggering these things, you know, that could sometimes be risk, right? Yeah, prompt injections. Those are a real thing. And I think, you know, a lot of people are still trying to tell you what that means and how bad that can be. And so I always try to caution people about that as well, right? You really want to be realistic on how far-reaching you're doing with this. That's why I like the internal use cases. Things like that is a great way to start.
1:30:01To get familiar with the technology. To get familiar with the constraints for that. Other than that, the voice model stuff. I'm really excited to try that. I really want to. I love the secret pirate mode that they demoed. I don't know if you caught that session. I didn't see that session. Obviously, there are six voices. but there's a secret seventh mode if you add in a prompt to speak like a pirate. Love it. That was an old, I don't know if you remember Facebook way back in the day, had that as one of the languages you could select. Yes. Yeah, yeah. So that reminds me of that. Yeah, yeah. Yeah, lots of fun to be had with OSI as well.
1:30:37Okay. Well, thanks so much for jumping on. I know it's very random, but also, yeah. People love to hear from builders. So that's awesome. I love hearing from builders. And most of the interviews were done as we were sort of leaving the Dev Day venue and going to the after party. And I caught Div Garg of Multion, who we've been talking around and circling around a possible episode on. He's definitely one of the leading voices and thought leaders on agents because he's building a browser agent that's a very prominent one. Unfortunately, I have to take an L on this one because the audio is not great.
1:31:10Div's mic wasn't working and I don't know what happened to it. I try to always check these things, but you're only going to hear the output from my mic, which is slightly worse. But I opted to leave it in because Div is actually building an agent with OpenAI stuff and had access to GPT-4 Vision. And I think that people building with GPT-4 Vision will be surprised at his answer to me on whether or not it's useful for agents. Good to meet everyone. I'm this founder of Multion, which is an AI web agent that can automate browsing per year. So we can book your flights, order stuff on Amazon, order dinner, whatever you can imagine.
1:31:43Yeah. And I was actually reflecting. So everyone who listens to this already knows what was announced. I was actually reflecting that they didn't have any browser-based actions. So what were your thoughts on just generally their approach to agents? So it'll be very interesting because I feel like browser actions are just so risky. And things can go wrong. So if a company or you're OpenAI, you wouldn't want to build that. And they're better off just relying on a third party who wants to own that. And that's also the strategy we are taking with them. Like OpenAI launched a ZP integration for APIs.
1:32:11But we want multi-entreered with the new API solution. I want to do things beyond APIs. I want to connect to my personal accounts where I just have my logins already or I already have the cookies and I just want to go and interact with my personal accounts or personal data very easily. And I think this is very fascinating for us where we can launch a multi-on integration with their new platform. And then you can just go and give it a command like, oh, can you book this platform for me on chat.gbt? And then to launch a browser and the browser, you can see what's happening and then go do the whole thing for you.
1:32:40and it'll be all seamless and then people can have a lot of fun just like trying out all these different capabilities and like automating their like daily workflows. You can like save this as custom integration so for different agents you can have different custom like multi-on prompts that are already like pre-saved and then you were like oh I want to now go order something on like DoorDash I want to order my favorite burger then like Chajibit can go and like suggest you order a favorite burger and then it's like you can like now order this for me multi-on and multi-on goes and like she does that and buys that so we solve the payment for you we solve identity for you and we are owning all the risky actions that you can play.
1:33:11So you're going to build a GPT version of Multi-On? Yeah, we'll have a Multi-On GPT. Okay. Will that be like a replacement to your existing thing or just like an alternative way to use your same APIs or something like that? So it's like the direction we're going for is we want to make our AI agent embeddable within existing applications. So we're launching an API. Okay. And we already have a chat-gpd plugin. And so this will be like sort of like We'll use the API to cover this sort of new GPT experience. So for us, we actually don't have to change anything. We'll be very streamlined, just integrate our API into chat GPT, and we can start using it.
1:33:45Yeah, yeah, awesome. What about, I guess, the Vision API? I think one of the things that have always constrained browser agents is the DOM, which is very heavy. So the alternative approach is to use Vision. Would you explore that? What are your thoughts? So for us, we actually had early access to the Vision API for more than a month. And we tried it on a bunch of websites, it's maybe like 5 % of the websites is actually really useful, which are more like image heavy because 95 % of the websites, you can just, even if you do OCR, that's good enough. Yeah, it's not in the dataset. We have really good like parsing.
1:34:13So most websites, we can compress less than 3K tokens. So we are not, we don't really have to like worry about how heavy the text is. So we had one interesting use case about the Vision API. We had a user who got it to work on Tinder and then like the, then like multi-incorporation. Hot or not?
1:34:32Left and right. and the user actually got a match. Yes. I think you have found the killer use case from Altion. Yeah. Like this... Oh my God. Okay. Interesting. Interesting. Okay. But only image heavy sites. That's surprising to me. Because you know the original Vision demo, they actually showed a screenshot of Discord and they have perfect OCR. Yeah. It's true. It should be good for you. It can be very interesting. But I think it's like even without vision we can just do like so much things. So like adding vision maybe like helps a bit but not it's not like really game changing for us right now.
1:35:11That's surprising. Okay. Well good to know. Anything else that you would highlight from today? I'm just like really excited about like open air trying to become like a marketplace. Yes. App store. Yes. So if this can take off they could potentially kill like Apple app store and become like the new thing there. And it's really hard to say like how things will go. They tried this with plugins before but this like this might actually work this way. But it was just really interesting to see how two years from now, how a lot of the development might look like, how the world looks like. I'm very excited about two years from now, I think everything will be so different.
1:35:43We might not even use computers or even mobile phones. You just have an assistant. You just talk to it, and the assistant goes and does everything. It'll be a fascinating world. Yeah. So one last question before we go. You have a nice side gig teaching at Stanford. Well, you were a PhD student, but you're still teaching or curating Transformers United. Yeah, so I dropped out from the PhD, but I'm still a lecturer at Stanford. Yeah, okay. So what paper should people read to catch up on this? What is top of mind in terms of research that is informing what we're seeing? Yeah, definitely. It's a good question.
1:36:16Things are moving so fast, and there's hundreds of research papers coming out literally in a few days. I'm really excited about developments that are happening at Meta. So a lot of this work is open source on the Lama stuff, all the MISTAL stuff. I feel like that's very interesting on the Transformers side. Do you believe sliding window attention was the key for Mr. Al? I feel so for them, but I feel like there might be other ways to do that. There's some secrets, right? There was probably some secret. Yeah. Okay. Well, that's all the time we have, but thank you so much. Thanks a lot. Thanks. Okay.
1:36:42And our next guest is Louis Knightweb, CEO and co-founder of Bloop AI and organizer of the AI meetups in London, where he is a very prominent and staunch member, unlike Raza, who has defected to San Francisco since our last conversation. Louis always has very interesting takes in person, and it was a pleasure to finally actually get him to come on the pod. But also, we recorded this while inside of a Waymo on the way to our after party. So, Louis, you are new to the pod, but we've been friends for a while. Maybe explain, maybe introduce yourself and how you come to the world of AI. Yeah, I guess.
1:37:18So we started Bloop, me and my co-founder, three years ago in a very different era for machine learning. And we both started the company because we wanted to help engineers navigate large code bases in a much better way. Yeah. And originally that was training our own models to do natural language to code search. And today we still do that, but obviously those language models are very small compared to the state of the art. Yes. And so they're just one part of a much bigger pipeline. I see you as a very astute technologist. You used to be a VC. You wrote the first check into Human Loop and you used to share an office with Human Loop to the point that I called it Human Bloop.
1:38:03Yeah. I think you that yeah i did yeah that's good we're considering renaming uh and you also run ai tinkers in london i do yeah london has a kind of a slightly different mix of of talent than say san francisco you've got a lot of agencies a lot of enterprises and so yeah we we just felt a need to start like a very startup focused event and that's why we we created ai tinkerer london yeah i think alex gravely will be very happy to hear about other stuff that you've been doing. And I've been to one of them and it's really good work. I might be the only one that's been to both. Yeah. I've been to both as well.
1:38:40Okay. So let's fast forward to today. A whole bunch of things was announced. What's top of mind for you? Yeah. So I think like context length is something that we spend a lot of time evaluating whenever something new drops. All of the kind of standard evals, you know, the kind of literacy tests things like that, they generally don't do a good job of measuring whether a model can actually use the context length that it claims it has. Yeah, context utilization is what I saw Will Depew today call it. Exactly. And so this basically started maybe five months ago over the summer when Claude2 dropped.
1:39:20And, you know, obviously it had 100k context and we were really excited about that. So we ran an experiment to see basically if we hid 10 pieces of information in the prompt and we increase the size of the prompt, you know, so you do it at 1 ,000 tokens, 4 ,000, 8 ,000, et cetera, up to 100 ,000, how many of the original 10 pieces of information can it retreat? And we essentially found that the accuracy drops off a cliff between one and 10 ,000 tokens. And so, and we repeated the same experiment with GPT-4 and, you know, we found similar results that 32k gpt4 can only find one of the 10 pieces of information but if you are only using a thousand tokens it can find nine of the pieces of information so what that tells us is that you know context utilization five months ago was was was not great with with all of the state-of-the-art models so with the announcement of 128k today and that's the first test you'll run that's the first test i'll run and they are you know having spoken to a couple of the team members who do eval today from OpenAI, you know, they're pretty confident that the model's got better ability to answer questions at those context lengths.
1:40:29So it's time to measure. Time to measure. Any other of the API features? Reproducibility, does that matter to you? I think, to me personally, no. I kind of like the creativity. I normally have my models at like, you know, 0.1. Yeah, exactly. And a bit of temperature. But I know lots of people on the blue team he'll be very happy i'm sure and then i guess the json features the there's so many like the multimodal features any of that appeal for for you personally even if json is is definitely a big one i think it allows you to to kind of standardize how you call different models yeah so instead of having to build you know the and it's not a massive thing to build but to build the the the kind of function calling integration and then if you want to try bi-anthropic, you've got to go and have a completely different way of interpreting the output.
1:41:19So if you can just stick with JSON across all of your different LLM providers, open source models included, that's definitely, which just allows you to evaluate different models more easily. Yeah. Yeah. Very excited about that. You are, so you compete in a pretty competitive space with the code assistance, code search, code assistance, right? There's Sourcegraph, there's Codium, there's other call the MNN, it's co-pilot and so on. You've never ventured into the agent side of things. Yeah. Is that a conscious strategy? Are you waiting for the right time? Are you waiting for the right APIs? I think, I mean, we're seeing traction at the moment with companies that have very large code bases, right?
1:41:59And it's not something we hear from those users that, you know, when we listen to their problems, it hasn't been like an obvious fit to try and build, like maybe an auto GPT type of agent. I'd still say, you know, we're very interested in agents. The pipeline we have at the moment, it's basically GPT in a big while loop with function calling, which, you know, like nine months ago definitely did count as an agent, maybe less so now. So, you know, it's just customer and problem driven. And we don't, you know, it's not a hammer for the nails that we've got. Yeah. So two comments on that. One, I think OpenAI has sort of put their flag a little bit in the definition of an agent.
1:42:40They had three things, right? They had custom knowledge. They had custom instructions. And then I forget the third one, custom tools. Let's just say. So by that definition, we're doing... Yeah, so we've been doing that since about February. That's the definition. Then the second observation I would say is you talk to developers, but what if the target customer for agents is not developers? It's the PMs, right? So we definitely see a lot of PMs. using the product or people that are defined as like reading more code than they write. So, you know, it could be designers trying to understand the implications of an interaction.
1:43:18It could be PMs trying to fact check a contentious time estimate from a developer or something like that. Low trust environment there. Talking from, I've seen some stuff. Egregious things, yes. Yeah. So basically, it's still not that appealing for you, but you'll keep a lookout for it? The staples? I think based on the definition OpenAI, you know, released today, we tick all the boxes. And I think we were one of the earliest adopters of that, if that's the definition. You just don't brand yourself with the agents? I don't think it's important to users. I don't think that's why people use the product.
1:43:57I mean, we're very solutions focused. I think a lot of our branding at the start of the year was about models. And, you know, we put GPT-4, GPT-3 right there on the front page. And now, you know, we've kind of reoriented to be more about solutions. I think that reflects kind of maturity of the ICP we're going after and where we are with sort of stage of company life. Yeah, yeah. Cool. Any other things that you personally, not Bloop related, are just excited by, interested by from today? Any interesting conversations with others? Loads are really interesting ones. I had a fascinating talk with some safety researchers.
1:44:39They were here? So there's a couple of PhD students who had looked at adversarial attacks through fine-tuning of models and found that basically it's such a hard problem to solve. If you enable fine-tuning, it's basically impossible or very difficult to make it so that you can't disable all the safety features. You can just train it to spit out all sorts of stuff. So that was pretty fascinating. I'm excited about the Waymo we're in right now. Oh, yes. So we should tell people we're recording in a Waymo. Haven't been looking at the road the whole time. Is this your first Waymo? It is my first Waymo, actually.
1:45:20Yes. Thank you for taking my Waymo virginity. But I've experienced this together. I've been a cruise stan the whole time until they ran over someone. So my take on cruise, like at sample size, 10 cruise journeys before they got shut down. and three of them resulted in something popping up on the screen saying that I had been in a collision. Did they use the word collision? Yeah, yeah, yeah. That's surprising. I'll show you after that. I got pictures of it. I took a fair amount of cruises and I didn't, yeah. And so it was the same situation almost every time, which was a car was in front trying to park.
1:45:54And I think they just maybe bumped fenders or maybe the crash detection. Oh, there was actual contact. I think in one of the cases, I think there was. In the other two, I didn't feel anything, But it came up saying, like, you've been in a collision and somebody comes over the intercom. Checks if you're okay. So, yeah, I mean, out of 10 rides and three of them ended like that. So I think, yeah, definitely some questions there. But this way moves pretty smooth. Maybe also we're in a better neighborhood for driving because we're going to Golden Gate. The time of day, that's a really good point. I noticed that all of the ones I took at night, all of the cruises I took at night were fine.
1:46:29And when I took one during rush hour, it was a completely different experience because the routes it would take, it had this really aggressive, maybe traffic management, something that was going on. So it'd take a long time to get from A to B. Yeah. Yeah. Yeah. It often puzzles me slash interests me that self-driving is almost solved. We still have some bumps in the road. Sometimes the bumps are human. It's solved in San Francisco where you've got wide open roads, nobody cycles. and that's that's not true some some people's like i live here excuse me some people's like okay i mean compared to like okay compared to london where you've got you know roads half the size built for horse and carriage and millions of cyclists and buses and all sorts so i think you know it's going to be a long time until we have that same experience that of a cruise or way mo today london well i understand london's a tougher neighborhood uh but still you know we're 80 there 75 80 there whatever right but like and it seems like the the stuff that we do in the rest of our lives in terms of ai automation is so primitive compared to this which is the car that we're sitting here right now and i find that weird i find like the relative ease or the relative like here-ness of this technology is very disparate like how come it didn't trickle down from self-driving to the rest of tech yeah it's it's interesting isn't it well i don't know how those pipelines are built i assume that's the secret sauce right but the flip side of that argument is like maybe it's very scary that we know like now many more people understand the the mistakes that these these types of systems can make because we're all getting hands-on with GPT and this system is equally as problematic and we're just oblivious to it because it's a black box.
1:48:23Almost at your drop off. Check the app for walking directions. Okay, Waymo. All right. Well, I think, yeah, that's probably all right. But thanks so much for giving a quick review. And thanks for having me. Yeah. Yeah. So that was Louis, whose opinion I think is very reflective of the people who are building code generation or code search type startups based on top of GPT-4. And as we headed into the Dev Day venue, we actually caught Shreya Rajpao from Guardrails AI. And there was an interesting comparison here in our conversation between how she views the LLM stack versus how OpenAI views the LLM stack.
1:48:59OpenAI actually had a closed door session where they gave some thoughts on how they felt that people should start from prompting and build up into a full software system. And they actually deferred a little bit from Shreya. Don't worry, all that's recorded. The videos will come out in a week, but you can listen to Trios' take. We're reviewing AI Engineer Summit. Yeah, we're reviewing the AI Engineer Summit, and it was a very, very well-organized conference. And a small thing that I was thinking about is that your swag for speakers. Is it on? Okay, it's on. Your speaker swag was, like, not surprisingly, I guess, but, like, really weirdly very nice.
1:49:35And it just kind of, like, showcases this attention to detail that I think, like, really kind of permeated the entire, you know, conference. like every single decision was very well thought through and you know kind of like to a degree of like quality that's very rare to see so yeah it was it was amazing i thought you guys did like an absolutely fantastic job yeah this one mostly goes to ben so i'm definitely gonna make sure that ben understands that i really appreciate the work that he does and this is why i couldn't do it myself you know i'm mostly the content guy but i don't he's the logistics and he's run conferences for eight years so that's why i keep working with him yeah i also kind of really enjoyed the 18 minutes you know really yeah yeah when i saw that i was like huh is this going to be you know it's going to be enough and like is that but it was like yeah yeah yeah yeah i i think the 18 minutes was actually the right kind of bite size it's optimized for youtube yeah i see interesting okay because it's not the in-person audience that matters i see i see i see interesting okay i need to promote my my video more uh yeah is yours up yet i don't think it's up yet it's not up yet yeah we're releasing we're dripping them out to spread it out.
1:50:37Sounds good. So yours maybe in two weeks from now. Okay, sounds good. Okay, so welcome back. Thank you for having me. I think you were guest number five. You were super early. So we're at the after party now. How do you feel about the whole day? I'm really excited. I think it was... Yeah, I think the excitement in the air with everybody just waiting with bated breath to see, I guess, what gets destroyed but also what gets really optimized. I think this is like very, it feels like you're really part of a movement. And as Shannon, we were like, you know, us like early people in this space, we got to stick together because like whatever happens to any of our companies, you know, there's such a like, there's such a transformative moment in technology that, Yeah, so you don't care, right?
1:51:20Yeah, we're all going to like look back on this time. But I had a blast. Like I really, really enjoyed the releases. Yeah. What got destroyed? What got destroyed? I'm minding for hot takes here. Once again, And I think my takes are unfortunately very measured. I wish I had spicier takes. Your takes are within the guardrails of common behavior. I think retrieval is the big one for me. I think it's kind of really exciting to see the retrieval baked in. And that's one thing where I'm very interested to see, does that pattern become common by model providers? A, by commercial model providers and also by open source model providers?
1:51:59And how much of retrieval do you have to do yourself? and what remains challenging about retrieval compared to just this really easy API to just have it done for you, right? Yeah, I think what they did was effectively build the basic patterns in. But for the more advanced stuff, you're still going to need Langechain, Lamaindex, all those. Yeah, yeah, yeah. So for the longest time, I believe that in RAG, it's the retrieval that's the hard part, right? Yeah. And then generation is really easy. As long as you're good retrieval, you can get really, really far and the generation only gets you like a little bit over.
1:52:31And so I'm really curious to see like, okay, how, once again, like how complex do you need it to be in order to start seeing good results? Yeah. Okay. Interesting. And what are your normal benchmark tests like? Do you actually have a set of tests that you run whenever you're like exploring something? Or some personal favorites of like use cases that you think are tricky for LLMs to do well? I think like a big focus of ours is on hallucinations. always kind of like checking out hallucination and like conflicting instructions, et cetera, is one. Terse responses is another, you know, like how well is it at like not, you know, you ask it a question and here's this 10 point list and you know, very, very verbose.
1:53:09Do you have a terse response as a validator? Yeah. Well, we don't have it like, we don't have it publicly, but like we do kind of like check it. Yeah. So I think like those are kind of some of the things. There was one, there's one example in one of the closed door sessions where they, one of the answers were two terse. Yeah. Where I think everyone were laughing when they were like can you write a blog post about this and the guy and the gpt said sure i'll do it tomorrow yeah yeah yeah yeah i think like those are i think those are i'm really really excited about yeah just check i'm really really excited about json generation i'm actually kind of surprised to see how long it took them to get like it's they're probably just doing constrained decoding under the hood right like constrained generation okay because they're now saying that guaranteed correct json rather than you know more correct do you get what i'm saying I was parsing through their words.
1:53:55They've never had an issue producing JSON. It's just that sometimes it doesn't fit the JSON schema. Right? Am I wrong? You would know better than me. No, I think there are also issues with like producing. I think the obvious thing is like... Unbalanced brackets? When it's on context length, I think that's like an obvious thing, right? But like weird things, when you have like really long strings, then quotes, et cetera, become kind of weird. So I think those are some other ones. Schema is obviously kind of challenging, et cetera. Yeah. I think there are, even with function calling, like function calling, at least I haven't played around with it yet today, but previous generations of function calling wouldn't guarantee that your schema is matched, which would be an issue.
1:54:33And I think they're still not guaranteeing it because I kept waiting for them to say it. I haven't read any of the public docs or anything. Do you know if they're guaranteeing that it fits a schema? Oh, that's a good question. Yeah, that's a good point. They never said they guarantee. Yeah, they never said they guarantee. They guaranteed correct JSON. They didn't guarantee if the JSON matches the schema. So, okay, you can call JSON loads. Yeah, yeah, yeah. I'm very curious to see, once again, if this is a pattern that all of the other foundation model providers adopt. And I don't see why not. I think for them to own specific decoding models is going to make a lot of sense compared to a lot of the hacky stuff.
1:55:11Yeah, cool. Any other favorites? Doesn't have to be guardrails related. Any favorite conversations, favorite demos? favorite i oh the gpts and the assistants i think you want to make one for yourself yeah i do want to make one for myself it doesn't add like yeah not very godreels related i do want to kind of play around with like how well it works with like some of the things we track but yeah it was just so fascinating to see the marketplace i am very very curious to see you know what the marketplace looks like like is it are people going to have like really really vertically specialized things on the marketplace like if you have a generic you know sales assistant or something right like Like how much our SQL generator, how much, how popular does that become versus like sales assistant for X vertical at Y stage of the sales process?
1:55:55Oh my God. Do you know what I mean? Like it's, it's so easy to do this now. Yeah. That like where, at what level of specialization do you need to be to kind of start seeing the results? And that is one thing I'm very excited to see, like how that, how that pans out. It scares me a little bit because it's basically, they say the future of programming is natural language or something like that. Yeah. And that's great, but it really is a new platform, a new operating system almost that they're creating. And I don't know how to position myself. Not that I have to, because my world is very developer-oriented.
1:56:26But this is a whole no-code world that you and I don't touch. Yeah, yeah, yeah, yeah. Whoa. Yeah, yeah. I really want to see, is there just going to be assistance for everything? I'm generally curious to see the impact of this on knowledge work. Like, you know, which, yeah, like how much of my work, like if I'm getting annoyed by something, is my first instinct going to be like, you know, let me just, you know, spend the five minutes to build an assistant for this? Like, is that how everybody's now going to start thinking? You know, and that's one thing I kind of really want to see. Yeah, that's exciting.
1:56:57Okay, last question. You spoke at AI Engineer Summit. Let's advertise your talk a little bit and point people to your talk. Yeah, so thank you again for inviting me to the AI Engineer Summit, one of my favorite conferences that I've attended, you know, this year. My talk was about the new paradigms for working with large language models, for building really production-ready applications when the technology that you're working with is underneath all of it, non-deterministic. Really fascinating thing, which was the OpenAI's talk about building production-grade applications talked about how essential it was to build guardrails as a way to make it to product.
1:57:32You're talking about the one from today. Yes, the one from today. Which people haven't seen yet. Which people haven't seen yet, but really, really cool talk. So I think it really validates what we've been saying pretty much since the beginning of the year, which is that you'll get to a certain point, but at that point, you need to start adding guardrails to your application if you need to get your users to start getting value out of what you build out. I have your chart and I have their chart. They put guardrails at the first layer. It's not at the end. It's actually right at the beginning for user experience.
1:58:06Yeah, that's right. Yeah, that was kind of interesting to see that they put it as part of the UX. I'm still kind of very candidly, I'm still kind of digesting that. Like I think of it as part of the infrastructure. And I don't know if it's as much UX as it is, you know, just like one of the components that you need in your stack. But I think a lot of what they said today completely validated, you know, what we've felt for the longest time. And also what I go really in depth about, like in the talk that I gave, right, which is that what happens once you have the bare bones application ready? What is the process of actually adding guardrails for what you care about?
1:58:41Like, what does that look like? You know, what are the risks that you care about? How do you verify that those risks are happening or not happening? If they are happening, how do you quantify them? And then how do you mitigate them? That was what the talk was about, which I would really recommend people go and check out. Awesome. Well, you did a great job. We're going to post the talk soon. And thanks. It's good to see you again. Thanks again for inviting me. And that was about all I managed to get before the after party. At the after party, there was actually an after after party thrown by Noose Research.
1:59:08So let's hear a little bit about OpenAI versus open source AI from Alex Volkov. Okay, so we are in the one day after Dev Day here with Alex. Hey. Hey. Very, very recognizable voice right now. We don't have to introduce you. Hey, everyone. And we are here to talk about the two parties that happened yesterday. There was one official DevDay OpenAI after party where I interviewed Shreya, who was just before this. And then there's an unofficial one for keeping AI open by Noose. So what was it like? Just compare and contrast. So let me maybe start with who Noose Research is. Oh, yeah. Most people haven't heard of Noose.
1:59:42It's written N-O-U-S. I mispronounced it now multiple times. It's Noose Research. It's one of the few organizations online that started from a Discord and then kept going up until a significant amount of people are working with them, affiliated with them, of folks who take open source model to its most extreme capability. So, collect data sets from open source, open source and more closed source, and depending on that, they release different licenses. And then they fine-tune open source models that were released to us from Lama, for example, and Mistral, which is a French company that recently released a 7B model that's the best.
2:00:17And they've been doing this since Lama 1, but recently it really kicked into high gear with Lama 2 releases because Lama 2 ended up being with a commercial license. So you could actually use this for actual, you know, products and services. And Mistral came out with like a full Apache 2 license with a BitTorrent link. I think you remember that. And so these organizations suddenly became like a very... very important currency in the world of like where the whole world of AI is going because they're lining local models and many companies love open ai but either cannot afford this or cannot risk the chance the open ai changes something like we saw with dev day and so many people are turning on to like okay if we want to run our own hardware how do we actually do this and you can run it you can run llama 2 and all these models on your own hardware but then you want to fine-tune them for your own purposes and so how do you actually fine-tune and now organizations like news research was probably the biggest one alignment labs shout out to austin and folks from from alignment labs skunk works and many of these like people come up and say hey we have the know-how and we only started learning about this like eight months ago six months ago themselves but now they're like the specialized more people that fine-tune models and actually release the best kind of models on the hug and face open source leaderboard yeah yeah and in in my knowledge all two models that i keep hearing about one is hermes and they recently switched the base model for hermes from llama to mistral because apparently it's better yeah her miss is like an instruction data set 900 000 instructions i don't really know where it's from maybe i don't want to know they also do some like fun models there's like a mystical model that they do some some stuff like that i think it's actually a little bit weird that they keep releasing models like they release like three models a week it's insane right and it's very hard to keep up like i'm like okay which one is actually the one that i should pay attention to yeah so first of all you're welcome to join thursday i and then we talk about all the models every week it's kind of interesting to that if i do like a recap for a month the beginning of the month most of the updates don't matter because like every every this i'm doing monthly and i i feel this like i'm doing this i'm doing this for historical posterity like yes five years from now people want to look back then they can look at my notes because i only have 12 a year yeah nobody's gonna look at your notes they're gonna have a gpt trend on your notes answering everything i have yeah i'm doing like every week and every week we're talking about like this model outperforms that model like significantly and we're noticing significant changes from week to week literally in the spend of a month we went from a 33 billion parameter model which is big and parameter count is not everything there is right you can have a smaller model with like larger longer training that actually will perform better than whatever but we're noticing smaller and smaller models doing outperforming bigger ones significantly zephyr from hug and face outperformed llama 70b and zephyr is like only like a 7b model on some things on some things for sure and so that's very interesting because like it's really hard to evaluate evaluation frameworks are bad everybody's saying that they're not representing of anything people can fine-tune overtune on them and so there's this whole kind of subculture of open source mostly on discord some of them on on x and twitter spaces and for some reason but i find it very like humbling and incredible they also hung out in thursday eye and so that's how i got to this that's how i got to meet like news research folks technium emozilla and organized the the counter party event last night together with some other eac people that we know from twitter as well including mark adjson so apparently he was supposed to i didn't see him uh but like i saw a photo with a bald head of a big guy so i was like is that mark i don't i don't know anyway but the opening eye party was at a art museum and then the news research party was at a club it was a club yes at Folsom Street in San Francisco, a club.
2:03:5710, 15 Folsom, I think. OpenAI was a very highbrow, buttoned up event. There was a live band, someone playing jazz. Which I think I mentioned this once. It was too loud. We want to talk. We don't want to listen to music. No, no, no. We're just old. Everything is too loud. And then it was like a lot of people, a lot of networking, a lot of people trying to get together, maybe do business together. Very awesome. Many people from OpenAI actually showed up. a lot of people. We stood in line, there was a long line for the magneters to step in, and everybody passing us around was like OpenAI employee that passing straight through.
2:04:33And then that ended around eight, which is like the standard San Francisco buttoned up. Oh yeah, that's when you go to bed. That's when you go to bed. And that's when the other party kind of started. And I think they just seized the opportunity because everybody's in town for the OpenAI stuff. Why not make a splash an announcement for like for open sourcing ai so literally the invite was keepaifree.com yeah which was the website and the email open.com yeah and you had to register you had to go in there and this was to me an incredible kind of show of twitter in real life so all of the folks who follow mark andreessen he recently stepped into this thing with like the techno optimism and stuff he started to boost the e uh effective accelerism yeah uh folks and so there's a lot of like signature stuff from that like uh ecosystem on twitter there's like don't thread on me with like you can take away my gpus there's like all these signs across the club the it's a very visual club as well so where the djs is a whole like a 3d projected thing so there's like a bunch of like art and like live things about keep ai open i found it like very very super cool i'll i have to tell you a tidbit i saw me and killian were there from open interpreter we saw two people with lab coats it was like what's the deal with lab coats so we went and asked and they just said hey we just like we came back from our work where we work on semiconductors we're actually like touching chips whatever we just like didn't change out of it and my head was like so incredible in the keep ai open gpu kind of poor party we have people who literally work on superconductors came from the work like they're working on chips yeah semiconductors or superconductors very different things i think semiconductors yes yeah yeah we uh we had the superconductor episode uh a while a while back i think people still recovering i'm i'm personally still recovering from that that was the whole thing for me yeah so is news research like vibes uh you know like what what is the mission apart from to keep publishing open source models i think you'll have to get some news people to actually speak like about the mission about the actual product but as as far as i understand this no matter how much the product side will be and there will might be there's so many people they're doing like so incredible stuff that people notice like and you know so no matter how like how much of the the business side will be they're like committed to fully open source as much as possible including data sets including models that are like trasmestos for example their model that's like trained on the occult and the physical metaphysical or you you can't expect open ai to let you talk with a model to answer with like mystical questions mystical stuff astrology halloween so you you're very like easy into the astrology and halloween they're talking about like you can ask this model about like resurrection right like all of the occult like craziness they've collected open air will not let you do that and so there's i think when i will not let you do by default because they have lawyers and they were doing it sued yeah recently they announced the protection shield thing so you won't get sued because of their model so they're them and tropic all these big companies it's very important for them to protect the outputs and the models here these folks are like hey if you want to build a model fine tune this we're going to teach you how jump on our discord we're going to help you with producing like the biggest models and then if you know there's going to be like a financial aspect to this as well if your company wants to run this we'll also help you do that yeah so it's the same as stability basically it's that's from what it from talking to him that's what i gather yeah cool anything else that people should know about the party news i found this whole day to be like a very singular ai day and we don't get many of this gpt4 i think was the biggest one previously it was like march like a single march 14th that's what thursday i started we started talking about this every week this was a singular day in san francisco this like started pre-game party with swix and some other folks that i i got to feel like a little bit of san francisco and then dev day was incredible we just heard from simon there was like a garage that they made into a venue event yes probably custom venue event on the fly which like just talks to how much uh they can pull off it felt to me that like this dev day event and then the following party it felt a little bit like almost like an apple thing yeah we're like it's going to be a yearly thing that people will like try to get in as much as possible one thing to note that in the other party there were many people who didn't get in to this party and so you know they were watching for like a party this office right here this office people watched here and people watched in in the live space that we we 8 000 people tuned into our spaces 8 000 people tuned in i didn't even have a chance i always want to know the number oh so it shows the relative level of interest and you know like so according to 22 000 this is 8 000 just relative interest yeah there's like two spaces as well robert scobar he he stole the thought he stole some audience from us yeah shout out robert and i think that like it was a singular day and I think the news research keep open source open, EAC, Mark and Driesen, all these things together also added to the top of this because it happened in the same day, one on top of another, in the same place, San Francisco.
2:09:28I find it incredible. I would definitely come back next year to it. I think you'll be back sooner than that. There'll be other things going on. Thanks. Awesome. Last but not least, we go back all the way to the Newton where I started this podcast where we checked in with Rahul Sanwaka. better known as Rahul Ligma, who just celebrated his one-year anniversary as one of the biggest memes and celebrities in San Francisco. But by day, he's also the CEO and co-founder of Julius AI. And then I'll match it up. What's up, Swix? Hey, good to see you. It is one day after Dev Day, and we all had a chance to process.
2:10:07How do you feel? What's your top takes? Dev Day was awesome. We got to see a bunch of really smart people who are building cool things with OpenAI, GPT, Dolly. The event was very well put together. The keynote was awesome. The energy in the room was crazy. And I could see real-time social media firing up with all these takes. Overall, I think it was a good day. Yeah, I interviewed Surya Dantuluri. Yeah, I think you know him. He was like, Sam just killed my startup. And it was almost true for him because he has a bunch of plugins and plugins are kind of deprecated. Yeah. Yeah. Yeah, the plugin thing was interesting because it's going to be deprecated, but they just accidentally turned it off yesterday.
2:10:51Yeah, so you freaked out a bit. It freaked out, and then they brought it back up. Yeah. Yeah. So, top features that you're interested in that you want to explore more? I think people are super psyched about the Assistant's API, but personally, if you ask me, two things that I am most excited about is Turbo. Yeah. The speed is crazy. Have you actually, have you met, you know, do you know any like rough measure? Because I don't think they actually ever mentioned the speed relative difference. I started noticing the speed difference in ChatGPT actually like a few weeks ago. Oh, I see. So they already slowly eased us into it.
2:11:25Yeah, yeah. And I saw like takes on Twitter that, did anyone notice ChatGPT get much faster? And I noticed it too. Yeah. But, so it's Turbo, it was exciting. But the second thing that's exciting is multiple function calling. Yeah. And then the JSON output formatting. I think as developers are building on the dev API. So that's the thing that's super exciting to me. You know, of course, there's version stuff. There's code interpreter as a tool in the API. Yeah. But I think what will bring the most applications is actually the speed. Because there are so many things. If you look at our numbers on Julius, people are not patient.
2:12:07They want an answer and they want to answer quick. And we see clearly if you can get an answer to them a few seconds faster, there's a clear difference in the conversion. So speed is going to be paid. What is conversion for you? Is that just paying? Oh, no, it's like from first message to second message. I see. So we do code gen and then we run the code and then the code has an output. The user asks a second message and we can just see the funnel. Yeah. Where if it's faster, the code runs faster. And the second thing is multiple function calling. I think you're basically telling the AI that, so I think the people misunderstand functional calling.
2:12:43It's essentially tool use. And if you can tell the AI, hey, you can give me multiple tools to use at once, I think that's going to unlock different applications than before. Because before it was just like, okay, this is a task. Tell me one tool and what's the input for it. but if the AI can now use multiple tools in parallel you can first of all have more specialized tools and then the AI has more specialized instructions for each tool it's just going to unlock a lot of cool applications that previously weren't possible there was a practical limit in the number of tools that you can give it, right?
2:13:17so we had this discussion in March, February March, April when they released the Function API that this is subject to context window the JSON schema itself does that change at all? I don't know if you I don't. But what I noticed, though, even before, was that more functions and more options just confused it. And that's what I want to play with next. It's like, okay, what's the breaking point? I see. Like, does more options, you know, confuse it? Does it make it? Would you use multiple function calls as well? Oh, totally. Is that just theoretical? No, no, no. I have a direct application for it right now.
2:13:52One of them is oftentimes Ghibli writes code. And then we run that code and we realize that, oh, from GPT's last knowledge update, that module in Python has changed. It has new functions, new APIs. So today, the way we do it is when the error happens, we tell GPT, okay, you can go look up new documentation and then fix that error. But with multiple function calling, the way we would do it is like, give me the code, but then also give me a documentation lookup. And then when the error happens, I can just quickly fix that without another GPT call. Yeah. And then keep moving. Nice. But I mean, in general, it's just like multiple to use to me.
2:14:32It's just so exciting as a developer. And I wish people were talking more about this. Yeah. I mean, people are still coming to terms with just like the base model and prompt engineering and all that. That's still important. But for engineers, I think you should explore these other advanced features. True. Yeah. Anything on the multimodality side that you're interested in? I mean, originally will be super interesting for sure. and we have this functionality in Julius right now where you can generate React and HTML opponents. Like V0. I think Matt was showing me a little bit of that demo. Yeah, yeah.
2:15:04We've been hacking on it a lot. I think the missing piece here is that, well, you have an engineer who knows how to react and they probably wouldn't find this useful. But if I can allow anyone in the world to just draw a mock-up on a piece of paper and then run that and have the vision Yeah. What they've demoed. Yeah. Yeah. Turn into like actual components I could use on a webpage. That'd be sick. And what's even more sick is like have the feedback loop where you take a screenshot of the page generated and then feed that screenshot back into Vision and then come up with more instruction and have that loop.
2:15:40Yeah. Wow. Like a self-improving webpage. Isn't that crazy? Yeah. I'm super excited. Yeah. Yeah. So in my mind, Julius is very data focused. By the way, I didn't introduce you. I didn't introduce you I was just going to do it separately Yeah But people know who you are Yeah You have a Wikipedia page Yeah You just passed your one year anniversary As Rahul Lingma Thank you By the way Any fun things happen On the anniversary What are the fun things Ilya said Ilya recognized you on the squad Oh Ilya was like Oh my god This is Oh you're famous or whatever And No these guys are so awesome Like they're so humble But anything happened The first one year anniversary Nothing really Like it's I mean you knew about it A week before I like to set Anniversary dates That's awesome Because it reminds people of the passage of time.
2:16:22It's like, wow, shit, has that been a year? Yeah. And then you're like, I think it motivates me more than memento mori. Yeah, in this case, sometimes you're out of date. But it reminds me to spend my years wisely, to do interesting things with the time that I have. Memento mori is kind of depressing, whereas this is… This is like, oh, yeah, did you know one year ago we had this thing? Wow, it's been a year. Yeah. Okay, cool. But Julius, you data analysis chat thing. Yeah. Basically, Code Interpreter++ is how I think about it. Exactly. And also, you just cross 100 ,000 users. Yep. You have delivery modes across your plugin as well as a chat box, like a dedicated web app.
2:17:02Yep. Okay. Anything else that people should know? Well, the origin is, you know, writing code is super fundamental to doing things. You could not only automate a bunch of tasks in your life, which is writing code, but also it's how you just, like, interact with the universe, right? You have code that brings you a Waymo car and picks you up and just drops you out somewhere. And I think allowing these language models to write code and do things for you is really powerful. And data announces this application that we're most excited about right now because that's what it's good at immediately. But just on Friday, we launched FFMPEG support.
2:17:39And there were people trying to upload videos, turn the videos into GIFs, or like take a YouTube video, turn it into a short summary. and all these different use cases that we didn't truly like hardcore into Julius. We just told it, hey, now you can run FFmpeg and you can run ITDLP and MoviePy and all these different things. Do these tasks for me. And then people were just like organically discovering those things. There's this guy TDM on Twitter, CTOJr. And he took some meme video and put it on my own tweet, overlaid on my own tweet, and then I tweeted that. And then that got a bunch of likes.
2:18:14And I was like, dude, Like this is the first one that gets a lot of likes on, you know, FMPEG on Julius. So that's a lot of meme potential. It's a lot of meme potential, but that's not what we're going for. Yeah. You know, it's just like letting people like do things. Your target market is like the S &P, the enterprise. It's actually individuals who have data on it and they just want to drop academics. A lot of academics, actually. Yeah. A lot of academics, a lot of students, researchers, any kind of CSV, Excel data, you can just dump it into Julius and have it analyzed for you. We have this video coming out in a few days where you can now actually train a nano GPT on Julius.
2:18:53So you can give it, hey, here's the GitHub repo for Carpet. So you, yeah, it has you have GPUs to train it on or you just train it in CPU? CPU. Yeah, it takes a minute. Yeah, that's true, that's true. Yeah, I mean, Carpetty will like that. Yeah, yeah, yeah. Okay, cool. So the thing I really want to sort of ask you as a founder on is you know, I think there's always this existential threat about OpenAI building your features in a way so like the number two default bot in the in the gpt app store yeah is data analysis yeah and people can build their own by customizing and adding code interpreter yeah although i think there's also opportunities for you so on the roadmap that they presented in a closed session they also said you can bring your own code interpreter yeah so like how are you thinking about that i mean as a founder or as founder as so who's the audience is it like other founders or is it?
2:19:43Yeah. Other founders and people just interested in how you're processing this. Yeah. I mean, I think it's a very interesting story of processing this live because the news just dropped yesterday. Yeah, totally. Well, so the story behind Julius is that we actually launched Julius three months after Code Interpreter was announced and a few weeks after it was rolled out to everyone else in the world. So we were number two. And even then we got 100 ,000 users because I think there's a lot of work to do to get something to work properly. And there's a bunch of examples of this on the internet. So if I'm talking to founders, what I'll tell them is, man, so many people give up before even getting started.
2:20:25And that happens. Don't do that. Sure, you can change your idea. You can find new things to work on. But the way I'm processing is that we launched after Code Interpreter came out. And there's 100 ,000 people who think Julius is better than Code Interpreter. Or you just tried it out. Yeah, I'll try it out and use it over Code Interpreter. And there's like a lot of work to do. Like, for example, the FFMPEG stuff we launched on Friday or the HTML stuff, you know, React component stuff, all these different things. To get them to work, it takes some effort. How I'm processing it, I mean, you know, that's what startups are all about.
2:21:03It's like risk, right? If you want to build a risk-free startup, you probably don't want to work on startups. Yeah, just go get a job. Just go get a job, exactly. So I'm having so much fun. The way I'm thinking about this is like, whoa, there's all these new different things I could do now. I could build. That's so exciting to me. And I'm pumped. Yeah. Yeah. Awesome. That's it. Any last words? Call to action? Call to action. Let's go build some cool things and get a bunch of users. Let's do it, guys. Yeah. All right. Awesome. Thanks so much. Thanks, Wix. I think that's a meme that we can all get behind.
2:21:36Let's go build things for a bunch of users with AI.
From the publisher
We left a high amount of background audio in the Devday podcast, which many of you loved, but we definitely understand that some of you may have had trouble with it. Listener Klaus Breyer ran it through Auphonic with speech islolation and we figured we’d upload it as a backdated pod for people who prefer this. Of course it means that our speakers sound out of place since they now sound like they are talking loudly in a quiet room. Let us know in the comments what you think?
Timestamps
the cleaned part is only part 2:
* [00:55:09] Part II: Spot Interviews
* [00:55:59] Jim Fan (Nvidia) - High Level Takeaways
* [01:05:19] Raza Habib (Humanloop) - Foundation Model Ops
* [01:13:32] Surya Dantuluri (Stealth) - RIP Plugins
* [01:20:53] Reid Robinson (Zapier) - AI Actions for GPTs
* [01:30:45] Div Garg (MultiOn) - GPT4V for Agents
* [01:36:42] Louis Knight-Webb (Bloop.ai) - AI Code Search
* [01:48:36] Shreya Rajpal (Guardrails) - Guardrails for LLMs
* [01:59:00] Alex Volkov (Weights & Biases, ThursdAI) - "Keeping AI Open"
* [02:09:39] Rahul Sonwalkar (Julius AI) - Advice for Founders
Get full access to Latent.Space at www.latent.space/subscribe




