In short
Podcast Notes: The AI Daily Brief - Code Interpreter is GPT-4.5: A Summer AI Technical Roundup
Episode Overview In this episode, NLW is joined by Swyx and Alessio from the Latent Space podcast to discuss key technical developments in artificial intelligence over the past month. Topics include the Code Interpreter, Llama 2, advancements in AI agents, the rising interest in AI companions, and open-source developments.
Key Participants
- NLW: Host of The AI Daily Brief
- Swyx: Host of Latent Space, AI developer, and entrepreneur
- Alessio: Host of Latent Space, venture capitalist focused on AI
Key Topics Discussed
- The Code Interpreter and Its Significance
- Overview: Initially presented as a ChatGPT plugin, the Code Interpreter is considered to represent a significant advancement in AI capabilities, akin to a new model (GPT-4.5).
- Functionality:
- Acts as a sandbox for running and testing code.
- Allows users to upload files and perform operations, enhancing its utility beyond coding tasks.
- Users have found it effective for non-coding queries, suggesting it may outperform GPT-4 in various tasks.
- The Shift in AI Safety and Regulation
- Context: Discussions surrounding AI safety have intensified since figures like Geoffrey Hinton left Google, prompting more regulatory scrutiny.
- Impact on OpenAI: Concerns about labeling new models due to potential backlash from safety advocates.
- Llama 2 Release
- Significance: Llama 2 is the first commercially usable, GPT-3.5 equivalent model, enabling businesses to deploy AI on their own infrastructure while allowing for fine-tuning.
- Open-source Debate: The model's commercial licensing has sparked discussions about what constitutes "open-source," with the community expressing varied opinions on its implications.
- Custom Instructions in ChatGPT
- Overview: The introduction of custom instructions allows for greater personalization within ChatGPT, but is seen as a delayed response compared to other platforms.
- Market Position: ChatGPT is perceived to be playing catch-up to other competing AI platforms that have already implemented similar features.
- The Landscape of AI Companions
- Trend: There is growing interest and investment in AI companions, which offers a glimpse into the potential for virtual relationships and emotional support through AI.
- Social Implications: The emergence of AI companions raises ethical considerations regarding loneliness and interpersonal relationships.
- Claude 2 and Anthropic's Positioning
- Overview: Claude 2 from Anthropic offers significant advancements with a large context window, making it attractive to developers.
- Market Reception: Although Claude 2 is seen as beneficial, it remains overshadowed by OpenAI, indicating a need for better branding and community engagement.
- Open Source vs. Proprietary Models
- Discussion: A shift towards open-source models is noted, with developers increasingly moving away from proprietary solutions like those from OpenAI.
- Future Outlook: There is anticipation for more public case studies showcasing successful implementations of open-source models in production.
Key Takeaways
- Code Interpreter as a Game Changer: Viewed as a pivotal development that enhances the breadth of tasks AI tools can handle efficiently.
- Evolving AI Landscape: The introduction of Llama 2 and Claude 2 signifies a competitive landscape with increasing pressure on established players to innovate.
- Cultural and Ethical Considerations: The rise of AI companions prompts a re-evaluation of human relationships and the role of technology in emotional well-being.
- Future Predictions: Expect a rebound in AI developments post-August, particularly with major events like Facebook Connect and ongoing hackathons.
Conclusion This episode emphasizes the rapid evolution of AI technology and its societal implications. The discussions highlight the importance of keeping pace with technical advancements while considering regulatory frameworks and ethical considerations in AI development.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Today on the AI Breakdown, I'm joined by the hosts of the Latent Space podcast podcast to discuss everything that happened in AI last month from a technical and developer. perspective. From llama to code interpreter to open source debates and beyond, this is your summer technical AI roundup. The AI Breakdown is a daily podcast and video about the most important news and discussions in AI. Go to breakdown.network for more information about our YouTube, our newsletter, and our Discord.
0:27Welcome back to the AI Breakdown. Today, I am very excited to be collaborating with the hosts of Latent Space. Latent Space is a podcast focused on AI development and AI engineering. It's hosted by Alessio, who is also a VC investing in AI and other frontier spaces, and Sean, better known as Swix, who is an AI developer and entrepreneur. Now, one of the things that makes the AI space interesting is that among the people paying attention right now, say the types of folks that are interested enough and focused enough to be listening to a daily AI news analysis show like the AI Breakdown, even among those who are not developers themselves, they are still actively paying attention to technical developments in the field.
1:04There's an understanding, I believe, that when it comes to getting an edge in what's coming next, GitHub is much more important than Twitter for understanding what's coming down the pipeline. With that in mind, Latent Space and the AI Breakdown are hoping to do these technical news roundups at a pretty regular clip, and we kick off looking back at a month that had some serious action. We're going to be talking about open source developments, the latest in the agent space, why AI companions are getting more attention, and of course, breaking down the significance of Code Interpreter and Llama 2.
1:32It's a great conversation. I think you'll learn a lot. So let's dive in. All right, what is going on? How's it going, boys? Great to have you here. Hey, good. How are y 'all? Great to be on. Good. I'm excited for this. Yeah, no, I am super excited. I think, you know, we were just talking a little bit before this, that the AI audience right now is really interesting. It's sort of on the one hand, you have, of course, the folks who are actually in it, who are building in it, who are, you know, or dabbling because they're in some other field, but they're fascinated by it and, you know, are spending their nights and weekends building.
2:09And then on the other hand, you have the folks who are, you know, what we used to call non-technical perhaps, but who are actively paying attention in a way that I think is very different to the technical evolutions of this field because they have a sense or an understanding that it's so fast moving that the place that they have to be paying attention to is what's changing from the standpoint of developers and builders. So what we want to do today is kind of reflect on the month of July, which had a couple of, I think, really keystone events in the context of what it means for the technical development of the AI field and where it leads, how people's frameworks are changing, how people's sort of sense of things have changed over the last month.
2:54And I think that The place to start, although we could choose a lot of different examples, is with an idea that you guys have spent a lot of time sharing on Twitter and in other places that the launch of Code Interpreter from OpenAI, which is nominally a chat GPT plugin, actually represents functionally something closer to the release of GPT 4.5. So maybe we can start by just having you guys sort of explain that idea, and then we can kind of take it from there. I'll maybe start with this one. Code Interpreter was first announced as a plugin, at least in the plugins announcement from March. But from the start, it was already presented as a separate model.
3:36Because at least when you look in the UI, you don't go into the Chattapult plugins UI and pick it from a menu of plugins. It is actually a separate model in the dropdown menu. And it is so today. And I think, yes, it adds on an additional sandbox for running and testing code. than iterating on that. And actually, you can upload files to it and do operations and files. And people are having a lot of fun uploading different binaries and hacking to see what the container is and trying to break out the container. But what really convinced me that it might be a separate model was when people tried it on tasks that were not code and found it better.
4:14So Code Interpreter is poorly named, not just because it just sounds like a weird developer tool, but basically it's kind of maybe hiding some progress that OpenAI has made, but it's completely not been public about. There's no blog post about it. Code Interpret itself was launched in a support forum post, you know, low key. It wasn't even announced by any of the major public channels that OpenAI has. And so the leading theory is that, you know, I've dubbed it GPT 4.5. I think like if they were ever to release an API for that, they might retroactively rename it 4.5 in the same way that 3.5 was retroactively renamed when Chai GPT was released.
4:51And I think, and since I published that post or tweeted that stuff, the leading reason, the leading rationale for why they did not do it is because they would piss off all the AI safety people. Yeah, no, I mean, it was sort of correspondent, obviously, like, a thing that's happened less just this month, but more over the last three months is a total Overton window shift in that AI safety conversation. Starting from, I think, about April or May, when Jeffrey Hinton left Google, there has been a big shift in that conversation. Obviously, regulators are way more active now than they were even a couple months ago.
5:26And so I do think that there are probably constraints in how OpenAI and any other company in the space feel like they can label or name things. And even just as we're recording this today, we just saw a trademark for GPT-5, which is sort of most likely, I think, just dotting the I's and crossing the T's as a company because they're eventually going to have a GPT-5. I would be very shocked at this point if there are any models that are clearly ahead of GPT-4 that come out before there is some pretty clear guidance from the US government around what it looks like to release more advanced models than GPT-4.
6:02So it's an interesting moment. Let's talk about what functionally it means for it to be that much better, better enough that we we would call it GPT 4.5. And maybe what might be useful is breaking that apart into how it is improving the experience for non-coding queries or inputs. And then separately, how it has made chat GPT as a coding support tool different as well. I think there's a lot of things to think about. So one, models are usually benchmarked against certain tasks and that works for development. But then there's the reality of the model that, you know, if you ask, for example, mathematical question to like GPT-3, 3.5, you don't really get good responses because of how digits are tokenized in the model.
6:51So it's hard for the models to actually reason about numbers. But now that you put a code interpreter in it, all of a sudden it's not a math in the tokenizer in the latent space question. It's like, can you write code that answers the math question? So that kind of enables a lot more use cases that are just not possible with the transformer architecture of the underlying model. And then the other thing is that when it first came out, people were like, oh, this is great for developers. It's like, I know what to do. I just ask it. But there's this whole other side of the world, which is, hey, I have this like very basic thing.
7:27You know how I'm a software engineer by background. You know how sometimes people that have no coding experience come to you and it's like, hey, I know this is like really hard, but could you help me do this? And it's like it's really easy. And sometimes they think it's easy and it's hard, but a code interpreter enables the whole space of problems to be solved independently by people. So it's kind of having, you know, Sean talked about this before, about some of these models being like a junior developer that you have on staff for you to be more productive. This is similar for non-business people.
7:58It's like having junior, you know, dev or like a intern analyst that helps you do these tasks that are not even like software engineering tasks it's more like code is just a language used to express them it's like a pretty basic stuff sometimes but you just cannot cannot do it without so for me the gpd 4.5 thing is less about you know is this a new model that is like built after gpd4 it's more about capability so if you have gpd4 versus 4.5 you're probably going to get more stuff done with 4.5 just because of like the code interpreter piece. So for me, that's enough to use the code name. But as you said, Sam Allman said, they're not training the next model.
8:38So if they said this is 4.5, he would have like, he would go back to Washington DC and be in front of Congress and have to talk about it again. Yeah. Well, one thing that I always want to impress upon people is we're not just talking about like, yes, it is writing code for you, But actually, you know, if you step back away from the code and just think about what it's doing, is it's having the ability to spend more inference time on harder problems. And it matches what we do when we are faced with difficult problems as well. Because right now, any LLM, and this is before Code Interpreter, any LLM, if you give it a question like, what is 1 plus 2?
9:17It'll take the same amount of time to respond as something like prove the Black-Scholes theorem, right? And that should not be the case. Actually, we should take more time to think when we are considering harder problems. What I think the next frontier and why I call it 4.5 is not just because it hasn't had extra training. It's not just because it has the coding environment. It's also because there's this general philosophy and move that I see among OpenAI. And the people that it hires. So in my blog post, I called out Noam, who I personally met. So it's kind of awkward to talk about it, I guess, a friend or a friend of a friend.
9:51But it's true that I have met multiple people now at OpenAI who have specifically been hired to work on more inference time optimizations as compared to training time. And I think that is the future for GPT-5. The reason I think about this working time is that this is the direction of AGI, that we're going to spend more time on inference. And it just makes a whole lot of sense when you look at GNOME's background working on the Broadus and then Cicero, all of which is just consistently the same result, which is every second or millisecond extra spent on inference is worth like$10 ,000 of that in training, especially when you can vary it based on the problem difficulty.
10:32And this is basically ties back to the origin of OpenAI, which originally started playing games. They used to play Dota. They used to play all sorts of games in sort of those reinforcement learning environments. And the typical way that you program these AIs during these games is when they have lots of branches, you take more time to search and figure out what the optimal strategy is. And when there's not that many branches to go down, then you just take the shortcut and give the right answer. But varying the inference time is the innovation here. One of the things that it seems, and what you just described, I think aligns with this, is I think there's a perception that more advanced models are just going to be bigger data sets with more of the same type of training versus sort of fundamentally different techniques or different areas of emphasis that go beyond just how big the data set is.
11:30And so one of the things that strikes me listening to or kind of observing how Code Interpreter works is it almost feels like a break in the evolutionary timeline of GPT because it's like GPT with tools, right? Alessia, you just kind of described it. It's like it doesn't know about math. It doesn't have to know about math if it can write code to figure out the math, right? So what it needs is the tool of being able to write code, and that allows it to figure something out. And that is akin to, you know, humans are evolving for millennia, not using tools, then all of a sudden, someone picks up a rock, and this whole entire set of things that we couldn't do before, just based on our own evolutionary pathway are now open to us because of the use of the tool.
12:13I don't think it's a perfect analogy, but it does feel somewhat closer to that than just again, like, it's a little bit better than 3.5. So we called it four, it's a little bit better than four. So we called it 4.5 kind of mental framework. No argument there. Another big topic that relates to this that was subject of a lot of conversation, not just this month, but has been for a couple of months, is this question of whether GPT-4 has gotten worse or whether it's been nerfed. And there was some research that came out around that with maybe variable sort of feelings around it. But what did you guys make of that whole conversation?
12:48I think evals are one of the hardest things in this space. So I've had this discussion with founders before. It's really easy. we always bring up Copilot as one example of like cutting edge eval where they not not only look at how much of their suggestions you accept but also how much of the code is still in a minute after three minutes after five minutes after it's really easy to do for code but like for more open-ended generative tasks it's kind of hard to say what's good and what isn't like if I'm asking to write the show notes for our podcast which has never been able to do how do you eval that it's really hard so even if you read through through the paper that uh ling jiao and mate and and james wrote a lot of things are like yeah they're they're worse but like how do you really say that you know like sometimes it's not cut and dry like sometimes it's like oh the formatting changed and like i don't like this formatting as much but if the formatting was always the same to begin with would you have ever complained you know there's a lot of that um and i think with llama edge effort can like go wrong in terms of like being too tight you know for example somebody asked llama 2 it's like how do you kill a process in like linux and llama 2 was like oh it's wrong to like kill and like i cannot help you like doing that you know and i think there's been more more chat online about you know sometimes when you do reinforcement learning you don't know what reward and like what what part of like the the suggestion the model is anchoring on like sometimes it's like oh this is better sometimes the model might be learning that you like more verbose answers even though they're the right the same way so there's a lot of stuff there to figure out but i think some examples in the paper like clearly worse some of them are like not as crazy but i mean open eyes under a lot of pressure on like the safety and like all the the instruction side and we cannot like the best thing to do would be hey let's version lock the model and like keep doing evals against each other.
14:46Like doing an eval today and an eval like that was like a year ago, there might be like 20 versions in between that you don't even know how the model has changed. So yeah, evals are hard. It's the TLDR. And I think basically this is, what we're seeing is OpenAI having come to terms with that, the origin of itself as a research lab where updating models is just a relatively routine operation. versus a product or infrastructure company where it has to have some kind of reliability guarantee to its users. And so OpenAI, I think internally, its researchers are used to one thing, and then the people who come and depend on OpenAI as a product are used to a different thing.
15:31And I think there's a little bit of cultural mismatch here. But even within OpenAI's public statements, we have simultaneously Logan from OpenAI saying that the models are frozen, and then his VP of products saying that we update models all the time so they're not frozen. So which is it? It cannot simultaneously be true. I think they're trying to figure it out. I think people are rightly afraid of them basing themselves on top of a black box. And that's why maybe we'll talk about Lama 2 in a bit. That's why maybe they want to own the black box such that it doesn't change out from under them. And I think this is fine.
16:06This is normal. But OpenAI, it's not that hard for OpenAI And I had to figure out a policy that is comfortable with, that everybody accepts, and it won't take them too long. And this is not a technical challenge. It's more of an organizational and business challenge. Yeah, I mean, I think that the communications challenge that you're referencing is also extreme. And I think that you're right to identify that they've gone from quirky little lab with these big aspirations to like epicenter of a national conversation or a global conversation about existential challenges, you know, and the way that you talk in those two different circumstances is very different.
16:45And you're sort of serving a lot of different masters, hopefully always guided by your own set of priorities. And that's going to be inherently difficult. But with so many eyes on it, and people who are, you know, the thing that makes it different is it's not just like Facebook, where it's like, oh, we've got a new feature, in the early days that made us all annoyed. People were so angry when they added the feed and we all got used to it. This is something where people have redesigned workflows around it. And so small disruptions that change those workflows can be hugely impactful. Yeah, it's an interesting comparison with the Facebook feed because in the era of ad tech, the feedback was immediate.
17:23You change an algorithm and if the click-through rates or whatever metric you're optimizing for in your social network, If they start to decline, your change will be reverted tomorrow. Whereas here, like we just talked about, it's hard to measure. And you don't get that much feedback. There's sort of the thumbs up and down action that you can take in OpenAI. But I'm sure most people don't give feedback at all. So OpenAI has very little feedback to go with on what is actually improving or not improving. And I think this is just normal. Like it's kind of what we want in a non-ad track universe, right?
18:02Like we've just moved to the subscription economy that everyone is like pining for. And this is the result that we're trading off some amount of product feedback, actually. Super interesting. So the one other thing before we leave OpenAI ecosystem, the one other big sort of feature announcement from this month was custom instructions. How significant do you think that was as an update? So minor. so it is significant in a sense that you get to personalize chat tpt much more than you previously would have like it actually will remember facts about you it will try to obey system prompts about you you had this in the playground since forever because you could enter in the system prompt in there and just chat tpt didn't have it and this is a rare instance of the chat tpt team lagging behind the general capabilities of the open AI platform.
18:55And they just shipped something that could have been there a long time ago. It was present in Perplexity AI. And if you think about it, basically every other open source chat company or open, every other third party chat company already had it before ChatGPT. So what I'm talking about is character AI. What I'm talking about is the various AI waifu, AI girlfriend type companies, each of which have characters that you can sort of sub in as custom instructions. So I think ChatGPT is basically playing catch up here. It's good for obviously the largest user base in the world of Chat AI, but it's not something fundamentally new we haven't seen before.
19:34Yeah, I think that it was so clearly a user-centric sort of end user, non-technical user in some ways focused feature that is incredibly obvious and important in retrospect, but not sort of a massive change to something that feels like it should have been there the whole time. But I guess building off of that, the sort of assertion or the notion that it was ChatGPT playing catch up, I wonder to what extent one of the sort of sub themes from this month was also a little bit about trying to better understand or having a better understanding of where the value is going to accrue in this space. So specifically, we had the first round of layoffs.
20:14This isn't something we actually had talked about talking about in advance, but we had the first round of layoffs from a couple different companies. And that, you know, we also saw in June, the first decline in monthly users measured in terms of site visits and mobile app visits. It's over. Other ways to measure it. Yeah. The bubble has burst. It's over. You know, I think that to some extent, yes, it was, this is a very convenient narrative shift for publications that have been breathlessly talking about its, you know, inexorable rise. So I think it maybe is a little overamplified by that. But do you think that sort of take this custom instructions, you know, or something like perplexity, perplexity has been ahead of chat GPT on a lot of different features, right?
20:54The sourcing interface is amazing, right? There's a lot of things that I think perplexity does so much better. But ultimately, is a project like that going to just be eaten up by sort of features being copied by sort of the board that actually has the, you know, the technical innovation underneath. I mean, I don't know. What do you guys think about that? Man, as a, you know, my full-time job is being a venture capitalist. I think that's one of the hardest questions that everybody is grappling with. I think there's a few things to think about. So one is how quickly do the open models catch up? So I think everybody agrees that long-term, like access to intelligence through these models would be available to everybody.
21:36the question is how much of a head start do the incumbents have in terms of um and by incumbents i mean like the ai incumbents you know like yeah open ai like perplexity all these companies that were there like two three years ago because then the big companies are going to be like oh well i got a lot more data and i got a lot more distribution and you can see that with microsoft right it's like who's trying to build the new microsoft word in the world of llms like nobody you know like the existing players like office is building this in notion is like building this in their product like they already have so much of your data like superhuman just rolled in their ai thing in their email product they already have so much more to put into the model to tailor it to your use case that then it's going to be hard for the new startups to get there but if the startups are like two three years of you know at start that's a different question but i don't know a Llama 2 doesn't bode well for a lot of them, right?
22:37Like the 70 billion parameter model is pretty good. So now all of a sudden you got Llama 7B. That actually, I think perfectly brings up a segue to the other major obvious thing that happened this month from both a technical perspective, but also just, I think, long-term from a user perspective, which was Facebook releasing Llama 2. So this was something that was anticipated for a while. But I guess where to even start with the significance of LAMA 2? I mean, how do you sum it up? If you're talking to someone who sort of isn't paying attention to the space, what does the introduction of LAMA 2 mean relative to other things that had been available previous to it?
23:14It is the first fully commercially usable, not fully open source, we'll talk about that. First fully commercially usable GPT 3.5 equivalents model. That's a big deal because one, you can run it on your own infrastructure. You can run it on your own cloud. So all the governments and healthcare and financial use cases are opened up to that. And then you can fine tune it because you have full control over all the weights and all the internals as much as you want. So it's a big deal from that point of view. Not as big in terms of pushing it forward, the state of the art, but it's still an extremely big deal.
23:52I think the open source part. So the day that it came out, I wrote this post about, you know, why Lama 2 is not open source and why it doesn't matter. And I was telling Sean, I'm writing this thing. And he was like, whatever, man, like this lesson stuff is like so, so tired. I was like, yeah, I'll just post it on on Agro News in the morning. And I think it was on the front page for like the whole day. And it got like 228 comments. And I was recording the flash attention podcast episode in the morning. So I got out of the studio and it was like 230 comments of people being very like, you know, upset one way or the other about license.
24:28And my point, and you know, I was, I started an open source company myself in the past and I contributed to a bunch of projects is that, yeah, Lama 2 is not open source by like the open source Institute definition, but we just don't have a better definition for like models, you know, like, because it's mostly open source. You can use it for a lot of stuff. So what's like the, and it's not source available because for a lot of stuff, you can use it commercially. So how do we find better labels? And my point was like, look, let's figure out what the better label is. But even though it's not fully open source, it's still like$3 million of like flops donated to the community, basically.
25:07You know, who else in the open source community is stepping up and putting$3 million of H100 to make us train this model. So I think like overall, net-net is like a very positive thing for the community. And then you've seen how much stuff was built on top of it. There's like the quantized versions with GGML. There's like the context window expansion. There's so much being done by the community that I think it was great for everyone. And by the way,$3 million is the low end. That's just compute. There's a reasonable estimate from ScaleAI that the extra fine tuning that they put on top of it was worth about$15 to$20 million.
25:43So that's a lot of money just kind of donated to the community. although they didn't release the data they didn't tell us any of the data sets they just say trust us we didn't train on any of your facebook information which is uh it is the first instance where the models are more open than the data and i think that's a reflection of where the relative shift in value might happen as a result of llama 2 and so i don't know you can take that in multiple different directions but i just want to point that out yeah i was gonna say so we first had the examples i made so we first had the open models open source models which is like red pajama so the data's open the training code is open the model weights are open then stability kind of did the same thing with stable lm which is like hey the weights are open but we're not giving you the data so you can you can download the model but you cannot retrain it yourself and then lama2 is like we don't give you the data we'll give you the models, but you can only use it for, for some stuff.
26:45So there's more and more restriction, but like Sean is saying, and we talked about this before, everybody wants to train their own model. Nobody wants to open source the best dataset for X, which maybe is what more open source people should focus on. It's like how to build better specific datasets instead of yet spending, giving Jensen Wang another$5 million of GPUs. But the model gets more headlines for now, you know? So that's what everybody does. Yeah. And I want to point out, it's a reversal of the open source culture. There used to be this sequence of openness that you could kind of pick and choose from, whether it's open code all the way down to open data versus all the way down to open weights.
27:26And, you know, there's some varied combination. I wrote this post a long time ago. I don't remember the five levels. But yeah, like it's very strange. And I think it's just a relative discussion of where the money is going. And I think it massively shows that compute is becoming commoditized, which, yes, there's a GPU crunch right now. A100s are sold out everywhere across the board. People are commenting all about it this month. And there's people hoarding compute like nobody's business. But as far as the value in AI is concerned, it looks like compute is relatively commoditized. It's actually data that people are kind of safeguarding jealously.
Read the full transcript
28:05Going all the way back to the history of open source models, in Luther AI, when they trained GPT-J and GPT-Neo as the first reproductions of GPT-3, they released the data first. Stable Diffusion, when they trained Stable Diffusion, they released Lion 500B first. And that's, I think, reflective of the normal sequence of events. You release the data, then you release the model weights. But now we're just skipping the data part. And I think it's fair. It's a way to guard yourself. I think one of our conversations, I think it was Mike Conover, when he was talking about comparing our current AI era versus the 2000s era in search engines.
28:48He basically said all of the public publishable information retrieval research dried up because all those PhDs went to work at Google and Google just sat on it. And this is now a fight for IP. And I think that is just a very rational way of behavior in, I guess, a capitalist AI economy. So one of the things that we were talking about before, starting with the code interpreter 4.5 and why or GPT 4.5 and why they might not call it that, is the emergence of this sort of regulatory, if not pressure, certainly intrigue. Do you think that there's potentially an aspect of that when it comes to why people are so jealously safeguarding the data?
29:27Is there more risk for being open about where the data is actually coming from? The book's three examples is probably good. So MPT trained their model on a dataset called Books Tree, which is 190 ,000 books, something like that. And then people on Twitter were like, well, this stuff is not, you know, in the free, you know, it's under copyright still. You just cannot. Public domain. Yeah, it's not in the public domain. You can just take it and train on it. But the license for some of these books is like kind of blurry, you know, on like what's fair use and what isn't. And so there was like this old thing on Twitter about it.
30:04And then MPT, you know, Mosaic first changed the license and they changed it back. And I think Sean, Sean Presser from Luther was just tweeting about this yesterday. And he was basically saying, look, as ML engineers, maybe it's better to not try and be the main ethics knight and just say, hey, look, the data is open and let's try it. And then maybe people later will say, hey, please don't use the data and then we can figure it out. but like proactively not using all of this stuff can kind of keep the progress back and you know he's more coming from the side of like a luther which is like doing this work in public so for them it's like hey you know if you don't want us to train all this is fine but we shouldn't by default not do it versus if you're meta you know they said that they trained llama on like stuff available on the internet they didn't say they trained llama on stuff that is licensed to train on.
30:59It's a small difference. The other piece of this that I wanted to sort of circle back to, because we kind of breezed over it, but I think is really significant. We did get a little lost in this conversation around open source definitions. And I don't think that's unimportant. I think that people are rightly protective when a set of terminology has a particular meaning and a massive global corporation sort of tries to like nudge it towards something that is potentially serving their ends versus actually being by that definition. But I also think that your point, which is that functionally relative to the rest of the space, it probably doesn't super matter because what people mean is almost more about functionally what they can do with it and what it means for the space relative to more closed models.
31:43And I think one of the big observations has been that the availability from when Llama One was fully leaked, the availability of all of that has pretty dramatically changed, one, the evolution of the space over the past few months, and two, I think from a business standpoint, how the big companies and incumbents have thought about this. So another big conversation this month, going back to sort of the venture capital side of your life has been the extent to which companies or startups are, or big companies are not wanting to sort of sign on with some startup that's going to offer them, you know, AI, whatever, because their technical teams can just go spin up sort of their own version of it because of the sort of, you know, availability of these open source tools.
32:33But I'm interested, I guess, in bringing the sort of open source, you know, in air quotes side of the conversation into the realm of how it has impacted how companies are thinking about their development in the context of the AI space? I think it's just raising the bar on what you're supposed to offer. So I think six, nine months ago, it was enough to offer a nice UI wrapper around an open AI model. Today, it isn't anymore. So that's really the main difference. It's like, what are you doing outside of wrapping the model? and people need more and more before they buy versus building. Yeah, I think it actually moves the area of competition towards other parts of productionizing AI applications.
33:22I think that's probably just a positive. I feel like the competitive pressure that Meta is putting on OpenAI is a good thing. One of the fun predictions that I made was in the next six months, GPT, OpenAI will open source GPT-3, which is not open source. It's so far behind the state of the art now that it doesn't matter as far as safety is concerned. And it basically keeps OpenAI in the open source AI game, which would be nice to have. Of the things that people have been building, you called out a couple context window expansion, but have there been any that really stand out to you as super interesting or unexpected or particularly high potential?
34:02One of our short term podcast guest uh the mlc team they worked on wrapping llama 2 to run on uh macbook gpus so i think that's like the the most interesting gap right it's like how do we go from paper token to like unlimited local use that's one of the main main things that keep even people like me from like automating a lot of stuff right it's like i don't want to constantly pay open ai to do menial stuff but if i could run this locally and do it even if it's five times lower i would do it. So that's a super exciting space. Yeah. I would say beyond that, there hasn't been that much. I mean, it's only a few weeks old, so there hasn't been that much emergence coming from it.
34:43I would definitely say you want to keep a lookout for basically what happens in post-LAMA 1, which keep in mind, it was only in February. The same thing that happened with Facunia, Alpaca, and all the other sort of instruction to you and sort of research type models, but just more of them, because now they are also commercially available. We haven't seen them come out yet, but it's almost a guarantee that they will. You can also apply all the new techniques that have emerged since then, like JSON former, because now you have access to all the model weights to LAMA. And I think that will also create another subset of models that basically was only theoretically applicable to sort of research quality models before.
35:30And so now these will be offered commercially as well. So like, yeah, nothing like really eye-popping, I would say. But it's been five minutes. Yeah, it's been a very short amount of time. And the thing about open source is that the creativity on lock is very hard to predict. And actually, I think happens a lot in the, let's just say, the less official part of the economy, where I've been focusing a lot recently on the sort of AI girlfriend economy, which is huge. I feel like it's not polite conversation that the amount of AI girlfriend and AI husbandos, AI wife who's in AI husbandos that people have been talking about.
36:08But it's real. There are millions of users. They're making a lot of money. And it's just virtually not talked about in polite SF circles. It feels like one of those areas that's going to be an absolute lightning rod when it comes to the societal debates around this technology. You can feel it. People are going to hone in on that as example A of a change that they don't like. That's my guess, at least. So I have a really crazy longer term prediction, maybe on the order of 30 to 50 years. But AI Girlfriend for Nobel Peace Prize, because what if it solves the loneliness crisis? What if it cuts the rate of terror and school shootings by half or something?
36:51That's huge. My wife and I have joked about how every generation, there's always something like they always think that they're like so far ahead and they think that there's nothing that their kids could throw at them that they just like fundamentally won't get. And without fail, every generation has something that seems just totally normal to them that their parents generation writ large just like has such a hard time with. And we're like, it's probably going to be like AI girlfriends and boyfriends. We're going to be like, yeah, but they're not real. They're like, yeah, but it's real to me. You know, or having debates with our future 13-year-old.
37:26Our kids are only four and two now. So it feels like maybe the right timeline. Yeah. I've heard actually of all people, Matthew McConaughey on the Lex Rebman podcast. What? Yeah, yeah. No, he was great. Shout out. Shout out. Shout out, Matt. They were kind of talking about this and they were noodling this idea of like computers helping us being better. so kind of like we have computers learn how to play chess and then we all got better at chess by using the computers to like learn and like experiment they were talking about similarly and interpersonal relationship maybe you know it doesn't have to be you shut off from from humans but it's like using some of these models and some of these things to actually like learn you know how to better interact with people and if you're like shy and an introvert it's like okay, I can like try these jokes or like these conversation points with a model.
38:17And like, you know, it teaches me, Hey, that's not okay to say, or like, you know, you should maybe be more open or, or I don't know, but I think that's a more wholesome view of it than like everybody just kind of runs away from society. And that's like 10 AI friends and doesn't talk to humans anymore. It's much less sexy to just say like AI friends, right. And even though like there's the, If you look at the possibility set, the idea that people might have this sort of, to your point, like conversational partner that helps them effectively work through their own things in this safe space, that doesn't necessarily lead to romantic attachment just because the movie Her came out, right?
38:59Right. It can just be a panel of experts. I do have plans to build a small CEO, which is my own boss, and just for me to check in. And I actually will flag out, just living here in San Francisco, you come across a lot of AI engineers who are interested in building mental wellness products. And a lot of these will take the form of some kind of journal. And this will be your most private thoughts that you don't really want to send anywhere else. And so actually, all of these will make advantage of open source models because they don't want to send it to OpenAI. And that makes a ton of sense. For people who want to try it out, I'll also give a shout out to circlechat.co, which is something I just came across from one of my friends here in the co-working space that I have, where it's one of those situations where you can actually try out, like having a conversation and having a group of AI friends chime in and see what that feels like to you.
39:50It's the first example I've come across where someone's actually done this. Super interesting. So Lama and Code Interpreter, I think, stood out pretty clearly as really big things to touch. I wanted to check in just as we sort of start to maybe round the corner towards wrapping up. Claude 2 and Anthropic, how significant was this? In what ways was it significant? Was it something that was sort of meaningful from expanding the capacity set for developers? Or was it sort of more just a good example of what you can do if you increase the context window? But that's something that might ultimately become table stakes later on.
40:25Yeah, I can maybe speak to this a little bit. It is significant, but not earth-shattering, clearly. I think it is the first time that Cloud as a whole has just been generally publicly available, used to be on a wait list. Yes, it has a longer context window, but to me, more significantly, it is Anthropic finding its foothold in the very competitive AI landscape. Anthropic's message used to be that, yes, we're number two to open AI, but we're safer. And that's not a super appealing thing to many engineers. It is very appealing to some corporations, by the way. I think having the 100K context window makes them state-of-the-art in one dimension, which is very useful.
41:08The ability to upload multiple files, I think, is super useful as well. And actually, I have met a number of businesses. I'm close friends with Sourcegraph, who are actually choosing to build with Cloud2 API over and above OpenAI, just because they are better at latency, better reliability, and better in some form of code synthesis. I think it's Anthropic finding its foothold, finally, after a long while of being in OpenAI shadow. Yeah, and we use Claude for the transcript and timestamps in the podcast. So shout out the 100K context window. You know, we couldn't do that. When we first started the podcast, we were like, okay, how do we chunk this stuff for like GPT-4 and all of that?
41:51And then I was like, just put the whole thing in here, man. And works great. So that's a good start. But I feel like they're always, yeah, second fiddle. You know, it's like every time they release something, people are like, cool. Okay. Some people like it. Most people are like, okay. I feel bad for them because it's like, it's really good stuff, you know, but they just need some help on the marketing side and the community buy-in. So I just spent this past weekend at the Cloud hackathon, which is, as far as I know, Enthalpik's first hackathon. I tweeted a pretty well-received video where I was just in the hackathon venue at 2 a.m.
42:30in the morning, and there was just a ton of people hacking there. There were like 300 people participating for Cloud. And I think it's just the first real developer excitement I've ever seen for Anthropic and Claude. So I think they're on their way up. I think this paves the way for a multi-model future. That is something that a lot of people are betting on. It's just the odds are stacked against Anthropic, but they're making some headway. I do think that you should always be running all your chats side by side against ChatDBT and Claude and maybe Llama2. so I immediately I have a little menu bar app that does that that syncs all the all the chats across and and yeah I can say I can legitimately say that Claude wins about 30 percent of the time as far as any time I give it a task to do I ask it a question which is not you know doesn't make it number one but it actually is very additive to your overall toolkit of AIs that you should use yeah it's certainly the first time that you're if you go on Twitter on any given day you will see people saying things like if you haven't used Claude you know for writing you have to try it now or so you know like people who are really who have made a switch who are have no affiliation who are very convinced that it is now part of the the suite of tools that people should really be paying attention to which I think is great yeah we shouldn't be at a stage yet where we're you know total totally in on one just one tool set I'll also mention I think this month, or at least July, was the first discussion of whether is too much context not actually a good thing.
44:06So there's a pretty famous paper, I forget the actual title of it, that shows a very pronounced U-curve in the retrieval abilities of large context models. And so basically, if the item that's being retrieved is at the start or at the end of the context window, then it has the best chance of being received. But if it's in the middle, it has a high chance of being lost. And so is 100K context a good thing? Are you systematically testing its ability to retrieve the correct factual information? Or are you just looking at a summary and going, yeah, that looks good to me? I think we will be testing whether or not it's worth extending it to 100K or a million tokens or infinite tokens.
44:45Or do you want to blend a short window like 8 ,000 tokens or 4 ,000 tokens and couple that together with a proper semantic search system, like the retrieval augmented generation and vector database companies are doing. So I think that that discussion has come up in open source a lot. And basically, I think it matches human memory, right? Like you want to have a short working memory. You know, I was thinking about it. The one other, obviously, big sort of company update that we haven't spoken about yet was around the middle of the month, Google Bar had a big set of updates. A lot of it was sort of business focused, right?
45:23So it was available in more languages. It was, you know, whatever the, the sort of from a feature perspective, the biggest thing that they were sort of hanging their hat on was around image recognition and sort of this push towards multimodality. But do you guys have any thoughts about that? Or was that sort of like not sort of on the, the, the high priority list as a, as an announcement or development this month? I think going back to the point before we're getting to the maturity level of the industry where like doing like model updates and all this stuff, like it's fine, but like people need more.
45:52you know people need need more and like that's why i quote interpreter it's like so good right it's not just like oh we made the model a little better like we added this thing it's like this is like a whole new thing if you're playing the model game if not you got to go to the product level and i think google should start thinking about how to make that work because when i search on google maps for certain stuff it's like completely does not work so maybe they should use models to like make that better and then say we're using bard in google maps search but yeah i don't know i'm kind of tuning off a lot of the single just model announcement so bard's updates i think that the multi-modality they actually beat gpt4 to releasing a generally available multi-modal model right you can upload an image and have bard describe it and that's pretty interesting pretty cool i think uh one of our earliest guests roboflow uh brad their cto was actually doing some comparisons because they have access to a lot of the vision models.
46:49And BARD came up a little bit short, but it was pretty good. It was close to the state of the art. I would say the problem with BARD is that you can't rely on them having reliable updates because they had a June update, I don't know if you remember, of implicit code execution where they started to ship the code interpreter type functionality, but in a more limited format. If you run the same questions that Bard advertised in the June blog post, that Sundar Pichai advertised in a video that he tweeted out, they no longer work in Bard. So they had a regression that was very embarrassing, obviously unintended.
47:28And it shows that it's hard to keep model progress up to date. But I think Google has this checkered history with its products being reliable. They also killed off Google Domains, RIP. And I think that's something that they have to combat, which is like, yes, they're trying to ship model products. I've met the bar people. They're good, earnest people. But they have struggled to ship products even more than OpenAI, which is frankly embarrassing for a company this size of Google. Outside of the biggies, are there any other sort of key trends or, you know, maybe not even key trends, but sort of bubbling interest that you guys are noticing in the developer community that aren't necessarily super widely seen outside?
48:09You know, one of the things that I keep an eye on is all the auto GPT like things, you know, in this month, we had GPT engineer and we had multi on who held a hackathon. and there's a few things like that, but not necessarily in the agent space. But are there any other themes that you guys are keeping an eye on, let's say? I'm sure Alessio can chime in, but I do keep a relative close eye on that agent stuff. It has not died down in terms of the heat. Even the auto GPT team, who, by the way, I work on, they're on the first floor, they're building that I work on. They're hard at work shipping the next version.
48:42And so I think a lot of people are engaging in the dream of agents. And I think scoping them down to something usable is still a task that has so far eluded every single team so far. And it is what it is. I think all these very ambitious goals. We are at the very start of this journey, the same journey that maybe self-driving cars took in 2012 when they started doing the DARPA challenge. And I think the other thing I'll point out, interest in terms of just overall interest. I am definitely seeing a lot of eval-type companies being formed and winning hackathons too. So what are eval companies? They're basically companies that you monitor the success of your prompts or your agents and version them and just share them potentially.
49:33I feel like I can't be more descriptive just because it's hard to really describe what they do just because they are not very clear about what they do yet. Langchain launched Langsmith and I think that is the first commercial product that Langchain probably the top one or two developer oriented AI projects out there and that's more observability but also will tend towards eval as well because they acquihired in AI eval projects as well so I'll just call out just the general domain of how to eval models is a very big focus of the developers here in SF. Yeah we've done two seats in companies doing agents, but they're both verticalized agents.
50:12So I think the open source motion has been auto GPT, do anything. And now we're seeing a lot of founders is like, hey, if you take that and then you combine it with deep industry expertise, you can get so many improvements to it. And then the other piece of it is how do you do information retrieval? So in general knowledge, like documents, everything is kind of flat, but when you're in specific like vertical, say finance, for example, if you're looking at the earnings from this quarter, like 10 quarters ago, like the latest ones are like much more important. So how do you start to create this like information hierarchy between documents?
50:50And then how do you use that? Instead of doing simple like retrieval from like an embedding store, it's like, how do you also start to score these things? That's another area of research from founders. Oh, I'll cut out two more things. One more thing that happened this month was SDXL. you know, text-to-image doesn't seem as sexy anymore, even though, like, last year was all the rage. But I do think, like, it's coming along. I definitely wish that Google was putting up more of a fight because they actually, at the start of the year, released some very interesting papers that they never followed up on that showed some really interesting Transformers-based text-to-image models that I thought was super interesting.
51:31And then the other element, which, you know, I'm just, like, very fascinated, by a lot of the, I don't know, like the, I hesitate to say this, but it's actually like the character and like the, character AI, let's just call it, character replica and all the sort of not safe for word versions of that. I do think that a lot of people are hacking on this kind of stuff. The retention metrics on character AI blows away, you know, a lot of the metrics that you might see on traditional social media sites. And basically AI native social media is something that is, there's something there. that I think people haven't really explored yet.
52:09And people are exploring it. No Shazir, basically the leading light behind the Transformers paper. Character AI is his company and he's always a few years ahead of it. So not to keep returning to this theme, but I just think it's definitely coming for a lot of the ways that we view things. Right now we think Co-Pilot and right now we think ChatGPT. But what we really want to speak to is a way of serializing personality and intelligence. And potentially that is a leading form of mind upload. So that gets into science fiction, actually. I don't want to go there, but I do see a lot of people working on that.
52:47Yeah, I mean, we just got a Financial Times report that says that AI personas from Meta, from Facebook could be coming next month. Oh, that's what the report was. There's one that's Abraham Lincoln, one that's like a surfer dude who gives you travel advice. So it's, you know, the sourcing is three people with knowledge of the project or whatever. And it's, you know, obviously no confirmation from Meta, but it's no secret that Zuckerberg has been interested in this stuff. Yeah. And, you know, the FT piece is actually, it's a good overview of why a company like Meta would care about it in very dollars and cents terms.
53:24And I want to state, like, the first version of this is very, very lame. Like when I first looked at character AI, It's like, okay, I want to talk to Genghis Khan if I'm doing a history class, but it's like what a 10-year-old would enjoy. But I think the various iterations of this professionally would be very interesting. So on the developer side of this, I have been calling for the development of agent clouds, which are clouds that are specifically optimized not for human use, but for AI agent use. And that is a form of character. It's a character with a different environment, with the different dependencies pre-installed that can be programmatically controlled, can give programmatic feedback to agents.
54:02And there's a protocol forming that some of the leading figures like AutoGPT and E2B are creating that lets agents run clouds. This would definitely terrify the AI safety people because we have gone from running them on a single machine towards running clusters of machines. But it's happening. So let's talk about what comes next. Do you guys have any predictions for August or if not predictions, just things that you're watching most closely? I think like for me, probably starting to see more public talk about open source models in production with people using that as a differentiator. I think right now, a lot of it is kind of like, oh, these models are there, but nobody's really saying, oh, I moved away from open AI, I'm using this.
54:50But in our, we run an early adopters community with about 1500, kind of like a Fortune 500 large companies leaders. And some of them were like, oh, we deployed Dolly in production and we're using it. We're not writing a blog post about it. So I think right now the perception is still everybody's using OpenAI and the open source models are like really toys. But I think we're going to get into September and, you know, you're not going to see a lot of announcements in August proper, but I think a lot of people are going to spend August getting these models ready and then going into end of the year and say, hey, we're here too.
55:24You know, we're using the open models. Like we don't need OpenAI. I think right now there's still not a lot of public talk about that. So excited to see more. Yeah, I'm a little bit, as for myself, this is very self-interested, obviously, but we had another agenda. You know, I wrote about the rise of the AI engineer. And I think it's definitely happening as we speak. I have seen multiple tags, like people tag me multiple times a day on like how they're reorienting their careers. I think people professionalizing around this and going from essentially like informal groups and Slack channels and meetups and stuff towards certifications and courses and job titles and actual AI teams in every single company, I think is happening.
56:05I just got a notification like two days ago that, you know, in Meta, apparently you can sort of name your name, a job title, whatever you want internally. And so the emergence of the first AI engineer within Meta has been announced. And so I think as far as the near term, I do see this career, this profession come into place that I've been forecasting for a little bit. And I'm excited to help it along. Awesome. Well, guys, great conversation. Tons of interesting stuff happening, obviously. Ironically, I think it's a relatively more quiet time in some ways than it even was. And my prediction for August is that we're going to see the extension of that.
56:46We're going to see sort of the biggest breath that we've had, at least from a feeling perspective, maybe since ChatGPT. But then we are going to rage back in September. You've got Facebook Connect in September. You've got sort of just the return to business that everyone does after August. But of course, I think the hackathons aren't going to stop in the Bay Area. So people are going to keep building. And it's entirely possible that something hits in the next four weeks that totally changes that. Be exciting to see. Looking forward.
From the publisher
Today NLW is joined by Swyx and Alessio, the hosts of the Latent Space podcast to discuss the key technical developments from the last month of AI, including code interpreter; llama 2; the latest in AI agents; growing interest in AI companions, and more.
Latent Space podcast -https://www.latent.space/podcast / https://twitter.com/latentspacepod
Swyx - https://twitter.com/swyx
Alessio Fanelli - https://twitter.com/FanaHOVA
ABOUT THE AI BREAKDOWN
The AI Breakdown helps you understand the most important news and discussions in AI.
Subscribe to The AI Breakdown newsletter: https://theaibreakdown.beehiiv.com/subscribe
Subscribe to The AI Breakdown on YouTube: https://www.youtube.com/@TheAIBreakdown
Join the community: bit.ly/aibreakdown
Learn more: http://breakdown.network/
Twitter: https://twitter.com/nlw / https://twitter.com/AIBreakdownPod
