In short
Latent Space: The AI Engineer Podcast
Episode Summary
The Four Wars of the AI Stack (Dec 2023 Audio Recap)
Episode Overview In this special episode of Latent Space, the hosts recap the notable developments in AI over December 2023, focusing on what they term "The Four Wars of the AI Stack." This recap introduces key conflicts in the AI industry—concerning data quality, GPU resources, multimodality, and RAG/Ops. The hosts aim to summarize significant trends, perspectives, and emerging technologies shaping the AI landscape moving into 2024.
Structure of the Episode
- Introduction
- Hosts discuss the growth of Latent Space and the unique format of this audio recap.
- Highlights the overwhelming response to the December newsletter edition.
- The Four Wars of the AI Stack
- Data Quality Wars
- Concerns over data sources, including user-generated content (UGC), licensing, and synthetic data.
- Emphasizes the critical role of data attribution and the implications of the New York Times lawsuit against OpenAI.
- GPU Rich vs. Poor War
- The divide between companies with ample GPU resources versus those without.
- Discussion on the implications of the Anyscale benchmark drama and its impact on industry credibility.
- Multimodality Wars
- Evolution from text to image and now to 3D video models.
- Examines the competitive landscape between major players like OpenAI and emerging startups focused on specific modalities.
- RAG/Ops Wars
- Analyzes the tension between database companies and framework providers as they encroach on each other's territories.
- Discussion on the role of RAG (Retrieval-Augmented Generation) in enhancing context and performance in AI applications.
- Key Insights and Notable Mentions
- End of Low Background Tokens: Transition from low-quality data to a focus on high-quality, verified data.
- Synthetic Data: Growing interest and investment in producing effective synthetic data to compensate for locked-up human data.
- GPU Inference Dynamics: Understanding the pricing strategies and economic models around GPU usage and inference costs.
- Emerging Technologies: Notable mentions of companies innovating within the AI space, including LangChain, Eleven Labs, and others.
- Future Predictions
- The hosts speculate about upcoming trends, including the importance of quality data, the impact of multimodality on user experience, and the evolution of tools for AI engineers.
- Discussion on the potential for frameworks to either expand upwards or cloud providers to expand downwards.
Key Takeaways
- Data Quality: As AI systems increasingly rely on diverse data sources, the importance of quality and ethical sourcing cannot be overstated. The legal ramifications of data usage underscore this issue.
- Computing Resources: The disparity between GPU-rich and GPU-poor companies reflects larger industry trends and impacts the accessibility of AI technologies.
- Multimodality: The evolution of AI is pushing towards models that can process multiple forms of data, which could redefine user interaction and application development.
- RAG Technologies: Tools that enhance retrieval capabilities are becoming critical as users seek contextually relevant responses from AI systems.
Final Notes The episode encapsulates a moment of reflection on the advances made in AI in 2023 while setting the stage for the challenges and innovations expected in 2024. As the field continues to mature, understanding these wars will be crucial for engineers and stakeholders alike.
For more detailed insights, access the full recap and show notes on [Latent Space](https://latent.space/p/dec-2023).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:28Hey everyone, welcome to the Late in Space podcast. of us as hosts on what's going on in AI. You know, both of us are very actively involved. And I don't think this year will be any different. This year, there's lots more excitement to come. And we're trying to grow late in space in terms of the types of formats and the amount of value that we deliver to our subscribers. So one thing that we've been trying and experimenting with is this monthly recap that I started doing around August of last year, where I basically just take the notable news items of the month and then I sort them and categorize them according to some order that makes sense and write them down in the newsletter.
1:04And this last December recap was particularly exciting because it seemed like it popped off in a number of areas, particularly with the AI breakdown. Our friend NLW featured it on his podcast. And I figured we can just kind of go over that as a way of setting the stage for 2024, but also recapping what happens in 2023. Yeah. And people always ask me if December is like a slow month, but I think you almost broke Substack with how many links we had in the thing. No, we actually did. So a lot of people commented to me about the formatting issues within the newsletter that I sent out. And I know that they are there, but I couldn't fix it because Substack was broken by us with how long it was.
1:41Oh, man. But we had this kind of like four main buckets called the four words of the AI stack, data quality, and I guess like data quantity as well, in a way. The GPU rich versus poor, which we have a whole episode about with Dylan Patel. Multimodality. We're actually recording tomorrow with Luma Labs about their new 3D model. So we went from text to image to 3D video. I wonder what's next. And we're going to release Hugging Face as well. I guess I've been thinking about calling it multimodality 101 because the first modality beyond text that you should really pay attention to is vision. Right.
2:17Yeah. And then the RAG ops war. I think that's a... I don't know what to call it. I don't know if you would have called it anything else. This is my... I don't know. But I think beginning of last year, that was like kind of the hottest space because there wasn't much open source model work. And I think over the last maybe like four or five months, everybody's still focused on fine tuning Lama 2 and like a DPO to improve these models, Mextral and all these things. And people forgot about our friends at LankChain, Lama Index, and some of the things that were maybe top of mind. Vector DBs, you know, it seemed like everybody was releasing a Vector DB.
2:51early in the year. Yeah, I think that I'll be very surprised if any new VectorDBs come out this year, with one exception, which is something I'm keeping an eye on, which is TurboPuffer. I don't know if you've seen them going around. Yeah, all the smart people seem to be adopting TurboPuffer as the first serverless VectorDB. Yeah, no, and we're going to have definitely Jeff and Anton on the podcast at some point. I know they're going to be fun. I should also mention, the reason I selected these four wars was a process of elimination of wars that I think ended up not mattering. So for those who don't know, inside of my writing, I often include footnotes that are in themselves, just essays in footnotes.
3:33And so I think it's also notable the things that people thought were hot that were less hot than expected. So it was agents, definitely less hot than at the start of 2023. And then this one is very controversial, non-selection by me, I think. Open source AI is not a battle in the sense that I don't think there's anyone against open source AI, everyone is on one side. There's no opposing side apart from regulators. But in my mind, when I think about for engineers, engineers are all universally in favor of open source models. So there's no battle here. Everyone just wants it to improve. So it's not interesting to write about.
4:06We just want more open source. Yeah. The only battle is people offering inference on it. Yes. Killing each other in the process. Yeah. So I classified that as a GPU rich versus poor war. But maybe there's a better way to classify that. and you can give me some feedback on that. It's a struggle to try to categorize the world. Code models as well. I was very struck by a conversation I had with Poolside, I So Can't from Poolside. So they haven't been on a podcast yet. They're kind of stealth still, but they had a very, very notable fundraise. I think they had like$50 million raised. I think even more, yeah.
4:36For a seed, spending most of it on GPUs. And my conversation with Iso, he was like, hey, you know, like Replit was like one of our podcast's early biggest winners. Replit didn't really follow up with, They announced their 1.5 model, but it's not really widely used beyond Replit. There's StarCoder, there is CodeLlama, but it's not really... For how important code is, it doesn't seem like as big of a battlefront as just general function calling, reasoning, these other kinds of domains. And so I thought it was just interesting to note that even though we, as a podcast, try to pay particular attention to developer tooling, to code models, We interviewed Cursor, Find, Replic, Codium, and Hugging Face.
5:18These all seem like very small compared to the amount of money being thrown, the amount of heat in the other domains. And I don't know why that is. Yeah, I think it's maybe the fragmentation of the tooling. You know, like most people in code are using VS Code, Cursor, GitHub, one of the three. So there's maybe not as much experimentation versus with text. People are just trying everything. It's hard to try a code model. You know, I see code models being released, but it's not super easy to just plug it into your workflow. So I think engineers like myself are just lazy. Like, hey, I'm having great success with whatever I'm using.
5:53I don't really want to go there. Special case form of code is SQL and the semantic layer data engineering type things. We also had two guests on there from Seq and Cube. And we also talked to a bit of Databricks, a bit of Julius. Yeah, and we have Brian from Hex. And Brian from Hex. Does he count? I don't know. Yeah, I know. Yeah, yeah, yeah. I guess the Hex Notebooks, yes. Hex Magic, yes. Rexys, it's a different beast. Anyway, but yeah, I think people who come to AI engineering for the AI might actually end up finding themselves in data engineering in the end. In traditional ML engineering in the end, they might have to discover that they're doing Rexys and all the stuff that gets swept under the rug in a demo becomes their job.
6:36And I'll probably say just because we didn't select a theme for last year doesn't mean it wasn't important. It just wasn't top of mind yet. And maybe I think that would be an emerging theme this year. Yeah, I think that's kind of the consequence of the low background tokens, like the end of the low background tokens. Can you explain what you think low background tokens? This is our November recap. Yeah, the comparison that our friend Jeff Huber at Chroma brought up is steel before the atomic bomb creation. So steel before and no radiation in it. After all the testing, a lot of steel had radiation embedded in it.
7:09So it was really precious to get low background steel, meaning with no radiation. And same with tokens. You can assume that any internet content from three years ago, it's just internet. It's like people writing, it's not models writing. Instead now, anything we're going to get on common crawl updates and things like that, you never know if it's human written or not. And I think that will put more work on data engineering, right? because even basic stuff like checking if a text says, as a model created by OpenAI, it's going to be important. So people have just been blindly taking all the data sets offered by Eleuther and Common Crawl and all these different things, assuming that all the data in it is good.
7:49I think now, how do you build on top of it? And we've seen the New York Times lawsuit against OpenAI. We've seen data partnerships starting to rise in different companies. I think that's going to be one of the bigger challenges. and maybe we'll see more of the work that Databricks has done to build the Dolly 5K instruction tuning. Just first party creation of data. It's like you got people sitting at their desk every day. If everybody wrote five Q &A pairs or things like that, you would have a massive unique data set for your model. Yeah, for people who missed that episode, that was one of our early episodes as well.
8:22And Mike Conover since left to start BrightWave, which I'm sure we'll have him back this year. Yeah, they're doing a lot of interesting stuff. I think the next episode will be very cool. So how do you want to tackle this? Do you want to just kind of go through the four wars? Yeah, let's do it. You created this like Wikipedia-like infographic for each of them. So yeah, I should say like the inspiration for this actually was during the Sam Altman leadership battle, people were making mock Wikipedia entries for the debate and for like who's on the side of the D-cells and who was on the side of the EX.
8:54So I like that format because it's very concise. It has to list the key players and it's kind of fun to think about who's on what side and think about what is important and what people are battling over. And I think it is important to focus on key battlegrounds as a concept because there's so many interesting things you could be talking about in AI and they're not all equally interesting. So how do you decide what is interesting? I think it's money, it's power, it's people, it's impact, that kind of stuff. And so, yeah, that's what I ended up doing. And fun fact, the way I did this was I actually edited the HTML on Wikipedia, and then I just screenshotted it just to get the formatting.
9:33Good old developer tools. Developer tools is all you need. So the data war, belligerence. Yeah. On one side, you have journalists, writers, artists. On the other side, you have researchers, startups, synthetic data researchers. I guess maybe we want to talk about what are the axis of war. So one of them is attribution, right? I think there's a varying spectrum of how comfortable people are about this data going into a model. So some people are happy to have your model trained on it, no matter what. Some people are happy to have your model trained on as long as you disclose that it's in the model.
10:12Some people just hate that you trained on their model. And some people, like the New York Times, want you to destroy any artifact that might have touched your article. So that's kind of what we're fighting on. I just want to make it clear that it's not just like you should never use the data or you should always use the data. I think people are just trying to figure out what's the right form of attribution and how do I get paid as somebody whose data ended up being in this training. I think we're giving everybody a lot of great tokens with Latent Space because we do full transcripts on everything and we're happy for people to train models.
10:45Oh, yeah, please train a Latent Space model. Yeah, we would love it. So that's kind of what we're fighting on. And anything that people should keep in mind about this war and like maybe some of the campaigns that are going on? I think the New York Times one is probably going to go to the Supreme Court. It is very, very critical. It is landmark war that will probably decide what fair use means in context of AI. And I recommend, I think The Verge did a good analysis of this. Platformer maybe did a good analysis of this. There are like four criteria for what fair use is. and everyone basically converges onto the last criteria, which is, does your transformative use of my copyrighted material diminish the market for my content?
11:24It's very hard to say. I would suspect that yes, in some capacity, in some amount, but good luck proving that in a court of law. And I think a negative ruling on open AI would seriously stall the progress of AI. And that's bad for humanity, but good for content creators and writers. So obviously, we want them to be adequately compensated and recognized for their work. There's no easy outcome here apart from the existing copyright system, which is also somewhat broken. And it's just a very, very tricky, challenging case, I think. It's funny because we had something. I was a community moderator at a website called Rap Genius, which was a lyrics annotation.
12:03And there was a similar thing in maybe 2014 where the music labels basically came to the website and it's like, hey, this is not fair use. you know like you cannot reuse the lyrics to the song and eventually the website made deals with the record labels to like be able to do this and then google was stealing the transcripts to put in like the enhanced thing and they proved it yeah yeah we did all the like a busy like the things on the i some i's we put the dots like the accent and that's how it made oh i thought it was uh i thought they just varied the spacing or they like use the different kind of spacing in the unicode I think it was the IE thing, but maybe, I mean, this is like almost 10 years ago.
12:42So Rap Genius proved it by injecting some data poison into their corpus, and then Google reproduced it faithfully. And so therefore, they proved that Google is scraping Rap Genius. Did Google have to pay Rap Genius money in the end? I don't think so. There was also another issue with Rap Genius that we had that got blacklisted by Google for like, there was like a lot going on. But anyway, this is not a Rap Genius special. Yeah, I mean, ultimately, like, I think that we do need quality data. I think that if this case is contained to the New York Times, the New York Times worst outcome is that they will substitute it with Washington Post and they substitute with The Economist or like the second or third ranked newspaper that is the most friendly to AI.
13:21And then the New York Times will realize that actually their words are not as not that much more valuable than other words. Then the value of the content comes down very, very dramatically. I think it will be interesting. But yeah, I do think it's overstepping their bounds to call for the destruction of LGBTs. that's probably for sure. Then the bigger problem I have is with Stack Overflow and Reddit, which I named as the side of the New York Times. They have effectively shut down their APIs in order to try to train their own models. Probably same as Twitter, actually. I should probably have put Twitter.
13:51I put Twitter on the wrong side, maybe. I don't know. Twitter is on both sides. Elon is on every side, the side of chaos. Yeah, what this is, is basically every UGC, users generated content company of the 2000s and 2010s, now has a giant pile of user content that becomes valuable data that used to be open for researchers to scrape and train models. Now all of them are locking in their walls, right? Behind their wall gardens and then trying to train their own models to boost their benefits. So this is a locally optimal outcome for them, but a globally suboptimal outcome for humanity. Because why should we care about the closed garden of Reddit, you know, the Reddit model, the stack overflow model, the X model, as opposed to like it being a part of a data mix of 20 % Reddit, 20 % Stack Overflow, 20 % X.
14:37That seems like a much better outcome for the world, but everyone is acting in their very narrow self-interest in trying to make their own model, which is probably going to suck. Right. So next, Bor, after you get data... We should mention synthetic data. Oh, yeah. So what happens when you run out of human data? You make your own. So I would say when I went to New Rips, that was the number one discussion out of every single researcher's mouth. There is a lot of research coming from both, I guess, the big labs as well as the academic labs on what good synthetic data looks like. I don't know if you've talked to any startups around that.
15:12I just talked to Luis Castricado the other day, and he is promising a very, very interesting approach to synthetic data generation. I think his phrase for it is pre-trained scale synthetic data, as opposed to what the news research and the other open source communities have been doing, which is fine-tuned scale synthetic data. And so he wants to create like trillion token data sets that are all synthetic. And I'm like, okay, that's interesting. But also at the same time, these are all just downloads from GPT-4 or something else. Lewis is very aware of that and he has a way around it. I don't really understand it, but he claims that that's a good way around it.
15:47Andre Karpathy at NeurIPS highlighted this paper from DeepMind where they were bootstrapping synthetic data that could be verifiably proven correct. So specifically in math and in code. where there is a correct answer. So yeah, that makes sense. You can solve the synthetic data problem that way. But what about beyond that? There's just no answer. And wasn't part of the issue also that the way that the phrases are constructed and all of that in synthetic data ends up kind of making mode collapse even worse? Because one thing is right or wrong, right? The other thing is every sample is written the same way.
16:23Or as a similar, since it comes from a certain model, kind of as a similar route. of structure? Yeah, so I mentioned this in the Best Papers discussion with John Frankel. So the basic argument is you already have a flawed distribution from a language model. You are resampling that flawed distribution to double down on that flawed distribution. There's no extra information from humans. So on principle, how can this work? And so the only conclusion there is you don't need it to emulate a human. You need it to emulate a useful assistant, however you define it. So I think the goal of synthetic data is less to emulate human speech because that is basically solved it is now more to spike the distribution in useful ways and that's a phrase i borrow from kandron from imbue but anyway so i think that synthetic data will be a giant theme for this year and not least because the human data is being locked up behind walls so like it's a very very clear trend this is probably the most amount of money after gpus will be spent here so one war i did not put here was the talent war right like the war for phds and smart people.
17:26But when you break down what the talent people do, one is they make models and they run inference on GPUs or they run training runs on GPUs. But the other is they clean data, they find data, clean data and format data. And so yeah, these are all just proxies for the kind of talent that is flowing back and forth. And ultimately, I think you have to focus on what they're working on, the visible output of what they're working on, which is data. Awesome. All right, let's talk about the GPU inference war. I think this is one that has been heating up. And we actually have a bunch of these folks coming on the podcast in the next few days.
17:57Are we calling it Compute Month? Yeah, we can figure out a name, but we have Modal, Together, Replicate, there's a lot coming up. But basically, the Mixed Raw release, the MOE model, was kind of the spark of the war. I think the price went down like 90 % in one week. Yeah, I wrote 2-2-2 times, but yeah, 1 divided by 2-2-2 is whatever the... Yeah, and then there was like the benchmark drama between together and AnyScale on whether or not which one was faster and whether or not the benchmark was really reflective of performance. Yeah. This is very surprisingly ugly in a way that I think usually people try to respect each other's work and play nice and say nice things when people release stuff.
18:39Even if it's a competitor, you say nice things or you don't say anything at all. AnyScale, for some reason, they released a benchmark on which, of course, AnyScale looks the best. why would you release a benchmark where you don't look the best but then basically everyone featured in that benchmark didn't like it of course i do think there's some like methodological things so so for anyone doing benchmarks you have to understand that there's a real real real difference between like a public benchmark that is meant for just limited testing compared to okay if you're load testing us or if you're seeing what a real enterprise customer would see you have to give them a heads up you have to get a different api key a different endpoint and you test the real infrastructure not the demo one this is very common for infra companies and i think any scale just neglected that and it hurt their credibility like any scale is not new at this game like they should have done that but what was interesting was this benchmark drama reached even beyond any scale and we're gonna have sumith on and he's gonna talk about like why he weighed in because sumith doesn't represent any inference for it he just works at meta but he felt like this was uh this is a very interesting debate and i think we'll see more of this you have been a data investor for a while like database companies always do this and i think now we're just seeing this kind of fight come into the inference space yeah i think the hardest thing is the end customer cannot replicate it you know so like if you give me like a postgres benchmark i can run postgres on my macbook you know and run similar ones i think with models it's just impossible so people tell you this is the benchmark and you're like okay i have to go sign up to every single cloud now to try it it's just not easy.
20:14And we talked about this in Benchmarks 101, which is same with model benchmarks, right? Just like, oh, this model is so much better than this. And then it's like, did you train on the questions? And it's like, what? Oh, I don't know. So, and again, it's hard for people to just like run the models and test them, you know? So like there's a lot more weight, I think, in AI on benchmarks that there is in traditional software because nobody buys Upstash of a Redis cloud or whatever just based on a benchmark. They try them and check performance and whatnot because they have real production scale workloads.
20:46Here, it's like nobody's really doing anything with these models. So it's like whatever any scale says, I guess it's good. But then customers are going to go try it and just decide for them what the right thing is. Yeah. And I think it's important to understand it is not just about cost. I think what the price wall represented was a race to the bottom on cost. And you're like, okay, Deep Infra, which is a company, the name of the company is Deep Infra. Deep Infra has promised to just always be the lowest cost provider. Like, okay, fine. That's a good value proposition. But you're not only optimizing for that in a production application, right?
21:18You're optimizing for latency. That's one thing. You're optimizing for uptime. That's something that you can only earn over time. You're optimizing for throughput and other forms of reliability. It starts to tail off beyond that. But there's three or four dimensions that really, really matter, right? If you're not table stakes on any of those things, you're out. You're just out. Actually, there was a really good website that was released just this week called Artificial Analysis. Did you see it? Yeah, so this is what the industry needs, which is an independent third-party benchmark pinging the production API endpoints of all the providers and giving a third-party analysis of what this is.
21:54I actually built a prototype of this last year. Yeah, I was going to say. But I didn't like maintaining it. I'm glad someone else is doing it just because I don't want to keep up with all these things. But still, I think it's a public service that somebody should do. And so I'm glad that they did it. I think they did it very well. So yeah, I think that is where the, I guess the inference drama is ending for now. I don't think, you know, I haven't seen any continuing debate there. The only other thing that, you know, I did some extra work on this for the recap, which is like, are they losing money?
22:25You know, are they pricing correctly their tokens from Mixtrall? And I actually managed to go into Dylan Patel's write-up of the Mixtrall price war. And I think I reasonably worked out that you can serve mixed trial and the lowest you can possibly charge if you take the most aggressive amortization of all your capex and all that is 50 to 75 cents per million tokens, which is what perplexity prices their mixed trial at. And perplexity is a very smart player. They're not even an inference infra provider. They're just doing this for fun. But they're like, we don't want to lose money on this. We will provide it at cost.
23:01This is what cost is to us. so that means perplexity provides it at 56 cents per million output tokens that means any scale which is 50 cents Octo AI 50 cents Abacus AI 30 cents and deep infra 27 cents they're all losing money because we think that the break even is 51 cents and even that is like a full batch size and kind of max no no no max utilization I assume 50 % utilization so like very like you talk to practitioners very very good is 60 % average is like 30 40 so i just i say 50 right you assume 50 percent batch like 16 100 tokens per second generation that's also very very high these are all very favorable numbers like probably the real number is closer to 75 cents per million than 50 cents per million anyway anyone charging under 50 definitely losing money so then it's like okay either you don't know what you're doing which in which case good luck or you know what you're doing and you're purposely losing money for something and what is that and i don't know but i think it's an interesting aggressive strategy to pursue if you are doing it on purpose so this is something that like the classical like walmart would have a lost leader like they they really really on purpose lose money on things so that they get you in the door to try things out like i don't know if that makes sense to you as a yeah it's like the well it's like all the you know the candies are placed at the cash register because maybe you just went to get the thing on discount and then you buy a KitKat or whatever and then make money on the KitKat.
24:29They all have the Pokemon trading cards at checkout now. So if you bring your kids or buy the discounted whatever for you, then you end up spending more. But to me, the thing is like, where's the checkout register where you upsell people with these things, right? Yeah, I don't know how you upsell. That's really the big thing. I don't know. I'm curious to see. I don't think Cloudflare still has a life. I wonder what they're going to charge for our workers. Yeah. They cannot serve mixed trial. Their GPUs are too underpowered. Cloudflare AI is like very good marketing for very, very underpowered inference, right?
25:03Yeah, well, I don't know. I think it all depends on like what is going to be needed, right? So they have that missed trial 7B right now. Yes. But they cannot serve mixed trial. Yeah, yeah, yeah. Okay. I wonder, but I think they don't want to get into this race right now probably. No. Yeah. So yeah, I'm curious, going back to the last leading, It's like, is there going to be a better model that comes next that they hope that you already integrated their thing with? You know, if you're using Together to serve mixed raw and then something else comes in that you're going to replace mixed raw with, hopefully you're still going to use Together and they're going to get better unit economics on it.
25:41I don't know. It's a good question. It's a good question. Thank you, VCs, for paying for all of our imprints. No, no, no. I think these are, you know, everyone in here are grown adults. they're smart investors I'm sure there's some kind of long term strategy here and I'm trying to figure that out like assume that people are smart and then what will smart people do yeah I think it's the same with Uber right it's like how could it have been so cheaper at the start you know like you look back at DoorDash all these things it's like like last year was a great year for Uber all my Uber friends are like suddenly very rich again one thing I will mention on like the engineering sort of technical detail side is The rise of Mixture of Experts is something that we covered in our podcast with George and now with Mixtrall.
26:28And it represents the first successful, really, really commercially successful Sparse model. And Sparse in a very interesting way, in a sense that the divergence between the amount of compute you need at training versus the amount of compute you need for inference continues to diverge. But also in a weird way where you need to keep all the weights of the MOE model loaded. even though you're not necessarily using them at all times. So basically what I think that is, is like, I think that that is going to impose different needs on hardware, different needs on workload, different needs on like batching optimization.
Read the full transcript
27:05Like Fireworks recently announced Fire Attention where they wrote a custom cruder kernel for Mixed Draw on H100. It's like super, super domain specific. And they announced that they could, for example, quantize from like 16-bit down to 8-bit with like no loss in performance. Like all these magical details emerge when you take advantage of like very, very custom optimizations like that. I think like the rise in MOEs this year is going to be, going to have very meaningful impacts on the inference market and how it's going to shape how we think in price for inference. It may not be that we have this sort of input token versus output token paradigm for long, particularly because we have things like different forms of batching, different forms of caching.
27:44And like, I don't really know what that looks like, but I'm very curious. I see a lot of opportunity here. If I was an inference provider player, like that's something I would be trying to offer to people as a way to differentiate because otherwise you're just an API. Yeah yeah no it was in a way counterintuitive because most of the struggles with inference as well are just like memory bandwidth you know so we have now models that scale worse a higher batch you know but I'm glad I'm not in that business I can tell you that as far as there's so much work to be done at like so many low levels of the stack you're already trying to provide value to the customer on like the developer experience and all of that but you also have to get so close to the bare metal to like make this model like writing a kernel imagine if you had to write you're like a cpu cloud provider and you have to like write instruction sets it's like just nobody would get in that business you know so i i salute all of our friends at compute providers doing this work and i mean together it's doing so much for like 3dial and like flush retention 2 and and whatnot so yeah so and and that's something that i would I will leave as the last part of this sort of war of GPU rich versus poor.
28:51The GPU rich people are the model trainers and the infra providers. They say like, we have the GPUs, come use our GPUs, and then we provide you the best inference, right? And that's what we've been discussing so far. On the other side, on the GPU poor side, are like all the alternative methods, right? The modulars, the tiny corps, the QLORAs, and all the other type of stuff. I even put consistency models in there because, you know, Any efficiency or distillation method where you reduce your inference or GPU usage by like 25 to 40 times, it's a GPU-poor-friendly approach. So I will also put Apple and MLX in there.
29:29And that's also like Apple is finally making moves in inference. And that will be a game changer for local models because then you just don't need any cloud inference at all. You just run it on device, which is fantastic. And then obviously RWKV and Mamba and Stripe Taina from together. like all those like emerging models i don't know there's something i've been worried about for a latent space how much attention should we give to the emerging architectures because there's a very good chance that one these things don't work out two they take a very long time to work out and then three once they work out they're like for limited domains and like not super usable so i don't know if you have opinions on that i can follow up with one conclusion that i've had but I want to throw that question open to you.
30:12So the one conclusion is RWKV and the state-space models, including Mamba, have historically just been pitched as super long context models. And I'm like, that's not something I need because I'm okay with 100K context. I'm okay with RAG and recursive summarization, all those techniques to extend your context, like rope and yarn and all these things. So I'm like, why do I need million context models? Why do I need 10 million, 100 million, 1 billion models? Like, well, why? So the easiest argument is, oh, you can consume very, very high bit rate things like video and DNA strands. And then you can do like SynBio and all that good stuff.
30:51And I'm like, okay, I don't know anything about that. Like what happens if like you hallucinate one wrong chain in your, you know, the DNA strand that you're trying to synthesize? Good luck. I don't know. That's why I've been historically underweighting intentionally. our coverage of state-space models and the non-transformer alternatives until Mamba. Mamba really changed things where basically for the same amount of compute, you can get a lot more mileage or a lot more performance for the same size of model. Now it's an efficiency story. Now it's a GPU poor story. It is no longer a long context story.
31:25It is just straight up, we are strictly more efficient than transformers. I'm like, oh, okay, I can get that. Does that change anything? I don't know. No, that makes sense. I think people look at the slope, right? Which is like, oh, you can get the context. higher and higher. But in reality, it's like, if you kept the context smaller, instead look at the anti-slope, so to speak. It's like same context, it's like a lot less compute. Yeah, so that was not clear to me until Mamba. And so I think that's interesting. There's a concept I've been trying to call the sour lesson. You know, the bitter lesson is stop trying to do domain-specific adjustments, just scale things up, and it's going to work.
31:59That's general intelligence. General intelligence dislikes any attempt to imbue inside of it special intelligence. like if you have like any if switch case or if statements or like if finance do this if something do that don't bother just just scale things up and it's going to do all of them simultaneously all better at once that's the bitter lesson the sour lesson is a parallel is a corollary which is stop trying to model artificial intelligence like human intelligence right the neuron was inspired by the brain but doesn't work exactly like the brain machine learning uses back propagation the brain does not use backpropagation.
32:34We keep trying to create alternatives to transformers that look like RNNs because we think that humans act like RNNs. We have a hidden state and then we process new data and we update that state. But maybe artificial intelligence or machine intelligence doesn't work like that. Maybe we just fail every time we try. So that's the sour lesson. Every time we try to model things. And my favorite analogy, I actually got this from, I think, an old quote from Sam Altman, who was like, you know, we made the plane, the airplane. It was inspired by birds, but it doesn't work anything like birds. And it works very efficiently.
33:10It's probably the safest mode of transportation that we have, and it works nothing like a bird. So why should artificial intelligence work like human intelligence? And that is the philosophical debate underlying my continued cautiousness around state-space models. I feel very vulnerable saying this because I don't think there's any justification once you look at the empirical results or the mathematical justifications for these things. But there is some grounding in philosophy that you should have when you think about, does an idea make sense? Is it worth exploring? Yeah. I think now there's a lot of work being put into it, right?
33:47And I think transformers have shown enough success that people are interested in finding the next thing. So before, it wasn't clear if transformers were really going to work. So people are kind of working on them. But yeah. Okay. Maybe in the 2025 recap, we're going to have. Yeah. I mean, we'll try to do one before that. So we actually have a link. I don't know if you know this. Shreya Rajpao from Guardrails. She's married to Karan from Hazy. Yeah. And so now he's started one of the other state-based model companies. I forget the name of it. So we'll see. I'm sure like this will be an emerging topic this year as well.
34:20So we don't have to wait till next year. Yeah. No, I think we're going to have maybe the sour lesson, you know, overview. I mentioned this in the Luther Discord, and then they were like, okay, so what is the spicy lesson, and what is the salty lesson? What is the sweet lesson? I want the sweet lesson. Sounds better. Cool. Talking about GPU port, let's do multi-modality. I feel that Stable Diffusion was the first GPU port model. Yes, absolutely. I don't know if I mentioned that. I just didn't mention it. Stability, I think in 2023, they shipped incremental things. I don't know if Stable Diffusion 2 was out there.
34:57But everyone's talking about XDXL Turbo, which is a form, which is the alternative to a consistency model, but looks like a consistency model. They ship video diffusion. They ship a whole bunch of stuff, but just wasn't as big as 2022 when they made a huge impact with stable diffusion. Yeah, I mean, it's hard to, it's hard to, it's hard to top that. But yeah, MidJourney has been doing great. Obviously, I actually finally signed up for a paid account last month. So MidJourney, yeah, yeah. I'm part of the$200 million a year that they're getting. What's confirmed is, I think like a, Business Week article or Economist or Information article that this team has now reached at least 200 million ARR, completely bootstrapped.
35:33I think their employee count is somewhere between 15 and 30 people. I don't know if you know exact numbers. I have heard rumors that their revenue is actually higher than that. That was what was reported. But it's between the 200 million to 300 million range, which is crazy. Especially if it's primarily B2C, which it looks like it is. yeah yeah it's like b to fiverr to b i think i think there's like a ton of you can see the majority specialist yeah you can like get in discord and see what people are generating you know and you can see a lot of it it's like product placement ads and a lot of stuff yeah and dolly 3 doesn't seem to have any impact on the dolly 3 got so much worse after the gpt4 really the all-in-one well first of all before you could generate four images and they had like very good vibes now the vibes are like boomer vibes oh no every time I generate something the images I have here are Dali 3 every time I generate something on Dali it looks like some dusty old yeah like I think it's a skill issue I think you're wrong no but that that was the great thing about Dali 3 right it's like it made the prom better for you yeah yeah yeah like before like literally like when it first came out I'm like hey make a coliseum with like llamas and it was like this beautiful thing i feel like now it's not i don't know again it's a model right so it's like maybe i just got unlucky yeah i'm in the wrong way exactly yeah there's there's a lot of players in this i don't even think i put like some of the players i'm really excited about like you know the image and team spit out to be to create ideogram you know that was a few months ago and they didn't put it here because i forgot it's too much i can't keep track of all of it i will just basically say that I do think that I used to, at the end of 2022, start of 2023, I was not as excited about multimodality.
37:22Obviously, I'm more excited about it now. I used to think that text to image was more like hobbyist kind of work, but$300 million a year is not hobbyist. It is not just like not safe for work because Midjourney doesn't do not safe for work. So it's real. It's a new form of art. It's citizen art. It's exciting. it's unusual and interesting and you can't even model this as an investor you can't even model this on an existing market because like there's just a market of people who would typically not pay for art and now they pay a little bit for art which is digital not as good as a human but it's good enough I use it all the time yeah I'm surprised I haven't seen a return of digital frames that were very popular during the NFTs boom people are like oh the very very first late in space post was on the difference between crypto and AI in this respect.
38:15So I called this multiverse versus metaverse. Crypto is very much about metaverse. Let us create digital scarcity and let us create tokens that are worth limited edition, that were something. And then you display it probably in your PFP as your representation of yourself. And what AI represents is multiverse, which is a very positive sum instead of zero sum, where if you like a thing, okay, I'll choose a different seed and I'll make a completely equivalent second thing. and that's mine. And that means very different things for like what value is and where value accrues. I mean, I still cling to the insight even though I don't know how to make money from it.
38:51Obviously, Midjourney figured it out. I think Midjourney like made the right approach there. The other one I think I'll highlight is Eleven Labs. I think they were another big winner of last year. I don't know. Did they announce their fundraise? I think so. Rumor is they're now Yeah, rumor is Rumor is I can say it. You don't have to say it because I only heard it from my friends. Rumor is they're now unicorn and they just focus on voice synthesis, which again, did not care about it at the start of 2023. Now we have used it for parts of latent space. I listen almost every day to an 11 labs generated podcast, the Hacker News Daily.
39:24I don't know what the room for this to grow is because I always think like it's, it's so inefficient to talk to an AI, right? The bit rate of a voice created thing is so low. It's only for asynchronous use cases. It's only for hands-free, eyes-free use cases. So like, why would you invest in like voice generation. I don't know, but like, it seems like they're making money. Right. Yeah. I mean, Sarah, my wife, yeah, she uses it while she drives to talk to ChatGPT. Just like, yeah. so ChatGPT uses their own TTS. Yeah. Yeah. Okay. But you can see the modality. You should bring Sarah in at some point, but what is our interview?
39:59We're doing a bunch of like home renovation. So maybe she's like driving to Home Depot and it's like, Hey, what am I supposed to get to replace the sink? You know, or okay. all these sort of things that maybe were like Google searches before. Now you can easily do eyes free, hands free. Yeah, a lot of people have told me about that and I just say when I'm by myself, I always listen to podcasts. So I don't have time for ChatGVT and ChatGVT, you know, probably the number one thing they can do for me is give me like a speed adjustment so I can listen to it. That's funny. Yeah, anyway, so like I'm curious about your thoughts on like how as an investor, I think this is the weirdest AI battlefront for investing because you don't know the time it's funny because there was a bunch of companies doing synthetic voices a while ago and i think the problem a lot of them got through like good arr numbers but the problem was like a repeatability or use case so people are doing all sort of random stuff you know and the problem is not it's kind of like mid-journey the problem is not that there's not maybe a market of interest it's like how do you build a venture-backed company where like a scalable go-to-market that like can go after a customer segment and like do it repeatedly.
41:09I think that's been the challenge. I don't know how 11 Labs is doing it, but you could do so many things with text to voice that is like, how do you sell it? You know, who do you call? Like, that's like the hardest thing, right? If you're raising like a series A, a series B, it's like, how are you going to invest this money in sales and marketing to get revenue back? It's kind of like the basic of it. And it can be challenging. That's why sometimes investors are like, you're making money and that's great for you. but like how there's no industry it's hard to like just tie together you know I would be interested because I feel like there's a category of companies in the early 2010s that did this meaning they offered an API with no idea how you're going to use it I'm thinking Twilio Twilio has a cohort of like sort of API first companies that are all like sort of Twilio inspired I think there's a category or a time in the market when it makes sense to just offer APIs and just let your customers figure it out and it's actually okay and then there's sometimes when it's not okay and i think the default investor mentality right now is that it's not okay if you don't know what your customer is doing i think truly is a funny example because i think in the middle 2010s uber was like 15 percent yeah but like i'm just i'm talking like move yourself back as to like to little seed investor to a series a investor they had no idea wasn't even around but i think the the thing now it's like text to voice is not new you know like that's really the thing it's like what's new now is that you can generate very good text to then feed into that model.
42:35So that changes why the market is interesting. But if you really think about it, the models today are a little better. They're maybe like 50 % better than they were three years ago. But the transformer models under the feed at what to say, they're like a billion times better. A lot of people use it for automated customer support, things like that. Before you had scripts they were reading, now you can have a transformer model converse with the customer so it makes it a lot more useful in cases but we'll see how yeah we'll see that changes okay the last thing i'll mention here why is this a war which is open ai and gemini and google are working on everything models versus each of these individual startups all working on their selected modality and so this is a question of like the big tech company is going to actually win because they can transfer learning across multiple domains as opposed to each of these things being point solutions in their specific things.
43:30The simple answer is obviously everyone will win. Right. Because the AI market is so huge. You know, there's a market for the Amazon basics of like everything, you know, one model has everything. And then there's a market for, no, like the basics are not good enough. I need the special thing. Do you have an opinion on when does one market win over the other? Or is it just like everything's going to win? Yeah, it's interesting. I think like it works when people wouldn't have used the product without the Amazon basics, you know? So like maybe an example is like a computer vision, you know, like, I mean, we have.
44:00Yeah, vision is so important now. Yeah. It's like, you know, before people were like, why am I bothering trying out to set up a computer vision pipeline and all of that? Now they can just go on GPT-4 and put an image and it's like, oh, this is good. I could use this for this. And then they build out something and maybe they don't use GPT-4V. They use RoboFlow or whatever else. That's kind of how I think about it. It's like, what's the thing that enables people to try it? you know so in a way the god model can do everything fairly okay it's like dali and mid-journey you know all these different things and maybe like the mixed role inference wars so like another example it's like i would have never put something in my app at like two dollars per million tokens but i did it at 27 cents per million token you know and now it's like oh no i should really do this it's a lot better so that's how i think about how the god model kind of helps the smaller people than build more business.
44:54Yeah, creates a category. Yeah, rag and ops. Yeah, last but not least, where to begin? We had almost all of these people on the podcast. They're honestly the easiest to talk to because they look like DevTools. And you are a DevTools investor. I worked in DevTools. I think they're also more mature as businesses. There's more of a playbook that is well understood by the customer. Like, yes, I need a new stack here. Maybe not. Okay, so my biggest problem with putting databases versus frameworks versus ops tooling in the same war is that they're not really a war. They work cohesively together. Except when one thing starts to intrude on another thing.
45:32And that's why I very consciously put together this sequence which is databases on the left, frameworks in the middle, ops companies on the right. What's the first product of Langchain? Langsmith, which is an ops thing. So now suddenly the framework companies are not so friendly with the ops companies because they're trying to compute the ops companies. And what are the ops companies trying to do? The ops companies are trying to produce SDKs that compete with frameworks. Okay, then what are the database companies trying to do? First of all, they're fighting between each other, right? There's the non-databases, all-adding vector features.
46:01We had some people approach us and we had to say no to them because there's just too many. And then there's the vector databases coming up and getting$235 million to build vector databases. You know, obviously you're an active investor in some of these things, so you cannot say everything. But just on databases alone, one of the biggest debates of 2023, where do you stand on the whole thing? That's the million dollar question. Well, one, in the start, there's kind of like a lot of hype, you know? So like when LangChain came out and LamaIndex came out, then people are like, oh, I need a vector database.
46:30They search vector database and it's like Chroma, Pinecone, whatever. But then it's like, oh, you can actually just have PG vector in Postgres and you already have Postgres. Did you know it could do that? People are like, no, I didn't because nobody really cared. So like there's not a lot of documentation. Same with MongoDB vector, Cassandra, all these things. Elasticsearch. You can actually put vectors and embeddings in everything. It's a different kind of index. And I think like Jeff and Anton also, what they always talked about even early on, it's like this is like an active learning platform.
47:01This is not just like a vector database. It's like, what do you do with the vectors? It's like what's most helpful. It's not where do you store them. So that's kind of the change. I think that was old Chroma, by the way. I don't know if that's the new current messaging. Well, but I think I'm just saying like to them, it's never about this is the best way to put a vector somewhere. It's like this is the best way to operate on the vectors. And the store is like part of it. But there's like the pipeline to get things out and everything. You have to build out a lot more. So I think 2023 was like create the data store.
47:34I think 2024 is going to be like, how do I make the data store useful? Because the vector storage has come out of its heist. So there needs to be something else on top of it. Unless they can come out with some kind of new distance function or something. They teased a little bit of what they're working on at the AI Engineer Summit, which, yeah, density and whatever other fancy formulas that Anton is cooking up. But yeah, I think I tweeted about this maybe like two, three months ago and I think I pissed off Chroma a little bit. But the best framing of what Anton would respond to here is what people are embedding within vectors is a very different kind of data from what is already within Postgres and MongoDB and all the others.
48:10In some sense, it's net new data. And that actually struck a chord with me because that's how I started to understand structured versus unstructured data. That's how I started to understand, you know, one of my kind of heroes is Mark, who's CTO of MongoDB. This guy was the former GM of AWS RDS. And for those who don't know, GM is like, you're the mini CEO of that business. And when you work at AWS RDS, you run a$1,$2 billion a year business. And now, and then he quits being Mr. ProSgress of AWS to join MongoDB, the enemy. When he gave that speech of like why he did he was like actually if you look at the the kind of workloads that's happening postgres is doing well obviously structured data always going to be there but unstructured data and document type data is just rising exponential rate even faster and like for him to say that uh means different things anybody could have said that anybody could have pointed made a chart that showed what he did anybody could have said that but for him to have said that i think it was a very big deal because he he's rich he doesn't have to work but he like believes in this so much that he was like, okay, I'll just join MongoDB.
49:15So I'm like, okay, there's a real category shift between structured data and unstructured data. I believe it. I don't think it's just that you can put JSONB inside of Postgres and be done. That's not a NoSQL database. Okay, fine. So what is this new thing of vectors? And how do you think about that as a new kind of data? And I think if there's a third category of something beyond unstructured data, I don't know what it is. Like context or memory or whatever you call it. Whatever you call this kind of new data, that might belong in a new category of database, and that might create the new MongoDB of this era.
49:47And it could be any one of these guys. Right now, Pinecone has the lead. I think they're a$750 million company. Valuation. Yeah. And then all the others are much smaller. So if this is really a new data category and there's room for a key player, then it's probably going to be one of these guys. By the way, I left out VV8 and I put QGrants in there. Do you know why? No. Anthropic and OpenAI both use QGent for their internal RAG solutions, which means that for whatever reason, we should probably interview QGent. They passed the evals when WeV8 and Milvus and all the others didn't, which is interesting.
50:23There's a lot that we don't know. Yeah, interesting. I think, going back to your point of LangChain building LangSmith, at some point, some of the vector databases are going to be like, why am I letting my customers use Lama Index? It's like, I should be the RAG interface since I'm owning the data. That's why I put them next to each other. Right now they're friends. Yeah, right now. Yes. I mean, if we think about the Jamstack era, you had Vercel started as Zite, which was just a CDN. And then you had Netlify, you had all these companies. And then Vercel built Next.js. And so they moved down from the CDN to the framework.
51:02And it's like, now they use the framework to then enable more cloud and platform products. Which way is it going to give this way? I think what we learned from before is that you rather own the framework and then have the cloud to support it than just have Netlify and not have your own framework. Just given the way the two companies are doing now. So for those who don't know, I worked at Netlify and I was very, very intimately involved in this. So we don't have to say any in private. No, no, no, it's fine, it's fine. It's well known that Versa won and Netlify has pivoted away to a different market.
51:35But is it over learning from an end of one example that you always want to own the framework. No, no, no, no. No, because then the counter example is the same, which is Gatsby. Yes. Where you own the framework, you don't own the cloud, and then you don't make money either. So it's kind of like, I think we still got to figure out where the gravity is in this market. I think a lot of people will say the gravity is in the model. A lot of people will say the gravity is in the embeddings and the data that you put into it. A lot of people don't know what they're talking about. So I think 2024 is supposed to be the year of AI in production.
52:06I think we're going to learn soon who bleeds into where. I think that statement is like the year of Linux on the desktop thing. It's just always going to be true. People are always going to be saying it. We're going to be here one year later and this year is the year of AI production. And it's always going to be incrementally more true. But what is the catalyst? What is the big event that you will point to and say, aha, now it's in production? I don't know. I think actually being that it's not in production. You know, like a lot of companies, it's funny, like one, they're just like an inherent timeline that large companies work within.
52:44GBD4 came out in like April. That's like eight months. It's like most companies don't buy things within eight months and like implement them. So I think like part of it, just like a physics time limit that like even people that have been really interested, you just cannot go through the whole process of getting them live to all of your customers. So I think we'll see more of that in good and bad, right? it's going to be a lot of failures and a lot of successes, hopefully. Yeah. Any other commentary on tooling, RAG, Ops, anything like that? I always tell people, as much as I'm interested in fine tuning, I think RAG is here to stay.
53:18Don't even doubt it. This is a necessary part that every AI engineer should know. Yeah. Well, I think, yeah, it's tied to the infinite context thing, right? I think the leftover question is like, do you want to have infinite context and hope that the model is good enough at parsing which parts matter to your query? Or do you want to use RAG and wrap very specific context injection? I think so far, most people will say, I'd rather do a context injection with just what I care about than put a whole document in there and hope the model gets it. But maybe that changes. I don't. There's no way it changes.
53:54Hey, you know, that's great for Lama Index. Yeah, yeah, no, it's great. It's going to make a lot of money, I guess. No, it's not clear that they're going to make a lot of money, right? Because they're just an open source project. I don't think they've launched a commercial thing yet. I don't think so because, yeah, Jerry was talking about it on the podcast, but it wasn't. Yeah. Yeah. So, I mean, we'll see what they launched this year. Yeah. I do have my notes. The year of AI in production. The year of Lama Index in production. Yeah. Okay. So that's the four wars. We also covered a bunch of other non-wars that we skipped over.
54:22I did remember that you actually just published a piece on the semantic versus - The syntax. The syntax. Do you want to cover that as an evolution? Yeah. I think like I kind of mentioned this a couple of times on the podcast, but basically the idea of like code has always been the gateway to programming machines and we spend a lot of time making it easier so you go from punch cards to like COBOL to C to Python just to make it easier for the person to read and write the code and through it we started adding kind of like this semantic functionalities in it so in Python you can do array.sort you don't need to know bubble sort you don't need to know any algorithm that you learn in school to do it and I think the models are kind of like 100xing this, which is like, now all you need to do is like create a signup form, you know, where people put a name email and send it to this endpoint.
55:11So it's going to be a lot easier for people that know the semantics of the business, which is, you know, your product managers, your business people, the layer that goes from customer requirements to implementation, basically, and have them intervene in the code. So, you know, how many times as an engineer, you have to like go change some button color or like some button size, like these small things that like you really shouldn't be doing. And now you can have people with natural language intervene in the code and write code that can actually be merged and put in production. I also wrote the bear case for it, which is like, we already have so much trouble getting engineering teams to collaborate and get all their changes together without conflicts and all of these things that maybe also having non-technical people trying to do things will be hard.
55:57and models, they just think about solving the task at hand. They don't think about, I've always told my engineers, it's like, you need to leave the code base better than you found it. If you're like writing something, it's like, just, we cannot always keep adding like a quick hacks, you know? And I think models are great at quick hacks, but sometimes it's like, oh, this is like the 16th bun that you've changed a style for, you should make a class for it. That's like the dumbest example. So I think if that happens, then I think I'll be a lot more bullish on coding agents. Until you can have non-technical people manually query models and look at results and then say this is ready to go, it's going to be hard to have autonomous agents do it.
56:38Yeah, so I actually had a tweet about it today because Itamar from Codium actually published Flow Engineering as his next evolution of prompt engineering. And they've been working on in-IDE agents. They call it agents. You can debate about the definition of an agent at the end of the day. My split of it is inner loop versus outer loop, which I think you understand that. Maybe I have to explain it to the audience because every time I talk about it to developers, they've never heard of it. So inner loop is everything that happens between a Git commit. Outer loop is everything happens after the commit is committed and it's pushed up for PR.
57:11So maybe that's too reductive, but that's something like that, right? Like inner loop happens within your IDE, outer loop happens in GitHub, something like that. Okay, so I think your conception of an agent is outer loop-y, especially if it's non-technical, right? Like the dream, like you mentioned sweep.dev in your write-up. And there's also CodeGen. There's also maybe Morph. Depends what Morph is doing. And there's a bunch of other people all doing this stuff. Even small developer was also like, you know, write in English and then create a code base. And I think it's just not ready for that.
57:44Outerloop is a mirage going to forever be five years away. And the people working on interloop companies have been the right bet. and you can work on inner loop agents. Actually, Code Interpreter is an inner loop agent in a sense of limited self-driving. It's kind of like you have to have your attention on it. You have to watch it. It can only drive a small distance, but it is somewhat self-driving. And so I think if you have this gradations in your outlook on autonomous agents and you don't expect everything to jump to level five at once, but if you have an idea of what level one, two, three, four, five looks like for you, I haven't really defined it apart from this concept of inner loop versus outer loop.
58:20But once you have to find it, then you can be like, oh, we're making real progress on this stage. And this other stage, too early for now, but at some point somebody will do it. Yeah. I think like, yeah, maybe level one, it's like, I think of it more as just the auto-completion in the IDE. Level two is like asking cursor, hey, how can I make this change? But then level three should be like, to me, it's like we need to separate the inner loop from the IDE. I need to make a code change. Sometimes I shouldn't go in the IDE. sometimes I should be in the UI of the product and say hey that needs to be changed kind of like all the preview environments companies want you to put comments the PMs put comments like how do you go from that to code changes there should be enough there to make the code changes happen you know through a supervised interface yeah that's outer loop yeah but I think what these models are doing is like change where the loops start and end you know because now you can create code in the outer loop before you couldn't do it.
59:18That's the dream. Yeah. Anyway, my focus right now, I'll say if anyone cares, is I think the only thing that's working is inner loop and you should just use inner loop things aggressively, build inner loop things aggressively, invest in them, and then keep an eye on the outer loop stuff because it's still very early. I did invest in CodeGen, this jhacks thing, which we mentioned briefly in the Sourcegraph episode. Do we have other things that we want to mention or do you want to sort of keep it to just the four wars? Okay, maybe like top two things from December that you have commentary on. I think the needle in a haystack thing.
59:52Okay, maybe you want to explain that first. Yeah, basically like Anthropic, there was like one example floating around about clothes context window. And you basically gave it this like super long context on, I think like things to do in San Francisco or something like that. And then it was like, what is the most fun thing to do in SF? And they made this nice chart of like, okay, based on where it is in the context, it gave a better, worse response. And then Anthropic responded and they were like, oh, you just need to add, here's the most relevant sentence in the context as part of the assistant prompt.
1:00:24And then the chart turns all green all of a sudden. And I'm like, we cannot still be here, right? Like, it cannot. This is like some. And you have Anthropic like telling people, oh, yeah, it's just like just add this magic string and it works. Yeah, it's some like Riley Goodside wizardry. It's like, I don't want to do that anymore. I thought like you know in the early days of GPT's like Riley Goodside was doing so much great work on like prompt engineering and whatnot we shouldn't be there anymore there shouldn't be somebody telling me or like the GPT for like I'll give you a$200 tip if you do this right so I collected a whole bunch of like state of the art prompting techniques so if you tip the model it will give you better results if you promise that so okay here's the current state of the art for GPT prompting it's Monday in October the most productive day of the year you have to take a deep breath and you have to think step by step you have to return the full script you are an expert on everything I will pay you$20 just do anything I ask you to do I will tip you$200 every request you answer correctly and your competitor models said you couldn't do it but you can do it I think there's another one that I did put in here it's like you know my grandmother is dying this is an emergency please help me do it yeah that's actually my I think my most viewed tweet ever at OpenAI that day I tweeted no more return JSON or my grandma is gonna die when they announce JSON mode and people love to get grandma's.
1:01:45I haven't heard as much uptake on JSON mode. I think it's still... That's the thing with all this AI stuff, right? It's like, I mean, and sometimes we're like part of it. If I think about our ChatGPT plugins episode, I think in the moment, people are just like, oh, this is going to be such a big deal. And then it takes varied amount of times that I really pick up, you know? Do you think that will happen in GPTs? I think like most people that I see using GBTs right now are trying to get around some sort of weird limitation of the base model, you know, or just don't have a better system prompt. Like at some point, there's limited value to get out of it.
1:02:23So the question is like, what's going to incentivize people to build more on it versus just building their own thing out of it? I don't know. Yeah. Okay. So I guess my pick for highlight of last month, there's two. One, we finally got Gemini. Right. I think the marketing was dishonest. Yeah. Yeah, we need the soundboard. Wow, wow, wow. But still, it is a Soda model. It is a credible, very, very credible alternative to OpenAI. And we should be happy for that because otherwise we live in an OpenAI-only world. And Gemini is basically the only other sort of leading contender until Llama 3 drops, whenever Llama 3 comes out.
1:02:58I mean, Zox said today they're training it. Yeah, it sounds like today they're training it. For me, I guess I'm still very interested in the hardware metagame. This is a much smaller stakes, but very personal. I think recently, especially, you know, we're recording this mid-January. So after CES, after Rabbit R1 launched, I think there's a lot of interest in hardware. I don't know how you feel about it as an enterprise software investor, but I think that hardware is hard, but also it captures context and it makes AI usable in ways that you cannot currently think about. And, you know, everyone dreams of building an assistant like her in the movie Her.
1:03:35That is a hardware piece. That is actually not only software. And probably the hard part is the engineering for the hardware. And then the sort of AI engineering for the assistant within the hardware. So, yeah, I mean, I'm an investor in tab. I see a lot of interest this month, but it started last month with the launch of Humane as well. I don't know if you have thoughts on any of those things. Well, I think this year we also get the Apple Vision Pro thing. So I think there's going to be a ton of experimentation. I think Rabbit got the right nostalgia factor. You know, it kind of looks like a toy-looking, yeah, Tamagotchi type.
1:04:08Game Boy Advance, something like that. I'm curious to see what you got beyond that. I think, yeah, I mean, obvious, like right where we have the studio building tab. And I think that's another interesting form factor. And I think if you ask them, I think in our circles, a lot of people are like, well, what about privacy and all these things? But he will tell you that we're kind of like a special group that most people value convenience over privacy, as you learn from the social medias of the last few years. So yeah, I'm really curious to see how it develops. I really like technology where you're slightly uncomfortable with it on a social level.
1:04:43For Uber, it was like dysregulation around taxis. For Airbnb, it was staying in strangers' homes. And now it turns out for OpenAI, it was training on people's content. Now it's becoming a matter of regulation. And OpenAI's data partnerships are a form of private regulatory capture, which is a playbook that is fantastic. I hope it was on purpose because whoever did that is a genius. So I'm like, okay, I do think that every great new company, especially on the consumer side, is provocative in that sense. They're doing something that is not yet kosher. And so I think the humanes, the tabs, anything that is working on that front where it's like, yeah, I'm not sure I'm comfortable with this, but maybe it could change.
1:05:24That is a really interesting shift. I'm excited from that point of view, but at the same time, most hardware companies fail very, very quickly. They have a very hot start and then everyone puts it in their drawer and then never looks at it again. So I'm very, very aware of that. Here's the core thing of it, right? Avi doesn't think it's a hardware company. Most of the cost of the$600 for tab is going towards GPT costs because it's actually processing context. And the whole idea is that context is all you need. In this world of AI applications, whoever has the most unique context wins. A unique context could be the quality data war.
1:05:57A unique context is I have Reddit info, I have Stack Overflow info, I have New York Times info. So if I have info on everything you say and do at all times, that is something that no one else has. And if it becomes a good store of that, then what can you do with that? So I'm most excited for him to expose the developer API because then I can come in and do all my software stuff. But he has to build the hardware layer and get acceptance for that first. Right. Yeah, no, I'm excited to see. I'm sure we're going to see a lot of people walk around with them. So I'm excited to see. Actually, so I think he doesn't like me because I asked for an off button.
1:06:31because I want to be able to guarantee you if you're having a conversation, I want to see it's off, right? It's kind of like, oh yeah, my phone is on silent mode, right? There's a physical silent mode button, but now he just wants it to be always on. That's a whole new market, like a soundproof, like soundproof storage for your AI pendant so that you can guarantee the person cannot hear you. Awesome. This was fun. Please, if you're still listening after one hour, 21 minutes, let us know what we did right, what we did wrong what you would like to see differently it's the first time we try this out but yeah awesome thanks for doing this cool
1:07:31Thank you.
1:08:01Thank you.
From the publisher
Note for Latent Space Community members: we have now soft-launched meetups in Singapore, as well as two new virtual paper club/meetups for AI in Action and LLM Paper Club. We’re also running Latent Space: Final Frontiers, our second annual demo day hackathon from last year.
Edit from March 2024: We did a followup on the Four Wars on the AI Breakdown.
For the first time, we are doing an audio version of monthly AI Engineering recap that we publish on Latent Space! This month it’s “The Four Wars of the AI Stack”; you can find the full recap with all the show notes here: https://latent.space/p/dec-2023
* [00:00:00] Intro
* [00:01:42] The Four Wars of the AI stack: Data quality, GPU rich vs poor, Multimodality, and Rag/Ops war
* [00:03:17] Selection process for the four wars and notable mentions
* [00:06:58] The end of low background tokens and the impact on data engineering
* [00:08:36] The Quality Data Wars (UGC, licensing, synthetic data, and more)
* [00:14:51] Synthetic Data
* [00:17:49] The GPU Rich/Poors War
* [00:18:21] Anyscale benchmark drama
* [00:22:00] The math behind Mixtral inference costs
* [00:28:48] Transformer alternatives and why they matter
* [00:34:40] The Multimodality Wars
* [00:38:10] Multiverse vs Metaverse
* [00:45:00] The RAG/Ops Wars
* [00:50:00] Will frameworks expand up, or will cloud providers expand down?
* [00:54:32] Syntax to Semantics
* [00:56:41] Outer Loop vs Inner Loop
* [00:59:54] Highlight of the month
Get full access to Latent.Space at www.latent.space/subscribe




