NeurIPS 2023 Recap — Top Startups

30 Dec 2023 · 2 h 42 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Latent Space: The AI Engineer Podcast - Episode Summary

Episode Title

NeurIPS 2023 Recap — Top Startups

Podcast Description In this episode, the hosts recap the highlights of NeurIPS 2023, focusing on notable startups and innovations in AI. The episode features interviews with several prominent figures from various AI startups, providing insights into their latest advancements, challenges, and perspectives on the future of AI.

Key Themes and Discussions

  • Startups at NeurIPS: Emphasis on the importance of startups in advancing AI technologies showcased at the NeurIPS conference.
  • AI Innovations: Highlights various innovations such as multimodal language models, code generation, and embedding techniques.
  • Community Engagement: Discussion on the significance of community and collaboration in AI development, particularly through open-source contributions.

Featured Interviews and Insights

  1. Jonathan Frankle (MosaicML/Databricks)
  2. Discussed the acquisition of MosaicML by Databricks for $1.3 billion.
  3. Emphasized the trend of everyone becoming LLM builders and the shift towards bigger models.
  4. Speculated on the future of multimodal models and their practical applications.
  1. Lin Qiao (Fireworks AI)
  2. Introduced Fireworks AI and their focus on inference platforms.
  3. Discussed the challenges in optimizing AI models for production use and the unique aspects of their technology.
  1. Aman Sanger (Cursor)
  2. Highlighted Cursor's funding from OpenAI and its rapid growth in the AI coding space.
  3. Explored the concept of caching in AI applications and its implications for AI performance.
  1. Aravind Srinivas (Perplexity)
  2. Shared insights on reaching 1 million app installs and the importance of real-time data in improving AI models.
  3. Introduced PPLX Online, an LLM API with no knowledge cut-off for developers.
  1. Will Bryk (Metaphor)
  2. Explained Metaphor's approach to building a search engine focused on handling complex queries.
  3. Discussed the significance of their unique algorithm in contrast to traditional search engines.
  1. Jeremy Howard (Answer.ai)
  2. Talked about the philosophy of open-source contributions and the importance of community in AI.
  3. Emphasized the value of diverse experiences in building effective AI systems.
  1. Joel Hestness (Cerebras Systems)
  2. Explained the unique capabilities of Cerebras's hardware in terms of speed and efficiency.
  3. Discussed ongoing research in sparse models and their implications for AI deployment.
  1. Jason Corso (Voxel51)
  2. Introduced Voxel51's toolkit for AI engineers and its focus on data-centric ML practices.
  3. Discussed the importance of exploring data sets for effective model training.
  1. Brandon Duderstadt (Nomic.ai)
  2. Provided insights on Nomic's products, including GPT for All and Atlas, and their focus on open-source language models.
  3. Explored use cases and applications of Atlas for unstructured data exploration.
  1. Luca Antiga (Lightning AI)
  2. Announced the launch of Lightning Studio, a new development environment for AI.
  3. Discussed the importance of reducing complexity in AI engineering workflows.
  1. Jay Alammar (Cohere)
  2. Shared thoughts on the evolution of transformer models and the potential future of LLMs.
  3. Encouraged engagement with the community and collaboration in AI development.

Key Takeaways

  • Community and Collaboration: The strength of the AI community is emphasized, with many startups relying on open-source contributions and collaboration for growth.
  • Emerging Trends: A clear trend towards multimodal models and the integration of various data types in AI applications.
  • Innovation in AI Tools: Many startups are focused on creating tools that enhance the capabilities of AI engineers, making it easier to build and deploy models.

Conclusion The episode provides a rich overview of the current state of AI startups showcased at NeurIPS 2023. It highlights significant advancements in AI technologies and emphasizes the collaborative nature of the AI community, paving the way for future innovations in the field.

For more insights and full access to Latent Space, visit [latent.space](https://latent.space).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:06Hello, hello. This is Swix back again with part two of our New Reps coverage. This time we're going to cover startups. And it's a special episode because this is the last episode of 2023. We are definitely looking back at the year with rose-colored glasses. This has been a fantastic year. We only started this podcast in February and it's grown so much. Thanks to all of you who've listened and give feedback and shared it with your friends. And we actually managed to invite a few of our former guests back on the pod together with some new friends and probably some new voices that you're going to be hearing next year.

0:38So this is not a hard-hitting interview series. You know, it's not that kind of interview. It's not that kind of podcast where we try to go too deep. Today, we're just going to go broad. And we're just going to check in on a bunch of startups that we like and monitor and we're present in Europe. So first up is John Frankel of Mosaic ML. We last talked to him in May for the MPT 7B episode. That's episode 13. And I have to say that was one of the best performing episodes of the whole year. So you're welcome to go back and listen to that if you missed it. And since then, they were bought by Databricks for$1.3 billion.

1:12And actually, during the interview, they were in the process of getting acquired. They just couldn't say anything about it. But it's definitely one of the biggest AI news of the year. And you can listen to what it's like or what's going through John's mind back then, as well as now, today, six months later. Hey, Jonathan. Welcome back to the pod. Thank you so much. This is an interesting place to have the pod under the overpass of interstate whatever it is. Yeah, interstate whatever in the city of New Orleans. Yeah, it's really good to see you. Since you were last on the pod, Mosaic got acquired.

1:41Yeah, thank you. I think you really deserve all the credit for this. No, you guys were sitting on that news, and we didn't know what was going to happen. But I did come away from your interview with a very, very high impression of, like, you guys are in a perfect place, perfect time. And it makes a lot of sense to join forces with Databricks. Yeah, they're kind of, I mean, I will say we really didn't want to get acquired. You did not? We didn't. I mean, we loved being independent. Sure. Like, we loved doing our own thing. Sure. But this just made too much sense. Like, you know, they do data. We do LLMs.

2:14We both do enterprises. We're all a bunch of academics. Like, it was just kind of, we couldn't think of a better match. And so it just, we kind of came to the conclusion, like, okay, I guess we can't not do this. Like, it's too perfect. Yeah, yeah. And you've done a bunch of other podcasts on the acquisition. So I don't, we don't need to retread. I'll send people that way. Just like what's new in Mosaic World? In Mosaic World, honestly, like we're just cooking. I think we've been a little quiet lately. Yeah. Or at least we look quiet from the outside. It is certainly not that we haven't been busy and it's certainly not that, you know, we're not doing cool stuff.

2:45Part of it is that, you know, getting acquired, there's a bit of administrivia involved. You know, we had to go through new employee orientation, get health insurance, you know, meet our amazing new colleagues. Part of it is like, you know, the field has moved toward bigger stuff and we've moved toward bigger stuff. So I think we'll have some exciting stuff to talk about soon. But my philosophy is always speak through the work. So I don't want to hype. I don't want to get people excited. You'll see the work and you judge for yourself. You talk about the industry moving towards bigger stuff. What trends are notable to you in the second half of this year?

3:16Everybody's figured out how to build LLMs. It's no longer a coveted skill of a handful of people. But now we've all become LLM builders. The field has kind of narrowed an aperture again. And yesterday when we were all figuring out how to train ImageNet, But, you know, now we're all figuring out how to build really big, really powerful models. And, like, that's now just an assumed skill. The rest is kind of what do you do with that skill? How do you build a product? How do you differentiate? What cool thing can you do that's different from everybody else? That's going to determine kind of, you know, what 2024 is going to be like.

3:47Yeah. I guess, like, a lot of people are banking on multimodal being, well, 2024 being the year of multimodal LLMs. I feel like that's a little bit too broad a brush. I don't know. Like, what's valuable in that front? I mean, so multimodal is going to be a huge deal. But it's already a huge deal. We have multimodal models. The Lava paper author I also interviewed on this pod. Yeah, Lava's amazing. I've been playing with it a bunch personally. It's awesome. And we've got Bard, and we've got Gemini, and we've got GPT-4V, and I'm sure there are going to be plenty more where that came from. I think the question is, as with all good things, cool promise is different than delivering value.

4:24Yeah. And I'm really curious, what do people genuinely do with this? in real production settings and the settings that will actually pay off the huge investment that's made to build these multimodal models. Right. I'm also kind of curious, like, are we going to start to see some big open source multimodal models? Like, you know, we've got lava, it's moving in the right direction. But like, is somebody going to, you know, build something that looks a lot like GPT-4V or something on that trajectory and kind of start another arms race in that direction? Like, it'll be interesting to see. I'm honestly pretty curious and I'm watching with bated breath for what everybody does.

4:56Yeah. I think in our chat earlier today, you said we kind of live in a diverse world where every company has kind of found its niche, maybe? If you want to go through that logic. Yeah, I'm kind of like, I think there are the optimistic and pessimistic scenarios for where we go. I don't know. I kind of think there's a boring scenario where everybody basically is building these giant LLMs and maybe language image models or what have you. And they're all kind of the same. It's just you've got the Google version and the OpenAI version and the Amazon version. And it almost feels like cloud providers in some sense.

5:28Like, you know, what distinguishes AWS from GCP? It's kind of, you know, where you are. Lightly different consoles and configurations. Yeah, it's like different interface. Maybe you prefer one or maybe like you've been using one for a while and like you're used to it. Or your IT person really likes this one because, you know, they used to work at that company or what have you. That would be a pretty boring world. But I think that's unlikely to be the case. I'm kind of, I'm looking at like, you know, all the cool stuff coming out of Gemini, you know, all the cool stuff coming out of OpenAI. And then like, I'm looking at Adobe, like, they're building, really Firefly, they're building like a different model with a creative perspective.

6:02Like, I'm kind of looking at this and going, maybe we'll have a wide diversity of models and everybody will be building models, like just by virtue of the fact that we need so much data to build any of these models. Everybody's going to play to their strengths and, you know, use every resource they have at their disposal. and Google has, you know, they have YouTube. I don't know if they're using it, but that's a cool resource. OpenAI has put a ton of energy into text data. Adobe gets creative people. And there are a few other companies where that came from. So I'm kind of honestly curious if we're going to just see really different models for different people.

6:35And I don't know, that's a pretty cool world to live in. That is great. We won't see this arms race. We'll just kind of see diversity. Yeah, and we shouldn't forget Bloomberg, which teased Bloomberg GPT, but that's a source of significant tokens, the financial world. Yeah, yeah. I mean, my whole business on the Mosaic and Databricks side is helping people leverage the data they have, so I'm excited about a world of diversity because not only do we have these crazy diverse foundation models at the larger scales, but everybody embraces whatever they have. Our friends at Replit do a code model, and Bloomberg does another finance model, and somebody does a healthcare model, and everybody draws on their strengths, and that's a cool world.

7:14Are you bullish? Every company training their own model? Sorry, that's a stupid question to ask you. I mean, I think, you know, I'll give you the honest answer because I think it's, you know, the business mosaic answer is, oh yeah, I'm super bullish. Everybody should train their own model. Come do it on Databricks right now. Like you should start from a base model that everyone shares, right? Like that's kind of useful. Maybe or like you work your way up. I think it's like there's a journey with playing with any of these models that may or may not end with training your own depending on where you go on that journey.

7:45You start by playing with an API, and maybe you do some retrieval, and maybe you do some fine-tuning, and then maybe you build your own model. But it's a journey, and I think there's a destination many people will get to that involves training their own model. Yeah, totally. What other trends are going on that you're liking or seeing or hating? Honestly, maybe this gets to the question of overall impressions of NeurIPS. I thought this was a pretty garden-variety NeurIPS in some sense, which feels weird to say in the age of, you know, chat GPT and everything else that's happened in the past year.

8:15Yeah. But this felt like the most normal conference I've had since like 2019. You know, I mean, we've had a pandemic in between and everything, but like the past couple years, I actually internally at Mosaic, I always do a long write-up of every conference and the trends I see. Okay. Like some public, some that are more relevant to what we're doing. And like a lot of the write-ups I've done over the past year or two have been like all about like the unease. sometimes it was just like my write-up for icml 2022 was all about people capitulating to scale and the five stages of grief and you know how different people were responding academics people at google brain back when it existed um you know all that stuff and it almost looks quaint to think about that it was insightful to say people have capitulated to scale in this day and age where you know, tens of billions of parameters looks mundane.

9:06But this kind of felt like, okay, the academics are trying to find their way forward. It's no longer just kind of coping and ignoring, but like trying to find their way forward. The industry folks are doing their thing. A lot more people keeping secrets than used to, but it's still like, you know, a lot of people also aren't keeping secrets and can talk about what they're doing still. So it kind of felt like, you know, equilibrium. I don't know how long it'll last, but this was a lot less of a frantic and stressful conference than I think I'm used to, at least in the past couple of years. You know, I'm in a new role in some sense.

9:36I'm on the business side now. I'm on the industry side and I'm trying to find my own path. But I felt like a lot of us have changed roles in some sense as the past couple of years have, you know, have taken place and everybody's moved around and figured out what they want to do. But we've all kind of found our place at this point. I feel like, you know, we may be in different places, but the ecosystem, the community has kind of sustained with, you know, a bunch of new PhD students and all that good stuff. like it's kind of you know i don't know it's nature healing in some sense from the insanity of the past couple years and a reminder that you know we're all kind of small pieces in a much bigger you know ecosystem and community yeah and it's still growing though uh apparently the latest stats was something like um 15 000 attendees oh my god oh my god i will say one big difference you know in the time right before the pandemic deep learning was getting so popular the conferences would sell out the day registration open like as a student you'd have to rush to register yeah or you wouldn't even get to go that i don't think is happening anymore it's here it's easier yeah and they're also live streaming stuff you know so yeah but it's kind of interesting that like i guess we've adjusted to the huge capacity and everything that's you know going on um but it's you know even so with it getting bigger it didn't feel that different to be honest maybe it's just that i joined the community when things were already big yeah but like you know there were some journalists here some vcs here but that's always been the case.

10:54Yeah, it's always been the case. You always have, you know, overrated, underrated papers. We'll maybe save the overrated stuff for later, but any underrated stuff that you want to highlight from this year, it doesn't have to be at the conference, but I just want to mind you for underrated papers that people should pay attention to. I'm going to flip this a different way. Okay. Because I'm not a fan of overrated or underrated, and I'm not, like, I'm not a fan of passing judgment on stuff. I just don't, like, far be it for me. I, like, one of my big gripes is, like, we shouldn't have best paper awards like and i say that having gotten one back in the day so i feel like i have the ability to say that not just out of bitterness but out of like recognition that it's dumb sure but test of time though test of time is great test of time is awesome yeah um you know and you know i look forward to everybody using lottery tickets in 2029 no um if you're working on lottery tickets you know there's a lot of other cool stuff out there but i think it's really i'll turn that question into like what areas should academics be thinking about.

11:49I don't know, what would I work on as a PhD student right now? Or what would I recommend a student work on? And all the biggest questions in the field come down to how you measure and how you evaluate. Those are just such fundamental questions. Until we know how to measure things, until we know how to evaluate anything, you can't really even do any science. We don't know what we're even talking about. And so I'm also thinking a lot about like synthetic data. Can we generate useful evaluation sets for all the little properties we want to find about an LLM? Creating data sets is really hard, but a model can help us do that.

12:17So I'm kind of curious, like, you know, Can we bootstrap the evaluation process with synthetic data, figure out good ways to help ourselves build good data sets? And then, you know, from there, maybe we can start to really take a bite out of the evaluation questions and get moving on the actual science of understanding what's going on with these LOMs. All that seems very academically viable. None of those require huge amounts of compute. They require creativity, ingenuity, but that's an abundance in academia, even when compute isn't. Yeah, I would say that that's actually one thing I've had a big delta on for this year.

12:48Yeah, tell me more. I'm curious. Synthetic data, I always thought it was, you're just kind of sampling from a known distribution anyway that you know it's imperfect and doesn't match human preferences. And it's Kanjun from Imbue that actually changed my mind on this. Tell me more. That is a smart person you're talking to. She's like, you actually don't want to match human preferences. You want to spike it in different ways and useful ways. And so you want to synthesize data in useful ways that don't necessarily match human preferences. And when she said that, I was like, oh, okay, I think I'm actually sold on this as a viable practice.

13:20I would actually make a completely different argument, but she's right. So I'm probably going to make a wrong argument now because Ken Joon is pretty much always right. And, you know, when she disagrees with me, it means I'm wrong. But, you know, the way that I look at it is synthetic data is not about like it's not about relying solely on the model. Like we as computer scientists love the idea that once you automate something, you fully automate it. Okay. It's really about like, how do you reduce the amount of work necessary to create something that's truly useful? And so synthetic data is not about, can we like whip up a data set automatically and then make a model better?

13:53It's about how can you use human time most effectively? And maybe labeling data or creating a data set from scratch is not the most effective use of human time. Maybe it's curating a data set that a model generated, you know, to pick the examples you like most and edit a few of them. When I think about the millions of different small properties of LLMs we want to study, like in some sense, the unit tests of LLMs that we want to develop, you know, that's going to require a bunch of tiny eval sets on specific, really niche things. It's really hard for a human to just write from scratch. Nobody has the time or patience for that.

14:24If a model can help you do it and you can curate, you don't end up in a full feedback loop. You have a human there, but you're just making better use of your mime. Yeah, it makes sense. I'll just observe that this sounds like weak labeling. And I talked to Raza Habib from HumanLoop, who actually pivoted away from weak labeling. Interesting. Tell me more. I don't know. I just think it might have just been too early. I'm still a believer. This is the thing about all of deep learning. You never know whether you're too early, and too early is often six months too early. It's no longer the, like, you know, Yoshua Bengio and everybody being 20 years too early.

14:57Or Schmidt-Huber. And Schmidt-Huber, of course. We have to salute, you know, Schmidt-Huber as well. It's not like being 20 years too early. It's like you might be six months too early and some crazy thing is going to happen. or like something will finally click. And there goes that. Yeah, yeah, totally. Cool. We're almost at probably your destination. The workshops tomorrow, you said, are like kind of the highlights for you for New Rips. Yeah, yeah. What should be my workshop strategy? You know, I don't, I haven't, I picked out a few, but. Oh, wander. How do you do New Rips well, basically? Wander and go to a lot of the poster sessions.

15:28Like the talks at workshops are always great. But, you know, often, honestly, the workshops are pretty eclectic in terms of talks. You try your best as a workshop organizer to put together a coherent program. but presenters are going to do what presenters are going to do, and you can't really stop that. But instead, I love the poster sessions because you get students who are working on really crazy creative stuff that isn't even ready for the conference yet. You're actually seeing things that have not been put out on Twitter yet, and that's such a nice change from Nurefs where all the conference papers have been out for months, if not longer.

15:59Wait, I observed the opposite. Things that have been on Twitter forever are now out of date, and there are posters because that's how long it takes to submit a paper. Yeah, yeah. But for the workshop poster sessions, it's the workshop poster sessions that are awesome because you're truly seeing stuff that was created this fall. Manage an archive yet. Nobody's talked about it. Probably makes no sense yet, but may evolve into something really cool. And so and you also like there's not as much competition to talk to the people. You can just kind of chill. So I love to like wander from poster session to poster session throughout the workshops because like that's my favorite part.

16:31I don't know, I can hear somewhat important people talk anytime, but it's like talking to the people and getting a glimpse of what might be ahead. Being able to say, oh my gosh, I remember seeing the poster for this paper that a year later becomes very important. And kind of asking yourself, is this nonsense or is this brilliant? And not actually knowing the answer, having 50 million people on Twitter having told you the answer. that's kind of I don't know it's it's fun it takes me back to like what the conferences were like for me you know when I was early in my career like you know it was just kind of some random people coming and chatting with me and you never really knew what was important and what wasn't but it was all kind of cool and fun.

17:09Use a formula hypothesis and you know search that way. Yeah so I'm looking forward to tomorrow if you find anything interesting just let me know and I'll go interview with them. I've been recording sessions with poster presenters all the time and I want to expose people who don't come to NeurIPS that this is what goes on. And there's so much, I found, so much talent that does a lot of work that you don't hear about them online because they're just not online or they just don't have the reach that I do. So I want to give them that reach. Yeah, I think there's like, I'll say two things kind of to close up.

17:40One is kind of that I feel like there's now so much hype attached to NeurIPS and iClear and ICML just by virtue of the hype that's attached to the field. I don't know, this feels pretty mundane and boring to me. It's really cool, but it's also just a bunch of academics walking around having boring conversations, getting coffee, and pretending to party. I definitely, my experience of... Pretending to party, I love it. No, I'll say that, I'll tell you my experience of NeurIPS last year. I don't know, these conferences have a reputation of being over the top with industry parties and things like that.

18:11And my impression was that was probably true in 2017. Like that year is known as the Neurips that broke Neurips. For various reasons, I wasn't there at that time. That was before I was even in the field. But my experience last year, especially post-pandemic, was a whole generation of students had like heard stories and these stories had been built up in their minds and they were trying to live out the fantasy of what they thought Neurips had been like. So these very boring happy hours, people tried to turn into ragers and it was hilarious. It was just adorable in some sense. So, you know, it's worth remembering like, you know, there's the fantasy and there's the reality And the reality is, you know, it's a boring industry conference where people are, or academic conference with some industry component where people are trying to make money and convince people to look at their posters and get a few citations.

18:53Lots of hiring. Lots of hiring. Lots of hiring. Lots of hiring. I think things have really settled into a new normal. And, you know, with all the hype and all the craziness over the past couple years, people feel like everything is just exploding and changing all the time. Like you see those LinkedIn posts of everything has just changed. I hate those. I hate those so much. I hate LinkedIn. And if anyone is a LinkedIn influencer, I hate you. But, you know, it's kind of like this felt like, OK, like maybe there's a steady state again. Maybe we can all catch our breath a bit. And it kind of felt like after a pandemic, after all the technical development that's happened in the past couple of years, like it's nice.

19:29We can chill like we can kind of breathe a little bit. And there's something really nice about that. Yeah, love that. Well, it's so nice to have you on again and chat and catch up. Thank you so much. It's good to see you. Thanks for jumping on. In case it wasn't obvious, that was not up to the usual standards of our recordings because that was a walking interview. I was carrying these portable mics all over NeurIPS. And really the only way to schedule podcast interviews with people, especially busy people like John, at NeurIPS is to show up with a portable mic, shove it in their face, and talk to them.

19:58And that's what the majority of the podcast conversations are for this episode because that's the only way I can see someone, grab someone, do you have 15 minutes, and talk through something. that's what happens and that's how we schedule interviews with a whole bunch of people that we would not get otherwise new reps is too chaotic to schedule anything else otherwise one takeaway from john's interview which i want to highlight apart from the whole it's the new normal conversation it's the focus on synthetic data generation this is a recurring theme that is continually coming up from my conversations with literally everybody in the space and how do you do it right how do you do it with the blessing of open ai by dance was recently banned from open ai because they were considered to be distilling from gpt4 which is not allowed under the terms of service i've heard that they're not the only company that is accused of or being thought of or rumored to be doing that probably the right approach is something that looks like deep minds approach which on monday of neurfs published a paper called beyond human data scaling self-training for problem solving language models.

20:59And the concept is honestly not that complicated. For the domains of math and for coding, they were able to computer generate data for training on. And they found that when training Palm 2 on that synthetically generated data, improved their results and performance on the benchmarks for those relevant domains. It makes sense that, you know, we can scale beyond human data on those dimensions. That's the trivially easy stuff. And the question is, how do you scale beyond the verifiably correct. If you listen to part one of our NeurIPS coverage, we talked about DPO, which is more efficient usage of existing information.

21:35So not exactly using synthetic information, but just as a sneak peek of 2024, we've actually already recorded an episode with Nathan Lambert, now of the Allen Institute on RLHF and RLAIF. And I think those approaches might scale beyond just the narrow domains of math and code. So next up is someone who's new to the pod, but not new to me. I've talked with Lin from Fireworks a bunch over the past few months, and they've definitely blown up in the inference space. So in some sense, you can think of Fireworks as a competitor to Together AI or Replicate or any other sort of inference-serving platform that you might think about, but they have a really good team, and they've been doing some very good work with Mistral.

22:14Lin and her team have an amazing track record, which you hear about in the interview, and their customer list is pretty stellar too. So it's worth checking out and checking in on the inference business with Vintel from Fireworks AI. Rewind, we can do all that because this will be edited. Okay, so who are you and what is Fireworks? Hey, Sean. We started Fireworks last year and me and a few founding engineers, we have been working at Mata on building an air platform and specific PyTorch for five years. When we started PyTorch, it was a framework for researchers. and we took the mission to build one framework for both production and research and streamline research production transition, operating PyTorch at a huge scale for MATA and for the industry.

22:59So by the time we left last year, it is running more than 5 trillion inference per day across 50 data centers for MATA. And we feel like this is a great impact we have landed. But when we look at the industry, it's really, really behind. And we founded Fireworks to really bring this expertise to help industry adopt AI in the fastest way, adopt the state of our best research into production in a very streamlined way. And why Fireworks is a name? Because PyTorch holds fire, and we want this fire to be everywhere. That's why we come up with our name, Fireworks. Nice, nice. Well, there's also Lightning and Lightning Labs is a kind of spinoff of that effort.

23:49Right, right. And basically, I think there are multiple teams working on better inference for PyTorch. Could you elaborate on how do you see the landscape of inference-as-a-service companies? I don't know if you consider yourself that, or infrastructure companies in general, I guess. Right. So I think when we think about inference optimization, there are different angles, right? I still think PyTorch team, when I was there and now, the PyTorch team, they are still doing a great job pushing for PyTorch performance optimization across training and inference through the PyTorch compile project. The goal here is to keep the simple PyTorch programming API, which is really good for researchers, and then take the heavy lifting of doing optimization in an automatic way.

Read the full transcript

24:45But then because PyTorch needs to support and sustain a broad community, so the workload is much more diversified when you think about optimization. And here at Fireworks, we take the same philosophy. We want to keep the simple API of PyTorch programming language and take the heavy lifting of the optimization, but more specific target at industry verticals, right? For example, when we started a company, we started from ranking recommendation. And we have a product around that. And then later on, our customer we engage with, they're asking us, hey, can we help on Gen AI? Because all the Gen AI models are PyTorch models.

25:28It's bigger, it's more complex. It's even harder to operate and optimize. So then we started vertical on Genii across large language model and image generation, other modality as well. But because we focus on vertical, so we can't afford to take a much more specialized optimization approach. And that is complementary to PyTorch Compile with PyTorch Driving for a broader audience. So that's where we are. And I will say because of our PyTorch expertise, we are the best. when it comes to performance optimization across the following areas. The performance for JNI models are pretty complicated because there's no one bottleneck on system resource consumption point of view.

26:15The bottleneck can scatter across CPU to GPU communication, the compute itself, memory bandwidth, and many other things. So we develop a very special scaling algorithm that allow us to tackle those bottlenecks independently instead of blending them together. So that's a very unique thing we are doing. The second is we build custom kernels across attentions, especially multi-query attention, mat-mool, or reduce. And those custom kernels outperform anything in the industry. Yeah, we also do many adaptive technology that just when we run the inference, performance will get better. The more you run the workload, same workload, it will start to adapt to the workload and become better and better.

27:12So across all this, that enables us to be in the leading position of JNI inference provider. Just to give people a mental image, obviously they can go to the website. you have a self-serve option that people can try out. You mostly have a library of existing popular open source models. You just started creating your own models, which we can talk about. I didn't know that. That's super exciting. You actually recently enabled Mixed Trial in one day after their release by reverse engineering the code? That's right. What's the high level of that? Yeah, so I think we did that twice, right? The first time when Mixed Trial 7B got released the same day, they released in the morning.

27:53And then in the afternoon, we launched Mission 7B. We were the first to get work. And this is basically, like, they release weights but no code. And then you have to implement code by guessing the... Right. For Mixtro, that happened last week. They only released the weights. And there's no code. And I think it's really fun for us, right? So because thanks to the technology we develop over time, we actually build a slew of componentized libraries that enabling new models is not every time built from scratch. So because all these models share similar kind of model architecture underneath with different components, and that's why we have the velocity of the speed.

28:47But it was actually fun to hack it. Dima, he goes by Dimitri Zgokov. Your CTO. Yeah, our CTO. He basically took the Lama model and tried to retrofit to the mystery of weights and it worked. It worked. We were like thrilled. Oh, it's actually working pretty well. But on top of that, it was just a base model. It's not an instruct-tuned model. It's not really usable for chat. And then overnight, we tune a chat model and deploy to pull bots and used by many other users already at high scale. And the feedback is really, really good. Of course, now we switched to Mr. Instra as the official version. But we still keep getting users feedback on the overnight tune chat model, sometimes even perform better.

29:38So yeah, so that's what we do when When it comes to the velocity of quality and velocity to high speed, we are the best company in the industry. Yeah. Mentioning speed, I should also mention that a lot of AI engineers listening on the podcast would be familiar with the Vercel AI Playgrounds, which you are the primary provider for, right? I mean, that's the one that's most visible because they name you. But I don't know if there's any other that you serve that you can name as you're the sort of inference provider. Here's just kind of a very highly selective list of the customers. Yeah, of course.

30:10It's not exhaustive. Yeah, we get the marketing rights. So we already serve Tome. They're doing really good PowerPoint generation. If you haven't used that, please try it out. It's really cool. Yeah, I used it for my keynote for my conference. Oh, that's fantastic. Yeah. I use like a magic trackpad to serve the Tome. And then obviously, whenever I need to generate images, I actually generate it from inside of Tome. So I was using Fireworks without knowing it. That's fantastic. We also serve the co-pilot kind of application. For example, Sourcegraph released Kodi. By the time this releases, we'll release our episode with Sourcegraph and Stevie Agui.

30:49Oh, that's great. We recorded one. We're good friends. We also are the inference backend provider for Poe. That is a very popular chatbot. And Poe is building... Wait, doesn't Poe just Anthropic or GPT? At the beginning. Okay, now they have their own models. They are going big on open source models to provide a variety of different experiences and bring different experiences and much better performance. And of course, from their point of view, cost efficient. There are many other big enterprises, for example, with DoorDash. They're using us. Did they say for what? Yeah, so we actually released ranking recommendation stack with them.

31:35to power their main business. Because when you go to that website, there are a lot of ranking recommendation stuff happening, including ads and kind of restaurant search recommendations and so on. One thing I wonder about is for something like a DoorDash, and I'm a bit newer to Rexxus in general, shouldn't those be pre-computed? Why does it have to be fast or live? It doesn't have to be live, right? Actually, there are a lot of dynamism, right? Because your personal preference may change. It's also quickly learning. And their distribution channel, their participating restaurant may change. Their menu may change.

32:16There's a lot of dynamism in the matching criteria here. And as I worked at MADL for a long time, to actually do highly adaptive ranking recommendation, personalized ranking recommendation, yield the best performance. when it comes to the relevance and revenue. Yeah. I'm just asking like offline versus online. I don't know how sensitive this is to latency requirements. Oh, yeah, yeah. No, so a lot of time, most people, like, of course, at big companies, people do online training. But for those enterprises, I haven't seen the need to go online training yet. So usually training is offline. but it's periodic.

33:04You have to refresh with new information and then you launch and deploy periodically. Okay. So I teased this earlier. I didn't know that you had your own models that you're also training. So you just released Clean Lava. Yeah. What's the story behind that? Right. So I think everyone knows GBTV and the space of multimodality, right? I think as I talked about in one of the interviews when I was at Meta for PyTorch at the end the moderator asked me hey what I think of the future my answer is multimodality because we live in the whole world that has so many different modalities across image, audio, text, video and so many other things and that is the mix of our world and real world experience So, yeah, we really think multi-modality will be a very important aspect.

34:00So, we take the very popular Lava model from Microsoft, but it has the kind of GT4 training data, so we replace that with our own training data and make sure it's commercially usable. Yeah, we're super excited about this. Yeah, I mean, it sounds like you'll be exploring more models as well and just putting all your platform and you're the fastest way to access them. We're here in New York. You're talking to a lot of industry folks. Any other top of mind conversations that you're just hearing a lot that may be surprising to people? So I mostly talk with many startups that is emerging. So number one, it's really refreshing to me, but not surprising, that there's so much product innovation that's happening.

34:47across the board. So much energy there, but a lot of those are built on top of JNI. Of course it's not surprising, but it's kind of validating fundamentally innovative technology can can reboot a huge part of the industry. So that's really really refreshing. The second is I think there are a lot more, hey how we think about working together, right? How we build a bigger, more interesting product for a broader audience together. I think those conversations are very, very interesting to me. Yeah, yeah. Okay, very cool. And you're also here to hire or recruit. Oh, yeah, absolutely. Maybe put out a call.

35:35Who are you looking for? What's the profile? Yeah, we are definitely growing very fast as a company. We are looking for system engineers. We already have a rock solid inference serving, but we are scaling it quickly and aggressively. Anyone with cloud infrastructure experience move really fast, join us. We are also looking for researchers who have a lot of experience and understanding data a lot, understanding quality a lot, can get to kind of quickly help our customer get to high quality. And whether through training our own models or fine-tuning the models and the building task-specific fine-tuning services, those are the areas we are pushing really aggressive on.

36:28And of course, we are hiring across the board of go-to-market people, all the way from marketing, solution architects, sales rep, and so on. Yeah, yeah. It seems like you're scaling very quickly. Thanks for coming on. Oh, thank you for having me. Cool. When I first met Fireworks, I was very impressed by their team. But since then, I've been more impressed by their execution. And my guess is that this will not be the only time that you'll hear about them on the Lanespace pod. So far in organizing and editing this podcast, I've been trying to bias towards reintroducing previous guests of the pod as a form of end-of-year check-in episode with friends.

37:03But so many of them actually mentioned Fireworks. You'll see later with Cursor and Perplexity that I had to put Fireworks first just because that many people have interacted with them, used them, and loved them or compete with them. I think it's a really interesting open question as to how much moat any one commodity infrastructure provider can have. The people who are not in the business say there's no moat, and the people who are in the business, like Lynn, see tons of moat in the software that they write, which obviously is proprietary to them. It's also interesting to see them start training and releasing their own models.

37:33And Fireworks released a lava variant, which we previously covered in our previous NeurIPS episode as one of the best papers of 2023. So I highly encourage you to check out that conversation with Haltian if you're interested. So I say all that to preface the conversation that we're going to have with the next two guests. The first is a return guest, which is Aman Sanger from Cursor.so. We had them on in August to talk about their amazing rise to power as the AI first code editor. They've definitely exploded all over my timeline. And at the time of the interview, I myself was a VS Code, CodeE, Codeium, Copilot, Codeium fan.

38:06And since then, I've actually switched my own workflow over to Cursor because of the better workflow that they provide. But still, there's a lot of open questions around their business. Just like Mosaic, during our podcast interview, they were actually sitting on a fundraise. And they recently announced their fundraise with OpenAI. So let's check in on Cursor. Okay, cool. So I'm back with Aman. Hey. Hey, how's it going? Hard to catch you. You're a difficult man to find. I guess so. You've been exploring Europe's and you also announced your fundraise since our last episode. Yeah, so we raised$8 million from OpenAI.

38:37They've been a fantastic partner, and I think it was a great decision. Open AI users are themselves. Yes, we have a lot of OpenAI users, and we're growing pretty fast inside the org. The thing that we like to say is, like, Cursor is the means by which research happens faster, right? Like, as we make programming happen faster and faster, as we make programmers much more efficient, we're making researchers more efficient. And the bottleneck for research is really just implementation. If you can come up with an idea and then actually have the code, have the experiment all written for you immediately.

39:09Research would just happen much faster. And so that's the goal that we're working towards. And I think we're tied a bit of the way there with a lot of OpenAI users. Yeah. What's the funniest or most interesting feedback you get from OpenAI people versus regular coders? Do they prompt differently because they work at OpenAI? So they actually probably have less feedback than some of our other users who are less familiar with language models because they know what the deficiencies are. They kind of know what's going on underneath the hood. You can probably give them interesting input on what people are trying and failing with.

39:39Yeah, that's true. We do give them a lot of feedback on a lot of their early alphas and whatnot. And so you've been tearing up the Twitters recently, putting in some effort. What are your sort of top messages that have been really resonating with people? I was a big fan of the KV caching tweet. It's surprising that not too many people, it seemed like not too many people knew about this before. yeah um so that this is when people learn about transformers uh it's actually not in the documented literature and the academic side of things that kv caching is a common industry practice yeah you only find out when you talk to industry people that yeah kv cache so like when you say kv cache it's really confusing because the kv cache like the kv cache can be cached right it's like almost like a double caching but the key idea here is well let's let's look at all the big closed model providers, right?

40:27They all have like these chat models. And with chats and with conversations, like the first end conversation messages are always fixed. And that means like the first, let's say, like end tokens are going to be fixed. And that means when I put the next token in, why do I need to redo all the work of recomputing the keys and values for those first end tokens? And a standard inference trick for this is you take those keys and values and you move them from GPU RAM to CPU RAM. You store them there for some period of time before they're evicted. And then if another request comes in with a matching prefix, the matching original conversation history, you just load those back into GPU RAM.

41:08And you save a ton of time on compute. Your time to first token goes... And then because you're saving on compute, you can increase your throughput. And this is a trick that you don't really see in any of the open source inference engines. So you don't see that, but people implement it on top of it. Yes. Well, I'm understanding, like, I think together, for example, I think it's implementing this. And I just talked with Lin Tao from Fireworks as well. Yeah. So one of the interesting, I always assume that it's because of personalization. Like, hey, in my system prompts, I have today's date. I'm going to have to update that once a day.

41:40Fine. Like, no big deal. Yeah. But maybe if people have more customized prompts. But, you know, you said there's some kind of cash eviction policy where if there's like a 95 % match, you use the cash. Yeah, I don't know what the exact eviction policy would be. You could probably use, like, assume you have, I don't know, like 100 gigabytes of space per device, probably a lot more, actually. You probably have up to a terabyte of CPU RAM per device or maybe per machine. You could just do something like least recently used. And then if you start to use up more space than exists on device, you just evict the least recently used request.

42:18You are a consumer mostly of the GPT-4 API. Yes. they don't really expose this in the API. How does this affect you? I think it's actually pretty important to understand what's going on underneath the hood to take advantage of these things. So we use dedicated instances. So they expose their capability to you? Somewhat, but the key thing is they expose very little, actually. Isn't that weird? Yeah, but the only way that you can really take advantage of this, and I kind of had another tweet about this, is you need to really understand what's going on underneath the hood. So you can then plan for when memory utilization is spiking based on how many tokens you're currently using or how much memory the instance you can speculate is more.

43:04When are you getting a lot of cache hits? So you don't expect to be using as much compute, which means you can then increase your throughput without worrying about things going, latency spiking or things going down. And I don't know if you've, I've taken this thought to quite an extreme level. Like you can use this to cache RAG stuff, like RAG results. Yeah. Just general prompts, right? You can, you can. So I did have another tweet about this where there's, no one's done this to the best of my knowledge. And I think this would be very, very hard to do. But you could technically cache the entirety of some corpus in something like S3.

43:41if you have a model which has smaller size keys and values. So this would be, instead of full multi-head attention, it could be something like grouped query attention, which is, I think, usually around 8x smaller, or even multi-query, which can be 64 to 256x smaller. And so then what that means is you can actually read the weights from BlobSorge if you have everything really optimized. You can read it into RAM a decent bit faster than it would actually take to recompute the key to KVCache. I think that will be very tricky to implement. And I think there are actually not too many use cases where it would be useful.

44:20I think the code bases, there's actually one where it could be. Yeah. My final observation on this is, OpenAI had the opportunity to offer caching to people with the assistance API. And again, they're charging you for the whole thing every single time you send a message to the assistance API. and like is it like i find it like is there is there some explanation is there is it just like a we can do it so we're gonna do it i mean it's tricky when you're not like using i don't know what they're doing underneath the hood but if you assume they're doing something like uh caching at a machine level like so i assume they're not they're serverless right so you have to load unload and that costs causes a cold start and that's a problem for for them yeah so it's like really trivial when you have server endpoints, server-based endpoints, or dedicated instances.

45:09It's probably quite tricky to get right. I mean, I'm not really confident as to what their decision-making was there, but I'd imagine it's much more difficult to get right. Got it. What was your second tweet that we prepped? One of them that I thought was interesting was generating a retrieval data set. Yes, synthetic data. Using synthetic data. I mean, the key thing here is there's a lot of using synthetic data to, like the outputs of models to actually train weaker models. And so a lot of people have done this with GPT-4 outputs. This is actually, I think, that requires, I guess, the claim that you can train on GPT-4 in outputs, and you'll still get pretty good models out of that.

45:47Yeah, it's well-established. Yeah, which seems reasonable, but we're actually relying on a weaker claim because all we're doing is, I mean, people can check out the tweets and see it in more detail, but GPT-4 is quite good at this task of ordering four candidate documents given a query as to the relevance of the query. There have been papers that show this list-wise re-ranking, and it works really well. So if you do that for enough documents, and you do it in an efficient way, which we kind of use a variant of ELO called ProofSkill to do, you can then get a really high-quality re-ranking data set, a really high quality ordering over, let's say, 100 candidate documents given some query.

46:31So we use GPT-4 in the loop for doing a bunch of different synthetic data stuff. This is one of them. And I feel like more people should be doing it for this kind of stuff. Yeah. Yeah, I think people are exploring synthetic data a lot at the back half of this year for choosing models as judges, models as synthetic data generators. Yeah, I think models as judges is like almost certainly going to work. If you like use chain of thought, it's a very easy task. I think this is a very easy task. This is how we do RLEIF? Yeah, yeah. It's interesting. RLEIF, I was looking at that paper again, and it seemed to really be good for, if you look at it compared to RLHF, it helped with harmlessness.

47:14I don't believe it actually helped in helpfulness. But it helped to achieve the Pareto optimal tradeoff, which is no decline in the other two. I think if you compare it to RLHF, it was pretty neck and neck. I don't think there's a statistically significant difference with helpfulness at least. But it is interesting. Like RLA-IF is just effectively getting better at censoring the model rather than improving its like almost like capabilities, right? Its helpfulness. While RLHF, like it'll do it as well as RLHF, but it doesn't offer anything additional there, which kind of makes sense to me. First impressions on your apps?

47:49I mean, very interesting. Lots of very smart people. I've had lots of very interesting conversations. I'll probably be back next year. I was kind of lukewarm on it coming in because everyone goes like, oh, it's a big conference. It's hard to navigate and all that. But then you run into a few papers, people, authors that are interesting. And then you're here. A bunch of other people I want to meet are all here. It's a nice way to get everyone in one place and just catch up on everything. The house parties are fun. Yesterday, it was just a lot of parties. I don't know. To me, it's very overwhelming.

48:20but I think the more exposures or epochs that you have on NeurIPS, the better. And I'm basically trying to do this audio experience to try to bring people in because there's many people who just never come. But they should get a sense of what's going on here. I find there are people here who you've never heard of on Twitter. They're not on Twitter. They just know more because they've just done the work. They read everything. Have you seen the Datacomps paper? I'll walk you over and show you. I was very impressed by their work. Like, these people, they just come out of nowhere, and once a year they do this.

48:52Yeah. And this is the place to find them. So that's why I'm here. Yeah, I mean, I completely agree. Yeah. Right? There's really such a good congregation of, like, very good researchers. Right? Yeah. Are you trying to hire them? Let's make a hiring call. Yeah. I mean, look, I think right now we're a very small, very strong team. We were five last time. Yeah. So we are seven now. Only six engineers, though. Yeah. So very small team. You're more millions than people. Yeah. Look, we're a very small training team, and we're looking to grow the team, but we're looking to grow it very carefully and slowly, because I think a lot of companies fall into the pitfall of hiring too quickly.

49:32So yeah, we're really looking for fantastic people. We're seeing incredible traction, incredible growth. There's a lot more really interesting problems to tackle, and people should check out our blog post on that, because I think it's very exciting. The fundraising post? Yeah, there's a fundraising post, and then we kind of link there. There's a problems post. If you go to any sphere.co slash problems 2023, there's lots of interesting work to do. Yeah. And I think we have a really good chance of being the team that can crack coach in. So it's a really exciting space. I think you'd be joining a very small, strong team.

50:07And so, yeah, if you're interested in working with us at Cursor, we'd love to talk. You can just reach out to aman at cursor.sh. Nice. SH. Oh, okay. Yeah, well. I thought it was so. We might try to get.ai or.com. We'll see. Cool. Well, thanks for jumping on. Yeah, for sure. Thanks for having me. So there again, you see one of the topics that I highlighted from my conversation with John Frankel, which is why I put it at the start, which is synthetic data generation in all its glory. And for Aman and Cursor, they're particularly interested in LLMs as rankers or LLMs as judges. And that seems to be generally a more blessed way than directly distilling the output of LLMs.

50:45And you can look out for our episode with Nathan in 2024 to go deeper on that. Another founder that recently raised that is the talk of the AI community, particularly with Guillermo Rauch and Toby Lutke recently endorsing the product, is Arvid Srinivas or Perplexity AI, which started off being, maybe we will construct SQL queries for you. And they went to maybe we'll construct SQL queries on our Twitter screen for you. and now they've blown up as a potential Google replacement, which is a huge increase in ambition, but they have the web app and the mobile apps to prove it. So here's Arvind with Perplexity.

51:21And so congrats on all your success of Perplexity. The two most recent accomplishments which I have seen, at least on my feed, is one, you hit a million people on your mobile app. That's huge. On both platforms, Android and iOS, independently. Is that because of your slick video editing skills? Actually, we have a good brand marketing designer. But, I mean, more than everything else, I think the app is really good, fast. We spent a lot of time on it. In fact, our first rollout of the app was not that great. It was slow. It used to crash. Users complained, and we listened to that and, like, recruited a good mobile team, much faster, more reliable.

52:00Any technical decisions that drove that? Like, is it React Native, that's slow, or something else? No, it's all native. we're not on one common React stack. And the reason to do that is that's the only way to make the apps feel fast. And I believe ChatGPT also does this. They don't use React Native. And then the other accomplishment is PPLX Online which you're showing on screen here. What are the headline things that people should know if they haven't heard of PPLX Online? Well, it's like the only LLM API that has no knowledge cut off. So if you're a developer and you just want to prototype products that need information from the web or has no knowledge This is the only way to do that and super fast pretty accurate you have two versions a 7b and a 70b So 7b is super fast 7b is like little slower But also like better quality and we plan to bring it up in the context of the Mixtro MOE as well That's been recently released.

52:51Yeah, I think you've been pretty transparent that there are fine-tuned a Lama 2 That's right. We're not in the business of pre-training Yeah, but like what do you fine-tune for between Lama 2 and what you have? Yeah, we fine-tune for like summarization the ability to take a bunch of sources and accurately give you a nice summary. And you are, I think, the only provider right now with online access or whatever. Yeah, that's right. But also, like, Grok has access to Twitter, which you don't have. And they'll release an API at some point. If they release, like, we'll be happy to use it, right? Our goal is to just give accurate answers on the web.

53:23Yeah. And Twitter is just one part of the web. Their vision is, like, Twitter is the everything app. We believe, like, that's the information out there that exists outside of Twitter that's also super valuable. In fact, you can even make an argument that information outside Twitter may even be a lot more valuable than information within Twitter because most of the links that get shared on Twitter are all from outside anyway. Yes. So it's only like what you miss out on is like a specific person's opinion. And usually journalists pick on that and write web articles, so it's all going to diffuse, right?

53:55Good ideas usually diffuse the rest of the web. So we're not really missing out much. It's a different source of data. Yeah, it's a different source of data. Also, it's all about what do you want? Is your source or citation already highly curated human artifact? Or is it some tweet? These are all questions worth asking. One thing that you do show off. So I was watching you demo just now. You have sentence by sentence citations. That's right. That's a design choice. Yeah. Because realistically, your source articles actually overlap the full paragraph. That's right. So why did you choose to impose sentence by sentence?

54:29That's how we write papers. I'm an academic. Every sentence you write in a paper needs to have a corresponding citation. As a user, it can be confusing. Like, when I click that link, maybe it's, like, the third paragraph. That's right. We can do better in, like, exactly navigating you to the right part of the link. But we're looking into all that. Yeah, of course. I mean, I do see you as, like, a search engine first with a very good language model team. That's right. Yeah. Right? Yeah. Answer engine. I would call it answer engine. Answer engine. Yeah. You are doing a really good job of that. I also noticed in your PPLX blog post that you also talked about the fresh LLM paper.

55:02That's right. Maybe, could you introduce that? And did you talk to the authors? Are they here in Europe? I did not talk to the others. It's not like we took a lot of inspiration from it. Okay. But it made sense to attribute the citation to them. Yeah, to intellectual backgrounds. What do you look for at NeurIPS, a conference like this? We're here for recruiting good, strong researchers to join our team, especially if they're more focused on shipping models to a search product that's used by millions of people. Awesome. We'll talk about your hiring call to action in a bit. I'm also interested in labs, like Perplexity Labs.

55:40Yeah. It seems like a place for you guys to experiment with serving models. That's right. Yeah. Everybody thinks you start as a rapper, and then one magic day, you just switch over from 3.5 to your own model. That's not how it works. In practice, your GPU is crashed or your nodes are not working or Kubernetes doesn't work as expected and requests are not having the throughput required. You optimize for latency, but then you are worse on throughput. So you're not able to handle spike requests. So all these things can happen, right? So you only know about these if you start small. and serve a playground where people come and test your own infrastructure and see how it holds up and then take the lessons from there and use it to serve it on production.

56:28So labs are sort of our playground for testing open source models and our in-house models that have been fine-tuned for factual accuracy and helpfulness. And it's a nice way for people to test open source models if they're curious about it, especially if they think about it as alternatives to ChatGPT. and then it's also a nice way for us to battle test our infrastructure. Same thing goes to the API. It's not like I believe these APIs are going to take over GPT 3.5 APIs or something, but it's a nice way for developers who want an alternative to explore, especially those who want to use faster, smaller models, like the 7B models, and it's also a good way for us to know how we can handle search requests and things like that.

57:09I want to push back on this. You said your playground is a way to battle test. But I think you probably get orders of magnitude more traffic on your main app than your side app. Look, we can't just directly ship to the main app, right? And you can never simulate real use. It's like a staging environment. Yeah, it's like a staging environment. But not just that, it's not just meant to be a staging. I don't want to downplay the importance of labs. Labs is sort of one of the fewest places on the internet today for you to go and explore and compare different open source models. And it also tells the user how fast our inference is.

57:42We give you all the metrics like tokens per second, the time to first token. It's also a very transparent way to communicate the speed of our infrastructure, which helps us also recruit good talent for infrastructure. But I think you're pretty opinionated that you are an app company first. Yeah, we are an app company. You're not an infra company. That's right. You just happen to have... We're not competing with Together AI or... Fireworks. Fireworks or OctoML. Yeah. There are too many of them, actually, honestly. What do you think they need to do to win as an objective third party? I think they need to raise an insane amount of capital and subsidize the cost so much and capture the market, or it's basically going to be impossible because you're all offering the same thing more or less.

58:20And NVIDIA is basically commoditizing it, right? Like with TRTLM and like Megatron and things like that. So most people's stacks are going to get standardized. So then why am I paying you? I'm paying you for the GPUs then. Right. That's a game you can only pay it, like it's an economy of scale thing. Which you're also buying your own GPUs and running your own stack. That's right. But we care about buying GPUs to serve our own product more than helping other people serve their products. Yeah. What have you learned being like a, I don't know, I feel like you're both an infra CEO and an application sort of product CEO.

58:52How do you balance that? Yeah, it's difficult, but you know, like you, one thing exists in service of the other, right? Yeah. Infrastructure exists in service of the product. You always have to remember that. For some people, product exists in service of the infrastructure.

59:32That's not how we are. and marketing power. But you are more AI native than they are? That's right. In a sense? That's right. You are a different search index. Like you have your own crawlers. We have our own crawlers and indexes, yeah. So like if I don't want Bing, then I use your stuff and maybe you turn out for your stuff. Yeah, that's right. That's cool. Okay, hiring. What are you looking to hire? What should people demonstrate when joining you? I think you have a very strong perspective on the kind of culture that you're building. Yeah, I mean we work pretty hard and like we want to get stuff done fast.

59:59So if you enjoy like fast shipping cycles and... Can you give illustrations? Like, what do you mean by that? You know, every two weeks, like, we have some announcements we make. So we work on very clear, precise projects that have, like, clear deliverables. And we kind of constantly want to keep improving the product. So as a machine learning research engineer, if you're excited about, like, training models and shipping them to production for such a useful use case like consumer search, and want to do it at the same velocity as us, like a startup rather than a big company that has to wait for several months to get something into production, that's a unique spot to be in, right?

1:00:38And you also want to be part of a growing exponential rather than something that's trying to defend its territory, right? Defend its territory. Yeah, like Google. Google's defending. So, attack. Yeah. So, you want to be in an attacking team. Have you heard, like, what does Google say about you? Like, are they interested in buying you? I think they're being pretty appreciative and respectful of the product, right? But, like, SGE is not great for some reason. Yeah. By the way, I don't think Google people are not talented. Like, they're probably more talented than we are. I think it's just that their incentives are not clear, and they might have to cannibalize their own business model.

1:01:18This is the classic innovator's dilemma, right? Exactly. They have a cash cow and they're trying to present that. You don't have ads, but you're serving subscriptions. And that's the main business model for now. As of today, yeah. That's it. Well, thank you very much. All the people that we talked to so far and some of the best founders I know, whether or not they're an AI, are fierce nerds. And Aravind definitely reminds me of the fierce nerds concept. But I don't think I'm the best person to tell that story. Maybe I'll tag in Sean Puri. Have you ever read that Paul Graham blog post called Fierce Nerds?

1:01:52No. What is it? It's an amazing post. I'm going to read you a couple pieces of it, but it's one of those like Paul Graham, I think is somebody said this earlier. They go, what's that guy? Andrew Tate. They saw some tweet that was really funny. It was Paul Graham was my Andrew Tate growing up, which is like so funny. So it's such a funny, it's such a deep cut joke, but if you get it, you're like, it just hits the spot. All right. So he wrote this post and he goes, most people think of nerds as quiet, you know, sort of like diffident people, right? Just sort of like, you know, passive. And in most social situations they are, they're, they're quiet and you know, they're not the star quarterback in the middle of the gym, right?

1:02:24They're kind of a fish out of water and a bunch of different things. He goes, but this is an illusion because when that only happens when non-nerds observe them because they're observing them in non-nerdy situations. So you see a nerd at prom, you just see them as a quiet sort of passive nerd. There's no alpha in them. But in fact, some nerds are quite fierce. Fierce nerds are a small but interesting group. They are extremely competitive, more competitive, I would say than competitive non-nerds because the competition is more personal to them, partly because they're not emotionally mature and they distance themselves from it, but also because there's less randomness in the types of competition that they engage in.

1:03:02Therefore, they're justified in making it more personal. I'll cut it off there. That's a clip from the My First Million podcast. And that's a story about how Dharmesh Shah, the HubSpot CTO, is a fierce nerd. And I really like that concept because one, it helps to validate that nerds can also win and why nerds can sometimes win more than regular people and obviously for more you can read that paul graham essay but i think arvin is a fierce nerd and i think perplexity is a fierce nerd company they do have competition though it's not like perplexity is the only company going after google not the only company going after search one of my favorite parts in in compiling these ensemble episodes is juxtaposing two competitors next to each other or people who disagree or have different worldviews.

1:03:46Like you just heard Perplexity. We just heard Arvind dunk on all the infrastructure companies, including Fireworks, which we just had on. Now, I'm not the right person to tell you who's right and who's wrong, but I know for a fact that they cannot all be right. And that's what's fascinating. That's what makes the market. So next in full disclosure is a personal friend of mine. It's Will Brick from Metaphor Systems. Metaphor launched end of 2022 with an AI search engine narrative as well. but their approach is more of a pre-trained LLM research engine as opposed to Aravind's answer engine. They're all very minor differences in the end.

1:04:18At the end of the day, people want to punch in a query and get results. And Metaphor's approach is different. They are going after the infrastructure play rather than the application plus infrastructure play and it's just nice to contrast them together and I'll leave the conclusions to you. What is Metaphor? Metaphor is a search engine over the internet but it's better than Google at handling complex queries. Okay, why is that? Why is that? Because we train a search algorithm from scratch to handle complex queries, basically. It's a totally different algorithm, yeah. Why are you at NeurIPS? I'm at NeurIPS because we want to learn about all the cool things people are working on and also because we want to hire some crazy good researchers to help build a future of search.

1:04:57Metaphor has a search engine. That's what you launched last year. And then you also released an API. And I've actually been using the API. It's actually really good for augmenting LLMs of search. I don't know how much to which you want to lean being an app versus an infrastructure company. Yeah, so we're leaning towards search infrastructure. So we really see ourselves as like we want people to build applications on top of us. We see the future as like everyone will use LLMs as an interface to everything. And we want to be powering the search that underlies that. I think we want people to build really cool UIs on top of our search.

1:05:27But the hard part and the thing that we're focusing on is really good search results. Yeah. Can you give examples? Do you have some really cool examples like tweets and books and PDFs and stuff? People really get excited about researchers working on something similar to them in the Bay Area or something like that. People have actually met. Oh, yeah, yeah. Competitive intel research as well. People have met people in real life based on searches because the results are so high quality. And they're not SEO spammed in any way. It's just exactly what you're asking for. It's cool to see that digital information to real-world interaction thing happen.

1:06:00I actually also interviewed Arvin from Perplexity, who I feel like is also in that sort of search domain, but he's less focused on search infrastructure. He's more focused on just being a search engine. I don't know if you compared yourself to Perplexity in that way. Yeah, I know. We get asked this a lot. I mean, Perplexity is doing a great job at combining LLMs with search results, and that does make for a better search engine. That is the future of the user interaction. but we're just like more focused on the search results themselves and really trying to handle the queries that you know google bing are not good at yeah so i mean we we want people to build llm style interactions on top of our thing as well wait so you say google bing are not good at do you think that people will use you in complement to google and bing or do you just completely replace that at least in the beginning like we're going to be used in places where google and bing don't work well so i mean if your application wants to know the weather or wants to know like that Taylor Swift song.

1:06:56Basically, if your application knows the right keywords to search with, then sure Google and Bing are going to be fine for you. But if you want to make these complex almost metaphorical queries with natural language, which are really the most powerful ones, then you should be using metaphor. Yeah, yeah. I was actually walking from your we were walking from your sort of sushi party that you just had at a recruiting event. I hope the food was good. It was pretty good. I love me a little bit of sushi. And I was actually talking to people about your sort of auto-prompting feature. Because a lot of people There was someone from MidJourney there, and they were saying how Dolly3 also does sort of auto-prompting or rewriting of the prompts.

1:07:31Yeah. Is there art to auto-prompting? How do you feel that? How do you feel about your auto-prompting feature, basically? Yeah, auto-prompt is like we convert. We use ChatWT, basically, to convert the queries that come into the search engine into queries that are formatted for Metaphor's models. Because Metaphor is trained to predict links given text. So the model really, like the best way to prompt metaphor is to search in a way where a link naturally follows, which can be confusing. So we have this auto prompt that converts into the right format. You can kind of think of metaphor as in the same state as like what GBD3 was in.

1:08:02I don't know if you guys remember, but, or if you remember. It's not instruction tunes. Yeah. Yeah. It's like, you know, two years ago, GBD3 was autocomplete. So you had to like prompt it in order to get the best output from it. It had a lot of power, but it just had this weird user interface metaphors in a similar situation. the problem is when you rlhf you like and we've tried this like it does reduce the power of the model and like it's just okay to to keep because like often we're using this auto prompt like it's okay to keep this model the way it is requiring this autocomplete type search and yeah would you call yourself a search llm like very very long ago the original pitch for metaphor that i heard from you guys was you're a llm that predicts links instead of instead of tokens oh well an llm is Yeah, I mean, LLM is like, it's modeling, like, yeah, usually like language.

1:08:46And we're not really, we're not exactly generating the links. We're, we're, we search over an index. Yeah, yeah, they're not hallucinated at all, right? They're actually from an index. Yeah, I wouldn't call it a search LLM. Okay. It's more like a, really a search engine. You might even think of it as a research engine. Yeah. There are a lot of different ways we're trying to explain. I mean, I think we're, like, using terms that were developed in an old era for a new type of thing. So we might have to invent new words or wait until they are created. Yeah, yeah. What else should people know about metaphor in general?

1:09:12What other interesting work are you guys doing? I think just the vision is super exciting, and I think people don't realize how exciting the vision is. Basically, the vision is to solve search. What does that mean? It means no matter how complex the query, Metaphor should be able to handle it. So we're talking AI researchers similar to you who are in the Bay Area, who've worked on Rust before, who went to so-and-so college, who would be a great candidate for this startup. Whatever it is, we should be able to handle it. And language models are powerful enough to understand language at the level of a human.

1:09:42So you should theoretically be able to make a system like this. It's just a matter of how fast can it be. And we want to make these things, like, do all those complex queries really fast. And imagine if you could do this. Imagine if this was, like, possible. And then you combine that with, like, you know, GPT-4, GPT-5. And that's how we want our customers to combine us with, you know, combine us with GPT-4, GPT-5. Suddenly now you have the ability to literally answer any information query, no matter how complex. Like, the entire world's knowledge is at your fingertips. That's, like, insane. We basically become all-knowing.

1:10:13Omnipotent. Omniscient. Omniscient and then omnipotent. Omniscient and knowledge is power. I skipped a step. I can do that sort of QED proof of why omniscient equals omnipotent. I am very excited about you guys. I've seen you grow literally from your living room. And it's definitely not over. What's it like having a meme-y celebrity CTO who keeps tweeting viral shit? I mean, I love it. Like, Jeff literally just goes, Jeff has figured out Twitter. He just knows how to go viral because he has really good takes. And we often throw up a party in response to his viral tweets. So you want to talk about the Andrew Huberman party?

1:10:50Yeah, okay. So he had a tweet that was like, Andrew Huberman has single-handedly destroyed the SF social scene because, like, everyone, whatever, is, like, sober at parties and goes with them early. And so, of course, we had an anti-Huberman party where, you know, everyone stayed late and we had, like, you know, a bunch of beer and, like, everyone. Well, my favorite was all over the apartment that we had the party in. You plastered quotes from Andrew Huberman about how alcohol is that for you. Right. Alcohol will destroy your brain and all these things. Look, everything in balance, right? We should have fun in life, but also be safe and everything.

1:11:19And he had another tweet about how he was going to go on a date, but the girl ghosted him, and that allowed him to focus on coding that night. So, of course, we had to have a ghosted NSF party where everyone came to code together because you're already going to be ghosted on Friday night. You might as well code together while you're at it. Yeah, I love that part of the social scene. And I think Metaphor is also really driving that somehow. So congrats for all you do. And it's just nice to check in with you. I've personally been enjoying the Metaphor approach to LLM search APIs. I've often said this in context of the capabilities of GPTs.

1:11:53So if you think about it, what are the capabilities of ChatGPT as it is today, as well as GPTs as announced on Dev Day, right? There's the LLM base layer, but then you tack on three core capabilities on top of it, right? One is retrieval meant to generation where you upload files and you do RAG on it. And second is a code interpreter where you do generate code in a sandbox and then you run code and you correct code. And finally, you execute it. And third is you have a search feature. And so we have a bunch of companies competing for the RAG functionality. You can check our episodes mutually with Harrison of Langchain and Jerry of Llama Index this year.

1:12:30There's a bunch of companies competing for the code interpreter capability. lead is obviously Replit, but then abstractly there's also Deno and Valtown and anyone who runs code is in that game basically. But what is surprisingly uncontested is open web search. And so far, I think it's perplexity and metaphor that are leading the pack in their different approaches. One, the PPLX API is an integrated LLM plus search API. And then two is metaphor, which is search only. And you kind of bring your own LLMs. For our next guest, we're actually going to go over to our last return guest, which is one of our most recent hits, which is Jeremy Howard, previously of Fast.ai, but now of Answer.ai.

1:13:10It seems that all people want is answers, and Jeremy doesn't have them, but he has questions. Outside of the Decibel event, recording. I realized I had to be the interviewer, and I was like, I should probably should buy one. And I had to pick a wine, and Sean told me, pick the most expensive one. Yeah, it's on Decibel anyway. Because Decibel, it's paying for it. So the one I'm having is from a$160 bottle, and it's really good. And I did the same. And I'm not having any wine. You're too young to drink. Yes. Could we go around and identify voices for people listening? Maybe, Tanisha, do you want to start?

1:13:39Sure. My name is Tanisha Abraham. I am the CEO of MedArc, which is a medical AI research organization. I also work as a research director at Stability AI. And I've been collaborating with Jeremy Howard for more than a year, like a few years, a couple of years maybe. And he's also the president of MedArc, and he's been heavily involved in my venture as well. And you have a podcast together, which I really enjoyed. Oh, yeah. Yes. Jeremy had me on his podcast, which was... Your first and only episode? What the hell? Yeah. It turns out that maintaining a podcast is hard. It's easy. Just shove microphones in front of people's faces.

1:14:15So I'm Jeremy Howard. This is my voice. And as of today, I'm Jeremy Howard of Answer.ai, I guess. Yes. And repeat guest on Lanespace. Your last episode did really well in terms of the number of views. Yeah, you guys are good interviewers. Well, also, you drop a lot of spice, which is what we like as podcasters. Oh, yeah. We also have Jess Leao on for the first time. Hey. Yes, hello. I'm Jess Leao, and I'm a partner at Decibel. Excited to be here. Excited to be providing the wine also. Standing in for Alessio. Oh, so good. Alessio dished us tonight, right? So you're the better replacement. Yeah, it's good because in a previous conference, Alessio was wearing my badge and replacing me, so now I can be Alessio for today.

1:14:55Or you just worked really well. A shorter version of Alessio, basically. So today was the Answer.AI announcement. Maybe you want to cover that. What should people know about it? What should people know about it? Oh, I don't know, man. You went from fast AI to the dark side now. No, it's not at all. It is the light side. It is actually. It is the light side. Fast AI, look, I spent the last week in San Francisco, and the amount of love I received for fast AI was overwhelming. I couldn't believe how many people told me it changed their life. which is just amazing but I have to say it's actually time to be rejuvenated the mission is the same bring AI to as many people as possible but now we can't do it on the back of my bank account I've been paying for everything we can't afford it anymore you've had donations and stuff you were very steadfastly against donations no donations, no revenue you of any kind, totally independent.

1:15:59But now, you know, I think we can do a better job by having a bank account with money in it. So thank you, Jess, for sending us money. Jess, what is it like when someone like Jeremy comes and like goes, you know, we need a bank account? You know, there are some people that you go through a pitch and then there's some people that you email and you start prepping the wire. And I would say that Jeremy fell into the ladder. I didn't even ask for this money. I was just going to have a chat with Alessio to get some advice. And then Jess turned up and Jess's other partner, John, turned up and was like, what are you guys doing here?

1:16:37They're like, oh, we'd like to give you money. So I was like, oh, okay. So that was good. They have good taste, right? Yeah. I've talked to you a bit, especially at the modular conference, which I'm wearing the badge of. Nice hoodie, yeah. The hoodie is really nice. So you're interested in fine-tuning. you're interested in fundamental research what like could you list out the main areas of interest maybe i mean basically the interest is in making ai as useful and valuable as possible yeah that's how we make it like as accessible as accessible as possible as widely used as possible help as many possible people as we can with this technology right so how do we do that it needs to be cheaper needs to be faster it needs to be easier to use and it needs to be more integrated into people's day-to-day lives into the stuff that they do this is like hard you know and so in the end I guess I was inspired by Thomas Edison's invention factory in the late 19th century where they had the same situation they were like oh look electricity has been invented okay what do we do with this it's a source of power I don't know and they're like oh let's create the record player and the light bulb and the refrigerator and you know it's like recognizing that now you have electricity you can make all these things that's hard, it requires really smart researchers who deeply understand the underlying technology recognize like oh there are some gaps here but they could be filled if we like use this different kind of filament or whatever and so you actually need like deep technical experts who also have the like curiosity and playfulness and spontaneity to like think like oh what if the world had this new thing in it I wonder if we could put that thing in the world now that we have AI Yeah you were very complimentary of like the open source so we last met at the open source meetup as well we met so many times and you're very complimentary of their approach towards just trying things, like model stacking, for example.

1:18:52Is that the kind of people that you're looking to collaborate with? I think partly. I'm deeply involved in the open source community and I want to continue to do that. All the best models outside of your open AI and stuff are all created by the open source community. at the moment through just trying crazy things. But it'll be a mix. I also want to work really closely with the best academics in the world. And I also want to collaborate with the people in parts of the world we've never even heard of who never get a chance because nobody gave them a chance. And so one of the things we're going to be doing a lot of is recruiting in really weird ways.

1:19:44you know, to find those people who are underappreciated. Would it be like a challenge, like a Kaggle type challenge? Yeah, like Kaggle-y kind of things and, you know, physically find ways, you know, or through like open source bounties and stuff like that. Like basically give people an opportunity to show that they can do amazing shit that nobody else can do. Yeah. It doesn't matter how old they are or where they live or what the color of their skin is or whatever. Yeah, I think what the FASTII community has shown is that a lot of people who don't have a traditional background that are really talented people.

1:20:22And I think, yeah, it's great that that was there for, that the FASTII community was there and that Jeremy continues to highlight those talents as well. Actually, let me give props to Tanishq as an example, right? So Tanishq is the CEO of a research lab of which I'm the president, Merak. And how old are you, Tanishq? I'm 20 years old, which is why I'm not drinking the wine. Who wanted to drink at this wine bar? You know, so like Tanishk's a great example of somebody that most people wouldn't hire as a CEO, but why the hell not? Like he finished high school 10 years ago. He finished high school at 10.

1:20:58You know, he had his first degree at what, 14? Like he's somebody who's... I mean, that's somewhat like he went after the traditional accreditation, or the pieces of paper that you would pursue to show yourself as qualified. So in a way, he's part of that status quo. In a way, but unfortunately people are ageist. Yes, they are. And I also know that I never actually did a computer science degree or anything like this. My start with AI was actually through the Fast AI course. And so it's been a long journey since then. Yeah. What would you ask him about Answer? Because I already know a lot of what's going on at the company.

1:21:43What is he not saying? Is he too humble to say? I think what he's not saying, he already has a great team of researchers. There are already two researchers at Answer.ai that are amazing researchers that I've had the chance to also interact with over the past year or so closely and also just more generally. I'm looking forward to seeing what Answer does. And I'm really excited to continue to collaborate with Jeremy. I think this will be even better for me. Selfishly, I'm very excited because I think it will be better for me to work closely with Jeremy as well. Even though he's in his own research lab, but I think the collaborations that will come out of this will just be amazing.

1:22:24So that's what I'm excited for. And Jeremy, last time you were on the podcast, you said that one of the most consistent pieces of advice that you always give is that people just need to show up, follow through, do the work, that stuff. Obviously, Tanish did that. Yeah, so Tanish is one of those rare people, right? I feel like Tanish is more special than that. What else did he do really well that is vulnerable? So, I mean, God, how old were you when I first came across you? Like 15 or something, maybe? Wait, what? It's been so long? Yeah. Because he only took FastAI a year and a half ago. No, no, no.

1:22:57He was a FastAI student back then. Okay. And, you know, he kind of got on the forum, helped answer questions, you know, asked interesting questions of his own. To stick with that for five years, that's tenacity, you know. And the last course we did was the hardest course we've ever had. It was the diffusion course. It was the first ever stable diffusion course. And none of us knew what the hell was going on. And, you know, he was the one who slogged through the math, figured out what the hell all those Greek letters were saying, and did the first math of stable diffusion video that, as far as I know, that ever existed.

1:23:46He did that with Wasim, right? Along with Wasim. So, you know, he, like, he slogs through difficult shit. And the thing that I notice now is like, you know, Tanishka's kind of famous or was kind of famous as a child prodigy. Yes, you did a TED Talk when you were 14. He was on Child Genius. He did a TED Talk when he was a kid. I was nine when I did it. He was nine. Okay. Okay. So I kind of thought like, oh, things are easy for child prodigies. You know, they're so smart that they just, it's easy. And I'm like, oh, no. Actually, Tanishka's nearly as dumb as me. and so he just works really he just works really hard and I'm like, Tadis, what does this mean?

1:24:27He's like, I don't know like, oh, okay we better figure it out and so that's been interesting to see that like actually child prodigies have to work really, really hard as well that's part of what makes them a child prodigy is that they're tenacious and they don't give up even over five years Does it look that way to you? Is that what you... Yeah, I think so And I think, again, part of it... You agree you're nearly as dumb as me? No. Say it again for the pod. I think Jeremy's trying to trick me here. But I think the Fasting A community has been so friendly that it's been a really pleasant experience to stay with that community.

1:25:08And I think that has also enabled my tenacity because I enjoy being in that community so much. So that's why I've stuck around in that community for so long. So without that, without the community that Jeremy has built, I don't think there's any way I would... It supports you. I had the same with Free Code Camp. So, you know, I think a lot of it has to do with... I'm going to cry. I think a lot of it has to do with building good communities. And Jeremy has done a really good job of doing that. And it's actually a lot of hard work to build a good community and to nurture and grow that community.

1:25:38And, you know, I've been in many communities and I've kind of observed, you know, how different communities in the AI field have grown. And Fast AI is like one of the best communities that I've had a chance to be a part of. So, again, props to Jeremy for doing that as well. I'm so embarrassed right now. I want to give you the perspective. You've been an AI investor for a while. Yeah. And how do you view this community and this moment here? The one thing I will say to the conversation that we're just having that I think is awesome is... We can move here a little bit. Yeah. People keep coming and drinking more wine right next to us.

1:26:11It's a mobile studio. Yeah, we're truly a mobile studio, middle of New Orleans. Let's go. One of my favorite heuristics as an investor is distance traveled rather than just your, rather than just like what do I see today in your resume or whatnot. Because I think if you just go by a certain pedigree or credential or whatnot, you miss a lot of people who have traveled a really big distance, who didn't have advantages to certain opportunities or came from different places or not from the U.S. like you name all the different you know all the different lists and I always try to look for those kinds of people because they're the ones that are always pushing the frontier and like really run through walls and I think this conversation is a good example of that right and I mean no one has a longer distance traveled to Germany I literally and and in the literally from Australia yes yes and and when we were and I think when we were meeting last week you were talking about this a little bit around looking for engineers and people at places that aren't necessarily where everyone else would be looking at, but that has yielded some of the best, deepest relationships you've had, right?

1:27:16Absolutely. I mean, companies turn resources into valuable products and services, right? What are the resources that we suck in? It's people and GPUs, you know? And money. And well, we need the money to get those GPUs and the people, right? Like, the GPUs are, you know, reasonably lucky. You can replace one with another, no worries. So it's actually the competitive advantage, the thing that makes you different, is the people. So this is the most important thing for us to achieve our mission is to build this team, you know, to build this really special team. And, you know, I think the way to do that, and the way I've always built teams is to say is to look at people and say like okay where is this person now and what would it have taken them to get there you know like so if somebody's like you know was kicked out of high school you know because they were dyslexic or because somebody was like grew up in the mountains of Bangladesh and didn't have a pc until they were 16 or you know somebody fought against you know uh you know a woman who grew up in an environment which he had to fight against like institutionalized sexism or whatever it's like these are the people to me i just kind of go like okay this person's gone from like negative 43 up to 99 yes overcome a lot that's a kick-ass amazing person whereas somebody who's gone from like 98 to 99 it's like Okay, it's still cool.

1:28:53But they're probably not the people who are going to change the world. Yeah. And so we want to be a small team where literally every person in it is somebody who can change the world. And the nice thing is when you're in a small team like that, it's just really enjoyable because everybody's just really great to be around, really inspiring. and so yeah that's why we're kind of looking for these extremely special individuals yeah cool so that's a hiring call explicitly you know if anyone's listening who fits that profile and really wants to work with you they should reach out right and now we have a website to send people to so I was going to wrap it up with just overall NeurIPS tips right like what is it like to be at NeurIPS this year if you've been here before?

1:29:50And also, like, what's your best tip for doing Europe's rights? Anyone can take it. I guess I'll start. This is my second Europe, so maybe I don't have a lot of experience with it, but, I mean, I've been enjoying it a lot so far. For me, I think it's about networking with people, and that's the best part of Europe's, because at the end of the day, AI moves so fast that half of these papers are already kind of outdated. Like, you know, we've already seen, like... months ago, right? Yeah, yeah. In order to get here, they have to be reviewed. Exactly. So, you know, we're already seeing the second version or the third version of a lot of these models already, and, you know, so, I mean, it's, for me...

1:30:30So archive is all you need? Archive is all you need, yeah, I guess, yeah. So for me, the value comes out of talking with people and meeting with people and networking, and that's why we're coming to events like these that to network and make these connections, and, you know, I actually meet a lot of collaborators and other researchers at all these conferences. And just to be clear, when you say networking, like it's not like networking in that sense of like getting ahead. It's a kind of a really nerdy kind of networking. So like earlier, Tanishq and I were at another reception where it's like, oh, there's Albert Gu.

1:31:03He's the guy that like two days ago released the Mamba paper. And we go after him and say like, oh, you know, we had a conversation about state-based models and why he's using that and what he thinks the opportunities and limitations are and is there still room for attention? So when we say networking, You know, we mean like geeking out on deep conversations about people's academic areas of interest. Yeah. I always follow up the question of like, okay, like what's your name, where you work? And then what are your interests? And then we tend to go from there. Yeah, just like what paper did you write last or?

1:31:33You know, I will say one thing. So even though the posters, there are a bunch that truly you go by. And even the people presenting are like, yeah, this is kind of out of date. The one hack that's really fun is a lot of those people are also already working on the next thing. and they can give you sort of an early preview of something that actually is not an archive yet. And so that I actually have always, my favorite parts of the conference are actually just walking around the poster session, shaking hands with people who are presenting and learning about what they're most excited about, what they're working on, what are some of the new things.

1:32:02So I find that really fun. And also in my case, since I'm a VC, my best tip is throw an event with a lot of good wine and let the people come. Excellent. Jeremy, you have any tips? I mean, like Tanishka, this is only my second Europe. But I've been to quite a few conferences in general, and my tip, number one tip for all conferences, is don't go to any sessions. Yeah, just stay outside and talk. Whatever they're saying, very, very slowly, and they're probably not an expert at verbal communication either, you can probably get the better version by just reading the damn paper that they're reading out to you.

1:32:38So don't bother with that. So I'm like, yeah, hang outside, you know, in the hallway, look on the app to see who else is around reach out to them and like try and like find a group of six or so interesting people to go and like check out the you know local Louisiana sausage special outlet with whatever reception hopping this is our fourth reception tonight oh my god fourth and best right Jeremy this is why we came to this one last so we can hang out here until the wine's finished so a lot of people hate on the official in europe's conference app hoover uh but i kind of like it because of one thing people can organize their own meetups and list it here it's awesome it's actually really good yeah so i'm brazilian and there's a brazil like little chat and it's so fun everyone's talking in portuguese talking all the time they're sharing all the things that and these are people talking about actually like interesting concepts in portuguese um so it's it's actually really fun i i love the app i didn't even know you're brazilian I am, yes.

1:33:41Liao with a little squirre. Liao, yeah. My accent kind of trips people, and it also trips people when I say something incorrectly and you can't really tell, but I'm really Brazilian. Yeah. Well, we should do a steakhouse next time. Oh, yes, please. Yes, please. That's one of those dinners. Done. Churrasca Rias, right? Yeah, Churrasca Rias. Exactly. My favorite was there was a meetup for people who are interested in sushi. That was the meetup. I love it, yeah. There was nothing machine learning about it. So at ICML, it was really fun. And there was one meetup that I went to that was just like swimming in the morning because it was in Hawaii.

1:34:12It was actually kind of awesome. And then people were like actually discussing like super legit topics in the ocean. I'm actually kind of sad I missed out on ICML. But like it felt indulgent to go to Hawaii for that. Yeah. Okay. Well, I just want to bring it to a close. The last thing I was going to say is, Jeremy, I don't know if you know, I picked your meme as the best meme of November 2023. It was Laundry Buddy. So what's up with Laundry Buddy? Why do you hate it so much? What did it do to you? No. It did nothing to me. For people who are out of the loop, what did you do? I couldn't have walked it back more.

1:34:49Jeremy did walk it back on Twitter. Are you going to make me revisit my shame? I just think it's a fun story. Just for your show, I'm going to revisit my shame. Some people don't know. I made a bold claim that Laundry Buddy was not the peak of open AI's path to societally beneficial artificial general intelligence. I was wrong. It is, in fact, very much on that path. It is well loved. To be able to know that the world's best artificial intelligence can help you figure out how to sort out your whites and your colors. whether to use powder or pods, and what to do if you get a stain and you don't have laundry nearby.

1:35:39It's special, it's important, and it's a part of my life that I will never want to be without. I love that. So the ChatGPT now has an official Twitter account, and they even got in on the Laundry Buddy meme, which is amazing to me. I actually spent a couple of hours this morning hanging out with Boris Power from OpenAI, who was in there batting for Laundry Buddy from the start. Wait, there's an anti and pro Laundry Buddy? No, I mean, he was just a particularly strong enthusiast. He had the grace to not even bring it up, unlike you. I had to. It was so funny. I cracked up so much. It was great.

1:36:20Well, thanks for chatting, Susan, and I'll return you back to your evenings. May your clothes be well-launted. thanks for having us cheers thanks that was jeremy howard together with tanishk abraham and jess leo tanishk and jeremy recorded a podcast separately so if you want to learn more about he's done long-form interviews in more detail than i can cover because it's a lot of biomedical stuff and that's one of the areas that we are not very knowledgeable on and for jess leo she was an investor in mosaic is one of the newest partners at decibel and led the round in answer AI. Next, we're going to go to some people on the show floor of the NeurIPS Expo.

1:36:55They're not people I had prior relationships with, but they're still doing interesting work nonetheless. And the first is we're going to check in with Cerebrus, which is not only producing giant massive GPUs, but also publishing interesting research. So here's my conversation with Joel Hesnes, Principal Research Scientist at Cerebrus Systems. That started working about a year ago. We started building out multi-box systems so that we could do cluster-level training, so larger-scale models. And so this last year, we've just been showing off what it's capable of. So early this year, we started with our Cerebrus GPT models that showed compute-optimal scaling for chinchilla-style scaling, but it's open source.

1:37:40All those models we released are open source. Based on that work, we got attention of a few different groups. One of them was the OpenTensor Foundation, and they came to us and said, hey, we want a great 3 billion parameter model that does something that's easy to deploy, like in a laptop or something. And we wanted to do very general language capabilities, long sequence length. And so we trained the BTLM language model for that. Concurrently with that, we also had an engagement that started up with Group 42 in the United Arab Emirates. So that's this poster, Core 42. They had interest to train large Arabic language models.

1:38:25So the first demos that we did for them were just Arabic models. But then they said, let's do multilingual Arabic and English. So we've been training the JACE 13 billion and 30 billion parameter models this year. We've released both of those publicly. The first version of the 30 billion just came out. And the quality of that model is, in Arabic, is better than any other public models currently. And then in English, it's competitive with models like Falcon 40B. So we're on a good track there. more releases to come through Core 42. We're excited to have that be open source and to contribute to the community there.

1:39:11Anecdotally, since we're already chatting, might as well keep going. The UAE also notably has the Falcon or TII Institute. Are they related? Are they competing with each other? What's going on? Initially, there was a little bit of competition. They're funded by different people, different groups. But there is a countrywide effort going on in the United Arab Emirates to consolidate a lot of their AI efforts. Yeah. And so that's why we're seeing very impressive and good pushes towards let's make it open, let's collaborate some more. And so there might be opportunities in the future for us to coordinate directly with TII.

1:39:52And we have looked at things like their data sets, like Refined Web. So there has been some exchange. Yeah, with the macro data refinement process. I don't know if you know, it was a reference to an Apple TV show. Severance, anyway. My fun fact. A little bit of editor's note, the TII Institute people were actually there at NeurIPS presenting a poster on RefineWeb, the data set that they did for Falcon 180B and 40B. So I asked them about the name. My last question is about the name. Is it from Apple? Is it from Severance? Yes. So what's the story? No, it's just like, in the end, we had someone look at the data every now and then, like, go through the thing.

1:40:31And that's like looking at the scary numbers. So, you know, this was the magazine. You know, nobody comments about this. I know. I was like, wait, I saw this in Severance. Yeah, I know. Right? Like, I was like, this is a good joke. Because it's exactly what you do when you do filtering. Exactly. If you haven't seen Severance, it's a great show. It's on Apple TV. Great watch for the holidays. Pretty short. And it's interesting. I guess you can call it AI related now. But it's cool that, well, so one of the things I often get asked about, because we have listeners in a lot of different countries, should every country have their own model?

1:41:02I think this is a really tough question because the volume of data in different languages is power law, ZIPs law distributed. So the number of low resource languages is massive. We're talking over 100 languages that are low resource. You just have too few tokens to do a lot with in the language modeling context. So it's much harder to deal with those. Now, we've actually seen a few different techniques at NeurIPS that are targeting those sorts of settings. And they're doing things like train a base language model in English and then do transfer process where you co-train with both languages. That makes a lot of sense.

1:41:46It makes a lot of sense. In that setting, you want to get the knowledge representation from one language and then try to adapt the style, grammar, syntax, I guess, the easier part. In Arabic, we're in a sort of medium resource language. There, I think it makes more sense to try to mix two languages if you want to do multilingual. And then it helps you do things like translation. And then higher resource languages, so if you're talking European languages, French, Spanish, German, those I think you can do probably from scratch in those languages. And probably pretty easy to do multilinguality also.

1:42:28Yeah. So, yeah, it's definitely a very interesting open direction we're pushing for. In fact, I'd maybe reference, we have a multilingual workshop on Friday where we've invited a bunch of groups to come and give talks about their experiences with training different language models. Cool. Well, people can check out the authors. I'm sure this is published in Findable Online. Yes. Cool. So we should probably get to intros a little bit. I mean, we're already recording. Who are you and what do you work on and what does your teamwork on? So my name is Joel Hessness. I'm a principal research scientist at Cerebra Systems, and I'm the lead of our core machine learning group.

1:43:12So I've helped us bring up our foundation language models first and helped kind of set some of the direction for expanding outward from there. So we started by expanding out a lot on the common language functionality, and now we're expanding into other places where transformer models can be used. So targeting things like multimodal and other workloads that are similar. So a lot of our effort has been bringing this up and coordinating with the broader Cerebrus organization to lower these applications down, get them compiled to run ad efficiency on our hardware. so there's been a lot of performance optimization, making sure numerics are correct for training large models, making sure things train stably, things like that.

1:44:05So yeah, we're focusing on scaling out right now, getting much larger clusters. We've sold a couple already. To G42. G42. Yeah, exciting things to come there, I think. Exciting things to come. So we're going to cover some of the other posters that you have here, But one thing I guess I, people are very unfamiliar with anything but NVIDIA. What should people know when working with a Cerebrus chip? Sure, yeah. I think maybe people might be familiar with our wafer. So Cerebrus uses a full wafer for our processor instead of cutting the wafer apart into pieces. If you cut it apart, you end up packaging it into a bunch of different cards.

1:44:47And then you package those into a box. Then you have to network them. And then you have to network them all together with a bunch of extra software. That's very complicated for large-scale applications. And so instead of doing that, we leave it together on a single wafer. That single wafer goes in a single big box. The performance is roughly equivalent. Our CS2 box is roughly equivalent to maybe 20 A100 GPUs. And you can program it like running on a single GPU. So it's just much easier to use. Nice. And is it cost-effective as well? I assume it is because you're saving a whole bunch of overhead.

1:45:24Right. So the manufacturing process has a lot lower costs because we don't have to deal with as many moving parts. Fewer points of failure. Reliability is quite good. And we aim to be price performance comparable to GPU systems. Cool. Awesome. That's the hardware stuff. We're also going to talk about the streaming things in a bit. But yeah, I'd love to, whatever you want to pick next as one of your works for this year. Just give an overview of some of our research directions. So our hardware is, it has native support for completely unstructured sparsity. What that means is we can send in, say, if we're using the weight streaming mode, which I mentioned, a weight that comes in, we can do a vector multiply with some activations.

1:46:17So you can use that in your matrix multiplies on the wafer, but you can do that on a per-weight basis. You don't need to load the whole thing at once. You don't need to load the whole thing to do matrix multiply. So what that means is we can do unstructured sparsity, just send in the weights that you actually want to use in the matrix multiply, and you can get a sparse matrix multiply. Isn't the decision for, like this is the argument, classic argument against that kind of sparsity, is that the decision actually takes longer than just doing the math anyway. like the branching the the sort of turing complete branching that's a yeah so uh part of the approach that we're using is a weight sparse approach which means the sparsity is in our in the model itself and so then while you're training that you'd prefer those weights to be uh the same sparsity structure for a while okay so there are techniques that train some kind of constraint some regular thing right yeah so so uh the the early works in this for things like the lottery ticket hypothesis where you'd find the john frank john frankl is like 10 feet from us yeah and uh there you find the mask uh by by doing some heavy duty training and then you rewind and retrain the model from scratch yeah now that's uh static sparse so that you have the same weight sparsity all the way throughout that works great on our hardware we have however added a bunch of new functionality that's sort of beta in our recent release that allows you to change the sparsity throughout training and so that's um that's something that's being used in recent research works like uh the rigging the lottery ticket uh hypothesis work so wriggle and then another one called set um a different approach to deciding how to change the sparsity.

1:48:04But those updates happen infrequently enough that it doesn't harm the performance on our hardware. That's cool. Awesome. So this is Sparse IFT is the paper that you published. Yes. So our Sparse IFT work looks at different ways that you can swap out layers for sparse versions using the same flops that might be able to get you better representation capability. So if you have pressure in your representation that's in your activations, for instance, let's widen the layer and sparsify it to give the model more activations. You can store more in those activations. Those end up staying dense. So our results here show that we can get something like a 2 to 3x performance improvement at 75 % sparse.

1:48:55Or you could flip it around and you can get, for the same flops, a better model by sometimes 3 to 5%. That's probably budget-wise. I guess you're choosing between pre-training and inference, just like many people, like what you're optimizing for. Yes. That's great. Awesome. And what else are you leading? So I'm also working on some of the pre-training efforts that we're doing that look at things like gradient noise to estimate good batch sizing and make sure that we're making efficient use of the compute. So there are techniques. So we have a poster, the efficient and approximate per example gradient norms paper.

1:49:37This is, yes. So this is at the, we have this published at the WANT workshop with NeurIPS. And the basic idea is gradient norm calculations are, typically if you wanted to do the gradient norm calculation, you'd want to aggregate all the gradients together and then calculate the norm. And you do that over your batch. So that's, it's helpful if you want to measure some training dynamics. But if you want to look at something like critical batch size to understand how well is my model training in terms of efficiency, you actually want to have sub-batches. You want to understand the grad norms of the sub-batches also.

1:50:19If you use that and then the large batch grad norm, you can calculate noise statistics, like signal-to-noise maybe. If you use this technique that was defined by one of my teammates, Gavia, we can do an approximation that allows us to run some statistics over activations and run some statistics over the delta gradient values coming back. And then you can take a dot product, an element-wise product of those now, it's much more compute efficient, to calculate for each example, this is an approximation of the gradient for that sample. And then you can arbitrarily kind of combine those back together to get estimates of gradient noise.

1:51:07Okay. So this is something where we improve the compute requirements. We use this in a few different contexts currently, but it improves the compute requirement for this from, for high-dimensional tensors from the dimension of the tensor down to linear time computation. Nice. And do you, is there like, I forget what this is called. It's kind of like an annealing curve or something where you use this technique at the start to initialize and then eventually you sort of wean yourself off it. So this is something you do want to track throughout training, especially if you're doing like phase training or if you're changing the data distribution or something.

1:52:00It's really helpful to have these statistics to decide, am I using an appropriate batch size that I'm getting good generalization with the new data? It helps you set learning rates and things. So this is something you want to track throughout training. It gives you an estimate of how big the batch size could be. Yeah, excellent. Very cool.

1:52:26One more? Sure. So then given that we have a sparse accelerator, we're also looking at applications where you can deploy sparse models. And part of our work is figuring out how to find those sparse models that you use in a deployment setting. And so we have other work that's related to like the sparse GPT work that's been recently released where we do some pruning after dense pre-training. And we do some retraining to get the capabilities of the model back before you would put it in deployment. How much of it can you get back? Actually, I'm not totally familiar. This is work from my team members.

1:53:11I know we can do so for for large very large language models that have have not been trained on a huge number of tokens you can do easily upwards of 50 % sparsity and fully recover the like the upstream losses from this retraining so this is this is a really big next step challenge for a lot of the organizations that we work with they're interested now I have they're able to pre-train a very large model with the hardware. Now they're interested in figuring out how to deploy it in an efficient manner. So we're working with a few different groups on this. So we're working with Qualcomm and another group called Neural Magic that does inference for these large models.

1:53:58Yeah. Amazing. I was going to ask if you need the same data set to retrain, but it looks like you train on the pile. So I guess that's a no. Yes, you can actually shift here. Obviously, different data distribution means you have to be a little bit careful about how you do the retraining. So I think there are a few different things we've learned about different learning rate warmups, different learning rate levels, I guess. Because if you're doing a big distribution shift, you want to allow the model to shift a little bit. And so you want a slightly higher learning rate. But like, for example, you prune LAMA 2, and we don't know what the original dataset was?

1:54:39Yeah, I mean, well, so we kind of know that Llama 2 is a little bit similar to something like Slim Pajama and Llama 1, but yeah, it is definitely a different dataset. We do know that Pile and Slim Pajama have a fair bit of overlap in some things, but it is definitely a different distribution. Yeah. So this is a lot of work that our Applied ML team, our Applied ML team is working on. And we're expanding that team currently, by the way. So Cerebris is hiring for anybody who's interested in listening. You can check out our website, cerebris.net slash join dash us. If you'd like to check it out. Send us your resume and we'll take a look.

1:55:24Yeah, thanks for spending some time with us. Before we go, what's one NeurIPS tip that you want to give to people if they're attending NeurIPS? How do you do NeurIPS right? How do you do NeurIPS right? Well, so it's grown roughly 5x in the time that I've been attending Neurip. So it gets more overwhelming every year. So pace yourself. And I like that they've kind of backed off a bit on the talks and in favor of poster sessions. Like just you got to go wander around. You got to talk to people. You got to check out posters and kind of let stuff sink in and ask questions. So, yeah. Yeah, excellent.

1:56:04Well, thanks so much for your time. Definitely. Thanks. That's it. I think Cerebus is doing very interesting work here. Most people know them for their hardware, but I think they're doing very interesting work on the software and LLM trading side. And I'd be interested to have them on again in 2024. So next, we're going to go walk down the floor to Voxel51, which is not a company I've actually come across before. But it seems to be an interesting pair together with the next guest as well. So this is another one of those situations where I get to put two competitors next to each other and let you decide as to how they differ and how they talk about themselves.

1:56:36Sure. My name is Jason Corso. I'm the co-founder and chief scientist at Voxel. I'm also on the faculty of EECS and Robotics at the University of Michigan. So Voxel 51 is a spin out of my lab. We make a toolkit for AI engineers that sits on top of things like PyTorch and TensorFlow. And I think of it like a model and data set debugger. The key problem that we face is not that we can go download data sets and then train models on them, or even with foundation models, go pull one off the shelf and then expect it to work exactly the way you want. The problem is really the code development of a data set to then go and actually use one of those models or train or fine-tune your own model.

1:57:16So 51 lets you represent the data that you're using or building alongside your models in a way that is extensible, visualizable, and flexible so that you can write single Python lines of, like simple single lines of code in Python to do queries of your data sets and your models. Like, show me the corner cases where Model A is outperforming Model B, and it's outdoors. Or show me, you know, intersections in my BDD data. Or let me visualize my embeddings that are either just vision or point cloud-based or multimodal. And then visually interact with them with lassoing on the 3D embedding. Is the concept of active learning still in vogue, or is it, like, not cool these days?

1:57:57Well, I mean, so 51 is a pretty flexible ecosystem of capabilities. The heart of it really is that data-centric data model of unstructured data. So we support images, video, and point clouds. You can, in fact, there's a blog that one of my colleagues at Voxel51 wrote maybe like a month ago on how to implement an active learning workflow on top of 51. So it's plausible. It seems like it will lend itself easily. Yeah, exactly. It's plausible. I mean, the challenge with active learning is, you know, will just more data help or do you need to write more data? Of course the write more data. That's kind of a, you know, that's the question, I think, right?

1:58:32Yeah. Is it primarily vision that you work on or is it just anything? Yeah, so my experience is in computer vision, mostly video understanding and imaging problems. So that's where we got started. However, the software is pretty flexible, so you can add your own data type. Like, you know, we're considering adding audio, adding text, IoT, you know, like temporal signals. But right now it's images, video, and point clouds. I've often heard it said that, you know, But the best researchers and the best engineers are really the people who get their hands dirty in the data sets. ALEX KOMOROSKEY - Oh, yeah.

1:59:02You have to get your hands dirty. So in some sense, the whole company exists. Because I was worried no one was getting their hands dirty enough. They were just expecting to take a data set, take a model, and then train it once, and then out pops your usable thing. No, that's not the way it works. This is a hard problem in building intuition, building a comfort or an ability to take a 10 million sample data set and find the 1 ,000 samples that are giving you this problem here is hard to do and that's what 51 really lets you do. Yeah, yeah. So why the name actually? I have to ask. Well, we had 50 bad ideas.

1:59:36And this is the first, the one that was like actually good. Well, that's the way we say now, but the actual original way we got started as a company was as a video understanding as a service platform. And so that's why, so the voxel in the name is in the space-time volume of pixels. Yep. And 51 was just to elicit ideas of Area 51. Can you find the right voxel? Is it there? That kind of thing. We've subsequently way pivoted away from that, as most startups will do at some point in their journey. Yeah, it makes the domain easier to buy. Sure. Exactly. So anything else people should know about your platform, like top use cases, top customers that you always brag about?

2:00:12Sure. Well, I mean, it is open source, right? So as long as you have the three key assumptions, local data, one user, one machine, there's no limitation on the machine learning that you can do with 51. when you want to violate one of those assumptions, like work on a team or work in the cloud or whatever, then we have an enterprise product that you would talk to us to purchase, basically. And that's kind of like a Google Drive layer on top of the open source one. Yeah, very reasonable. Yeah, the only, I mean, we sell to, a lot of companies do use it. I'm not going to name them here, but you can go to the website.

2:00:42There's a logo wall of those we can name. But it'd be great if you're listening to give us a GitHub star. Yeah. That's our, like, we're here at NeurIps to get users. Stars for swag. Stars for swag. You got it. Excellent. You published a guide to doing CVPR, right? I did. We're here at NeurIPS. What would be your guide for doing NeurIPS right? So how to do NeurIPS right? I think there's some key things of doing large conferences right. One is, like, don't expect to do too much per day. Yeah. Right? So what I've always done, even when conferences were, like, a quarter of the size or less, like, for any one day, identify five to ten papers in the morning that I just want to go, want to understand for that day, right?

2:01:20So then I will make sure, though, to spend time with that poster presenter at the oral talk. To me, that's the key. And then at the end of that day, I do tend to write a summary for my own brain, my own notes of what I did, like what the key points were for those papers. That's definitely one winning strategy for a big conference like this. All right. Any other advice for people building, you know, any papers that you're excited for this year? Well, I mean, advice, I don't know. If you don't know your data, then you don't know what you're doing, is the way I would probably say it. And indeed, getting close to your data is part of the model building process.

2:01:56Just to say it again, I think of it as a co-development process of data sets and models, not of a model training problem. Yeah. I actually had a really interesting chat with someone from Cerebrus, actually, where they talked about how they were doing evals on their loss per region on a data set as they were training their large language models so that they could increase the exposure on a specific subdomain if they saw that specifically loss was not progressing as well in that particular subdomain. So it's kind of like online training and watching their models evolve while they're training. Yeah, I guess it sounds like on specific subsets of the data, which is really important.

2:02:32Cool. Well, thanks so much for your time. Thanks very much. Nice to chat with you, Sean. Coming from data engineering, it's pretty interesting to see this space develop. It's interesting also that a lot of them emphasize open source, which we'll see with the next speaker, which is Brandon from Nomec. Who are you and what's Nomec? Yeah, hey, everyone. My name is Brandon AI. I'm a co-founder and CEO of Nomec. Nomec is a company that does many things, but we have two main products right now. One of them is GPT for All, which is an open source ecosystem of low resource language models. So it lets you do things like run, you know, Mistral 7B fine-tuned on OpenOrca on a MacBook or, you know, some esoteric GPU, things like this.

2:03:09The second product is a tool called Atlas. It lets you explore massive unstructured data sets in your web browser. Since we're here at NeurIPS, a lot of people seem to respond to calling it massive clickable T-SNE as a service. Yes. I was actually thinking, is it T-SNE or UMAP? Yeah, so it turns out if you squint closely enough, they're the same algorithm up to a choice of low-dimensional kernel. So we optimize the T-SNE objective function. One of our uses of IP is we have the world's fastest optimizer for it. So if you take, say, the NVIDIA Rapids UMAP implementation, which is kind of the fastest version of this in the wild, off the shelf and run it on Wikipedia, on the biggest machine on AWS.

2:03:45It's going to take you a couple of days to actually get that map. And we can do it in about four hours. So yeah, it lets you make the maps part of your iterative daily workflow as opposed to having to wait a week to get them. Nice. We'll throw a video on this on the show notes, but maybe you could sort of narratively show what you're showing. Like you showed a TikTok example and a Twitter example, right? So these are really for visualizing massive multimodal data sets. Yeah, so the fundamental thesis behind the tool is that the shape of data that people have has fundamentally changed as a result of generative.

2:04:14Instead of having these big Excel spreadsheets of tabular things, you now have vectors plus metadata. And we need to rethink visualization and the implications of that for the visualization stack. You are kind of seeing at the database layer they're starting to penetrate with vector DBs and stuff. But I think there's going to be radical kind of implications for that change all the way up the stack. And so you can use it on, getting back to your original question, Twitter data, TikTok data, images, sounds, text, Anything that you can stuff into a vector, which is pretty much anything these days, you can map and you can understand.

2:04:44Yeah. Can I bring my own custom embeddings and see the impact of that? You can. So there's two ways to get data into the platform. One way is bring your own embeddings. And then you just pip install Gnomic from Gnomic Import Atlas and then atlas.map embeddings. You supply your embeddings. You supply metadata on top of them. And then a couple minutes later, you'll get a web link back to a map where you can click on it and fly around it. If you just have raw data, we have a bunch of out-of-the-box embedders that we develop and we work with partners to develop that you can use to map it out the box as well.

2:05:12Yeah. And this is not open source, but GPT4All is. So there are aspects of the platform that are open source. The entire thing runs on a graphics engine that we developed called DeepScatter. It's the only tool out there that can render a billion point scatter plots in a web browser. And to do that, you have to, again, kind of fundamentally rethink how graphics in the browser works from the ground up. That is available source, but unfortunately, it's not fully open source. It's okay. Yeah, you don't have to apologize for anything. I do have to. You know, I wish we could open source everything, but we are unfortunately subject to capitalism, and so we cannot.

2:05:43But in the limit, I would love to open source everything. I also maybe heard you in another introduction talk about this as Looker for language models. Elaborate more about that. Yeah. Do you have a query language? What are you thinking about as the overall vision? Yeah, so I want to bring it back to the analogy of the new shape of data disrupting the stack, right? So the first place we see it hitting is at the database layer. We see vector databases, and there's a million of them nowadays. is, I think that that change is going to propagate all the way up the stack. And we are interested in what happens to the BI analytics visualization layer.

2:06:13And so really what we're thinking of this as is sort of like a tableau for unstructured data or a looker or Power BI or something like this, where we've built the entire visualization system with embeddings as a first-class citizen. And so that enables a lot of different actions. Some are already in the platform. Some I can't tease yet, unfortunately. but having embeddings as a first class primitive enables a lot of like very very useful things that you're not gonna be able to get unless you have that so what do people use atlas for like just maybe list out some more use cases that might not be obvious from people just thinking about visualization yes we'll start with the most technical one we'll go the least technical a lot of ml engineers use it to understand and evaluate their models and training data yeah so we just did some work with hugging face on their obelix data set which they use to train their Idefix model, doing some evaluation and training data analysis, looking at what areas of their...

2:07:03We actually interviewed those guys. I was in Paris, and I talked to Leo and... And Victor. Yeah. Yeah, those guys are sick. But yeah, so we worked with them on this, and we discovered a couple of things in their training data they should have actually cleaned out of it. There was a bunch of end-of-sentence tokens to be replaced that made it through stuff like this. Some really garbage content. Did you do anomaly detection? Or is that up to people to code themselves? Yeah, so the anomalies usually manifest as the little moons on the outside of the map. Oh, sure, okay. And then you can just hit them with the little lasso tool and stuff like this.

2:07:32But one of the things about the Hugging Face map that I found fascinating was because we supply a topic model out of the box, you can look at things like, are there topics where the loss tends to cluster together? And for the Hugging Face model, there was this high loss mode in the poetry topic, which I thought was super interesting. And so I've got two theories for it. One is that poetry includes the distinct subversion of common linguistic patterns. And so, of course, language models will be bad at it. But the more, perhaps, optimistic theory is that poetry captures something that's fundamentally human, that the machines have not grasped yet.

2:08:05The pragmatic version, I think, is probably what's happening. But I like to be optimistic. IdaFix is a visual data set. And you are multimodal. Yeah. Okay. So they have poetry in there. Yep. Interesting. It's sort of interleaved webpages of, like, it'll be an image and then some poetry. So that's the more technical side. And then coming down to the less technical side, you know, a lot of our customer base at this point is like consulting type companies. And they find the product really useful for connecting domain experts with large data sets. So generally what will happen is you'll have these domain experts, be it like a doctor or someone in regulation, someone with subject matter expertise that'll be handed this massive set of documents from a client and be like, I don't even know where to start.

2:08:41I don't even know what's in this. And so a couple of the consulting partners we work with actually now have a KPI that's like time to Atlas, where it's like how quickly from the data set hitting the company does it get to Atlas so that we can send an analyst to the map and they can start to explore it. And so we're really excited about enabling sort of traditionally non-technical people to explore and analyze these massive data sets with this no-code interface. You know what you should do? You should hook up with Google. Doesn't Google have a big set of publicly available data sets? Yeah, so we've actually done a couple of collaborations with Google Cloud on some of those data sets.

2:09:13We can maybe link the blog posts or something. Sure. Okay, awesome. Just NeurIPS in general, you've been here a number of years. What do you look for when you come to NeurIPS? Any tips that you have for people coming to NeurIPS? Oh, that's a good one. Yeah. Big tip is just, like, if you see someone cool, like, they're probably nice. So chase them down and, like, have them talk to you. Shove a microphone in their face. Yeah, yeah, yeah. No, I love it. But it was, like, my second NeurIPS or something, I saw Oriol Vinyols walk by. And he had just done, like, the StarCraft stuff. And I was like, okay, this guy is sick.

2:09:44He's doing some really cutting-edge stuff. so I ran up and asked him for life advice and he was so down to earth and shatter with me for a bunch of time about modeling and life and how to think about my career and stuff and so yeah, if you see a hero, shoot your shot Yeah, very, very cool Any papers that you're keen on this year or maybe really affected you in previous years? Oh, that's a good deal. This year I think QLora's here, which I think is like a very very interesting Yeah, it's a very interesting set of implications for the low resource world. Can you elaborate? Yeah, so one of the things we think a lot about at Gnomic is the accessibility of AI technology and one of the things that's become very clear to us and I think everyone this year is like there's the GPU rich and the GPU poor and so I think methods that make it so that anyone in the world can interact with this technology like QLaura are just like so so so valuable and so I think any research into like low resource training of models and low resource deployment of models is just going to be so good for everybody especially like the open source community, I really love to see it.

2:10:45Yeah. You just reminded me. So we forgot to talk about GPT-4-All. Yeah. Very, very early win, I think, in the overall space of things. But now, more recently, in my mind, Llama CPP has come out to be its own platform. Yep. Old Llama is emerging as a thing. There's a bunch of ways in which people run models locally. How should people think about GPT-4-All in the context of all that? Yeah. So one thing that a lot of people don't realize is that a lot of the core contributors to Llama CPP actually work at Gnomic. Yeah. And so I guess the operant advice here is just play nice with open source. GPT for all is this thing that's going to be free forever for our community.

2:11:21We're going to keep trying to improve it as our Discord recommends and as people call for. But if we can do things like go and contribute to other open source projects that are high impact, we're going to. And so the hope here is that as economic pressures apply, open source stays collaborative is really the goal for us, I think. Okay, cool. Well, that's it. Any other last words? What are you looking for? How do people find you? Yeah, you can follow us on Twitter at nomic underscore AI. You can also find our website, nomic.ai. Hiring engineers, researchers, any specific profiles? Super interesting people.

2:11:57Yeah, come chat about interesting things in our Discord, really. We can just our website and stuff. But really, the best way to get involved is make some maps, do some open source work. A lot of the people that we hired in this last spree of hiring were big open source contributors. and so like yeah just give back to the community and then you know we'll try and find you and boost you yeah awesome well thanks so much for your time yeah take care i think the way that nomics embracing and supporting open source ai is encouraging and i think more companies should learn from that but they're definitely far from the only open source ai company out there lightning ai is one of the oldest i guess if you can call that old in the space and i happened to catch luca the cto at their booth and at neurops they were there to launch lightning studio which their new development environment.

2:12:41Hey, Luca. Welcome. Good to see that you guys are launching a new product today. Yeah, sure. It's super exciting. It's the result of many months, if not years of work and realizations. So maybe let's establish a baseline. Most people will have heard of PyTorch Lightning. What was the evolution of Lightning AI? Yeah, so PyTorch Lightning is a very healthy community of people using it. We are 5.5 downloads, about 80 million downloads in, sorry, 5.5 million downloads, of course, per month, about$80 million in total. And it's one of the frameworks that comes from the era of traditional, quote, unquote, deep learning, that is one of the main actors in the Gen AI space.

2:13:22Because, for example, Stable Diffusion was trained using PyTorch Lightning, a bunch of models. PyTorch Lightning powers Nemo from NVIDIA. Yeah, they're a custom chip design language model. So basically PyTorch Lightning has evolved and grown into Gen.AI. And with the release of 2.0, 2.1, we've tried to make it better and better for use cases in which you have very large models and you have a hard time not going out of memory. And do distributed. PyTorch Lightning has always been very focused on distributed trainings. one of the things that he did the best but when models get very very large I think that's where we improved a lot this year we also launched fabric lightning fabric which is a it's a framework it's a companion framework to fighters landing where you get all the constituents of the lightning trainer but now you can write your own trainings so for people doing very optimized stuff, very bespoke, I don't know, you know, collectibles, they want to place them where they want, they want to fully own the training loop, or they're doing stuff like reinforcement learnings where it's not the traditional training loop, you can still do it with a trainer but it's a bit more difficult, then Fabric lets you just write your for loops.

2:14:50But we'll still abstract away strategies, precision plugins, the login, the aggregation of metrics, and all this stuff. I like to think about these frameworks as frameworks that reduce the surface area for mistakes. Because mistakes nowadays, well, a few years ago, mistakes, exactly, right? They cost a lot of time to a PhD student. Right now, they cost a lot of money. So you don't want to make too many mistakes there. and Torch Matrix is the third project that we have that is very healthy and is following a lot of the metric computation. Again, you don't want to compute accuracy and aggregate it across a multi-machine job in the wrong way, right?

2:15:33Because you'll get wrong indications and it's really easy to do it incorrectly. Yeah. And this year we started doing... Yeah, I should mention these are mostly open source. Yeah, these are all 100 % open source. I think Fabric in particular was pretty popular. Yeah, so Fabric has powered also our language model repositories, LitLama and LitGPT. Basically, back when Lama was originally released, I and me, of course. The weights were leaked. Huh? The weights were leaked. Yeah, exactly, exactly. But at some point, there was a model being published by Meta as well. It was GPL licensed, so we didn't really like that.

2:16:19And so we said, why don't we take NanoGPT, because I was working with NanoGPT at the time, and turn it into LAMA. And that started the whole thing of minimal implementation, single file. You have everything there. You have no layers to go through to understand how your layers are. And that became something that became very popular within many organizations. So it's still very popular. So the LLM efficiency challenge, the starter kit had LeadGPT in it. And LeadGPT today supports many models, many different models. But it's very easy to get to the bottom of the implementation of every single thing.

2:16:58So it's very hackable. My philosophy is make it hackable before you make it fast. Because more people can contribute to it and we have contributors being very successful. there have been initiatives of models being pre-trained using that, like TinyLama and 360AI, I think, a few days ago came out, and they said they used LeadLama to pre-train their 7 billion parameter models. So it's great. And a lot of those learnings went back into public and back into PyChurch planning. And this is how we're kind of growing organically towards supporting Gen.AI use cases. There's an example of one of those learnings from those outside usage of LitLama, I guess.

2:17:44Sorry, can you say? What's an example of one of those learnings that you got from 360 contributing back? Well, 360 is very vanilla in the sense that we just learned, I think, the day before yesterday that they used us. So it's great. We're very happy about that. From TinyLama, they did some optimizations on top of our code. and they trained a 1.1 billion parameter model on 3 trillion tokens. I think they're still doing that. I don't think they're done. And then some of the improvements that they made, and we upstreamed it to our, like, for example, I think chunk cross entropy, some kernels that they were using.

2:18:31And then we were happy to see that even our data set that we optimized because it chunks your data and it can stream very quickly, work for them. So it's kind of a mutual thing that we're doing. And also, all the quantization support. For example, right now, Fabric and Byterge Lightning support bits and bytes natively. And it's basically one of the few solutions where you can use quantization on any kind of model and not just the model that the original authors decided to support. Yeah, it's kind of flexible. But here today, I think the main thing we're doing today is launching our platform. Yeah, you just launched Studio today.

2:19:14Yeah, exactly. I mean, Studio, again, is a result of many months and years of work. It basically makes you build AI at scale, but it feels like it's your laptop. So to me, it's kind of the first time I've seen a platform not leaking the abstraction of orchestration on the cloud and so on. literally there's nothing to learn. You put VS Code in the browser and then you add all the... You can even connect from your local VS Code and code there. You have the whole machine. It's a whole machine. It's a cloud development environment. Exactly. And it's built around reproducible environment. Exactly. But when you go in there, it's not that you need to build your Docker container.

2:19:56You just go in there, you present it with a machine, you can start working immediately. If you pip install something and then you decide to switch instance type, your dependencies will carry over. Or if you decide to duplicate my studio, everything that I set up on that studio from the environment to the data, the code, the checkpoints eventually that I put there, you will find them. And so you will spend zero time setting up your environment. So are you snapshotting memory? How does this work? Well, that's secret sauce. You're not using containers, you said. Yeah, well, I mean, we do, like, if you think about it, then it's not too complicated fundamentally, but it's very complicated to actually get the perfect experience out of it.

2:20:43Like maybe describe your design constraints. What are you optimizing for? We're optimizing for velocity. So we don't want people to spend time thinking about things they shouldn't think about. Like when you're coding on a machine and you now want four GPUs, you should just be able to get four GPUs and keep working. without thinking about, oh, now I need to go to a console, spin things up, board my environment, attach drives. These are all things you shouldn't think about. And again, it goes back to limiting the surface area for mistakes, right? Because you can do what you're good at and not do what you shouldn't ask for.

2:21:23It's like the fabric philosophy that's expanded to the dev environment. Exactly. We're very excited. You can do small things like in Colab, accept that your data is persistent and you can switch off and switch on and everything will be there. Yeah. Or you can even train large language models. Yeah. What are the larger customers doing? You know, what are you doing for them? Because I feel like this might be targeted towards the smaller customers. No, actually, we run with, we work with very big financial institutions. and we're actually pre-training models ourselves. So the scale at which you can operate is pretty large.

2:22:08It looks like something that you can do small stuff with, which is true. It's super smooth there. But if you need to launch a job on 100 GPUs, you can just do it, provided that you have the machines. But we manage reservations, so we can target reservations. Or you can attach your own cloud account and negotiate your quotas with your cloud provider, and we'll just orchestrate on your cloud account. Yeah. Any cloud providers you would shout out as particular... I mean, people know the big three clouds, but any other providers that you would shout out as very good partners to work with so far? Right now, we'll be focusing on AWS.

2:22:42We'll expand, of course, because... Yeah, everyone needs everyone else. Yeah, exactly. Apparently, Oracle's doing very well. Yeah, yeah. We talked to Oracle. We talked to most of the cloud providers out there To us, it's more a matter of sequencing. We have a very good relationship, of course, with AWS right now. They've been supporting us for the launch and so on. But surely we'll get into getting the best machines for our customers. And in the near future, we'll also support on-prem clusters in terms of orchestration, like Slurm as an orchestrator or as a scheduler. People make feelings about Slurm.

2:23:21Well, yeah, but in this case, you don't have to deal with it, right? We take away the pain and you still can orchestrate on top of that. It's still not out, but it will come in the near future. We're already doing that with some companies. Yeah. So I want to talk about the workshop that you're doing on Friday, your efficiency challenge. Yep. Was it motivated by a paper? I saw it like a cramming paper. What is the maximum you can do with one day of compute, something like that? Yeah, so we noticed because they, Mark Sarofim and the other organizers, ended up choosing LGBT as one of the models for StarCade.

2:24:04And we were happy about it, of course. And so we said, yeah, what we can do together. And we ended up, and we really like the principle. So we believe smaller models can empower people a lot, getting control and understanding how to extract value from AI. And so I think there's a dire need of consolidation, getting smaller, getting more efficiency, and getting the result you want in the shortest time as possible. And that's how the velocity will increase and how eventually open source will get there on par, if not beyond what's available in the closed source world. So we are fully supportive of that.

2:24:52The way we ended up contributing is we maintained a public leaderboard. And it was a nice experience because we integrated with Discord. There was a Discord channel. this is for the efficiency challenge discord exactly the efficiency challenge discord and we set up an agent that was running on a few of our machines and people could submit through a DM to the bot so that the bot would then spin up a job run things in a queue get back the results from evaluation and then essentially get a ranking on where they were. And that, I think, helped a lot motivating people to compete against each other, but in a very constructive way.

2:25:43And to be honest, in the first month, it's been very, very bumpy with that. It was all new infrastructure, and we were doing it in spare time, so it wasn't the best of the experience. So together with the community that wasn't there, they helped us figure out what was not going well and I think at the end we had more than a thousand submissions that were successful. Many more submissions that didn't complete submission problems like user code problems but there were more than a thousand submissions that were actually fully evaluated on that leaderboard. So the challenge is over but I don't know if you've done the analysis on anything to learn from the winning entries.

2:26:27So we've been, the rest of the organizers, yes, they have put a lot of effort in the next three weeks, four weeks, to reevaluate everything, run the first ones from scratch, and they've done an amazing job. And some of the code that we wrote for the public leaderboard ended up being part of this evaluation infrastructure. I was very, very busy with the launch. Of course. So I didn't participate there. I'll try to talk to Sebastian. On Friday. Yeah. Yeah. Okay. What will be the details from the winners? Yes. But I must say that all the community has been super nice. They were super constructive.

2:27:12I remember when Mistral first came out, there was a huge shred of people just getting in there, analyzing it, trying to find it. It was so much energy that we definitely want to push it forward. and we'll create public studios with evaluation frameworks on them. And we want to enable this kind of... So studios are shareable, of course. Yes. Yes. Not only shareable between you and me, but also community-wide. Yeah, yeah, yeah. So there will be a lot of things. You can go in there, use your pre-credits to just run the evaluation on your model. You can do that. Yeah, great. So last question. I've been asking everybody this.

2:27:48You've been coming to NeurIPS for many years. What are your NeurIPS tips? Oh, wow. Yeah. Go to posters, which I cannot do. Okay, so I've been to like, you know, there's been three poster sessions so far. The popular ones are just crowded. There's just no way. Yeah, I think there's not, yeah. I don't like to go to popular ones. Yeah, okay. The less popular ones and just talk to them. So I had so many super engaging conversations. Even in topics, like even something is not apparently something that you should focus on. Yeah. Your brain will oxygenate itself a lot. And typically after these conferences, I always come back with a full of ideas.

2:28:30And so I would say get enriched as much as you can by interacting with people, having very honest conversation with them. Yeah. I had to share. I had an off the record conversation with one of the presenters who said, like, yeah, I don't like this paper I'm presenting. I don't believe in it. I was like, wow, that's really honest. because like they submitted it months ago right and since then the world has moved on yeah well but that's part of the you know the struggle I don't I don't know how it must be being a postdoc or or PhD students or even a master students yeah nowadays in AI it must be like so stressful yeah back in the day it was a lot easier you know the prices have got bigger so yeah for sure yeah okay well thank you so much for your time and congrats on your launch yeah i think we'll you know meet each other on the platform maybe yes i definitely will try it out thank you a lot bye a few of my ai engineer and ml engineer friends checked out lightning studio and they were pretty impressed so i'm personally interested to check it out next year but last and not least i want to give the mic to jay alamar of cohere and llm university but more importantly of the illustrated transformer and is now writing a new book.

2:29:45We're here with Jay Almar, educator of many things. I've learned so much from you. Literally, it's one of those moments where at New York you just kind of see someone walking, I'm like, is that Jay? And then I had to get your attention a few times. But it's so nice to finally meet you. It's great to meet you and great to be here and sort of meet all kinds of brilliant folks. I've watched your stuff and sort of been watching the revolution and how you're helping sort of crystallize people's thinking about this new domain of AI engineering. Yeah. I think the title is very helpful as categorizing that class.

2:30:20Yeah, trying to do for my audience what you do for just general ML education, which is, I think, something that you've really done an incredible job of. Yeah, no, it's wonderful. It's what the community needs, definitely. As machine learning and AI sort of goes out of research and goes into industry. Exactly. It's a different persona, different background. One of the reasons I'm doing this recording is the kind of people that follow my stuff don't come here. Maybe they shouldn't, right? Some of this is too in-depth. But I'm curious, you've been to many NeurIPS. What is your general take of the vibe?

2:30:54What are people talking about? What's top of mind? There's a lot of LLMs. That's interesting to see. Suddenly a lot of interest, right? Yes. Yes, that is, let's say, maybe possibly a new development in NeurIPS. that's the area that's growing. And a couple of interesting keywords or groups or directions are diffusion, even diffusion for text models. There was a paper on that yesterday. What's the point of diffusion for text? Don't people want to stream things out? Well, I mean, if you think of, like, on the application side, autoregressive generation has some problems. So if the model makes a mistake with token 5, you're stuck with that problem.

2:31:34Yes, that's what Tree of Thought solves, which the Tree of Thought guy was here. Yeah, it's one, let's say, one avenue. But it's like maybe if the model does not fall within a mistake in that way, you can unlock new sort of different applications. But also all the image generation stuff, that's really where. Yeah, I'll make a plug. I actually had a, so there's a lot of house parties that happen after New Rips, which is fantastic. I ran into a guy from Mid Journey for the first time. They have this new storytelling section. And they are actually exploring text diffusion because of storytelling, because you have to generate a coherent story just like you would an image.

2:32:10True, true. So I would buy that as a use case. That is fascinating. And then with the agent stuff, like if you're interested in the future of agents, which is there's a lot of the reasoning stuff. The reasoning research in NeurIPS will most likely sort of inform the upcoming what's going to happen in agents next. So chain of thought, three of thought, that domain of research for me is very fascinating because it's going to be applied very quickly. The React paper comes out. It's in line chain. Everybody's using it. Everybody has a sense of what agents are. But that really shows you the potential of what they're going to be in the future.

2:32:46We're still in early days on agents. Any other top-of-mind sessions? Did you go to the Chris Ray run this morning? I thought that was pretty cool. I have that one and a couple of the other keynotes. I'll be re-watching them. But mostly I'm just talking with people, the recording video that's been my and sort of trying to orient myself it's an overwhelming amount of content and people and posters and talks and so i've been yeah looking at visualizations of you know these are the papers at neurops this these are the ones that could be interesting to you yeah people have published like tsn things of them and it's good but like it's not as good as just kind of seeing the vibes and i actually think the conference organizers do a good job of curating like what the you know oral session papers should be true you know like i i've generally found them I'm generally very insightful.

2:33:32I just found out about DataConf from one of the oral sessions. I don't know if you've seen them. They're effectively a new ImageNet. Oh, nice. Which is like, oh, that's cool. New benchmark. And yeah, I mean, to me, I'm taking it all in. So it's impressive how many people do so much work and you've never heard of them. That's true. That's true. Yeah. And they're conversant in all the techniques, all the papers, all the stuff that you see online. Yeah. They're just not online. That's true. And they just do research quietly. And then once a year, they show up here. Yeah, and sometimes you meet somebody here, and they would mention they worked on that other paper, and it's a paper that you're very familiar with.

2:34:07And then you go into their Google Scholar or something, and you're like, I've been reading this person's work for years, but the name never really specifically popped up until you meet them in person. So that's why it's definitely an interesting experience. No particular order, but have you had those, any underrated person you would call out as like, hey, everyone should pay more attention to the work that this person's doing, One thing that comes across, which is why workshops are good, and we can get into that sort of later. David Bao's work on interpretability and editing language models and editing their knowledge was one thing that sort of really stood out to me after I met David and sort of heard about Bavits' work.

2:34:49Editing by editing weights? Yeah, they have a method of editing the model, exactly. Is this the one where they played Go? This is Rome. This is where they convince a model using that method that the Eiffel Tower is in Rome and not in Paris. And then they have subsequent methods of, let's say, if you make 100 edits like that, the model degrades. So they have subsequent work on, okay, this is a better method to do many more of that. But also things like, I've seen like Logit Lens and sort of where in the model is this token being suggested? Is it layer one or is it layer five? Or that localization is interesting work.

2:35:30Yeah. So you do all these interviews on your YouTube. We'll send people there. Is this part of your work at Cohere? A little bit, yes. What is your deal? Yes, that's true. So a bunch of them go on the Cohere YouTube channel and the Cohere socials as well. So yeah, my work at Cohere, I get to learn in public, basically. I love that. So Cohere builds language models for embeddings, re-ranking, and generation. And through selling them, I get to see how industry is solving problems with them. And that, to me, is very fascinating. To see the technology coming out of research and then how it goes into industry and how people use them, how people sort of need to be educated on the best ways of using them.

2:36:11That view, to me, is something I'm lucky to have. Yeah, it's a good job to get, to be honest. If you love that stuff, you might as well get paid to do it. You probably don't know this, but I actually have written a book on learning in public, and I am a big advocate of getting developers and engineers to learn in public. Well, you do it so well. Yeah, this is my way of doing it. My final piece is, you know, you've written a lot of foundational work on, like, transformers. A lot of people are talking about the state-space models and what happens after the transformers. Do you have personal views on that?

2:36:39Not yet. Not yet. I'm on the lookout. So, yes, there are always new ideas that, you know, there's maybe poster number 502 here that nobody paid attention to. Maybe in six months we'll see that, oh, it crushes everything else. So that is always something you can never sort of expect. My favorite fact about the Transformers paper is it itself was not accepted. It was like a poster-only paper, right? I don't know the story behind that. It was a big deal for machine translation. Yes. But it's like, okay, yeah, there's a cool translation paper. It's one of many, right? Yeah, we already have Bert. One new attention method.

2:37:12We had Badano attention and we had Luang attention. and now we have also one more. But then BERT comes out and is like, okay, this is more than translation. Then GPT comes out and is like, oh, this can generate text. I'm still missing a good survey paper on everything that happens since attention is all you need. The evolution towards the modern decoder-only paradigm. And I feel like someone needs to write that. Everyone's too busy inventing new things to stop and write what happens. Because it's a massive thing. There are a few people who, because there's a lot of work on different kinds of attention for the transformer specifically.

2:37:46How to improve it for this problem, for that problem. But one thing that I'm doing is rewriting the Illustrated Transformer with the ideas that have stood the test of time since then. So it's like six years after, which ideas? So people are using Rope, Flash Attention, Alibi, and then Rope and Alibi, let's say positional encodings, localized attention. So some ideas that people are continuing to use over and over. Group query? Multi-query and group multi-query, yes. And then sliding window? Not yet. So it's in Mistral, but we maybe need to see it in more work. My conspiracy theory about Mistral is that so the Mistral paper heavily features sliding window attention and everyone is like, bullshit.

2:38:27Like, come on. I mean, because you see it. You saw that for two years after the transformer, everybody was proposing new ideas. And if you put this in the transformer, it does better on this. But then what stands the test of time? The vanilla transformer really stood the test of time. and did better than even a lot of these quote-unquote enhancements. But these ones, let's say, stood the test of time. So this rewriting is going to be part of the book I'm currently... Are you writing a book? Yes, I'm writing a book for O 'Reilly called Hands-On Large Language Models, including this as a chapter.

2:39:01So if you want an updated, illustrated transformer, that's going to be a part of it. When you launch a book, you should come on and do a full episode with us. That's amazing. Yeah, exactly. And then just general Neuribs tips. as an attendee? If people were coming for the first time, what would you advise them to do? I really love the visualization. I posted about this by Hendrik Strobel and Ben Hoover of the T-SNE of all the papers, but also it's clustered. So if you're interested in language models, that is clustered. I use that for my planning. It's so useful. These things are absolutely incredible.

2:39:33I got to meet Hendrik. They have a lot of very interesting ideas there and help to sort of orient yourself. And I've also seen work kind of like it, but where you can do semantic search on, so you can say, you know, agent papers. And it doesn't need to match the actual keywords. With GoHere, we have a demo on RAG on NeurIPS papers as well. So you can ask a question, you're like, okay, I'm interested in LLM and efficiency. It'll say, okay, this paper, this paper, this paper. And it's retrieval augmented sort of generation. So these are the three tools, but I think we need a lot more of these tools to make sense of this.

2:40:09Yeah, I need it for the meetups too. You know, in your conference app, there's all these like meetups for very specific things. That's true. I started one for Singaporeans because I'm a Singaporean in tech. And yeah, there's just a, there's a bunch of very, very specific, like running meetups, like nothing to do with tech specifically, but like, you know, this is also a social event, right? Like that you're meeting. Okay, yeah. You wouldn't happen to be at MNLP. No, why? Some people did that because it was like last week and some people went to MNLP in Singapore and then flew back here. Yeah, it's a tough call.

2:40:39Yeah, I'm not going to do that. That's rough. That's rough. Well, thanks very much. It's a pleasure to have you on. Pleasure to meet in person. It's so good to meet. Love your work. Thank you. Keep doing it. Any calls to action for people? Well, I'm JL Amar on Twitter and YouTube, and we have LLM University, LLM.University. I collaborate with Luis and Muir Amar to educate about... Some of the best YouTube, very short, but very comprehensive, authoritative. I'm very lucky to collaborate with these folks. it's incredible but yeah thanks for doing all that thank you appreciate it okay and that's it for our New York's coverage and for Latent Space Pod in 2023 we are still doing a listener survey so if you are listening through here you're definitely a big fan we definitely want to hear from you what you like about the podcast what you want to hear for 2024 we've got a couple of really good episodes already recorded for the start of 2024 so we're going to start the year strong and come out to the one year anniversary of latent space so thanks for all your support have a wonderful end of the year and we'll see you soon dj hit the outro

From the publisher

We are running an end of year listener survey! Please let us know any feedback you have, what episodes resonated with you, and guest requests for 2024! Survey link here.

We can’t think of a more Latent-Space-y way to end 2023 than with a mega episode featuring many old and new friends recapping their biggest news, achievements, and themes and memes of the year!

We previously covered the Best Papers of NeurIPS 2023, but the other part of NeurIPS being an industry friendly conference is all the startups that show up to hire and promote their latest and greatest products and papers! As a startup-friendly podcast, we of course were ready with our mics to talk to everyone we could track down.

In lieu of an extended preamble, we encourage you to listen and click through all the interviews and show notes, all of which have been curated to match the references mentioned in the episode.

Timestamps & Show Notes

* [00:01:26] Jonathan Frankle - Chief Scientist, MosaicML/Databricks

* see also the Mosaic/MPT-7B episode

* $1.3B MosaicML x Databricks acquisition

* [00:22:11] Lin Qiao - CEO, Fireworks AI

* Fireworks Mixtral

* [00:38:24] Aman Sanger - CEO, Anysphere (Cursor)

* see also the Cursor episode

* $8m seed from OpenAI

* Tweet: Request-level memory-based KV caching

* Tweet: GPT-4 grading and Trueskill ratings for rerankers

* [00:51:14] Aravind Srinivas - CEO, Perplexity

* 1m app installs on iOS and Android

* pplx-online api 7b and 70b models

* Shaan Puri/Paul Graham Fierce Nerds story

* [01:04:26] Will Bryk - CEO, Metaphor

* “Andrew Huberman may have singlehandedly ruined the SF social scene”

* [01:12:49] Jeremy Howard - CEO, Answer.ai

* see also the End of Finetuning episode

* Jeremy’s podcast with Tanishq Abraham, Jess Leao

* Announcing Answer.ai with $10m from Decibel VC

* Laundry Buddy, Nov 2023 AI Meme of the Month

* [01:37:13] Joel Hestness - Principal Scientist, Cerebras

* CerebrasGPT, all the Cerebras papers we discussed

* [01:56:34] Jason Corso - CEO, Voxel51

* Open Source FiftyOne project

* CVPR Survival Guide

* [02:02:39] Brandon Duderstadt - CEO, Nomic.ai

* GPT4All, Atlas, Demo

* [02:12:39] Luca Antiga - CTO, Lightning.ai

* Pytorch Lightning, Lightning Studios, LitGPT

* [02:29:46] Jay Alammar - Engineering Fellow, Cohere

* The Illustrated Transformer



Get full access to Latent.Space at www.latent.space/subscribe

More from Latent Space: The AI Engineer Podcast

All 247 episodes
NeurIPS 2023 Recap — Top StartupsLatent Space: The AI Engineer Podcast · 2 h 42 min
Listen in VO