MosaicML's Naveen Rao on Making Custom LLMs More Accessible - Ep. 199

12 Jul 2023 · 31 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

NVIDIA AI Podcast - Episode 199 Summary: MosaicML's Naveen Rao on Making Custom LLMs More Accessible

Podcast Overview The NVIDIA AI Podcast explores how cutting-edge technologies are transforming various sectors, featuring discussions with innovators and experts. The episode titled "MosaicML's Naveen Rao on Making Custom LLMs More Accessible" focuses on MosaicML, a startup dedicated to democratizing access to large language models (LLMs).

Episode Highlights

Guest Introduction

  • Naveen Rao, CEO and co-founder of MosaicML, discusses the company’s mission to improve prediction accuracy, reduce costs, and streamline the training and deployment of large AI models.

Key Issues in AI Adoption

  • Two main barriers to widespread adoption of AI models:
  • Coordination Difficulty: Training large models requires managing multiple GPUs, which can be complex.
  • High Costs: The expense of training models can be prohibitive.

Importance of Accessibility

  • Rao emphasizes the need for businesses to have control over model behavior, respect data privacy, and be able to iterate quickly to develop AI products.
  • The aim is to distribute capabilities widely to prevent the concentration of power that new technologies can create.

MosaicML's Offerings Products and Services

  • Inference API: A curated set of open-source models allowing for easy access and customization.
  • Managed GPU Service: Simplifies the process of training and fine-tuning models within the client's cloud environment, emphasizing data privacy and IP ownership.
  • Training Templates: Options for users to start from pre-trained models or build their own.

Target Audience

  • Primarily ML engineers and organizations looking to enhance their machine learning capabilities while maintaining control over their proprietary data and models.

Comparison with Competitors

  • Rao discusses how MosaicML stands out by:
  • Moving rapidly with community input and new model releases.
  • Offering economic feasibility, enabling training of models for significantly lower costs compared to competitors (e.g., training a GPT-3 model for under $400K).

Company Development and Challenges

  • Founding: Established in January 2021, navigating market entry challenges, and aligning pricing with customer value.
  • Growth Phase: Currently expanding their customer base and refining their service offerings, experiencing high demand for their solutions.

Recent Announcements

  • Open Source Models: Released the Mosaic Pre-Train Transformers (MPT) family, including a 7 billion parameter base model with variants for chatting and instruction following.
  • Notably demonstrated the model's capability to generate creative content, showcasing its advanced abilities.

Broader Implications of AI

  • Rao discusses the democratization of AI, stressing that AI should not remain in the hands of a few. Ensuring diverse perspectives in model training is crucial to mitigate bias and enhance the richness of AI applications.
  • The conversation touches on the evolving nature of AI, emphasizing continued innovation, efficiency, and the potential for meaningful impact across various sectors.

Future Directions

  • Rao highlights ongoing efforts to enhance training processes by optimizing data selection and improving infrastructure for better model performance.
  • The focus remains on making AI technology more accessible, cheaper, and faster for businesses across different industries.

Conclusion Naveen Rao's insights underscore the importance of making AI models accessible and controllable for diverse organizations, while also addressing the inherent challenges and opportunities present in the evolving AI landscape. MosaicML aims to empower users with innovative tools that facilitate the development of tailored AI solutions.

Further Reading and Resources

  • MosaicML Website: [mosaicml.com](https://mosaicml.com)
  • NVIDIA AI Podcast: [NVIDIA AI Podcast](https://ai-podcast.nvidia.com/)

This episode represents a significant conversation about the future of AI, emphasizing the importance of accessibility and innovation in this rapidly changing field.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:10Hello, and welcome to the NVIDIA AI podcast. I'm your host, Noah Kravitz. My guest today is Naveen Rao, the CEO and co-founder of Mosaic ML. Mosaic ML is part of NVIDIA's Inception program, which helps startups grow by giving them access to cutting-edge tech and NVIDIA's expertise. Mosaic ML is on a mission to help the AI community improve prediction accuracy, lower costs, and save time. They're doing it by providing tools that make it easy to train and deploy big AI models using your own data in a secure environment. Naveen is here to tell us about why he founded Mosaic ML and why the company's work is vital to the future of machine learning.

0:50And I think we've also caught him on the heels of some product announcements, so we'll have some new stuff to talk about as well. Let's get right into it. Naveen, welcome, and thank you so much for joining the NVIDIA AI podcast. Noah, thanks for having me on. I'm really excited to be here. Great. So let's get into it. Maybe we can start by you telling the audience a bit about Mosaic ML, why you started the company, what the company does. And then, as I mentioned, you've got some news, so we can get into that as well. Yeah. So before the news, I think one of the things we saw very early on, and when I say very early on, the space moves so fast that was only a little over two years ago.

1:28Right. was, you know, we believe that the capabilities of large language models, large models in general, we're going to explode. And we're really going to bring new capabilities, new experiences to humans. It's a new tool. However, if we want to be responsible about that, we found that it's not a great world where only a few have access to these methods. We don't think anything's bad about people investing a lot and building the best possible models. I think that's actually great for the research community. It's great for pushing things forward. But the reality is these new technologies have the possibility of concentrating power in ways like we couldn't do before.

2:06Every technology transition does that. And the best way to mitigate harm in those scenarios is really to distribute those capabilities. It's actually very similar to how software development kind of blew up in the 80s when the PC became cheap and everyone could learn to code. I learned to code when I was a kid during that time. And I think we're seeing something similar now. We started this company really to bring accessibility to the latest state-of-the-art methods to a greater number of organizations. What were the big barriers? The big barriers are really difficulty. When you want to wield 500 GPUs and make them train a big model, that's really hard.

2:45There's a lot of software. There are things that go wrong. And you need packages to make this work well. And that just didn't exist in the tooling. The other part of it is cost. Really, it's a very deep question of how do I make a model train to the same performance with the least number of computing flops and the lowest cost in time. And really, that's where we spent a lot of our energy at the beginning. And, you know, we've actually shown, we've become the sort of de facto standard for showing costs. You know, GPT-3 models for less than 400K, diffusion models for less than 50K. We published lots of detailed blogs on this.

3:22And I think this has been taken up by the community and I'm really happy that we're able to enable a great number of people with these technologies. So what are your products? What are the services that you offer to the community? Who's your customer? How would they go about using them? Yeah, great question. So our customers are ML engineers. You can think of our tools as supercharging their capabilities. And we can do it within the confines of their cloud tenancy. So everything stays totally private and we never see data. So we really believe in data privacy. Data privacy is actually, I think, intrinsic to creating the world we want within AI.

4:00IP ownership, data privacy. So we sell to the ML engineer. We have a few different products that we sell to them. One is an inference API. Really, it's similar to OpenAI or any of the other inference APIs out there. However, the models that are served behind it our curated set, curated by us, of open source models. Those models, as they improve, we add the newest ones and people get access to it. So we kind of move at the speed of community there. That's actually something we just launched this week. We launched it on Wednesday. All right. Congratulations. Thank you. And that same service is actually available in the private scenario I described within the cloud tenancy of a customer on their own custom models.

4:39So they can take open models, they can customize them for their applications, or even train their own models from scratch. and we can serve them. Then one step back from that is fine tuning and training. We actually have a managed service for GPUs, again, within any cloud. We can run on all the major clouds and some of the new smaller clouds that are coming up and manage the whole process soup to nuts. You start with your data. We actually have templates for models that they can start from either pre-trained or completely naive and customers have a choice and we make that process go as smoothly as possible for them.

5:13So for those who are maybe a little less familiar with the details about, you know, they've been following the AI space. And obviously over the past, you know, six months in particular, there's been an explosion in the mainstream and the idea of a large language model is a little more familiar to a lot of people. But when you're talking about a customer, a business customer, an ML engineer who wants to train their own model versus using, you know, OpenAI or one of the others that's out there for the public and for customers to use, why would they do that? And what goes into that? And, you know, you touched upon the costs of training a big model and keeping it running, maintaining it.

5:57Why would a business customer choose to train their own model as opposed to using one of the publicly available ones? Yeah, I mean, I think the real world is always a mix of different solutions. And there are scenarios where, you know, you want to just use something that's out there. I think if you want control over the model behavior, like really tight control, you want to have your own model. It's sort of like insourcing software development. That's one aspect. The other aspect is data ownership. If you train a generative model on a data set, it can memorize parts of that data set. That's just a feature, actually, of the models.

6:35That's a good thing. And what it means is that wherever those weights of the models go, the data has gone. So you kind of need to control where the weights go. And that only happens when you own the model. And so really respecting data privacy, if you believe that that's intrinsic part of this new AI world, then you do need to think about your critical data and where it goes within a model and how that model is served to the endpoint. And so we give our users control over that process. And also just quick iteration. If you want to incorporate new data, you have some sort of a data pipeline that's built for your application.

7:14Again, you need to control how the model behaves. Right. So what are some of the things that Mosaic ML offers, you know, in comparison to some of your competitors? Some of the cool perks, maybe some of the interesting use cases that Mosaic ML offers. Some of the perks, I mean, we, like I said, move at the pace of the community. As new things come out there, we attempt to make it as economically feasible as possible and easy to use. we work across different sorts of models like diffusion models and large language models so we actually train from scratch our own diffusion model like i said it costs about fifty thousand dollars in compute costs which is pretty accessible i think i mean still no but it's it's uh it's a five figure number not a six seven or eight figure number so exactly that was a big thing for us try to get under 100k yeah and uh i think what that enables is people to try to customize these models for their application.

8:08And I think that that's really what we offer is sort of that we give you clay and you can choose to make a teacup. You can choose to make, you know, whatever little widget you want. And I think that's what's special about us. We're not end application company. We're a platform company. So we provide tools to enable a whole host of applications within different spaces. How long has the company been around when it was founded? It was founded in January of 2021. So just over two years. Just over two years. So during that two years, what were some of the challenges or even just interesting developmental milestones, technical problems, things that you had to figure out solutions for?

8:49Oh boy. So many. I think. How long do we have, right? Yeah. I know, right. You know, I think a big challenge before was how do we go to market? You know, how do we sell this stuff? How do we capture value that we're creating with our pricing and align our pricing with our customers, because I think the best business models are the ones where people pay more because it aligns to their value, right? And so, and it's our value prop to them. This is not simple. I think we experimented with a bunch of things that were really complicated, as every startup does, and we've settled on a few simpler ones.

9:24What's really interesting to me is that when you really, when you go through this process and you iterate a few times, you actually come to the realization that just about every AI application is kind of reselling compute in some packaged for it. An inference API is reselling a GPU time. Training is reselling GPU time. And it's really, how do you add value to that process and price align to it? And really that was the biggest learning. I think how we make things work in different clouds and actually the variability between different clouds, they're using, these are all NVIDIA A100 GPUs. Same basic GPU computing element, but how that's deployed and the software wears around, it's actually quite different for cloud.

10:02So we had to engineer solutions to make that each cloud performance, even though they're still running that core or GPU. How big is the team now? 55 people, something like that. And are you in growth mode? Are you kind of in, you know, go to market and handle customers and kind of see how that's happening? What's the stage of the company? Yeah, we're very much growing our customer base. I mean, we went to market a few months ago, so around the beginning of the year, end of the last year. And yeah, it's kind of exploded, which you might have imagined. Sure. Yeah. Wonderful. I mean, honestly, the biggest, the biggest barrier to growth is us in terms of closing deals and servicing those deals.

10:41So that's a good problem to have. And of course, we're trying to solve that. Our stack is highly automated, which has made things really nice. Like, you know, the dark side of large scale stuff is that, you know, GPUs fail sometimes. nodes fail, networks fail, and dealing with those failures gracefully and keeping the process going and hiding those details from the user who frankly doesn't care. That's actually a lot of the work that's gone in recently to make this whole process stable. And actually this morning, we released our open source models, our state-of-the-art open source models, better than any other open source model out there, very long context.

11:19We believe it's a great starting point for customers to build off of. And that's a culmination of all the work we've done in terms of engineering our stack to be reliable, the research work we've done and making long context work and stable training, you know, recipes work, all of this stuff. And I think that's why we're very proud of the stuff we released this morning. That's great. What models did you release? We call it the Mosaic Pre-Train Transformers, so MPT family. And we actually just released a 7 billion parameter base model with a few different variants. One variant is tuned for chatting.

11:54So it's really easy and nice to chat with it. We did put an interface up on Hugging Face that people can chat with. It's a little overwhelmed right now. We do have an inference service that will be, we already launched it this week, but we're going to be moving our models to for higher scale very soon. So people can get a better experience there. We also have a instruction following, instruction fine-tuned version. One of the more exciting versions, I think, is what we call the storyteller. That one is a long context. What context is, is essentially the prompts, right? We actually fed it the entire contents of the great Gatsby, 67 ,000 tokens.

12:33And we asked it to write the epilogue and it did. I posted it on Twitter. You posted it, okay. So I was gonna ask what happens, but people, what's your Twitter handle? Naveen G. Rao. Okay. So if you've been wondering all these years, what happens later? Naveen G. Rao on Twitter, we can get an answer for you. You happy with the answer? I'll just ask that. Yeah, it's actually kind of creative, right? Sure. We didn't know it spit out, but it actually produced some interesting stuff. In fact, I think we put some of it in our blog as well when we released the model. So you can check it out there too.

13:07But it's just a fun use case, right? I mean, I think this, the way we see it is, again, we're bringing, we're building clay, right? We're bringing that clay to the market and what people do with it. I'm excited to see, you know, new capabilities. It's actually kind of similar to NVIDIA's roots. It's like you bring computing that can do new stuff. What does it do? Well, people unleash all kinds of creativity like AI itself. Absolutely, yeah. So that makes me think something you said kind of at the beginning of our chat here. And it's kind of been a theme of, I mean, it's been a theme of the world, a theme of the advancement of computing and AI certainly in recent years.

13:40and we've talked about it on this podcast over the years, the democratization of all of these tools and going from GPUs, great for AI, people have GPUs in their machines at home for gaming, for other purposes. Oh, I could use those to start messing around with AI and you had these people doing all these things in the past years. And now obviously, the explosion of these LLMs into the wild, people having access, doing all kinds of stuff with them. You said something at the beginning about it being important that power held in these models not get too concentrated and that, you know, we continue to uphold this democratization of access to AI as being important.

14:26Can you talk a little bit more about that, kind of what that means to you and then what it means to, you know, how Mosaic ML goes about its daily business? Yeah, so AI, the reason it's so, I don't know, concentrating in terms of its power is that it literally can become a lens on which we view all data. You know, people have talked a lot about bias in AI and models. The reality is every model is biased. Every human is biased. And that's actually not a bad thing. It's how we make sense of the world. And that's how AI makes sense of the world. Those biases, you know, I subscribe to a viewpoint that there's no absolute moral truth.

15:05They are relative to what we feel as society, right? Right. And what we found as humans is the best way to do that is through some kind of democratic or distributed process, right? We vote for laws. We vote for people to represent us in our government. And that represents a majority. Yes, there's always people that are unhappy about it. But I think AI is actually very similar in that you need many people putting their biases into models that provide their lens, their perspective. And, you know, the market will decide what perspectives are useful, what are less useful, whatever, right? But I think I don't want to create a world where there are no models that don't disagree with me.

15:45I have a view on things. I have my own biases, of course, like we all do. But we need to enable everybody, even the people that disagree with us, to create those lenses on data. I think that's very important and core to what we do. And second part of your question, how do we go about doing this on a day-to-day? I mean, really bringing the costs down, making things in the 100 ,000 range, that's a lot more accessible to enterprises. I want to enable many enterprises with their own LLMs that are built for their purpose. You know, ChatGPT, all the API vendors out there, what was a magical thing about it is that it just lit up everyone's imagination.

16:23And my high school-aged kids told me about it. And I was like, you know, that's what I do. And they're like, oh, that's pretty cool. So finally they think that's cool. Right, it's a good moment. It's a great moment. But so it drove this awareness. But now I think when you want to build an AI to help you with healthcare, it's very different than entertaining someone via chat. If you want something that will advise you on your retirement plan, also very different. You know, if we want models that can model proteomics or genomics completely, not completely different architecture of model, but different training.

16:57So we need to move these capabilities into a number of different places. And that's what we do every day is we actually talk to organizations. We try to understand what their problems are and see how our tools can help solve them. I'm speaking with Naveen Rao. Naveen is the CEO and co-founder of Mosaic ML. Mosaic ML is in NVIDIA's inception program, which works with startups, giving them access to NVIDIA tech and expertise and advisement. And as we've been talking about, Mosaic ML just released, as we're recording this in early May, just released. It was a 7 billion parameter open source model as part of a family of open source models.

17:35We've been talking about the different use cases for Mosaic ML's tech. I want to shift gears for a second, Naveen, and ask about you. You mentioned learning to code as a kid in the 80s. And I know you were at Intel for several years along your path here. How did you get into coding and tech and working in tech and eventually into machine learning? Oh, so you're going way back. We don't have to go back, you know, further than you want to. That's okay. I'll date myself a little bit. So my older brother and my dad are both big geeks. My dad's always been a tinkerer. I think we just sort of got, I guess, blessed or cursed or whatever you want to call it with that same desire.

18:17and uh my brother was really into programming in the 70s even and we actually had a we had a personal computer when i was three years old uh in 1978 the texas instruments 99 foray yeah ti-99 yeah so i actually learned to code in in elementary school i learned to code logo you ever right yeah i remember logo with the turtle yeah yeah you could drive the turtle around you can make it draw stuff it was pretty cool and uh then quickly learned basic after that And really, it was almost like a game at that point. It was something fun. I could automate stuff and people thought it was magic. Yeah. Back then, like, oh, I can print something out over and over again.

18:54It's a little program like that. I actually started to build games and stuff when I was a kid. Yeah. And, you know, fast forward to college and, you know, I was a computer science and electrical engineering major. Loved building stuff with computers. Back then, it was the very beginnings of the internet. And really, one of the things I was fascinated by, kind of started from sci-fi, was artificial intelligence. the concept of a machine that is synthetically intelligent. And I even did research in the 90s on neuromorphic machines, Carver Mead's work, and really fascinated with this premise. And the thing that really grabbed my attention was the fact that our brains run on 20 watts of energy.

19:32The latest NVIDIA GPU runs on 700 watts for card or for chip. So I think it's still something that's a big gap, but it was just mind-blowing to me that like everything we are is 20 watts of energy right so um i you know i came out to silicon valley was in the startup scene all that kind of stuff learned how to build chips uh learned how to write software really well as a professional did it for 10 years and i was like okay maybe it's time to go back to that interest and i i did this kind of my family thought it was crazy because i had a very nice career and uh i quit my job and went back to school to get a phd in computational neuroscience science.

20:11Okay. And that was really driven by this desire to, I want to look back on my life at the end and think I did something. Yeah, absolutely. I nodded and said, okay. And then I thought, wait a second, computational neuroscience. Let's unpack that a little bit. What does that mean? It means understanding the computational underpinnings of how the brain processes information. So it's related to biophysics, it's related to physiology, behavior. It's kind of the middle of all of these things. And so it's really like, how does information come to the brain? How's it represented? Then how's it processed to, to affect output, which is movement.

20:46Yeah. So I worked in a motor control lab where we did neural prosthetics, literally decoding signals from neurons that were listened to in real time. Actually, my lab was the first one to do, uh, human neural prosthetics. People that were paraplegic, uh, or quadriplegic actually had an implant and they think we could decode their, their thoughts around movement. Right. Yeah. robotic arm. Pretty cool. How long ago was this? I started there in 2007. Okay. Have you kept up with the field? I mean, to some degree, you know, it's like there's only so many hours in a day. Oh, sure. Yeah. I asked because it was a few years ago now, but we did an episode with a startup working on similar things.

21:23They sort of described it as a USB port for the body. The founders had a buddy who needed a prosthetic leg, I believe it was. And they kind of saw what he got and thought, we can build something better than this, you know, and sort of went on with it. So that field and the idea of bridging kind of the body with AI, to use the term very broadly, you know, has always been fascinating to me. Oh, absolutely. I mean, really the premise and everyone in our lab kind of felt this way is that we can really crack how information is represented in the brain. We can actually start to make communication faster and more rich between humans.

21:59And what's interesting is that when you actually start to analyze, this is a total tangent, but we start to analyze how information comes out of a person, right? We speak, we move, we move our faces, it's all motor output. It's actually a lot of information. It's really hard to do better than it. Right, right. Just pure bandwidth wise, right? Evolution has created something quite amazing with humans. But, you know, after I finished that, I actually was a research scientist at Qualcomm researching back to neuromorphic architectures, how we can use some brain inspiration of computation and actually build better computers.

22:36And that's when deep learning kind of started to take off. I knew it from the academic field. I saw Jan Wakun speak years ago at my school and everyone thought he was crazy back then. But bless him and the others for keeping at it because really they knew that there was more here. And, you know, GPUs were getting dense enough and parallel computing was getting good enough that that's when the moment happened. And 2012 was kind of that moment. And we recognized very early that we need a new computing architecture. GPUs actually weren't it at that time. And my previous company was called Nirvana Systems, which was acquired by Intel.

23:16That's how I ended up at Intel. First AI chip company in this new spat of AI chip companies. And we were competing with NVIDIA. And really, I think I like to hope that we pushed NVIDIA towards architectures that we see today, which enable all the cool stuff we're seeing like large language models. So to put you on the spot, since you said evolution, we've been talking about the idea of the brain inspiring these computational models. Where do you think we're at? It feels to me as sort of an observer, but I've been, you know, observing, having the good fortune to talk to people like you for more than half a decade now on this show and working in the industry longer than that.

24:01But it really feels like we're in this just very intense, pivotal moment. Or I don't know if pivotal is the right word, but the pace is accelerating. And it's hard to kind of know how much of it is because of the hype around Gen AI specifically and the chat models making this technology accessible to so many more people in a way that just brings it into the consciousness. Where do you see the technology, the impact of the technology on the kind of greater world? What's the moment that we're in? Is it sort of the more of kind of the tail end of the past several years of the increases in compute and the democratization of tools?

24:42Is it kind of more the beginning of this, you know, entirely new era in, I mean, information and computing? Where do you kind of, if you sit and take stock of where we're at right now, what do you see? I think we're at the very beginning. You know, it's like humanity just got wings. There was something that has changed, right? And it's funny because like, I know we're living it and we were like, oh my God, you know, things are moving fast and like 10 years, it's taken 10 years to get to this point. And so it feels slow, but it's like, the reality is 10 years ago to now is actually just a blink of an eye.

25:16And there's been a constant building of technologies upon each other that work, right? And I think that's the critical piece that people seem to miss when they're like, oh, it's all AI hype and it died. And now it's sort of coming back. It's BS or whatever. The reality is we hit upon something where you can build large-scale optimization systems. And now that's gotten better. And we've learned how to make them even better, more efficient. Now, like my company has really been focused on efficiency of computing flops. Really, how do we get to that 20 watts, right? That everybody can realize. And I think that's a largely a, it is a physical problem in terms of a physical substrate, silicon or whatever.

Read the full transcript

25:58So there's a lot to be refined there. But there's also a lot to be refined on how these learning systems work and how they scale. Human brain uses each computing flop it has extremely efficiently. That's an algorithmic problem. So I think we got to look across the whole stack still. And we're just now figuring it out. I mean, a large language model takes millions of watts to train. It's nowhere close to capabilities of a human, even though now it starts getting, I don't know, sometimes indistinguishable. So I think we're very much at the beginning. So I wanted to ask you, you mentioned before we hit record here that you recently published a blog about H100 performance.

26:35Yeah, so the GPU has been something that's been intrinsic to the new capabilities we have. And we were actually very excited to see costs come down and time come down. And we were very keen to see the new stuff coming out from NVIDIA. and we benchmarked it. And we actually found both of those things to be true, about 3X faster. And that means everything that takes three weeks will take one week. Great. And the other important thing is you can do it cheaper. So we found 30 to 50 % cheaper performance per dollar. So even performance normalized, and to cost normalized rather, we saw a pretty big gain here.

27:17And that's fantastic. Cost savings. So again, back to the whole idea of accessibility, this is it. This is Moore's Law is part of it. New architectural features are part of it. Software stack is part of it. So all of it coming together is really driving the price down to a point where it's going to be just much more accessible by the community. That's great. And so for Mosaic ML in particular, you mentioned in a go-to-market phase, customer acquisition, growing what you're doing, and you just dropped a bunch of exciting announcements. So asking what's next feels a little, I don't mean to push you too far ahead, but, you know, as you look ahead to the rest of this year, whatever the timeframe is, what's the roadmap look like for your company?

27:56Yeah, I think enabling people to get to the best possible model and the shortest amount of time and least money is kind of what we, what our mission is, right? And so we look at different parts of that stack. We are optimizing inference. That's what we put out there and making that cheaper, you know, enabling more applications. We're optimizing the training process, but what we found is a big part of optimizing a training process is actually how you pick what data goes into your models. This is something that we are looking forward on how to do really well because ultimately it's gonna start with a really like raw data.

28:33Data is observation. How do we morph that into something that's suitable for learning and then feed it to a stack that's extremely efficient and then serve it. Actually, if you saw the announcement, again, things moving very fast. Last week from our customer, Replit. Replit is a distributed IDE company. They built a state-of-the-art code completion model on our platform. They started with their own data, along with some open data. Two people on their side using a Databricks pipeline and then our tools built a state-of-the-art model in less than a week. That's phenomenal. It's just amazing. Awesome, right?

29:11That is the story for us. We love it because it's like, that's what we want to do. Everybody can do this, right? Wild. Naveen, for folks who would like to learn more, obviously the company has a website. Is there a research landing page or their social media accounts? We mentioned the epilogue of The Great Gatsby. But for folks who want to dig in more to the technical side of what you're doing, perhaps looking to working with Mosaic ML, where should they go online to find out more? Yeah, mosaicml.com is our main website. Our blogs actually have a lot of detail in what we do. We are transparent in how we work and all of the details are there.

29:53So anything to do with our model release, our inference releases, any other technology, our streaming data loader service, all of these things are all detailed in our blogs. And you can also take a look at some of the solution pages, but for a technical audience, the blogs are really where you want to go. Perfect. Well, Naveen, it's been a pleasure. Thank you for coming on and talking about what you're doing and kind of looking ahead to the big picture, even on the very close heels of some big announcements for your company as well. Appreciate you being game to talk about the future in that way.

30:26And we wish you all the best of luck with Mosaic Amal and everything else you're doing. Wonderful. Thanks for having me on. Really appreciate it.

30:38Thank you.

31:10Thank you.

From the publisher

Startup MosaicML is on a mission to help the AI community enhance prediction accuracy, decrease costs, and save time by providing tools for easy training and deployment of large AI models.

In this episode of NVIDIA's AI Podcast, host Noah Kravitz speaks with MosaicML CEO and co-founder Naveen Rao, about how the company aims to democratize access to large language models.

MosaicML, a member of NVIDIA's Inception program, has identified two key barriers to widespread adoption: the difficulty of coordinating a large number of GPUs to train a model and the costs associated with this process.

Making training of models accessible is key for many companies who need to control over model behavior, respect data privacy, and iterate fast to develop new products based on AI.

More from NVIDIA AI Podcast

All 115 episodes
MosaicML's Naveen Rao on Making Custom LLMs More Accessible - Ep. 199NVIDIA AI Podcast · 31 min
Listen in VO