In short
This Week in Startups: Episode E1754 Summary
Podcast Title: This Week in Startups Episode Title: How open-source & distributed models can win AI with MosaicML’s Naveen Rao Host: Jason Calacanis Guest: Naveen Rao, Co-Founder and CEO of MosaicML Release Date: Not specified
Episode Overview
In this episode, Jason Calacanis interviews Naveen Rao, the co-founder and CEO of MosaicML. They discuss the dynamics of open-source vs. closed AI models, the societal implications of AI advancements, and the pitfalls of centralized regulation in AI development.
Key Themes and Discussions
- The Role of AI in the Future
- AI is seen as the next inflection point in technology, akin to the advent of language.
- Questions are raised about ensuring everyone has a place in an AI-driven world and that increased efficiency meets growing demand.
- MosaicML's Mission
- MosaicML focuses on democratizing access to large-scale machine learning and generative AI capabilities.
- The aim is to allow organizations to train their models using their data while maintaining control over their intellectual property (IP).
- Open Source vs. Closed Models
- The discussion highlights a debate between open-source AI models and proprietary models from companies like OpenAI.
- Rao emphasizes the benefits of a market solution over regulatory solutions, suggesting that a diverse range of contributors can lead to better outcomes.
- Data and Ownership
- Rao expresses concern for the lack of compensation for creators whose data has been used for training AI models.
- He uses examples from companies like Disney and Reddit to illustrate the importance of controlling proprietary data and the implications of not doing so.
- Training Models with Data
- Rao explains how organizations can use MosaicML to train models on their datasets, discussing the processes of prompt engineering and fine-tuning.
- He provides insight into the costs associated with training AI models, noting that effective use of data can lead to significant competitive advantages.
Episode Highlights
- AI's Impact on Employment and Education (40:42 - 48:49)
- Rao discusses the rapid changes AI will bring to job markets, emphasizing the need for society to adapt.
- Education systems must evolve to incorporate AI tools rather than relying on memorization.
- Regulatory Concerns and Open Source (54:37)
- The conversation touches on the challenges of centralized regulation in AI.
- Rao expresses skepticism about regulatory bodies determining who can train models and the implications this may have on innovation.
Key Takeaways
- MosaicML's Advantage:
- By enabling organizations to use their data for training, MosaicML seeks to ensure that companies retain their competitive edge without becoming reliant on larger, centralized AI providers.
- Open Source is Crucial:
- The discussion underscores the importance of open-source models as a means to democratize AI, allowing a diverse range of users to contribute to and benefit from AI advancements.
- Future Implications:
- The rapid pace of AI development presents both opportunities and challenges for the workforce, necessitating proactive adaptation in employment practices and educational frameworks.
Episode Sponsors
- Vanta: Offers compliance solutions for startups.
- Trovata: Provides cash management services.
- Microsoft for Startups Founders Hub: Offers cloud credits and resources for startups.
Follow-Up For more information about Naveen Rao and MosaicML, visit their [website](https://www.mosaicml.com) or follow Naveen on [Twitter](https://twitter.com/NaveenGRao).
---
This episode presents a crucial examination of the evolving landscape of AI and the importance of open-source solutions in ensuring equitable access to technological advancements.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00AI, in my view, is the next evolution of what humans can do. You know, language was a big technology that humans used to pass knowledge that that exploded what humans could do and the influence we could have on the world. AI is going to be that next inflection point. But how are we going to make sure that everyone has a place in that world? How are we going to make sure that the demand that's created by the increase in efficiency is commensurate or does it collapse? This Week in Startups is brought to you by Vanta. Compliance and security shouldn't be a deal breaker for startups to win new business.
0:34Vanta makes it easy for companies to get a SOC 2 report fast. Twist listeners can get$1 ,000 off for a limited time at vanta.com slash twist. Trovata. Starting up is hard. Trovata makes managing cash easy. Start automating your cash management at trovata.io slash twist. Use code TWIST for 30 % off one full year of premium features like AI forecasting. And the Microsoft for Startups Founders Hub helps all founders build a better startup at a lower cost from day one. Startups get up to$150 ,000 in Azure credits, access to free OpenAI credits, free dev tools like GitHub, technical advisory, access to mentors and experts, and so much more.
1:23There is no funding requirement and it only takes minutes to join. Sign up today at aka.ms slash This Week in Startups. All right, everybody, we are really focused on AI and this crazy revolution that started really with GPT-3 and 4 having a moment for open AI and people starting to realize, hmm, this stuff is going to impact everything. Since that time, Dolly and Stable Diffusion showed what's possible, and now people are incorporating generative AI, the ability to generate some type of content intelligently from a prompt into every single product. Whether it's Notion or Microsoft Office or Gmail, we're going to have AI companions, co-pilots in every piece of software.
2:18but you're left with a big question as an entrepreneur and enterprise do you do this on your own and do you own the ip and do you control your destiny or do you partner with existing platforms that are out there with me today is naveen rao of mosaic ml he's the co-founder and ceo correct naveen that's correct got it so you heard me sort of teeing this up uh you've been in machine learning for quite a time you had a company before your current company that you sold to uh intel i believe and so that's correct that was nirvana maybe you could explain what you're doing today uh with mosaic and why yeah so what we're doing today is really bringing these capabilities of large-scale machine learning which is generative ai in my in my mind to many organizations um i i think one of the things we've done even with my previous company was trying to really bring these capabilities to more people to create the world we want.
3:20I see success as people that disagree with me being able to build models equally as good as me. I think that's how we're going to make this world work. And it's become front and center now with debates around regulation of AI and putting some sort of government licenses and this and that. But I think really this is solved more in a market as a market solution where many people can build the stuff. Many people can imbue these models with the biases that they see fit. And, you know, we'll let the market decide where where things should be, not some sort of centralized regulatory agency. So your company allows my organization to take our data, put it into a language model.
4:06This is called Mosaic ML. Yeah. This is the software that will let me train my own model and then host it easily on AWS or whatever cloud I choose, I suppose. That's correct. Yeah. In fact, we even collapse the experience across different clouds. I mean, you can start training on 512 GPUs and AWS and then move it to 1000 GPUs and Azure. We actually make it very easy to move things around and be modular. And really that enables people to use their resources more effectively, but also not have a lock in from the provider and really just kind of own their IP. I mean, I think it all stems from the fact that respecting data privacy, I think, is important.
4:51Data is, you know, arguably an expression of your company, of your IP and building solutions that respect those bounds and enables people to build on top of that data and own that thing that is built. i.e. the model is very important and i think that's what we we aim to do at mosaic so this uh harkens back to the i don't know the thorn in my paw that i have been screaming about since the beginning of this which is hey what did you train these things on and are those people being compensated we now have a handful of lawsuits and letters that have been either filed or sent twitter to um microsoft about the training use of their data reddit coming out and saying hey this is our data if you want to use it we're going to need a fee and of course people trained on it without uh their those two sources permission and then of course you have getty images uh versus stable diffusion you have open source uh those the community uh or open source contributors versus co-pilot my github's uh composer essentially or um you know co-pilot for developers so if i was disney and i own marvel every comic book ever written by marvel every film every piece of dialogue written every treatment ever written things that made it onto uh the screen things that didn't make it onto the screen things that were spec scripts that were written that were never done you can be sure they've got 10 scripts for every one they actually produced disney could take that marvel or star wars corpus put it into mosaic compose a library never have to worry about having put that train a day training data into a public entity that would then go use it for future or maybe even claim some ip ownership of it and then they could let their writers on the marvel series ask questions hey tell me about this character dazzler was she ever part of the x-men what is she known for what does her dialogue sound like can i get some backstory whatever something the writers are fighting against but putting that issue aside this is a pretty compelling case for a company like disney to start this work now and to keep open ai and google's hands off this data, correct?
7:13That's right. Yeah. I mean, I think that's, that's one sort of really flagrant example, like kind of, uh, incentivizing content creators to keep creating content. Uh, there's even more, I'll call it mundane things where, Hey, I'm a, I'm a company that has, you know, huge data sets on the behavior of my customers. And I want to build something that gives me a competitive advantage in my space. Right. And I don't want to share that with my competitors. I want to build a model and express that competitive advantage directly, but I don't know how to build models, right? I'm not an expert at doing that.
7:48I can't hire a team to do it. They can use our tools to go and do that and leverage their data for a competitive advantage. But yeah, I think it all comes down to this sort of similar way of thinking is that we have data that gives us a kind of an economic incentive to keep gathering data, to keep creating new content. and then we have some way of expressing an advantage in the market and uh you need to be that could be proprietary data about consumer behavior or it could be you know i picked the the most iconic ip of all time star wars and marvel or at least in recent history that have you know generated massive massive profits for those companies if you're a sass or services company that stores customer data in the cloud, then you need to be SOC 2 compliant.
8:37You knew that from a third party and you need that third party to close big deals. And if you want to get compliant easier and faster, you need to use Vanta, V-A-N-T-A. Vanta makes it so easy for you to get and renew your SOC 2. On average, Vanta customers are SOC 2 compliant in just two to four weeks. Prepare that to three to five months without Vanta. And Vanta can save you hundreds of hours of manual work and up to 85 % of compliance costs. This is a total no-brainer. And Vanta does more than just SOC 2 compliance. They also automate up to 90 % compliance for GDPR, HIPAA, and more. You can't afford to lose out on major customers.
9:15We all know that. Listen, it's a hard year. Last year was hard. You can't lose those major customers because you don't have your compliance dialed in. Just work with Vanta. Get your compliance automated and tight, and tight is right. Lock down those big deals. Here's the best part. Vanta is going to give you$1 ,000 off. That's 10 hundies get$1 ,000 off at Vanta.com slash twist. That's Vanta.com slash twist for$1 ,000 off your sock too. Walk us through for non-technical people who are listening, you know, maybe founders of companies or capital allocators, exactly how I would start this process of taking my data.
9:50And I'm going to use myself as an example. It'd be easy for people who listen to this show to understand. We have thousands of meetings and notes from thousands of meetings. we have thousands of applications of startups asking us for funding we now have those in notion we have zooms calls with transcripts and summaries that we've done with ai and there are external data sources like crunchbase or linkedin that have signals that a startup has done well i.e downstream funding it's not perfect data sets but there's pitchbook there's crunchbase there's how many employees they have listed on linkedin not perfect but directionally correct so i want to see if i can moneyball startup investing you know i don't know if it's necessary for us since we have incubators and accelerators that let us place a lot of bets but i do want to do this uh at some point just for the giggles and to see what comes out of it yeah how would i start that process i have a non-technical team let's say of investors and a you know technical and that we can talk about technology but not developers how do i start this process of taking a corpus of data let's call it a thousand meeting notes or five that five thousand meeting notes uh and applications of startups how would i what would i do walk me through it yeah so i could first preface this with uh our tools are meant for technical audiences like ml engineers data engineers uh but there there are sort of conceptually three major ways to modify the behavior of a language model right so we'll start with language because we're talking about text.
11:25So when you're in a smaller data regime, when I say smaller, I mean that which was fits in a book, like less than a hundred thousand words say, uh, then we can use things like prompts, prompt injection, um, to, to actually change the behavior. So actually the model we released, uh, less than a month ago called MPT seven B, uh, has essentially an infinite prompt window. So we, we tuned it to have 64 K tokens. Uh, the token is about three quarters of a word. of a prompt. Explain what tokens are for folks and prompts just so we really can explain what's happening here with machine learning in this process.
12:05Yeah, token is really the important part of a word. So if I said evergreen, we think of that as one word, but it's ever and green are tokens. The vast majority of words, the token and word are the same thing. So that, boy, girl, those are all single token words. So that's why we kind of give this ratio of about 0.75 of, you know, tokens to words. Yeah, words to tokens rather, sorry. So when we, the way these models work is that we tokenize the language into these, into these blocks that are meaningful, a word or some piece of the word. And then that becomes the input to the model, which we call a prompt.
12:43That prompt sets what we call a context. context it's like uh i'm telling you hey jason we're going to talk about uh evergreens so now if i say something like naked seed that makes sense to you right so because i set the context really this is what a prompt does is it sets a context for a model and then it can sort of recall knowledge that has been trained upon from that context this is why prompt length is actually quite important so uh what we enabled was a very long context window that you could actually feed a whole book we in fact fed it the entire great gatsby and we asked it to ask it to write the uh the epilogue uh made up epilogue and it did so and it can do that across the entire context of the book it's imagine imagine like reading the whole book keeping it all in your mind and then writing writing it out that's what the model is doing so i could take smaller i could take the i could take take a successful company like amazon or netflix have some research on that company like a research report that was written plug that in and say of these thousand meetings i've done do you see any companies that would correlate with this company in some way and i could use the prompt of a 50 000 word gartner report or goldman sachs report on amazon or something from 1999 or 2000 i could take bill gurley's reporting on you know amazon from the 90s or 2000s mary meekers and start using that as prompt engineering for looking for patterns in startups huh something like that correct that's right uh prompting and context windows are i'll call them the weakest form of learning in a sense where uh you can take information put it in the prompt and have the model you know do some analysis on that the problem is sometimes there are weird conditions where let's say there's conflicting evidence from where the model was trained versus what was inserted in the prompt, you might get, you know, kind of undesirable behavior in those cases.
14:41So that backs us up to one more version of how we modify the outputs is what we call fine tuning. Fine tuning allows us to kind of condition the model to act in certain ways. Like if I ask a model, you know, racist questions, maybe we want to say, Hey, I don't, I don't, I don't want to talk about that subject. So we can condition it to do that. It's very similar to a human. Like if I put a human in a call center, I know they're talking about customers. I'm like, hey, don't talk about our competitors' customers, right? Or don't talk about our competitors. Just talk about our products. Don't talk about politics, right?
15:15Don't use swear language. These are things that I might tell a human, right? And so we can actually use fine tuning to condition the model to give us outputs that are like that. You can even imbue new knowledge through fine tuning as well, but I would argue that doesn't work quite as well. The real way I think to describe or modify the behavior of a model in a very profound way is using pre-training and data mix. So pre-training is where we take a model that doesn't know anything and we train it on a bunch of data. And the way we do this is actually an optimal mix of maybe some domain-specific data along with some general data.
15:51It's actually kind of similar to education, right? I have a child, I'm going to put them in school. I'm going to teach them about history and politics and math and science. And at some point later in life, you start to specialize, right? It's actually a similar process with, uh, with LLMs. And so in this example where I'm dumping in my startup data, um, what would be then the next steps, uh, for me to get value from it? What would I do once I've got the model set up? Yeah. So I think the first thing is to analyze how much data you're throwing at it. So if you're under that 100 ,000 word limit, then you're probably in the regime of tokens, sorry, prompts into tokens.
16:31If you're in the call it 100 million, so range, we can start talking about fine tuning. If we're now in the billion range, we can talk about pre-training and layering in this data. So that's really the analysis that we kind of walk our customers through. Typically, it's like which method you want to use depends on how much data you have to throw at it. Typically. So in your example, you said 5 ,000 transcripts, something like that. That's probably in the prompt regime. We're probably not doing anything beyond that. We may be able to do some light fine-tuning to actually condition the model to act in certain ways.
17:08Like, hey, I want this kind of information pulled out. Like, I want to know something about the quality of the founders. I want you to focus on that as an output. Right? Got it. I can condition the model to do that with fine-tuning. Got it. so i could say hey what's the problem you know because typically if you backed out of a deck the deck structure was architected to convince investors investors were optimizing for big problems solving big problems with high margin businesses with people who had great backgrounds who could execute so you could actually like the deck having a competitive landscape or the total addressable market or the problem and the solution those are the things that investors would go to first we know this because when you send a doc you send or some of these tracking software is a little bit creepy but it will tell you how long people spent on each page which pages they zipped right over like advisors who cares you know uh you know right there's a lot of stuff is uh you know thrown into decks just for performative reasons but the problem and solution and the background of the founders are paramount the number of customers and the pricing paramount the business model so you fine-tuning would be essentially that process of trying to tell the model this is important that's not as important um that's right and and you can even link it to outputs right we call this process reinforcement learning with human feedback rlhf and uh actually what you do then is you say well the inputs are all this all this deck material say and um then And these companies did really well and those companies didn't.
18:44Right. You could actually start to link it to an output and you can start saying, hey, show me companies that you think are going to do well. Right. And it can actually kind of pick up some patterns. And it would be different for a C stage investor. It would be they got to a billion dollar valuation. For a late stage investor, it might be this company went public or got bought for over a billion dollars. And so you could actually have two different outcomes could be defined as success. Totally. you know for y combinator or r accelerator at launch like you know success might be the company gets past 200 million dollars because we're investing in low single digit valuations when companies are just starting out and they're just ideas so you have a totally different uh approach there um so you're in competition with some of these open source projects is your solution open source and you know maybe you could speak to who's going to win ultimately uh having the great language models is it going to be the person with the greatest data the person with the largest open source community fine-tuning the open source projects to analyze that data who in your estimation is going to win the day or will it be parity where you know having a web a cdn a content delivery network sure there's on the margin some that are faster than others and you can probably have debates with the sys admin all day long but the fact is in 2023 you throw up any any of the top 10 cdns your site's going to work really well there's parity right right yeah yeah i mean uh so our models are open source we open sourced our 7b model a little over a month ago and or a little want to give people great starting points to get going.
20:32I think for us, what we're learning is our customers are on a journey here. This is all very new, right? It's new for every company right now, and they all want to do it. It just comes, people come at it from different starting points. Some people are like, okay, I'm a hundred percent in, I'm going to budget$10 million. I'm going to go do this. Okay, great. We can help you with that pre-training side of things. Some are like, well, we're dabbling. We want to understand how we can add value to our customers. Can we start with a smaller byte. So we want to meet them where they are. And open source models are a great way to do that.
21:01They can start with the open source model. They can fine tune it. It's relatively cheap. And then eventually start integrating to the application and then customizing even further. So we want to, we have the whole breadth here and open source is a very big part of that. I think who's going to win out of this is the one who can serve their customers the best. I don't think those principles are going out the window. Everything that you've talked about over many years, those things are still real, right? I mean, at the end of the day, you've got to give your customer something they want. If that means taking an open source model and fine-tuning it, and that's good enough, great.
21:36If it means that you need to pre-train something, that's fine too. If you're ChatGPT and OpenAI, like, yeah, they got to go build their own thing because that's their competitive advantage. I don't think that's true everywhere. But what we are seeing now is that the game is ratcheting up pretty fast. Like if I have something that I put in front of customers that interacts with specific kinds of data. Getting really good at interacting with that data probably means you need to own how that model works. If you don't, your competitors can buy that thing too. If I'm integrating an OpenAI API, maybe that's a great way to get started, but I don't have much of a competitive advantage because my competitors can go do exactly the same thing.
22:16You basically have decided to be on par with everybody. uh and everybody will get to the same place whereas if you have your own proprietary model and you're tweaking it and tweaking it everything you do past that open source moment where you use the open source software you own and those are accrete to your product or solution not to and this is distinctly different than just hosting picking where to host your server when you pick to host on amazon or google or rackspace azure whatever your the act of hosting on azure doesn't make azure or google cloud or aws better but the act of hosting on chat gpt or bar does make those models better correct and that is a subtle point where it can i guess yeah yeah yeah yeah so the physical infrastructure it's it's interesting i mean nvidia clearly had a huge bump recently they are the backbone of all of this both training and inference right now i mean we we we're actually encouraged we encourage many different types of hardware vendors to come to us and we want to run on their stuff nvidia is great they build really good products and we're running on top of them uh then the clouds are sort of the channel through which you get gpus right uh they also have some types of differentiation i mean network interconnect and you know reliability failover all this kind of stuff.
23:41You know, we find that it comes largely down to availability and price is the biggest differentiator along with some of these other more minor things like network capabilities. And really customers want choice right now. They want their cloud as a relationship that they do a lot of things on because they run their business on it. And they want some choice here. It's like, hey, you know what? I don't want to be like in a vice with one vendor because of this relationship. I want to have some choice. And so that's where the multi-cloud thing actually became a pretty good value prop from their perspective.
24:14Not every customer, but some of them. In that case, I'm building, AWS is giving me a great price, but Azure just dropped the prices massively. They're trying to win our business. I've got to keep running this model, growing it. It's not cheap to run these models. Can you give us an idea of what my, the job I gave you of my 5 ,000 meetings, what do you think this all costs to you know run these models at scale to add you know a thousand you know uh new startups a month to it and and really keep growing what is this going to cost well i think there's some misnomers out there that some people believe it's like you need to be at 30 billion to to build a model that even matters that's not true and uh but i think like in the level you're talking about let's start with your you know thousand documents when you're talking about hundreds of thousands or even maybe tens of millions of words it's really pretty cheap this is on the order of 100 bucks 100 bucks we could get we could do a lot in terms of fine-tuning um a thousand bucks you can do a really a lot so it's really not that hard uh but when we start talking about pre-training building models from scratch i'll give you the numbers our seven billion parameter model was trained on one trillion tokens, 1 trillion with a T.
25:29So that's approximately 750 billion words. A very long book has 100 ,000 words. So you can kind of do the math. It's a lot of content. That model took nine and a half days on 440 NVIDIA A100 GPUs, and it costs about$200 ,000 to build from scratch. Just to build that one model one time. Correct. And you run it again. You have to run it again. And you don't own all those. You rent those. You timeshare them. Correct. On other platforms. In the cloud. And these platforms in the cloud. Trovata is a cash management platform that helps you keep tabs on your runway, which is super important when you have to answer investor questions.
26:16And this is just going to gain control over all your financial data for you. Trovada scales from seed rounds all the way up to your IPO. And it makes it easier than ever to manage multi-bank liquidity with a single source of truth. You know, everybody now is putting their accounts and your money, you're splitting it across multiple banks. So you're protected with that FDIC insurance. Startups shouldn't be managing their lifeblood, aka your cash position in a spreadsheet. No. don't opt for bulky solutions that take months to implement when the banks now have super fast API connectivity. With investors like Wells Fargo and JP Morgan, Travada has pioneered the largest library of corporate bank APIs.
26:58Tons of unicorns like Carta and Fanatics trust Travada to gain visibility into their multi-bank data. It's the cash command center that helps you analyze, report, and forecast cash like a pro. Recently, they also launched Travada AI. This is the first genitive AI for fintech that uses GPT to automate cash reporting and business intelligence, while keeping your data private and secure, of course. So here's your call to action, go to travada.io slash twist to get started for free. Use the code twist for 30 % off premium features for one year, like AI forecasting and reconciliation. Educate the audience on the utilization rates of these in the cloud right now because we're hearing hey a100s h100s whatever there's a line around the corner we saw nvidia have this huge spike a couple billion dollars and yeah um you know unexpected orders came in uh so they're doing fantastic but uh if you want to run one of these models for your company are you waiting in line to get access to them do you have to reserve them is there like a line out the door to just use them or can you just use them anytime you need to well yeah it we are in a gpu crunch no doubt about it and that's not going to alleviate for a while i'm happy to talk about why that is as well it's a lot of time in that in the semiconductor industry um but right now um if you're willing to sign longer term contracts you can generally get them so we we as a company actually have blocks of gpus that we buy and we can bring to customers we call that a 1p deployment a first party deployment where we basically create a tenancy for our customer with GPUs that we already have contracted.
28:38So that actually works great. They basically pay us for a block of time and we can run those things and get them access and do it very efficiently and effectively. The other way we deploy things is within the tenancy of a customer. So a customer has a relationship with AWS. They believe they can get AWS to give them GPUs. We can run our software stack inside of their tenancy without ever seeing their data. That's something that people like because of the security and privacy. But as he said, the shortage of GPUs starts dictating how people go here. And so we actually do have a large number of GPUs.
29:12I don't necessarily want to comment on how many, but in the several thousands range that we can bring to bear. Now, the reason this is an an issue is that we're seeing scale scaling up these neural networks matters right for a 7 billion parameter model i needed on the order of four to five hundred gpus uh that wasn't true two years ago people weren't doing this and all of a sudden it's like what you would do on four or eight or 16 gpus now you're thinking i need you know 400 and so the the demand just went through the roof the new h100 that's the latest gpu from uh from nvidia uh is going to help a bit in the sense that each one is faster than the previous generation.
29:52So you don't need as many. But I think what will happen is it's sort of like goldfish. You grow the pond you have. As the capabilities of the hardware gets better, people just want to use more of it. And I anticipate us using routinely 1 ,000 GPUs for customer workloads. So that crunch is going to continue. Now, why do we have a crunch? I mean, can't we just crank out more silicon, right? uh what i think the the world doesn't realize is that there's really three places in the world that can build state-of-the-art silicon uh tsmc taiwan semiconductor that's the biggest one and that's where nvidia has a deep relationship samsung is another one that nvidia also fans on and then intel uh as a fab and intel primarily focuses on cpus for their um for their fabrication capabilities.
30:42And then beyond that, there's something called high bandwidth memory, HBM memory. It's packaged within the same physical package as the GPU. The process of packaging and getting memory together and making it all yield is actually the biggest bottleneck. There are two places in the world that make HBM memory, Samsung and SK Hynix. So this is your supply chain for these things. And there just ain't a whole lot more left of it. And to build out capacity means you got to build a whole new building you know and that's part of what the chips act was trying to do here in the united states is to create some redundancy have some of these on the north american continent and maybe have less dependency on regions that could be impacted by uh geopolitical events taiwan fill in the blinds right yeah um and so ramping those fabs up is underway but it this is a non-de minimis task it is a significant task to put one of these to stand one of these up this is a couple year process yeah i mean two to three year process and you know on the order of 10 billion dollars of investment to build a state of the art fab if not more these days so it's not small uh and it takes time i think that's the other part is that like if you want to react to a change in demand which there has been a big spike in demand the reaction time is a minimum two two and a half years just to build the capacity then you got to deliver that capacity so it's another year beyond that it's like a three-year minimum kind of thing the software and the models are getting so good um hugging face has like uh um a leaderboard of the models you're in the top 50 models and you have all these different players trying to make language models open source them and make them better so is it not true that these base models are going to be built and a lot of the demand to use them is not going to require they're going to get so good that maybe you're just not going to require to do as many new models or is it just induction where people like well i can make a new model i should run a model for my vertical etc yeah i think what's happening now is where one can't get the resources they're just going to take the other approach of I'm going to use an existing model.
33:00Great. It's a practical approach. They are giving up performance knowingly. If they had the capability to build their model, they can and will. So I think the demand is not going to tap out because of this. Like if I want to build a better model to be competitive in my space, I will. And if I can get the resources, I'm going to go do it. I might be strapped by resources, not be able to get them. We've taken a fundamentally different approach to a lot of companies in that we focused on efficiency of compute from the get-go. So meaning that, can I do more with less? Can I build a big model and make it cheap?
33:36So the reason that our model is state of the art and only$200 ,000 is that we put a lot of engineering time and research time into making it very efficient. We use that GPU completely. It's like, when I kill that animal, I'm going to eat everything, kind of a thing. And we're going to continue to do that. So we're getting more and more efficient with it but honestly the demand is going so fast that even with our efficiencies which bring nearly an order of magnitude of efficiency compared to what it was a couple years ago uh there's still not enough gpu compute and i think we are an absolute requirement to make this happen still not enough uh people are going to be seeing that they can build a better model and get an advantage and there's going to be an economic incentive to do it it's just they can't buy the gpu talk to me about the difference between specialized models and the general models and how this is going to play out because you know reddit bloomberg twitter quora these are very unique data sets not only are they unique data sets they've already have built into them some amount of categorization i.e a subreddit i.e a topic on quora that this is a legal topic versus a health topic i know it's the model can figure that out itself but the fact is these are very structured sets of data that have been built for decades that are really unique in the intent in building them stack overflow would probably fall into this so how does this i guess balkanize uh or manifest itself in the next two three four years is reddit just going to have a reddit gpt and quora already has their own gpt basically an interface maybe twitter has their own bloomberg created their own a small investment i have a small company i have skipped skift.com created their own based on their reports of travel companies they're kind of like a verticalized b2b travel publisher plus the transcripts of all their interviews plus all the research and the companies that they cover and they've made their own narrow language model uh how does this all pan out in the coming years so i'll give an intuition first uh before i go into the answer here i think the way to look at it is you know if i want to be of i want my kid to be a famous violinist what do you do you don't start them at 20 years old you start them at four right um arguably if they're if they're going to be a virtuoso and violin they're probably not going to be a finance virtuoso because they're going to put a lot of time and effort into making their brain very specialized toward that task.
36:14Even with everything that biology has given us in our brains, we still need to specialize to be really good. So right now we're at the very beginnings of this. Yes, you can talk about, I can build a model that can do a lot of things. It's going to be a jack of all trades and actually not a master of anything specifically. And that's okay. There are tasks where I want something general, right? If I want to take over multiple tasks that maybe people do or find mundane, I maybe want something general for that. And those general models will work for that. But when I really start getting into healthcare, being a co-pilot for a doctor or a nurse or a co-pilot for an investor, then I need some really kind of specific knowledge.
36:57And it's very difficult to make a general model do really well in specific tasks. That's one. The economics of building a general model that could potentially be very good at every task start getting kind of out of hand. I mean, just the training of itself gets very expensive. Then, because that model itself has to be so large, serving that model just has very unfavorable economics, especially when we're talking about compute being so scarce, right? Training and inference compute is basically the same kind of chip. So now I got to start thinking about, well, all right, if I want to actually serve this model to my customers, I need to think about the economics of serving that model.
37:33So I think we're going to be in a world where there are going to be some large general models and they serve some set of use cases and the cost to serve them is justified. There's going to be a whole tale of multiple expert models that are much smaller, that have much more favorable economics, maybe are very good at particular tasks and less good at other tasks. If I'm building something that's going to do customer support, I really don't want it to philosophize by why Rome fell. It just doesn't need to do that. It needs to talk about my products. It needs to get the user, you know, to fix their problem ASAP.
38:07That's it. Right. I don't want it to do anything general. So I think this is what we're going to see is this world where everything's kind of, um, uh, coexisting and solving different problems. We're already seeing that, uh, now, I mean, I talked to the founder of a company called perplexity AI, which is doing like kind of, you know, a search and, um, you know, uh, finding knowledge across different sources using LLMs and they're, they're doing a whole bunch of different things. They use every possible model they can to best serve the task that they have. So sometimes they use a general model to do filtering and they use specific models to condition the output the way they want them.
Read the full transcript
38:44So I think we're going to see this world where everything kind of coexists, which is going to be a bigger market. Our bet is that people constantly building experts on their domains is going to be the bigger bet. And the other one will be a consumer thing. Maybe it'll be different. I don't know. But I think in this world that's coming, we're going to see just a proliferation of all of these capabilities out there and the markets are going to be enormous. So it's almost like not worth sweating the details right now. All right, everybody, our friends from Microsoft are here. Tom Davis, a senior director at Microsoft for Startups, and you're a former founder.
39:20So Tom, tell me, the Microsoft for Startups Founders Hub, what is it? And what are you offering startups? Run us through the bullet pointed list of all these incredible benefits. There's lots of them. So we start with up to $150 ,000 worth of credits for Azure. That is not just traditional Azure, but also the Azure Open AI service, which is all the rage at the moment. You get benefits as well for productivity tools, so Microsoft 365 with Teams and Office in there, developer tools, GitHub, Visual Studio, but also third-party benefits like LinkedIn services as well. You can get access to Bubble.
39:58But as well, we have a special benefit with OpenAI, up to$2 ,500 with OpenAI. So you can leverage the latest and greatest models that are coming out from OpenAI. And when you want to go into production and reliability on services, you can shift across to the Azure OpenAI services that you get with$150 ,000 worth of credits. Amazing. Well done. And if anybody wants to sign up for that, do it now while you are in front of your computer, aka.ms slash this week in startups, aka.ms slash this week in startups. Well done, Microsoft, and well done, Tom. It's part of our mission to democratize access to innovation.
40:33So the more we can do for startups, wherever they are, whoever they are, the better it is for society in general. What do you think, as an insider, the impact is going to be on employment? So we'll go big picture now. we got into the details of these models congratulations on being one of the top 50 yeah it seems like you've got it dialed in there's going to be tons of use for this but what people are sweating is hey uh and i got my own feelings on it but i'm curious yours do you feel like even in your own i think you have 60 70 people in your startup do you feel like you need to hire as much uh or do you feel like as the ceo founder co-founder here your time is better spent taking the 67 brilliant people you've already assembled and just trying to make them 30 or 20 more efficient using ai tools where do you spend your time hiring the next incremental person or making the existing team better at what they do uh we're i still spend a lot of time hiring okay great people are still very hard to substitute i mean these models can do something that at a 20th percentile human i need 99th percentile players got it right so um i think there are things that 99th percentile players can do that very few other humans can do so we i spend my time on that so elite is still elite in your in your worldview the elite are not impacted by this trend well okay let's go down the the uh let's go down the rabbit hole at least not yet okay but but i think but i think to your point right even if i can make those elite players 20 more efficient that would mean i would imply that i need 20 fewer of them right right i mean making steph curry that has it right if you made steph curry two percent more efficient it would just destroy the league like this he's already too efficient yes right so if you think about all stars can you imagine making lebron james 20 more efficient i mean what happens to the league you know it's insane yeah no it's insane and i think throughout my career and i was here before the uh dot-com bubble and i've been a tech maximalist i felt that tech made the world better efficiency made the world better sure i've changed that a little bit and because of this new world and it's and i'll tell you why uh it actually has nothing to do with the technology but more about the pace of change uh what worries me is if i make the 50th percentile player 30 percent more efficient across the board i have the the change in demand won't be as fast as the change changes supply.
43:05And I think that's going to create this window of time for 30 or 40 years where we haven't figured it out as a society. And I don't know what the answer is. And that's the thing that worries me, to be honest. I do think the tech is going to happen. I think it enables humans to do more and to strive to solve bigger problems. AI, in my view, is the next evolution of what humans can do. you know language was a big technology that humans used to pass knowledge that that exploded what humans could do and the influence we could have on the world uh ai is going to be that next inflection point but how are we going to make sure that everyone has a place in that world how are we going to make sure that the demand that's created by the increase in efficiency is commensurate or does it collapse right uh so i i don't know the answer but uh that's kind of why i've taken suffice to say you're worried at this velocity if i could if i can summarize it correctly here and reflect it back to you which is an important thing to do in discussions uh the speed at which the efficiency is going to impact you know the 50th percentile below could be so um violent so fast it could happen so quickly that those people uh the demand for those people who don't make the jump could be uh so low that uh they they can't catch up in time and then they've got some number of years of their careers where they are sideline marginalized or otherwise not needed which is scary that's right it did happen quickly with things like the typing pool in i don't know if you're old enough to remember but law firms or you know many businesses would have a photocopying room a mail room and a typing pool and what and a filing room right and the filing room eventually gave way to box or you know google drive or whatever uh the typing pool everybody just typed their own stuff and mail became email and docu sign and and those rooms the photocopy room as well went away they don't exist in a modern office whereas those were half of a modern office's floor space previously but that took how long 10 years maybe 15 yeah something like that i'm talking about something that could change in three years right and um and i think the other part of it is that you know if you look at the the mail room and the filing room right perhaps that increased efficiency sum total for the business five percent let's say and it took 15 years we're talking about increasing efficiency by 30 percent in three years that that shift is so fast that like okay so you yeah or even a year right it could just be like boom you just can take on a new tool and all of a sudden it goes away so what happens then right so we we increase efficiency and delivery of goods and production of goods but now there's fewer people that can pay for it so so so what happens right as a society and this is where i think some people's minds go to ubi and then other people's minds go to entrepreneurship and it does really depend on i think your framing or worldview if you're an entrepreneur i think your mind goes towards well start a business uh or find more customers lower the price for whatever service you're doing you think about radiologists uh who you know look at um you know x-rays or computer you know generated x-rays mris etc it's pretty obvious that ai will absolutely do a better job in you know for most of that job in the next year or two if it's not already done so then what happens to those folks well we could do more mris we could lower the price of an mri we could let people take more mris or cts or pet's all these different tests what if we lowered the cost of those tests so that when your doctor was making a decision she wouldn't have to say i don't know about that it's worth it it's like who cares if it's worth it yeah it's it's not nine hundred dollars it's a hundred so go do it yeah we can do ten times as five yeah no and that's the world i i want right and that's where i would restore my faith in tech maximalism right where we can do that and we actually just do better at everything uh we did it already there's an example cheaper faster remember remember food insecurity yeah this concept of food is great and now what do we have obesity we in the 80s when we were growing up i don't know how old you are but you know but in the 80s you know we had live aid and we were trying to feed africa that was like oh my god this was the cause celeb of you know africa has no food and you know now if you don't have food in the modern era it's because some dictator in all likelihood has blockaded food from reaching you and the biggest drug in the world right now is our zampic and wagovi because we have an obesity problem in abundance we could have an abundance problem we could have an abundance problem in healthcare in our lifetime too many doctors too many nurses too many beds too much available you're going to be too healthy because we just figured everything out kind of like the abundance angle again great i i want that to be the case um and i i do i do go in my own mind to entrepreneurship um i just don't know if everyone's wired like that um is this what worries me every human has the motivation to become self-reliant radically self-reliant yeah and hunt for their meals as opposed to punch a clock and get their meal ticket uh it's a it is a that's right that's right what do you think that happened with education and this is the one i think is super fascinating because i'm learning so quickly right now i agree it's just using chat gpt as my default browser when i open my browser it's my default window now and i i'm retraining myself to use chat gpt4 as my first line and man i'll be on a podcast and i like when i'm talking to you i might like when we were just talking i said what why college jobs would be most impacted by ai and i saw radiologists on the list and i was like you know just for brainstorming i was like yeah that's an obvious one uh and that's how i did that throughout the conversation was ai I got to radiologists before I would.
49:29Pretty amazing. Interesting. Yeah. No, I, I think for education, right. It's, it's going to be, I have, I have kids. Um, I have kids in high school and you know, there's a, a traditionalism in education that again, kind of goes on a very long time scale, right? People are like, oh, you liberal education, you need to do this and you learn that. Um, I take a different approach where it's like, all right, you know, chat GPT is here. my kid told me he submitted a paper written by chat GPT. And I said, look, I don't want you to be dishonest. I'm okay with it. As long as your teacher's okay with it.
50:05So if you tell your teacher and she was okay with it, I'm perfectly fine with it. Because that's what your world is going to look like. And learning how to wield these tools and make them really effective is going to be how you differentiate yourself. So I think education should be more about like exposure to these tools and and solving problems uh directly and as opposed to sort of uh memorization of knowledge which was sort of human 1.0 right we had writing and human 1.0 was like okay if i can memorize stuff i i know something others don't that's gone right i i have google i have chat gpt i have all that stuff right now and i can i have access to everything that every human yeah has ever has ever written in a scientific paper yeah and you can get to it quickly i you know the thing i i find is interesting about kids and i'm i'm big on this montessori and like base level learning and i like this regio learning where you follow the kids instincts and if they're really into something obviously the aperture for learning goes way open you know if it's about some you know orcas my daughter was into you know killer whales for a little bit and she you could teach her anything with killer whales you could she would do math physics as long as it was with a killer whale you know as the thing we're weighing or the thing and the force of the killer well like she's going to be really into it so that's awesome but just personalized learning and unlocking student creativity as but two measures here you could take any personal lesson plan and i could say hey take this lesson plan for history and you know um um let's have a an approach to it that includes superheroes and it's like what how do we include superheroes and it's like oh well yeah they they did actually use captain america to you know uh study like uh you could you could make a captain america going through different world wars and or spider-man doing it would actually make total sense actually to that person and then they would be drawn into it spider-man teaches you physics great what could be better right like yeah the the personalized stuff to me is amazing for kids brains um and that was what the vulcans were doing you remember in star trek when the vulcans would go into those little pods they would there was like a one episode of the star trek series where like uh i don't know if it was one of the reboots but you know like they just put spock in a pod and he's sitting there in a pod and the computer is just throwing information at him and he's learning like i just see that as the future is like the ai knows what you know what you don't and is going to present the next lesson plan that you're most open to and will be most accretive to your life that's wild when you think about it i don't know i find myself optimistic it's like we're hacking uh human human learning process you know yeah are you optimistic right now watching this because the pace you've been in this for a while but the pace was very slow and then it suddenly breakneck which you know elon and some other people did predict that this will be slow until it's cataclysmic and what do you think pretty accurate pretty accurate and why why is that so accurate if you think it is well i okay so cataclysmic i think is the wrong word i think it is breakneck i am very positive um there are things i still worry about but i'm still positive um and i think we are what's happening now that i think is a bit annoying is that cataclysmic kind of rhetoric is being used in self-serving ways you know in anti-competitive ways example i think we're still in this phase open ai closed ai well i mean i'm i'm perfectly supportive of them being close they should be able to have their own competitive advantage if they want totally fine with that what i don't like is talking about regulatory agencies issuing certificates of you may now go train a model.
53:58I mean, come on, really? We're so much at the beginning of this whole journey. We don't even know the value of a model. We don't even know how we think about the data that went into the model. We don't even know the use cases for most of this stuff. Let's let the flowers bloom a little more, and then we'll start understanding the bounds of where the incentives break down when people don't own things and all that kind of stuff and get to there. I don't think the end of the world is nigh. I really don't think we're that close to it. Um, I think people are using that fear right now to, they're using that to do regulatory capture.
54:34Yeah, exactly. Pull up the ladder behind them. I do have to say, I do find it questionable that opening. I became closed AI, the whole premise. And I've told Sam this, and I've said it publicly like the whole premise was, this is too dangerous for people not to see what's going on. And then they said, well, it's too dangerous for people to see what's going on. So how do you make that, crazy shift it's almost like this uh we know better than everybody else but if you go on hugging face and you look at the 50 open source models they're doing it open source so why what's so unique about open ai that they get to make this decision uh and in fact google engineers you must have seen this in leaked documents said and i'll just quote we've done a lot of looking over our shoulder at open ai who will cross the next milestone will be the next one but the uncomfortable truth is we are in position to win this arms race and neither is open ai while we've been squabbling a third faction's and quietly eating our lunch i'm talking of course about open source plainly plainly put they are lapping us it's crazy yeah uh why why is google so scared of open source organization to compete with them well because it's hard to compete with with the whole community right there's this unleashed creativity from many many people you just can't you can't compete with it that's been traditional in software right i mean you've seen this linux all this like you can try to centralize it and that maybe that's that's the activation energy to get it over a hump but to compete with it is very hard and i think that's why everyone's scared of open source but i think back to a philosophical point i 100 agree with you it's like why do you have the um the mandate to to dictate what this technology will look like that's the part i have a problem with and i think the way to solve that is actually uh distributed capabilities many people having these capabilities right it's like um yes there there are economics involved and it's it's expensive but we can make those economics a little bit more favorable by by time slicing it actually looks very similar to semiconductors uh we actually call ourselves the lom foundry very similar to yes there's a large investment required to build a foundry But once you do that, the incremental cost of making a chip actually isn't terrible.
56:45And by enabling many to build these chips, you build, you know, Apple builds their own chips, Qualcomm builds their own chips, you enable this whole ecosystem. And I think that's how we solve this is through almost a market solution, not centralization. And it's like, you know, sort of paternalistic centralization. That's the issue I have. And maybe that's just the that's the entrepreneur in me. I hate when someone tells me that I'm allowed to do something or not. i mean there's i think it's great that we're having conversations about how fast this is moving and the impact it's going to have on society because usually everybody's very late to that party and so the fact that we're doing it in real time for the first time it felt like they play catch up with you know social media they play catch up with regular the regulatory framework for a crypto but here we are we're looking at ai which will have certainly a bigger impact than crypto did uh obviously uh and it will yeah i think it will probably have a bigger impact than social even though social has impacted governments media people's health their psyches this should be a bigger thing and it's actually good that we're having the conversations if some people want to regulate it for nefarious reasons or to pull the ladder up i think we can see that happening um but at least we're aware of it and it's a great that you're building something that democratizes it a bit and levels the playing field so companies individuals non-profits whoever can start building their models uh in a more open source free way and then portable right i mean the portability is also super important that no one person owns this hardware stack and that's not going to happen right it's not like there's any hardware advantage that's going to accrue uh here there's going to be many competitors to nvidia in the coming years you think or you think they're going to run the table so i it well i think what nvidia has done well is just they've executed really well against those competitors that have tried to come up the would-be competitors and um that's why they've continued to maintain advantage which is again they did exactly what they should do and it's i i would argue it's because of jensen's leadership being you know here this is what we're doing right very tops down very strong-handed um but i think there are going to be competitors and for the simple reason that we need more supply yeah um what worries me however is that the supply even if there were 10 nvidia's out there we may not have that much more supply simply because the bottlenecks back in the supply chain are further back it's not you know memory and packaging right yeah so uh but yeah i i'm encouragement i encourage anybody to build a competitor and this might you know what might be interesting about it is if the if it turns out the hardware stack throttles this a bit and that could be a built-in throttling then we don't need the government to get involved it's like hey we're going to be able to make so much progress here uh you know without exactly the hardware stack dramatically increasing so all right listen amazing job continued success you're hiring uh you mentioned all-stars 99th percentile uh how can people learn more or how can they see what jobs are open who you're looking for et cetera.
59:57Let's get you a couple of employees for coming on the show. Yeah, absolutely. Uh, go to our website. We have, uh, several listings on there for careers. And, uh, you know, basically people who, who are, who are builders, innovators, um, this is what we need. We're a small team and we rely on, you know, highly creative individuals who are amazing at what they do and it can implement them fast. And if you feel like you're one of those and you want to make a difference come talk to us mosaic ml.com slash careers all right we'll see you all next time on this week's service bye-bye
From the publisher
This Week in Startups is presented by:
Vanta. Compliance and security shouldn't be a deal-breaker for startups to win new business. Vanta makes it easy for companies to get a SOC 2 report fast. TWiST listeners can get $1,000 off for a limited time at vanta.com/twist.
Trovata. Starting up is hard. Trovata makes managing cash easy. Start automating your cash management at Trovata.io/TWIST. Use Code TWIST for 30% off one full year of premium features like AI forecasting.
The Microsoft for Startups Founders Hub helps all founders build a better startup, at a lower cost, from day one. Startups get up to $150K in Azure credits, access to free OpenAI credits, free dev tools like GitHub, technical advisory, access to mentors and experts, and so much more. There is no funding requirement, and it only takes minutes to join. Sign up today at aka.ms/thisweekinstartups
*
Todays show:
MosaicML Co-Founder and CEO Naveen Rao joins Jason to discuss the open-source vs closed AI debate, the profound impact of AI on society (41:06) AI’s rapid pace of change, and its implications for the future of employment and education (40:42). They wrap the show by breaking down the potential problems with centralized regulation (54:37).
Follow Naveen: https://twitter.com/NaveenGRao
Check Out MosaicML: https://mosaicml.com
*
Time stamps:
(00:00) Naveen Rao joins Jason
(2:54) MosaicML and its purpose
(5:10) Obtaining datasets and incentivizing creators
(8:30) Vanta - Get $1000 off your SOC 2 at https://vanta.com/twist
(9:37) The process of using your data with MosaicML
(11:55) Defining tokens and prompts
(16:53) Fine-tuning the AI model and reinforcement learning
(19:27) The competition with open-source models
(24:26) The cost of running AI models
(26:08) Trovata - Use code TWIST at https://trovata.io/twist for 30% off one year of premium features, like AI forecasting
(27:35) How the GPU crunch has affected cloud models
(32:13) Why demand will not cease
(34:21) Specialized models vs. general models (39:12) Microsoft for Startups Founders Hub - Apply in 5 minutes for six figures in discounts at http://aka.ms/thisweekinstartups
(40:42) The impact AI will have on employment
(48:49) The impact AI will have on education
(54:37) Thoughts on OpenAI becoming ClosedAI
*
Read LAUNCH Fund 4 Deal Memo & Apply for Funding
Great recent interviews: Brian Chesky, Aaron Levie, Sophia Amoruso, Reid Hoffman, Frank Slootman, Billy McFarland, PrayingForExits, Jenny Lefcourt
Check out Jason’s suite of newsletters: https://substack.com/@calacanis
*
Follow Jason:
Twitter: https://twitter.com/jason
Instagram: https://www.instagram.com/jason
LinkedIn: https://www.linkedin.com/in/jasoncalacanis
*
Follow TWiST:
Substack: https://twistartups.substack.com
Twitter: https://twitter.com/TWiStartups
YouTube: https://www.youtube.com/thisweekin
*
Subscribe to the Founder University Podcast: https://www.founder.university/podcast




