Open Models Change The Economics of AI

12 Sep 2026 · 57 min · 27 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

How open-weight/“open models” are changing AI economics for enterprises and developers, driven by lower cost, faster model release cycles, and tooling that makes models easy to run (Ollama local/cloud). It also covers hybrid routing between open and frontier closed models, plus security/geopolitics and GPU/hardware realities.

Guests

Jeffrey Morgan, co-founder and CEO of Ollama. Background: previously built Docker Desktop at Docker; later focused on developer experience for running open models locally and in the cloud. Ollama claims: 9M developers, 178K GitHub stars, used by 85% of Fortune 500.

Key claims

Cost is the biggest short-term pain open models solve; it then enables enterprise customization. Fine-tuning cycles are returning because open model release cadence is speeding up. Enterprises shift token consumption to open models (AT&T: 40% of token consumption). Open models dominate coding agents and are expanding to non-developers via agents like OpenClaw/Hermes.

Notable examples

OpenClaw token growth; context windows from 128K to 1M+; security testing where closed models refuse pen testing but open “security researcher” models on Hugging Face can. Ollama “day-zero” launch playbook: harness/SDK, packaging, and inference/hardware benchmarking.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Role of Open Models in AI Economics

0:00 to 0:44

Learn how open models address cost issues in AI and their potential for customization.

“Cost is by far the largest pain point that open models can jump in and solve, but every business has a vision of getting better control over AI and customizing it for their business.”

State of Open Models in Enterprise

1:16 to 2:24

Discuss the shift towards open models in enterprises and their usage trends.

“Well, I think the biggest thing we're seeing is a shift to open models, especially in enterprise.”

Impact of Cost on AI Adoption

2:24 to 3:22

Explore how cost influences businesses' decisions to adopt open AI models.

“You know, cost is something they can solve in the short term, but it then enables them to then go and customize these models for their unique use case.”

Growth of Coding Agents and OpenClaw

3:22 to 4:28

Analyze the rise of coding agents and the significance of OpenClaw in AI workflows.

“to automate a huge chunk of work over a long span of time to non-developers too.”

Challenges in Custom Training of AI Models

4:28 to 6:00

Understand the difficulties and advancements in custom training open AI models.

“We went from a context window of 128K to a million plus with open models.”

AI Safety and Security in Model Adoption

6:00 to 7:26

Discuss the importance of safety and security in adopting open AI models.

“What used to be more of a six-month cycle.”

Advantages of Open Models for Security Testing

7:26 to 8:06

Learn about the benefits of using open models for security testing applications.

“And it's really exciting for these businesses.”

Operating at Scale: Challenges and Strategies

8:06 to 11:20

Discover the operational challenges and strategies for deploying AI models at scale.

“But even the out-of-the-box models, they do come with safety training, but they're a little better understanding if you're doing this for a good use case versus one that's more of a negative or malicious use case.”

Future of Open Source AI and Market Dynamics

11:20 to 14:00

Examine the potential for market changes as open source AI evolves and diversifies.

“you get the model, you're lucky if there's the ability to access the model a few weeks in advance.”

Open Models vs. Closed Models in AI

14:00 to 19:06

Explore the dynamics between open and closed models in AI and their implications for developers.

“There was a really good talk from some of the Anthropic team, the platform team, and they talked about three big things.”
Show all 27 chapters

The Role of Local vs. Cloud Models

19:06 to 21:35

Understand the advantages and use cases for local versus cloud AI models in business.

“It's kind of an interesting idea because as the models get more powerful, ideally you just want to delegate to like your smartest model to figure out when to go to like an open model.”

The Future of GPU Usage in AI

21:35 to 23:26

Discuss the evolving GPU market and its impact on AI development and deployment.

“It's incredibly exciting because you can run that on not the lowest memory MacBook, but the second lowest memory MacBook you can buy from the store.”

Challenges in AI Model Deployment

28:07 to 28:46

Explore the complexities involved in deploying AI models effectively.

“And it's definitely a lot of spending time thinking through, you know, which model will get run where?”

Maximizing AI Model Efficiency

28:55 to 30:26

Get insights on using low-cost AI models effectively within startups.

“Suppose you were like a startup founder and you were just starting out now and you're building some AI company and you haven't raised a lot of money.”

Maximizing AI Model Efficiency

30:28 to 30:43

Get insights on using low-cost AI models effectively within startups.

“And so this new class of Flash models, where they're good enough for 80 % of the tasks, they're really fast, and they're ultra cheap, this new class of model that I think will enable some of those use cases.”

The Promise of Open Models

30:43 to 33:10

Understand the advantages of open AI models and their applications.

“If you think back to the coordination layer we were talking about earlier to be able to coordinate these flash models together to do different tasks can also yield great results that a bigger model can.”

Geopolitics and AI Model Origins

33:10 to 34:34

Delve into the geopolitical implications of AI model origins and security.

“Will there be an open model that becomes a god-tier model?”

The Origins and Evolution of Ollama

34:34 to 37:05

Learn about the founding story and evolution of Ollama in the AI space.

“phases where they sound more robotic, they sound more friendly.”

Navigating the Startup Landscape

37:05 to 42:00

Hear about the challenges and pivots faced by startups in AI development.

“So you applied to YC with a very different idea, right?”

The Evolution of Ollama's Development

42:00 to 43:56

Learn about the pivotal moments that led to the development of Ollama.

“Back then we were thinking of it as like the segment for LLMs.”

User Adoption and Market Dynamics

43:56 to 45:55

Discover how Ollama transitioned from hobbyists to enterprise users.

“And before that was two years of just frankly overthinking the customer, the product, and just not getting something out there.”

Monetization Challenges and Strategies

45:55 to 47:40

Explore the challenges of monetizing Ollama and the strategies employed.

“That is incredibly helpful to a Fortune 500 IT developer team because they don't have to ask for permission to use it.”

The Importance of YC and Community

47:40 to 49:51

Understand the value of the Y Combinator community for second-time founders.

“And one of them was a privacy-focused AI product, which Ollama really started with that in its open source incarnation.”

Learning from Past Experiences

49:51 to 51:54

Learn how past experiences shape current startup strategies in AI.

“You know, we went back and forth on this for a lot, which we shouldn't have.”

Adapting to the AI Landscape

51:54 to 54:30

Discuss how the AI landscape changes traditional startup rules and strategies.

“And this is the quality that's necessary.”

Curation in the Open Model Ecosystem

54:30 to 56:00

Examine the role of curation in managing the diverse AI model landscape.

“And I think from building a team too, it's that, you know, with AI now, there are just problems that you don't need to staff as heavily, whereas you did 10 years ago, right?”

The Value of Model Curation in AI Development

56:00 to 56:58

Learn how model curation simplifies the development process for software creators by aggregating various AI models.

“They want to build their next company, their next application.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Cost is by far the largest pain point that open models can jump in and solve, but every business has a vision of getting better control over AI and customizing it for their business. And that's really their North Star. Cost is something they can solve in the short term, but it then enables them to then go and customize these models for their unique use case. Early 2024, there was lots of interest in fine-tuning your own custom models. Then it sort of went away, or there will just be wasted effort or get stomped by the next model release. Seems like it's coming back now. You have a front seat to all of it.

0:33Do you think we're going through like another cycle or is it here to stay this time?

0:43Welcome back to another episode of The Light Cone. Today, we're talking to Jeffrey Morgan, co-founder and CEO of Olama, the easiest way to run open-source AI models locally and in the cloud. Olama is used by 9 million developers, has 178 ,000 GitHub stars, and is used by 85 % of the Fortune 500, which means Jeff knows a lot about the state-of-the-art of AI, what models score highest on benchmarks, and what developers actually download and keep using. Jeff, welcome to the light cone. Thank you for having me. We're down to here. Like, what is the state of the art? What are you seeing out there? Well, I think the biggest thing we're seeing is a shift to open models, especially in enterprise.

1:25And that's from a mix of US and Chinese origin models. And it's predominantly driven by coding agents and also AI assistants, more co-work use cases like OpenClaw and Hermes. And because you sit in the token flow of like so many tokens, you have really good data on what models people are actually using and how it's changing. What are the trends that you're seeing? Yeah, you know, Ollama started as a way to run open models on your MacBook or other hardware, NVIDIA, AMD, Intel. And earlier this year, we launched Ollama's cloud. And what we're seeing there is that it's predominantly Chinese models right now of Chinese origin.

2:01but they're being accessed by businesses all over the world, especially US and Germany is actually a big source of where open model tokens are being accessed. Is it all about cost? Is it our enterprises coming because they just want to get the cost down or is there anything more to it? Cost is by far the largest pain point that open models can jump in and solve. But, you know, every business has a vision of getting better control over AI and customizing it for their business. And that's really their North Star. You know, cost is something they can solve in the short term, but it then enables them to then go and customize these models for their unique use case.

2:39Is there a particular large enterprise that you can name that has done this? I think there was a great article in the information yesterday from AT &T. And it ends up they've already shifted 40 % of their token consumption to open models. And that's right now predominantly through U.S. and Europe models. but they're also evaluating the Chinese models. What kind of workflows do they run? Predominantly coding agents. I think what we've seen just from the extreme growth and per developer or per user token usage has predominantly been from coding agents. And then earlier in March and April, we saw OpenClaw take off and subsequently the Hermes project, the Hermes agent project take off, which has then opened up that ability to automate a huge chunk of work over a long span of time to non-developers too.

3:28whether it's like finance or support or marketing or sales. You had this actually very cool graph on the takeoff exponential for OpenClaw. Yeah. So earlier this year, this is a graph of token usage by developer on Olamis cloud, the average amount of tokens they use per week. And we kind of start, this is per developer. So this graph looks like it should be an aggregate of Olamis growth, but this is actually the per user. Exactly. This is on an individual user basis. How many tokens are they using a week? And so there's kind of like two big inflection points. One is that initial run up at the start of the year, which was driven by coding agents.

4:08So we saw Kimi, the GLM models, Minimax launch. Finally, we had open models that could power coding agents. And then in April, we saw this incredible growth from OpenClaw, really, which was then not just developers, but the rest of the world could take a hard problem, give it to an open model and let it go complete the task, which obviously consumes a ton of tokens as it's figuring out what tools to use, what data to go fetch. We went from a context window of 128K to a million plus with open models. And so all that enabled this explosive growth. So it went from roughly 5X, so under somewhere around, I don't know, 15 million tokens before all these co-work type of use cases.

4:48I think that's about right for that open claw jump we saw in April. Obviously in aggregate, it's in the 10 to 20X, if not more. As a whole through Ollama's cloud, we saw 150X since the start of the year. And so it's just the big interesting thing here is this huge surge in demand for open models, right? Whereas open models, I'd say in 2024, 2025, from the large models being served, were mostly being served as custom models. So you'd take an off-the-shelf model like DeepSeq or Kimi, and you'd fine-tune it for your use case um like for example you know cursor had famously done and from there you know you could serve that at scale but seeing out of the box open models being served that really only took off at the start of this year even since we started this podcast these things sort of come in cycles maybe like i feel like early 2024 there was lots of interest in fine-tuning your own custom models then it sort of went away and it's like or that will just be wasted effort or get stomp by the next model release seems like it's coming back now you have a front seat to all of it like what do you think we're going through like another cycle or is it here to stay this time i think the release of these models the cadence is only speeding up which makes it ever more harder to you know stay on top of that and have a post you mean the the the latest closed frontier models are releasing faster than i think on the open source side it's getting faster and faster you You know, this summer, we've already seen three iterations of the DeepSeq Flash model as an example.

6:17What used to be more of a six-month cycle. Yeah, so the gap is like closing, essentially. It's closing. And I think that makes it even harder to custom train models. On the flip side, I think the tooling is getting better. And so it allows teams that want to fine-tune their models to stay on top of it. I mean, we're also entering this moment where AI safety is becoming more and more of an issue at the Frontier Labs. So the frontier may well slow down to figure out its alignment and containment issues. And then meanwhile, the open source models and open weight models are continuing to grow and get better.

6:51Yeah. And, you know, we saw the announcement and the release of the GLM-5.3 model and its capabilities from a cybersecurity standpoint, you know, being extremely impressive. It creates a big opportunity for whether it's startups or existing businesses in the security and governance space to really jump in and help. Because I think if you look at that AT &T article I was speaking about, the blocker to adopting open models is largely around security and safety. But from our experience talking to customers, whether it's in Europe, whether it's here in the U.S., if you can solve the safety problems, by and large, adopting the Chinese origin model labs is completely on the table.

7:26And it's really exciting for these businesses. There's sort of this interesting moment right now where a hugging face had to use open weight models to actually even detect the hack from the frontier. A common question we get is like, well, where can I use open models that are leaps and bounds of an advantage over using a frontier closed model? And one of the key use cases is security testing and making sure that your software is secure. Yeah, because if you try to get Claude to pen test your product, it will just refuse to do that. Correct. Yeah. By and large. Whereas there are literally obliterated security researcher models that you can find on Hugging Face that allow you to do it.

8:04That is true. There are ones that are custom trained to be even more liberal to go and attack these problems. But even the out-of-the-box models, they do come with safety training, but they're a little better understanding if you're doing this for a good use case versus one that's more of a negative or malicious use case. A cool thing about Olama is that because you guys are such a key distribution channel for these models, my understanding is that typically the model developers are contacting you before the general release to coordinate launches and stuff like that. And you often get sort of previews of what's about to happen.

8:43And you get these incredible growth spikes when a new model drops. I wonder if you could tell us a bit about what it's like to operate this thing at scale. Yeah, absolutely. And like I said before, the models are coming out faster and faster and faster. And so we've developed a playbook to what does a successful day zero model launch look like? And there are a lot of things to get right. There's making sure that it's supported in your favorite inference engine, which is generally sometimes a multi-week process to make sure it's fast, to make sure it's accurate, to make sure it's up to spec with the reference.

9:12There's also finding the right use cases and harnesses for developers and users to make use of this model. Generally, these models have new capabilities. This morning, DeepSeq launched their first multimodal model from a large model LLM standpoint. They had previously had some smaller OCR models, which unlocks a whole bunch of use cases. But what they also changed was the DeepSeq harness, which launched recently, so that it could support this capability. So step two is then to find the harnesses, make sure they're prepared to actually run this model and to do that effectively. But every model is different.

9:44They all have different challenges and architecture changes and tool calling mechanics. And getting all this right is really hard. I think the key thing to do at the end of the day is to run benchmarks against the final product ahead of release and make sure that, you know, it's running as the research team at the lab specified. So I think one very important part that you play in the whole ecosystem is you kind of create a very legible standard to be able to make sure you have the best way to use the harness for each new model, each new tool call, and all of it is consistent across all the different models, which is pretty hard to do.

10:21Yeah, I think there are three things really that we try to package together. One is harness, for example, whether it's an off-the-shelf harness or an SDK to help use the model, some existing harness that is designed for this. And what's great now, there's so many great open source harnesses. The Codex harness is open source. OpenCode's a great one and one of the most popular harnesses from Olami users. Step one's getting that right, but then you've got to package it with the model, making sure that it's available, it's reliable. if it's in the cloud, that there's enough capacity for it, because day zero tends to be the largest growth day, obviously.

10:56And then importantly, there's the hardware and the providers. And that's actually where there's a lot of collaboration to be had, whether it's like an inference provider optimizing the model or, you know, with some of our partners, whether it's NVIDIA or, you know, for example, working with the Apple Silicon stack, making sure that it actually runs the model really fast because if the model is capable, but it's really slow, that's not a great experience. So getting those three things packaged into a box and generally you get the model, you're lucky if there's the ability to access the model a few weeks in advance.

11:25A lot of this stuff comes together in the last 24 hours before the model gets released. And so it's generally a fire drill. So you kind of almost like an operating system in the old world where you needed to really integrate very tightly with all the drivers, all the hardware, and then at the application level to make sure that all the apps were really tuned up well. And you are the glue for all of it, right? I do think the OS, which is generally a cliche analogy to use is a good one because you've got the drivers for the hardware and the providers and the inference layer, but you also have the application runtime and making sure that the harness works.

11:58Gluing that together, it's a very combinatorically large problem to solve. And so doing that well is really difficult. And so, but, you know, over time you develop pieces that, you know, allow you to quickly develop that, test it, release it, and kind of have this common runtime that can match any harness to any model. And that's kind of the role we're really playing for developers. You guys are very hardcore engineers and you fine-tune things all the way to from the origins with Apple Silicon all the way now to DJX. How do you build such a deep technical bench with that? I think a lot of our team, you know, we aren't AI researchers by background.

12:38We're from VMware and Docker and from, you know, other networking companies. And so the classic compute problems are kind of reinventing themselves in the inference land, whether, and so largely, you know, that's where we like to focus our time. But I think at the end of the day, you know, it all comes down to the developer experience. What is it like when the developer makes a call to the API and gets tokens back? What happens? And there's more and more happening in that layer right now. And getting that right needs all the layers of the stack to work well together. And I think, you know, one of the challenges with open models has been that hasn't been happening at the rate of what a frontier model lab puts out, where they have the classic five-layer cake that Jensen mentioned, which is like the apps, you have the model, you have the infrastructure and inference, you've got the chips, and then you've got the energy.

13:28And they've got all that ready to go for developers on day zero. And that's really the thing that we're trying to reproduce for open models. Of course, we're not gonna do every layer of the stack, but we can help orchestrate that. You layer the layers. Yeah, and maybe in open models, there's more than five layers. Like that model layer actually has a lot to it, right? There's the model weights, but there's also a lot of the orchestration components that each model is uniquely good at. There's this developer API layer. There's so many opportunities to build between the model and the application layer that are kind of hidden in today's, you know, five-layer cake stack.

13:58That's interesting. Do you think any of those like hidden layers might get unbundled and become their own companies or providers? Absolutely. There was a really good talk from some of the Anthropic team, the platform team, and they talked about three big things. One was knowledge. How do you connect your company's data and context to the model? One is coordination. As you know, when you make a request to in your cloud app or codex, it goes off and spins out a bunch of sub agents, some of those in the cloud, some locally. There's a coordination problem. And then lastly, there's an execution problem, which we talk a lot about as sandboxes.

14:32But these agents, more and more of them are moving to the cloud. There's a huge compute problem to be solved there. And if you look at what's happened classically in the cloud business, open source has meant that there are best of breed companies for each of those problems. Whereas you might have had kind of the, if you look back to the original generation cloud products, you have like the Herokus of the world, you have Google App Engine, where all those things were bundled together. But what developers ended up preferring is best of breed products for each one. It's the novelism, which is all things are just bundling or unbundling.

15:04Exactly. Like the frontier model. So Frontier Model Labs want you to be totally in their walled garden of managed agents and their context, their memory layer. And then meanwhile, Little Tech and all the founders out there and all the open source developers don't want to be caged in. So we're going to make all this other stuff. And it'll be an interesting moment to figure out what ends up winning. I mean, it'll probably be some mix of both. I think so. And you've got this abundance of open model tokens that's being created. There are just dozens of open model providers that are able to serve these tokens.

15:35The new scarcity, the problems now are what's above the tokens, right? How do you orchestrate an agent from A to B? These are problems that are tons of new systems and engineering problems that are just really hard to solve for an individual dev. There's no way they're going to build all those layers. Well, the interesting thing now is because the coding agents themselves are getting a lot better and ostensibly this is the worst the models will ever be. The classic reason why there was a moat here was it was just too hard to have really well-maintained software that was properly tested that actually satisfied user need.

16:11And what if that goes away? Like, we're literally at this moment where actually maybe the age, you know, you'll just have a cron and it runs a markdown file on some TypeScript and it'll just, you know, there's no lock-in anymore, right? Like you could be using OpenAI's memory system one day and then actually, you know, you could have an agent be constantly syncing that against your own memory system. And actually it just works. It's fine. Like it works over MCP. Like there's all, you know, there's plenty of like runtime testing and then there's no lock in. Yeah, I think a lot of the, you know, what we know of as harnesses today, a lot of those pieces will go down into the model.

16:52but you know as you kind of push a lot of the core loop of the model and hooks down to the model itself there are these pieces that kind of come out like memory is a great one and the the general guidance we're using is if it's a stateful problem like there's storage involved that's something that you know in the end can't go into the model because the model's trained and it's you know as we know training runs are now happening on like a monthly basis but it's still not up to date with the latest data so generally storing data is like a huge problem space that i don't think will ever make its way down into the model layer and managing credentials and that kind of stuff security credentials safety um open models don't have all the safety tooling that closed model providers give you out of the box but that's super important especially for businesses to adopt i'm curious where you see sort of the the end or future state for enterprise on the balance between sort of frontier closed models and open source model um especially on spend like my it feels like you know initially it was just the everyone's just like allocating all of their budget to anthropic or um open ai uh my sense now is yes open a open source is clearly growing but like so are so is like anthropic spend so like the things seem to be growing together like does that continue or do you think there's like a steady state where it's like i don't know like half the budget's going to be on the closed source frontier model and half's going to be an open source or something different the super majority of tokens and this is our take it will be open models within a business call it 80 90 that doesn't mean 89 of the budget will go to models in fact i think what the open model community is doing incredibly well together is lowering the cost to make it more accessible and so maybe you'll only pay 10 to 20 percent of the cost towards open models, but your token usage will have.

18:40Most of your tokens will be going through the open models. Right, which will enable a whole bunch of use cases on top because you have this abundance of tokens. You're not thinking about taking away token access from your team. You're giving more and more access. I think for the hardest tasks that's reserved for these frontier labs where a lot of the best researchers are. And then from there, there's a whole bunch of problems in the middle, right? Where maybe it's a combination of open and closed models working together. I think the steady state is that most of the software's open, most of the models are open.

19:07It's kind of an interesting idea because as the models get more powerful, ideally you just want to delegate to like your smartest model to figure out when to go to like an open model. But the labs who are in the models presumably don't want that. And I think, look, I think that all the labs are aligned in many ways to one thing, which is how do you serve the customer? And I think it will be up to the customer to decide if I have a router where some of the scheduling and harder orchestration happens through a frontier model, so be it. but a lot of the kind of line item work can happen through open models and the collaboration of the two together.

19:42I think we've seen a ton of projects, whether it's from Sakana AI or Open Router that have combined the two and it's seen really good results. This is not too dissimilar from a human organization, right? Like you have, you know, like a law firm, there's like a partner and then there's like a bunch of associates and like the partner farms out the work to the associates. It's like the same. We saw the same thing with cloud computing where it was really a blend of proprietary software. some of them provided by the cloud providers themselves. For example, AWS had DynamoDB, which was kind of their proprietary scale-out database.

20:12But then a lot of customers use that in conjunction with PostgresDB. And in the end, what we see is customers will use a combination of the two. I think it's a very common pattern. I mean, this is also the same design for why the Apple Silicon is actually more superior, the special accelerators for different kinds of workloads for, let's say, image processing, it's supposed to be audio that's it's been or even go way back in in the pc era you had like your standalone audio card right graphics card and all that speaking of apple silicon should we talk about local models because you're you're in a bit of a unique position and because you have large businesses both in cloud hosted models and locally hosted models that will run on your laptop what are you seeing in those two worlds and what what do you think is going to happen i think it's incredibly exciting because it's similar to the closed versus open model question.

21:02It'll be a mix in our mind. And that's what we hear from customers as well, where for easier tasks, you could run them locally and with lower latency and of course, lower costs when it comes to the per token costs. Ultimately, you're buying hardware up front and you'll use that in conjunction with these cloud models. What's exciting about this next generation of hardware, which we've had for a few years now is just how good they are at running the 20 billion parameter to 40 billion parameter range of models, sometimes up to 120 billion parameters. Yeah. Quinn 3.838B is now as good as Opus 4.6 for coding.

21:37Is that right? That's what the benchmark show. Yeah, that's wild. It's incredibly exciting because you can run that on not the lowest memory MacBook, but the second lowest memory MacBook you can buy from the store. So it's incredible. And are you seeing your users do that? Like, how are you seeing people use the local models versus the cloud hosted models for in practice? From the side of which models they're running, we're seeing a really solid mix of US and Chinese trained models being used for local. And we have the incredible models from, you know, the original llama models, of course, but also the Gemma models from DeepMind.

22:09These are great choices for local. But when it comes down to use cases, coding agents by and large are most effective with the large cloud models. You're solving really hard problems. You're writing code tests. It's really difficult versus some of the document processing workflow use cases that run extremely well locally because they don't have as difficult of a task in the end-to-end, you know, problem you're trying to solve. And so that's where we see this hybrid execution model where some of the easier, more straightforward tasks run locally. And then you have a router that can help decide, hey, we need to go to a large cloud model for this.

22:43And I think what that means for customers is that you're really dropping the costs even further when you're going from open models, not just because they're cheaper to run in the cloud, but because now you can run them effectively for free on the hardware you're buying for your business anyways. And what we're seeing ultimately from the cloud coding agent models is it is predominantly Chinese models being consumed today. And for the local models, it's a really strong blend of US, Europe, and Chinese origin models. Yeah, these two graphs are pretty stunning in comparison. Like basically for local models, the US and Chinese models are neck and neck, we're like tied.

23:18And for cloud hosted models, like the US is like recoloring the X axis. It's like 100 % Chinese models. Basically, we need more US labs to make large models. Is that what this graph is showing? Effectively. And, you know, with the launch of the Nemotron Ultra model, we're seeing kind of the first wave of that. And it's really exciting. NVIDIA as a company is so interesting because their moat is not like trying to start new software businesses or sell, you know, tokens. they seem to be quite interested in just releasing a lot of open source and helping the ecosystem. And then the fact that they do that then helps them stay ahead of the game on the hardware side.

24:00I think so. And, you know, ultimately NVIDIA, what's so incredible is there's helping power an ecosystem around open models, whether that's the hardware, the models. You know, we've seen the new DGX station computers that they're working on. I want one. Which have a GB300 on your desk. How do we get on that list? That isn't deafening loud. Do you know the price point on that thing yet? I don't know it off the bat. It's got to be like Gary's buying it. Well, I looked it up. I mean, you can probably run a frontier model for like, I mean, very slowly for like two, three hundred thousand dollars. Is that right?

24:36I think it's even more competitive than that. And you can run more than a frontier model at high speeds at a price point that isn't very far off what you can buy from a classic workstation computer. Oh, no way. If you think about, you know, quite a few of the customers we talked to, some of them are banks, for example, or industrial businesses. They already have these NVIDIA workstation GPUs in every single engineer's desk. Some of them, tens of thousands of them. And so this is entire. Oh, I want one. They've been selling them for all the CAD work for a while. Well, this is for the kind of original RTX A6000s.

25:09But this is the next generation of that. Yeah, yeah, yeah. I'm going to have to email Jensen. This was at GTC. So what we see here was at GTC. And we were one of the first people, along with Elon and a few others, to receive the DDX Spark as well, which sits on your desk and provides, you know, 128 gigabytes of unified memory to run that kind of 20 to 120 B model range. But obviously, that's just the beginning of a whole new range of hardware that can run the biggest models. So did you say you can buy a bunch of these and chain them and actually run a 400 B model? You can, absolutely. They have this really fast network link.

Read the full transcript

25:43And so you can stack them on your desk, almost like a miniature data center rack. The thing that people do with the Mac minis? Absolutely. Now this is like the production version of it. I think that's what's so exciting is you're seeing both from Apple and NVIDIA, this incredible leap to next generation hardware that's built for these models and is effective at running them. So you'd say like, this is the platform to get, like you could make Apple Studios work, but like if you want something that just can work, get DGX Spark. From our testing, both are very competitive. Got it. So I think a lot of it will come down to what you can get.

26:18And then also the tech stack you're looking for. I think there's an incredibly mature tech stack through the MLX project with Apple where they've done some amazing work to run LLMs on the Mac Studio, but also the smaller Macs. And of course the DGX Spark stack's just incredible. We're super excited as partners with NVIDIA for that. And it's going to be a cool renaissance for personal desktops. I think so. And, you know, it's funny with Olama's journey. We started local. Clearly, the coding agent demand is in the cloud. But that's going to come back locally in our minds because the hardware will catch up when you have a GB300 on your desk and you want the fastest coding loop.

26:58That's as fast as running your tests or as fast as making code editor change. We all remember the GitHub Copile experience of having the autocomplete come up in a few milliseconds, 100 milliseconds. That experience will make its way back to the desk, which has been a journey of starting local, going to the cloud. And then we think that'll come back local and you'll end up using the two together. Speaking of like what you can get in order to run a llama cloud, you need like a shit ton of GPUs. What are you seeing in the GPU market? I think what we're seeing is ultimately the prices are changing very quickly and the supply and demand volatility is very high there.

27:31And so, you know, I think ultimately for if you're a startup, getting access to some of the B200, B300 GPUs you need to run these latest models is very hard. Thankfully, there's a great set of inference providers building on top of that. And so we're seeing this extreme demand. Are you able to get all the GPUs that you need? Are you constantly like growth limited by how many GPUs you can get your hands on? What's the current state? We're lucky in that we've partnered with quite a few providers to work together to pool a bunch of GPUs together, which allows us to stay on top of our demand. But that's a lot of work.

28:09And it's definitely a lot of spending time thinking through, you know, which model will get run where? How fast should it be? Which region is it in? What will the latency be for the customer? There's a lot of hard problems to solve in that stack. And I think what's really exciting about products like Open Router, Olama, the Open Code project is for an end user developer, they can sign up and get access to this without having to go negotiate prices on a B200, B300. You know, think about their 24 month forecast in order to get access to some of these GPUs. YC's next batch is now taking applications.

28:46Got a startup in you? Apply at ycombinator.com slash apply. It's never too early and filling out the app will level up your idea. Okay, back to the video. Suppose you were like a startup founder and you were just starting out now and you're building some AI company and you haven't raised a lot of money. And so you like want to like use as many tokens as possible, like inexpensively. Like what would your advice be to that person about like how they can get like huge mileage with like a limited budget? there's this new class of models like deep seek flash is a great example and i think there'll be quite a few more where it's ultra low cost per token it's also low cost per task which is a really important metric and that class of models in my mind will be the first ones that come down to this idea of like unlimited tokens we all remember chat gpt you didn't really have to think about how many tokens you were using you would just use it every day you had unlimited ultimately i think we return to that but it's going to take a lot of work in the model the architecture to be custom trained for high volume token usage.

29:49And if you think about the start of the year, we really want open models got to the frontier of intelligence. We've bridged the gap where maybe like less than three months behind between the frontier closed models and the open models. But the next problem to solve is extreme efficiency. Seeing, for example, the GBT Luna model become very, very price effective for customers has been a huge boom. We talked a ton of customers where that kind of pricing enables widespread adoption within a team. I think we're see that with open models. We already are seeing that with open models. I think the DeepSeq Flash model is leading that charge.

30:19If we go back to the model breakdown on Alamas Cloud, the highest growth area is definitely the DeepSeq model. And this is largely powered by the DeepSeq Flash adoption. And so this new class of Flash models, where they're good enough for 80 % of the tasks, they're really fast, and they're ultra cheap, this new class of model that I think will enable some of those use cases. Yeah, those are going to be like the workhorse models to do like all the grunt work. Exactly. Yeah. You won't have to be thinking about how many requests am I making, how many tokens you'll be much more inclined to consume as much as you can because, you know, it's able to solve the hardest, not the hardest, but, you know, difficult problems.

30:58If you think back to the coordination layer we were talking about earlier to be able to coordinate these flash models together to do different tasks can also yield great results that a bigger model can. So by having these cheaper models, not only are they more accessible, they can run faster and you can access them in higher volume, but you can start to chain them together and build new problems that are solved by orchestration on top. And that's a really exciting area for new startups, for existing inference providers, for some of the larger businesses today that solve workflow problems. Ultimately, being able to chain these models together is going to be super helpful and you won't have to think about the underlying costs.

31:35Yeah, I guess, you know, when we first started talking about AGI, even on this podcast, there was sort of debate about, you know, and I think a lot of AI researchers would come out and say, like, there's just going to be a giant God model and it's going to do everything. But, you know, I think so far, like, it hasn't quite worked out that way. Like, obviously, you still have, you know, if you have to literally hack the NSA, maybe you need mythos or something. but for the majority of use cases, like you're talking about orchestration and you're talking about like smaller models, you know, that the task composition actually probably gives you a bunch of ways to make it more repeatable.

32:13It's more trustworthy. Like it actually does work at a cost that is like possible. So, you know, if it was going to be a God model versus like lots of, you know, smaller special purpose or even just like simpler models, it's turning out to be the latter so far. I think for most customer use cases, there's a level at which a model becomes good enough and then they can continue using that level of intelligence. Maybe the model will get faster, it'll have better architecture, it'll have new capabilities, but they won't have to reach for the God model. But I do think there are use cases where the most powerful models unlock them and that'll continue to be a thing.

32:53It'll be really exciting, sometimes scary as well on what they can do. But for the run-of-the-mill use cases where open models really shine, I think that's where we're hitting a point where you're not solving necessarily the hardest problems within the business, but they're hard enough where it's now unlocked by open models. Will there be an open model that becomes a god-tier model? I think it's possible. And we're seeing really exciting developments from Zipu AI and GLM where some of the tasks, they are frontier. And we all saw with the Kimi model how for web development, it became the best model.

33:27And that sent this new shockwave across the market, which is it's less about a gap and it's more about a head-to-head competition, which I think makes all of this even much more exciting. This is a bit of a sensitive question, but what do you think about this and the geopolitics around it? I think a lot of the geopolitical angles around this start with where the model's from. And the more we spend time with customers and users, a lot of it's actually how the model's run, where it's run, how it's run. Is it run a secure environment? And that starts to matter a lot more. But I do think, look, it's super important that a customer in the US can use a model trained in the US.

34:10And we have two kind of classes of customers we speak to. One is they don't really care where the model's from. They care about where it's run. But for every one of those, there's a customer that's saying, I really care about where the model's from. Because it's data. the way it's not even just a security issue as much as how does the model speak? You know, we all go through and communicate and we all go through, you know, a lot of the models go through phases where they sound more robotic, they sound more friendly. And a lot of that matters too. But I think the highest order bit is obviously making sure that you have a model that end to end, you understand where the data's from, which is great from the Nemotron models, that you can go and introspect what made this model.

34:50Because if you're putting in a mission critical task, which people are absolutely using open models for mission critical tasks. There's a post online about how Lama powers the analytics of a power plant to detect surges in Finland to make sure that the lights stay on. That's where these models, the model origin really matters. For like critical tasks like that, how do you ensure that a Chinese model, even if it's hosted in the US, isn't basically like booby-trapped to cause problems. The Manchurian candidate problem. Exactly. Have there been any known cases of the Manchurian candidate yet? I think not that I can think of off the top of my head.

35:28I feel like I would have heard about it. What you don't see a lot on some of the press articles is how robust some of the IT and security teams are at the businesses that we know of, the top Fortune 500 businesses. They're really used to this already because open source software, if you think the average application. It has thousands of dependencies. This isn't a new problem. And all it takes is one dependency for there to be a major security issue in the entire application. Yeah. Supply chain poisoning is insane. It's a thing. It's been a thing for decades. And it's not new in that sense. It's a little more opaque because you can't dig into the model.

36:05It is deterministic. But it's deterministic. And if you screen the model properly with safety checks, by and large, at least what we're hearing from customers, is that can be solved. Do you want to talk about the origins of Ollama? You guys came up through the Docker ecosystem, and a lot of people watching would love to be in the position you're in, where you have this sort of enduring brand moat that looks like it will extend for really until the end of time. It's a very powerful situation to be in. You basically found yourself on top of a giant oil well. For those out there wildcatting, can you tell us that story?

36:46You were actually working with Jared in 2021. Yeah, my co-founder and I previously built Docker Desktop while at Docker. So we really got an understanding of what makes a great developer experience. But I have to say the first few years of Olama as a company was really in search for what's the right problem to solve with this muscle we've built of trying to design a great experience for developers. So you applied to YC with a very different idea, right? For sure. Do you remember what the tagline was when you guys applied to YC in Winter 21? I think it wasn't well defined. I think we realized, let's go back to building a really great desktop experience for containers and Kubernetes.

37:28I remember what I wrote down on the application. It was a Kite-matic for Kubernetes or a Docker desktop for Kubernetes. Yeah, which was effectively Docker desktop. They had a great Kubernetes one. I think, you know, it's one of the challenges as a second time founder that, you know, Michael and I have told ourselves, we tried to over-engineer the idea in many ways. And I think even the two to three years after, like Olama, we did YC in 2021 and Olama wasn't launched until July of 2023. After we raised our Series A, after, obviously, after we had done YC, that journey was one of really in search for a customer problem that could delight a developer.

38:07And in some ways, it was almost a good thing that we tried different ideas and pivoted until 2023, because that's when Lama came out and started the open model wave. Which is why it's called Olama. Not necessarily. Okay, no. Oh, really? Llama means generally from our experience, whether you think of local llama, the subreddit, llama really just stands for open models. You know, as we were looking through the name, it wasn't necessarily from an existing model. It's more LLM. Exactly. It's like the animal plus LLM. Yeah. I think having that character was important. So we're like, what's a good name for a character, a face you can put to the name?

38:48Because open models are scary. You need a good animal mascot sometimes. Docker had one, GitHub. He still hasn't taken my advice to have llamas come to actual olama events. Oh my God. I'm curious what your Series A pitch was because you raised from Benchmark, like fantastic investor, but all of this, the future we're in now hadn't quite taken off in 2023. So what was like the pitch and the vision back then? Yeah. And we partnered with Benchmark in 2022. So it was Dolly days, but pre-ChadGBT. When it came down to the pitch, I think we weighed so much on like, hey, we're trying to build this great developer experience.

39:23We're solving this security problem. And we had known Peter, a partner at Benchmark, from our previous lives building at Docker because he was the Series A investor in Docker. And so a lot of it was weighted on the people and also why we exist. I think the what, I mean, solving SSO for Kubernetes, which is a real problem, wasn't really our passion. I think we were really lucky to find a partner that could see us for what we stood for and what we were trying to do versus the point in time problem we were solving at that point. I see. So you raised the A sort of pre-pivot, didn't you? Correct. Okay.

39:57I didn't realize that, actually. I was just looking at this cloud tokens by model family graph. And basically, if you just look at this graph, it looks like the Olama story begins in February 2026. And it explodes thereafter, which is so funny because, of course, it actually goes back to 2021. how was it like to be like sort of lost in the wilderness for like many years working on stuff that was like kind of working but like not really taking off and then all of a sudden to have things like just like explode like how did how did affect you and your co-founder psychology and the team and the employees what what was the experience like it was definitely scary and and for a few reasons you know one is like when you're when you're trying to solve a problem for devs or for a customer and you're just getting on the phone with them over and over again, and it's not totally clicking, that's, you know, it's less about the, are we in the headlines or is the project taking off the product we're building?

40:58It was just, are we truly actually solving a problem for somebody? And I think being lost in the wilderness, like what's your North Star, that customers are generally a great North Star, but not seeing the North Star is even scarier, right? Because often, you know, what problem you want to solve, you just haven't figured out problem. And I think the, you know, Michael, my co-friend and I, we started this company because we built a company in the past and we ended up being acquired by Docker very early. It was just the founding team. And our North Star was saying, we want to go solve a great experience for developers with something they find really hard.

41:27But man, in the two years where we're just finding that problem, it's really scary. You know, we had a team of more than 10 people, which made that really hard. And I'm so thankful to that team for staying by our side as we went through different ideas you know and and what's not really obvious is we went from this security for kubernetes to then like security for developers on the desktop which is like the pivot that we've never spoken about and then we kind of took that form factor when models came out we said well it was a leap but it was we knew kind of the the kind of problem and the feeling a developer wanted to have but lms finally made it realize like it was it was crystal clear at the point when we tried running the llama model and it was really hard and we're like okay this is a problem and it's really impressive when you get it working and it's kind of just a zero to one moment i'm curious for the story of that because there are actually like many pivots in the olama story but probably like the most critical one was like the pivot to olama to doing like locally hosted llms like how did that come about were you just like tinkering with ideas on the side and when you found the idea was it really obvious to everyone in the company that that was the thing to do or was there like like like a big debate and it wasn't until it took off that it became clear we you know sat in a room together i remember we were in toronto because we had a team split across toronto and palo alto and now we're predominantly in palo alto and we were saying throw everything out like if we had to start from scratch and we were just joined yc right now what would we do and you know we had seen two big problems because we had talked to some users and lms we tried using open source lms ourselves lms in general one problem was could you build a gateway to access any model and host that and make that really seamless.

43:06Back then we were thinking of it as like the segment for LLMs. It's a good way to think about it. Which I think has become really this big router idea, which is only at the beginning. It's a massive opportunity. And the other problem was we were a bunch of ex-VMware, ex-Docker folks were like, we know how to make things run. And so like, let's help make things run with open models. And then we kind of - Yeah, systems. And so we kind of tried to really introspect our team, which I wish we had done sooner because security is a very different team and sale than developer tools. And just by doing that, we gravitated towards saying, let's just try this thing.

43:39Let's give ourselves two weeks to launch the first version of Ollama. And then Ollama 2 came out and we said that was right at the end of the two weeks. So we said, okay, we're launching it. And we just had a bias to action. If you think back, like in two weeks, all of that happened, going from idea to shipping it to getting to more users than we had ever had with our previous stuff. And before that was two years of just frankly overthinking the customer, the product, and just not getting something out there. The first time I actually heard about olama was on reddit i didn't realize it was you guys i was on like that logo i just was interested in like running local models and it was on like the i think the local llm subreddit or whatever and everyone was just raving about olama and how great it was i was like oh it's a yc company yeah i found out later actually because you you were called a different company you weren't in our internal system yeah i remember catching up with jared and saying oh hey by the way there's all that security stuff we have this Olama thing now I think you were catching up with Jared and then I bumped into you on the stairs and I think you had your t-shirt or some swag or something and I was like you guys are Olama meeting a rock star or something it would have been really hard to time this but I wish we had taken that leap much sooner I mean the best time to do it was during YC but it was impossible ILMs didn't exist Lama didn't launch yet I guess you were also one of the first GitHub project that very quickly got to 100 ,000 GitHub stars, right?

45:01Do you remember how long? It was like very quick. Yeah, I can't remember exactly how fast, but it was much faster than Docker and Kubernetes. To your point, things kind of just started working and started taking off. And you're really, as a founder, just beside yourself because you can't totally explain why. I think it's the best way to explain product market fit. And there are different levels of product market fit you know we only started monetizing earlier this year with olamas cloud but just to see people fall in love with the product it's such a zero to one moment that um i wish we had done it during yc but in some ways it wasn't possible i also think it's just kind of wild to put into perspective like you sort of went from being in sort of like the cranks on reddit like interested in running their own like rigs at home to like 85 of the fortune 500 in like two years or something like that that's like a pretty that's the homebrew computer club to uh broad computer adoption like speed run that took 10 years for the pc yeah it took like 18 months 12 months yeah and that's one of the things that surprised us the most because i think look i think open models the original user is very much hobbyist just tinkering oh my god this is even possible but very quickly because you know two things one is they're free to get started with and you could run them anywhere.

46:16That is incredibly helpful to a Fortune 500 IT developer team because they don't have to ask for permission to use it. And so it just happened, what was really good for a hobbyist user translated very quickly to a developer within a business. It just happened to be a case where that was it. For example, databases, we saw some of this too, where a database that started for devs like MongoDB very quickly also moved to enterprise. but because LLMs are stateless, it made for such an easy transition. Now moving to the cloud, there's a lot more in play. There's an economic question if you're a customer.

46:53There's obviously security. Where's the model running? But what's beautiful about open models that both hobbyists and IT developers loved is you could just get started. You didn't need permission. Can we talk about the monetization angle? Because this is interesting too. So like in 2023, Llama 2 takes off. All of a sudden, you've got all these users, 100 ,000 GitHub stars. is like you've clearly found something, but it's basically like Reddit cranks who are using it. You're making no revenue and there's no obvious path for how you will ever make any revenue from all these like cranks on Reddit.

47:23It was two years before you actually figured out a business model for it, which funny enough is exactly the position that Docker was in. Like, how did you think about it during those two years? Were you worried about it? Was the team asking like, what's the business model going to be? How did you think about like figuring out how to make money from it? I think there's always two ways that we saw open models being able to monetize in a way that's great for the company, great for the developer, and great for the customer. And one of them was a privacy-focused AI product, which Ollama really started with that in its open source incarnation.

48:00But we always felt that there was this moment where, you know, you weren't using LAMA with the LAMA models, for example, with tool calling right away when they came out. So there were use cases where it was still reserved for the frontier models. And again, at the risk of overthinking it, we kind of saw that there wasn't the level of product market fit with open models that closed models had. And in some ways, philosophically, we want to align with when that happens, we want to be there to capture that. I think it happened this year with coding agents running with open models because you had the largest consumption of AI being matched with finally open models being able to service that.

48:35There are a lot of opportunities along the way to do it privately, securely. Again, a lot of the Fortune 500 have already adopted Lama. But we really asked ourselves, what would be the most important problem we could solve for a customer? And the local piece, while an important part of that story, never felt like the whole story, which was how do you access open models for the hardest problems? And so in some ways waiting, we knew we had to wait a little bit for the market to mature. At the same time, what are the risks of waiting? Well, you build a culture, if not careful, and we had learned a lot of this from our Docker days, where you don't think about monetization.

49:11It's not a priority. I think from our previous battle scars as a team, we kind of had, we knew about that. But I think the other component, which is really important, is making sure you keep in touch with your customers. One of the biggest risks of having an open source project that takes off is you consider your user base and customer base, your customer, just a blob on the internet, which is a really risky way to think about customers because you want to meet them, figure out their needs. What are they doing? What do they want to do in six months? What's their story? And I think that's the thing I wish we had done a little more in the last few years.

49:41And we're doing a ton of that now. One thing I'm curious, when you went through YC, you guys were second time founders. I'm curious what got you to decide to do YC, actually. You know, we went back and forth on this for a lot, which we shouldn't have. We should have just said, of course, we're doing YC. But by and large, starting a company is a really lonely experience. Even if you have a great co-founder, and Michael, my co-founder was the co-founder of my first company. He was my college roommate at University of Waterloo. But it's still lonely. And I think just having a set of peers, even though we did it during the pandemic, just talking to Jared and like five other groups of founders every week really helped you feel less lonely.

50:22And I think that's such an important part of it. And then, of course, when we finally moved down here and there was no more COVID, the network was just incredible. And the fact that we could meet founders building on open models, building on any kind of AI, we kind of knew that was going to happen because we had known so many founders from the University of Waterloo who had done YC pre-COVID. And they were like, it's really about getting together. And that was a big part of it. And we knew that was there. And, you know, I think that's what made it a no-brainer. But also just, I think there are a lot of mistakes you can repeat that you don't have to.

50:59And what I love about the YC community is how transparent founders are with each other about those. And, you know, I still keep in touch with the founder of Docker, who's an investor in our company. And we're able to talk about some of these challenges we saw in the previous generation of companies that, you know, we don't necessarily have to repeat or things that worked and we can bring, you know, into the future. Yeah. If you just don't repeat one of those mistakes that, you know, sometimes is the mistake that would have killed the company. Potentially. Yeah. Yeah. Our famous saying, you know, a bunch of our team is from companies that ended up working great.

51:31Docker's doing phenomenal now. but whether it's you know some of our team was early early at VMware and it there are always ups and downs and I think just having a group of people around the table who have a collection of those and also what worked actually what worked is actually even more important and just being able to like have that muscle memory is a big part of it oh man I was just thinking about this because we obviously hang out with work with a lot of 18 year olds or 19 year olds and then sometimes they're always asking like well what should I do and then I'm starting to realize like one of the more important things is if you've never worked on a team that shipped really amazing technology to like a lot of people or like like just real clear product market fit like do that once like even if it's a month even if it's like three months you would learn more in those three months because then you know what good looks like and then without that it's like i mean it's not like it's impossible like people at yc do figure it out because you know but it's that much harder Like the difference between having seen something that actually works from like beginning to like some form of like, this is what the bug database looks like.

52:36And this is how we release. And this is the quality that's necessary. And here's like the bar that we hold each other to. Having seen that, it just like multiplies the chance that people succeed. So it makes sense that, you know, starting off with a co-founding team that has seen a lot of that, pretty powerful. Yeah, I think it provides you a set of values you can work around, especially when you have so much power in your hands with AI. There are just parts of it that AI can help you with, but it won't hold you accountable to it. And how does software work? And to look at Ollama, for example, I'm sure there's versions of Ollama running in the wild from two years ago.

53:08How will your software work when somebody falls in love with it and continues using it for two years? Is it still going to be working well? Hopefully they update to the latest software or it's a cloud service. But I think you build that muscle memory. And we definitely have that from a lot of our more senior engineers on the team who are at VMware or NYSERA, for example. But at the same time, I think there are a lot of lessons we learned in the previous generation of DevOps and infrastructure that aren't valid anymore in the AI world. Oh, yeah. Tell us about it. What have you found? What is not valid anymore?

53:37I think a good example that I classically used is there was this generation of companies called Platform as a Service. in the 2010s. The Herokus of the world was a great example of this. I mean, Docker started out as that. Docker started out as a platform as a service. And there's this concept that if you're a layer on top of something else, that you're in kind of a vulnerable position as a startup, which is absolutely not true in the iWorld. And in fact, going up the stack can sometimes be even better because you're closer to the customer. In an infrastructure world, that's also the case. And that was like an analogy that we had to like, so many of these muscles, we actually had to break building Ollama.

54:14Another one was, you know, these LMs are never perfect. And like in the systems world, you want everything to be exactly as it's designed to run. It's tested. It's validated. But LMs by definition are not. That's a feature, not a bug. Exactly. It's a feature. You want it to be a little non-deterministic, I suppose. And I think from building a team too, it's that, you know, with AI now, there are just problems that you don't need to staff as heavily, whereas you did 10 years ago, right? If you think about what does your customer support pipeline look like? What does it look like to deliver a cloud service?

54:48It's a very different world with AI because how do you build a service where no engineer knows exactly how all the code works, which is obviously the case now? And so there's just new lessons we're learning going from some of our team from infrastructure 1.0 in the 2000s to cloud in the 2010s to now the AI space. There are a lot of rules that break. I mean, you're probably actually doing an incredible service to both sides of the ecosystem and that the end users get this very clean thing that just works, especially the tokens just come out and they're very clean and the API makes sense and it's rational and logical.

55:24and then on the flip side like i mean if you don't have a layer like olama i've directly experienced this where it's like oh yeah the underlying inference provider has a weird error for you know if you put this parameter in this way or it expects jason and you know it's not documented it's just like this insane minefield like you know the agents can kind of figure it out but like you're going to like bang your head into the wall for like a couple hours before you know the agent figures it out and in the meantime you're like this is a terrible experience you know and so you're like in there probably helping the inference providers fix all these fundamental bugs too yeah it's part of the the job we do and i think one of the big opportunities in the open model landscape is curation and taking a fragmented universe of models and inference technology and cloud services and harnesses and like making that actually just work is a really valuable problem because the end developer, to your point, they just want to build their software, right?

56:21They just want to build stuff. They want to build their next company, their next application. And I think that's where we come in, but it's where a ton of great services also come in. And we saw, you know, Open Router, obviously, is a good example of that from a wide model selection. So the developer doesn't have to sign up for, you know, 100 different providers. They can just go to one. They can pay in one place. I think we've seen with OpenCode, you know, you can have one harness that integrates with any model. It's a really powerful experience for developers is just looking to try the next model to see if it solves their use case better.

56:50So this curation and, you know, when there's a abundance of models and providers, now there's a scarcity in bringing that together into something that works. Thank you so much for joining us. That's all we have time for. Thank you guys for having me.

From the publisher

Ollama (YC W21) is used by 9 million developers and 85% of the Fortune 500, giving co-founder and CEO Jeffrey Morgan a unique view into which AI models people are actually using and how that’s changing.Right now, the biggest shift he sees is toward open models, driven by coding agents, falling costs, and capabilities that are rapidly catching up to the frontier labs. On Ollama Cloud, that shift has driven a 150x increase in token usage since the start of the year.In this episode of the Lightcone, Jeff joins us to talk about the future of open models and the story behind Ollama, from two years of searching for the right idea to building one of the most widely used AI developer tools in the world.

More from Y Combinator Startup Podcast

All 148 episodes
Open Models Change The Economics of AIY Combinator Startup Podcast · 57 min
Listen in VO