In short
Podcast Episode Summary: Fireworks Founder Lin Qiao on How Fast Inference and Small Models Will Benefit Businesses
Podcast Information
- Title: Training Data
- Description: Conversations with leading AI builders and researchers to explore the evolving technologies of AI and their implications for technology, business, and society.
- Episode Title: Fireworks Founder Lin Qiao on How Fast Inference and Small Models Will Benefit Businesses
- Hosts: Sonya Huang and Pat Grady, Sequoia Capital
Episode Overview In this episode, Lin Qiao, the founder and CEO of Fireworks, discusses her journey from leading the PyTorch team at Meta to founding Fireworks, a platform aimed at accelerating AI inference for businesses. The conversation delves into the challenges of AI inference, the importance of simplicity, and the future of open-source versus closed-source models.
Key Topics Discussed
Introduction to Fireworks
- Fireworks Mission: To reduce the timeframe from AI training and inference from years to weeks or even days, democratizing access to AI beyond major tech players.
- Platform Overview: A SaaS platform focusing on generative AI inference with low latency and cost, tailored for enterprises.
Lin Qiao's Background
- Experience at Meta: Led the PyTorch team, which involved substantial rebuilding to meet complex AI requirements.
- Transition to Fireworks: Aiming to leverage her experience to create a platform that accelerates AI deployment.
AI Inference Challenges
- Current Landscape: Startups often rely on closed-source models from companies like OpenAI but face latency and cost challenges as they scale.
- Importance of Inference: The shift from training models to deploying them effectively is crucial for businesses.
Open Source vs. Closed Source
- PyTorch's Dominance: PyTorch is preferred for its simplicity and effectiveness in research and production, making it difficult for companies to switch frameworks.
- Fireworks' Differentiation: While many platforms are framework agnostic, Fireworks focuses on optimizing PyTorch for better performance and easier integration.
The Concept of "Simplicity Scales"
- Design Philosophy: The idea that complexity can be managed by creating a simpler user experience while handling the underlying complexities.
- Conservation of Complexity: All applications have inherent complexities that must be managed; Fireworks aims to keep this complexity away from users.
Predictions and Future of AI Models
- Convergence of Open and Closed Models: Qiao predicts that quality differences between open and closed-source models will decrease, emphasizing the importance of customization for business applications.
- Small Models Phenomenon: Smaller, specialized models can be more efficient and easier to customize compared to larger, monolithic models.
Customer Journey
- AI Adoption Trends: Customers typically start with powerful models for experimentation and seek better solutions as they scale, focusing on latency and cost reduction.
- Engagement with Enterprises: A growing number of traditional enterprises are now seeking Fireworks for innovative AI solutions, reflecting a shift in their approach to technology adoption.
Future Vision for Fireworks
- Totality of Knowledge Access: Long-term vision includes creating simple API access to a comprehensive range of models and knowledge, including both public and private APIs.
- Function Calling Model: A strategy to enable seamless access and routing to different models and APIs, enhancing the user experience.
Competition Landscape
- Nvidia and Other Competitors: Discussion on the competitive landscape suggests that while Nvidia has a strong position, competition is expected to grow.
- Returns to Scale: Acknowledgment that as model innovations stabilize, the focus may shift towards optimization and application of capabilities.
Episode Highlights
- Discussion on Latency and Cost: Many startups face challenges with responsiveness, pushing them towards Fireworks for better solutions.
- Insights on Fine Tuning: The complexity of fine-tuning AI models and the need for automated solutions to streamline this process.
- Lightning Round: Quick insights on favorite AI applications (Fathom for summarization), predictions for AI models, and admiration for Meta's commitment to open-source.
Conclusion Lin Qiao's insights provide a valuable perspective on the intersection of AI technology and business needs, highlighting the importance of speed, cost-effectiveness, and user-centric design in the evolving landscape of AI inference solutions.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00We thought replacing other frameworks as library with PyTorges to be simple. It's just swap the library. How hard can that be? But we realized it's just a six month project. It turns out to be a five year project for us. To support entire matters, AI workload, building on top of PyTorges, because we have to rebuild the whole entire staff. From scratch, from ground up, because we have to think about how to load data efficiently, how to do distributed inference in PyTorch efficiently, how to scale training efficiently, and then we will end up reviewing the whole entire inference in Chinese tag on top of PyTorch.
0:45When we left, it was sustaining more than 5 trillion inference per day. So that's a kind of massive scale by two -surfer years. And the firewalls mission is to significantly accelerate time to market for the whole entire industry. Compressing it from five years to five weeks or even five days is time to market. So that's our mission.
1:24Joining us today is Lin Tiao, founder and CEO of Fireworks. Lin is an AI infrastructure heavyweight who previously led Pi Torch at Meta, which is the backbone of the entire global machine learning ecosystem. She's taken her experiences at Pi Torch in order to build fireworks in inference platform for gender -to -vana. We're excited to ask Lin about the market trends behind AI inference and how she plans to support and even accelerate the market shift to compound AI systems at fireworks. We're thrilled to have Lynn, CEO and founder of Fireworks with us today. Thanks for joining us, Lynn. Thanks for having me.
2:00We're really excited to talk about a lot of things with you today from PyTorch to the small model stack that you're building to what you're seeing in terms of enterprises building production deployments. But before we get there, can you maybe say a sentence or two on what you're building at fireworks. Yeah, so we started fireworks in 2022 and fireworks is a SaaS platform, first and foremost, for Genoa AI inference and high quality tuning, especially using our small model stack. We can get to very low latency for real -time applications, very low cost for sustainable basis growth, and customization, automated customization full tailored high quality for enterprises.
2:42So that is fireworks. Wonderful. I want to maybe start with the PyTorch story. PyTorch is kind of the foundation upon which the entire AI industry runs today. And you and Dima and some of your other co -founders were integral and leaders of that project at MENA. So think about PyTorch as the programming language for digital brains. things. And it's designed for researchers to very easily create those digital brains and experiment with it. The challenge of PyTorch is it's very fast with people to create various different deep learning models, the digital brains, but the brains don't think fast enough.
3:27So that's a challenge I took out to address while I was at all PyTorch. And you mentioned, you mentioned before or that most of the companies that are trying to build something similar to what your building and fireworks have chosen to be framework agnostic, whereas you very much made a big bet on PyTorch. Can you say why make the big bet on PyTorch and what benefits that brings to your customers? That is really based on what I see when I operate PyTorch and matter also across the industry. And I clearly see a fair no effect that PyTorch is because it's starting that as a tool for researchers, it starts to dominate the top of the funnel for model creation.
4:10And then the next stage of the funnel is people doing applied production work, they take those research models and test them out for production setting and try to validate that hypothesis and then feeding to production. So that's clear for no effects that's happening. And as PyTorch is de -platform research, it takes over the top of the funnel. It's really hard for people to rewrite into other frameworks for production, and naturally it just flows down towards the bottom of the funnel. And that's how PyTorch become dominant. And I'm start to seek more and more models, especially in more nation models, are all built in PyTorch and run in PyTorch and production, including the January A .I.
4:53models. That's why we only bat on pytorch and we don't want to distract us or to support other things. So researchers like it and it flows downstream from there. What are researchers like so much about pytorch? Simplicity. Simplicity scales. And that's kind of lesson learned through the journey of pytorch and met up and also building on the community. It is a relentless journey to focus on simplicity. And we have a constant seeking how to make the user experience simpler and higher and more complex in the back end. For example, when I started this journey and met up, there are three different frameworks.
5:34Coffee 2 for mobile, onexport, server -side production, pytoch for research is too complicated. And the mission is to not reduce three frameworks into one to simplify, but it's actually a mission impossible. After I consolidated all three teams and there's no consensus to how to simplify and build this one stack. And we took a very idealistic approach and take the PyTouch front end and take the Kapa2 back end and we said we're going to zip them together. It seems simple, but it's very hard to do because these two frameworks are never designed to work together. And the integration complexity is even much higher than build a framework from scratch.
6:20So two complex. And then we said, forget about it. We're going to all young high torches. Keep its beautiful simple front end and rebuild the back end. So we build Torch script, that's pad torch 1 .0. So that's really like the key focus on simplicity wins over time The other interesting thing is we thought Replacing other frameworks as library with PyTorges could be simple. It's just swap the library how hard can that be But we will like realize hits we thought it's just a six month project. It turns out to be a five -year project for us To support entire matters as AI workload building on top of a high torch, because we have to rebuild a whole entire stack, from scratch, from ground up, because we have to think about how to load data efficiently, how to do distributed inference in high torch efficiently, how to scale training efficiently, and then we end up rebuilding the whole entire inference in Chinese stack on top of a high torch.
7:26When we left, it was sustaining more than five trillion in France per day. So that's a kind of massive scale by two five years and the fireworks mission is to significant accelerate time to market for the whole entire industry. Compressing it from five years to five weeks or even five days is time to market. So that's our mission. Maybe when you look at the open source standards, there's a lot of people that are trying to do it on using VLLM or Tensurati LLM. How do you think about how fireworks compares to to what's in the open source. I really like both projects, and because my heart is in open source based on, yeah, I hear pipe torch experience.
8:07I will say both projects are great projects for the community. I think our biggest differentiation is, first of all, fireworks off the shelf is faster than both of the offerings. And second is, we're building a system where not just a library. And our system can auto tune towards our developers or enterprise workload to be much, much faster and to be much, much higher quality. And that cannot be achieved by just a library. And we're building all these complicity going back again to our January pie torch. which we are providing a very simple API but hiding a lot of automation, the complexity of automation, complexity of auto tuning behind the scene.
8:59For example, when we deliver our inference with high performance, high performance here means low latency and low cost, we handwritten Kura -Kunos. We implemented distributed inference across nodes and this aggregated inference across GPUs where we chop models into pieces and scale them differently. We also implemented semantic caching where given the content, we don't have to recompute. And we capture application workload patterns specifically and we build into our inference tag. We have many other optimization we are, we have been specific design for different use cases. That is not like general purpose or horizontal.
9:56So that is being encapsulated. We also have complex optimization for quantitative. You can think about a quantitative, just one technology or a heart that cannot be. but you can quantize so many different things. You can quantize KVCache, you can quantize ways, you can quantize communication across GPUs, across nodes, and EO different performance gain and quality trade -offs. We also automate like colloquial optimization. There are many things we're doing behind the scene to deliver a very simple experience to the app developers. So they innovate, this concentrate their cognitive bandwidth to innovate on the application side.
10:37You, I like your comment earlier about simplicity scales. And as you're talking through everything that you've built to make this such a simple and delightful experience for your customers, it reminds me of the idea of conservation of complexity, you know, like the amount of complexity required to deliver any given task can be neither created nor destroyed. It's just a question of who takes the complexity. That's right. And it feels like yours is a business where you have embraced an enormous amount of complexity to make life simple for your customers. And actually my question is about your customers.
11:08So where in the AI journey of your customers, where are their AI journey, do they say, wait a minute, we need something better. And then what brings them to you? Yeah, so we've seen pretty consistent pattern that last year people all start for OpenAI. Because they are in the heavy experimentation, exploration mode, many start up there, they have some creative ideas, application part ideas, and they want to explore product market fit. So they want to start with the most powerful model where OpenAir provides. And then when they, they feel confident they hit product market fit, they want to scale their business and then problem comes in because as I mentioned most of the genetic application there be to see consumer -posomeric developer facing it requires very high responsiveness.
12:04New law didn't see as a critical part of probability without that they don't it's not a viable part people are not patient enough to wait for half a minute for response that's not gonna work so they are seeking actively seeking law they can seek and then another key factor is they want to rebuild a sustainable viable business, they cannot bankrupt quickly. And the weird thing is... Not this market they can't. The weird thing is if they hit a viable product, that means they can scale quickly. And if they're losing money at a small scale, they're gonna bankrupt quickly. Right, so bring down the total cost ownership is critical for them.
12:46So that's why they come to us. So it sounds like, I remember you had this inside a year or so ago, we spoke about, you know, training tends to scale in proportion to number of researchers that you have, whereas inference tends to scale in proportion to the number of customers that you have. And in the long term, probably going to be more customers of AI products than AI researchers out there, and therefore inference is the place to be. It sounds like you're kind of, the customer journey sort of begins as people are going from training into inference. What sort of applications, what sort of companies are at that point where they're starting to really go into production.
13:23There are so many ways to answer these questions, the very interesting question. So first of all, my hypothesis when I start a company is we're going to take our startups first because they're most tech -abest. They will be ton of startup build on top of a GNII. Then we will go to dishonative enterprises because they are tech -forward. And then we'll go to traditional enterprise because they are like tech conservative. They want to observe and adopt when the technology and the product ideas are mature. So that's kind of my hypothesis. And it totally blew my mind. What's happening right now? Because we have a lot of inbound of startups.
14:08We are working with degenerative enterprises. We're also simultaneously working with traditional enterprises, including the health insurance company, healthcare companies, and banks. And especially for those traditional enterprise, usually I adjust my pitch to be very business oriented because hey, that's kind of my, maybe my bias and kind of want to strike a meaningful conversation with them, but they quickly dive into very low -level details, technical details with me. And it's very, very engaging. What are the people doing like at a traditional enterprise? is who are the people that you're engaging with?
14:45So we, really is it innovation? Innovation person, AI person, or is it more business line leader, somebody who owns a production application? Yeah, so I think it's start to shift. We are more engaging in start with CTOs. I feel like this business is shifting towards innovation driven business transformation. and that's why we encounter more CTOS than like the CLOs or other CSOs. So that's kind of an interesting shift. But yeah, across the board, I think there are multiple fundamental reasons that why that's happening. That's my hypothesis. One is, all the leaders realize current genoa wave is similar to the cloud first shift or similar to the mobile first shift.
15:34It's going to remap the landscape of industry. startups are growing really fast and the incumbents feel fattened if they are not innovating fast enough they will be obsolete they will be irrelevant but also across the incumbents they are heavily competing with each other they're competing how fast they kind of transition their business to be to quit more revenue to be more efficient using genNI so that's one phenomenon the second phenomenon is genNI is different from traditional AI I always say kind of this This is very different. Tradition AI is given a lot of power to hyper -scalers. Because traditional AI is you always have to train from scratch.
16:19There's no concept of foundation model you build on top of. And that means you have to go off and curate all the data. And the data rich company usually they are hyper -scalers. And you need to have a lot of resource investment to change your own models and so on. So that is before JNI, it's less affordable, it's concentrated in hyper -scalers. Post JNI, because of the concept of foundation models, people build on top of foundation models, and you don't train from scratch. It's not meaningful. It's all the same data. It's all the internet data you can crawl. It's a small less similar model architecture.
16:58It's a waste of resource if you train from scratch. Instead, you find two. You tune based on your small data set, high quality small data set, so it becomes a small data, small model problem. And it makes it so much affordable to everyone to assess this technology. And that's why everyone is jumping in to embrace it. How many of your customers are using you for fine tuning versus just using base model? And what do you think goes into building a great fine tuning product? It really depends on the problem they're trying to solve. We actually see the open source model is becoming better and better.
17:36The quality difference between open source model and the closed source model are shrinking. And my prediction is going to converge at the same model size. If you go open and closed, open merge. The open and closed will converge. Do you think there will be a time lag where closed is always six months ahead? Or do you think there will just be neck and neck? For the same model size, especially like between 7 to 70 billion, or even within 100 billion model size, the quality will converge. That's my prediction. We'll see after a couple of years and we will come back to this podcast and see how it goes.
18:15So the key here is customization. right, given like if this trend is true, then the key differentiation is how we customize those models towards individuals use case towards individuals workload. And is it easier to customize an open source model than a close source model or time -fitting model? So I want to say it's easier. It's just open source model tend to have a much richer community. And there are a lot more people working on building on top of those models. For example, a lot more three model. So it's a very, very good base model. It is a very strong instruction following. If all the instruction is very well, so it's very easy to use it to align the model to solve a specific problem really well.
19:01And for example, we have been investing in function calling strategically as a direction. We can talk more about that. It's an old topic by itself. But we find like fine to find to a function calling model on top of Lamato is so much easier compared with Fintening based on mixture models or previous Lamato and and other models So that's just kind of the base model open source base model is becoming very very strong in Instruction following in logical reasoning in many other base capability So it's very easy to work it to become a high performance model for service specific business task That's a power of small models.
19:42If we think about just open source software, open source, FHRSHR software, 20 years ago, open source was thought of as a fast follower sort of thing, you know, Red Hat being a canonical example. And then more recently, open source is not the fast follower, it's actually the innovator. You're needing about mango or confluent or some of these other great open source businesses that have been created. Do you think there are areas in the world of models where open source is actually and a lead close source and it is actually going to be ahead of the proprietary models. So I think that dynamics is very interesting right now because the proprietary close source model provided that they are batting on very few models, right.
20:25So open AI's LM models are like maybe three models, right. Or you can think about that as one model because what is a model model is the model architecture and data, 20 data, right. That defines the model. So I'm pretty sure they use all the models they have more or less similar training data, model architectures more or less similar so it's kind of scaled all parameters and so on. It's not just open to AI I think, and to opaque or mistro and so all these kind of model builders they have to concentrate the effort to focus on a specific like model segment And that's kind of the best RID. That's a bit small.
21:08But open source push a different dynamics because it enables so many researchers to build on top of it. So that's kind of the small model phenomenon. It's smaller and it's easier to tune, easier to improve quality, easier to focus on specific problem space. So it enables thousands of flowers to blossom. Thousands of flowers blossom. So that's the direction we believe in is to solve an enterprise problem, a thousand flower blossom is much better for enterprise. Because you just have so many problems and I bad to add, given any problem space, there is a solution for you and we further customize towards your use case in your workload.
22:02And what you get is bad quality, much lower latency for real -time application, much lower cost for business sustainability and growth. So we believe in that direction. Maybe to that point, have you seen your customers of our are able to match the quality that they got with OpenAI when they move over to the fireworks stack and like how are you enabling the what I call the small but mighty stack to compete. Yeah, so you really think pants on domain. So for some domain, actually people don't even find two. They use an off -shelf model as is. And it's already very, very, very good. For example, in the domain of coding call palette, code generation, transcription, translation, OCR, it's just phenomenal.
22:53Those models are really, really good. So that's kind of off -shelf and ready to go. But for some areas, it requires business logic, because every company is defining what is good differently. And then, of course, like, off -the -shelf model will not work, off -shelf, because they don't understand the business logic. For example, classification, different company want to classify, hey, you know, some marketplace want to classify whether it's a furniture or it's a dish or it's something else that's completely depends on their domain or summarization, you think summarizing is a very general task, but For example, insurance company want to summarize into a very special template Right, so and they're there are specific business tasks on Yeah, on many other things, we just kind of work with across the board, various different problems, and those requires fine tuning.
23:56And I want to call out fine tuning sounds simple, but it's actually not simple at all. So the end to end requires enterprise or developers to collect data for the end to trace. After they traced into label. After the label, they need to pick and choose which fine tuning algorithm to use. They're supervised fine tuning, they're STPO, they're slow or preference based fine tuning. As in, they don't label absolute good result. They basically say, I prefer this over that. They need to pick whether they want to use parameter efficient fine tuning like Laura or for full modified tuning. And for some tasks, they need to tune hyper parameters, not just the model weights itself.
24:49So among these many technology, they have to kind of figure out when to use water so it's very deep. And usually those app developers haven't even touched AI yet, and this law for them to pick up. And then once they tune and they test it, it's still improving some dimension. if it's still not good in some other cases and then they need to capture those failure cases and analyze, should I collect more data and go to this cycle again? Or it's actually a product design, right? It's very interesting. Some failure cases, not really failure cases, it's just they haven't designed what the product should react.
25:26For example, people are building in a system to auto generate content when people type. And if you're in a table and you're in a cursor in a cell and what does auto generate mean? Do you auto extend? Would you type in the cell? Do you generate more rows? Or you do nothing? So it's such a product design. So that requires a PM to be in the loop to think about the failure cases. So with all this complexity, what we want to do is take away the rudimentary stuff, take away the complexity of figure out which tuning approach to use, how to automatically label data, how to automatically collect data from production, want to take away all this and keep a simple API for people to use, but leave the design part to our end user.
26:17For example, how the product should respond should completely in their realm to figure out and solve, so that we want to kind of create that separation. And then we started working in this direction and hopefully we'll announce our part of there soon. I love that you're kind of liberating people to not have to think from the tech out and to actually think from the customer back and sort of use all the stuff that you've built to deal with the underlying technology and really focus on to your point the design patterns and the usability and making sure that they're actually solving an important problem in the compelling way.
26:51What is your vision for the fireworks platform? And to Pat's point earlier on conservation of complexity, we started this podcast talking about how you're conserving complexity for your customers on the inference stack. You just now talked about how you're conserving complexity for your customers in terms of the fine tuning work flows. What are the other pieces that have to come together? And what is your ultimate vision for what fireworks the product is? If everything works five years from now, what will you have built?
27:21So what we, like the North Star for fireworks is this simple API access to the totality of knowledge. So right now we're building towards that direction. We already have more than 100 models. We provide across large language models, image generation models, audio generation models, video generation models, embody models, and multimodal models as an image as the input to extract information. So that's kind of one side of the foundation model coverage, but put all the foundation model together, it still have limited knowledge, but because it's training data is limited. It's training data, it has a starting time, and in time, all the information they can crawl on the internet is still limited because there are a lot of knowledge that's hidden behind APIs.
28:15Hidden behind the public APIs that you don't have access to or you just cannot get real -time information. There are a ton of private APIs hosted with the enterprises. No way anybody will have access outside of the companies. So then what how do we get access to the totality of the knowledge for the enterprises is to have a layer to blend across many different models and public private APIs. So that's kind of the vision and the tool to the vehicle to get there is function calling. It's the function calling model. Basically this model is capable of understanding here at APIs you want to access and for what reason, it can automatically be the router to most precisely call out to those APIs, whether it's accessing models or accessing non -moder APIs in the most accurate way.
29:21So think about strategically that's extremely important to build these simplified user experience because then our customer, they don't need to figure out. They don't need to scratch their head and figure out All I need to find tune to be able to access those APIs and how to even do that myself is kind of a tall order for me to do. So we want to basically, you can think about that because many people are familiar with a notion called mixture of expert. So open AI is providing mixture expert and mixture expert becomes a very popular model architecture. The cost app has a router, set down top of few very big experts and each is specialized and it's something.
Read the full transcript
30:05And our vision is we want to be the mixture expert that access hundreds of those experts. And each of those experts are much smaller, much agile, but with high quality of solving specific problems. In that vision real quick, those experts live in fireworks, in AWS, in hugging face, like where do those experts come from they get put together with fireworks is the overarching framework? Yeah, our ambition is those experts living fireworks. That's where we want to curate, curate towards curate models we serve towards that. And that's where today we already have more than 100 models. And it will take some time to kind of build this layer in a very solid way, but working to release our next generation of function calling model.
30:59It's really, really good. A little preview on that. It has multiple layers or breakthroughs. We're going to announce it together with demos and example, and people can leverage and build on top of. Very cool. Do you see any viable competition for Nvidia on their horizon? That's a very interesting question. First of all, I think Nvidia is operating a very lookative market, and any lookative market invites competition. This is just the economics here. And also from the whole entire industry point of view, in general industry doesn't like monopoly. So that's kind of another trend like pressure coming from the industry.
31:55So I think it's just, it's not a question whether there will be competition to Nvidia's question of when. Do you think it's coming soon? I think it's coming soon. I think it's coming soon. So I think I mean, obviously, we can look at Nvidia's competition in multiple segments on the general purpose competition segments GPU that MD is coming up. That's interesting. I think also in a specific AI segment where the AM model space is stabilized. There's no more innovation. The problem is well defined and this is the model. Then custom -based will have its own role. So I think I will look at the market that way and I do think there will be competition commissum.
32:46Can I see about that by the way because you guys are in this part of the market where you are model agnostic to some degree and it's really about the optimization of those models when it comes to put them into production. Do you think that the returns to scale on the frontier, the models that are out on the waiting edge. Do you think the returns to scale are starting to slow down? Do you think that we're going to go into a phase where capabilities have started to mature or asymptote and the race is more about the optimization and tuning and application of those capabilities? I think both will happen at the same time.
33:24One is it will start to stabilize and plot two in the model app capability point of view and will heavily customize. Our strategies is heavily customized towards the use cases and workloads. So that's one direction. And the second is I want to caution that, because that matter was a thing for certain period of time that that is a model for ranking recommendation. And we should heavily index on that assumption. But then after a few years, it's not a case. There's significant amount of model innovation in seemingly stabilized modeling space And that pushed the S curve for a matter. I think same phenomenon will happen in the gen I space that a new model architecture Will happen and we're kind of overdue See meant we've talked about competition from other other vendors other direct competitors What about OpenAI?
34:32Does OpenAI keep you up at night? Like they drop prices on their APIs all the time. They're making their models. They are also trying to win the better faster cheaper race. Like how do you think about them? And how do you think about ultimately, what you're gonna build that's different from where they are going? Right, so again, I feel like for the, they are actually going smaller, right? And cheaper, I think for the same model size, For the same model bucket, whether it's closed source or open source, the quality is gonna come out. That's again, that's my prediction. And the real meaningful thing is to push the boundary here is have a customization or automated customization, tailored towards individual use case and the individual workload.
35:23I'm not sure if OpenAI has the appetite to do it because their mission is AGI. if they hold their mission, which is a great mission actually, but it's kind of solving a different problem than solving an enterprise problem, which basically means there are a lot of problems, a lot of specific problems that is really good for the small models to customize towards. And that's where we want to focus on our energy and build on top of open source models, assuming the trend that they are going to converge in quality. I love that. Our prayer reloft, less time you were here, made the point that, you know, in prior technology waves, internet, mobile, it was the people that did all the hard work driving down the marginal cost of running this stuff that actually enabled all the application development on top and all the end use cases that we get to enjoy every day.
36:14And I love that you were taking that exact approach with AI where it's still so cost prohibitive for most people to run a reduction and by just dramatically bringing down that cost curve of actually in the whole industry blossom. So it's really wonderful. Should we close out with some rabid -fire questions? Yeah, let's do it. OK, you can go first. No, go for it. OK, let's see. Favorite AI app. We do a lot of video conferencing, and the notes taker for video conferencing is a game changer for us. Whatever it is, there's so many different varieties, but I just love that. Which one do you use? I think we use fat on.
36:52Yeah, I was able to use that. It's really good for training and also summarization. So the infinite short hour time. Thanks. Well, we're the best performing models in 2024.
37:07My prediction is there will be many given the rate that every week, every week, there's a new So this is all good news for the whole entire industry. It's really hard to predict which one. But the one production I'm pretty confident is the model quality is well keeping proving and keeping increasing. And the world of AI, who do you admire most? I always say matter. It's not one person. But that matters commitment to open source. I think matters the most brilliant in the journey of JNII by continuous open source, a series of long models. And continue to push a boundary, continue to kind of shrinking the quality differences.
38:07So okay, so what matters doing is basically decentralized power, fun hyper -skillers, to everybody who has a dream, to innovate, foundation models, Janie models. I think that's really brilliant. Well, agents, performer, disappoint this year. I'm very bullish on agents. I think it's going to blossom. That's all we got. All right. Thank you. And thank you. And he's really fun to have this conversation. Thanks for having me. Thank you for joining us.
From the publisher
In the first wave of the generative AI revolution, startups and enterprises built on top of the best closed-source models available, mostly from OpenAI. The AI customer journey moves from training to inference, and as these first products find PMF, many are hitting a wall on latency and cost.
Fireworks Founder and CEO Lin Qiao led the PyTorch team at Meta that rebuilt the whole stack to meet the complex needs of the world’s largest B2C company. Meta moved PyTorch to its own non-profit foundation in 2022 and Lin started Fireworks with the mission to compress the timeframe of training and inference and democratize access to GenAI beyond the hyperscalers to let a diversity of AI applications thrive.
Lin predicts when open and closed source models will converge and reveals her goal to build simple API access to the totality of knowledge.
Hosted by: Sonya Huang and Pat Grady, Sequoia Capital
Mentioned in this episode:
Pytorch: the leading framework for building deep learning models, originated at Meta and now part of the Linux Foundation umbrella
Caffe2 and ONNX: ML frameworks Meta used that PyTorch eventually replaced
Conservation of complexity: the idea that that every computer application has inherent complexity that cannot be reduced but merely moved between the backend and frontend, originated by Xerox PARC researcher Larry Tesler
Mixture of Experts: a class of transformer models that route requests between different subsets of a model based on use case
Fathom: a product the Fireworks team uses for video conference summarization
LMSYS Chatbot Arena: crowdsourced open platform for LLM evals hosted on Hugging Face
00:00 - Introduction
02:01 - What is Fireworks?
02:48 - Leading Pytorch
05:01 - What do researchers like about PyTorch?
07:50 - How Fireworks compares to open source
10:38 - Simplicity scales
12:51 - From training to inference
17:46 - Will open and closed source converge?
22:18 - Can you match OpenAI on the Fireworks stack?
26:53 - What is your vision for the Fireworks platform?
31:17 - Competition for Nvidia?
32:47 - Are returns to scale starting to slow down?
34:28 - Competition
36:32 - Lightning round




