In short
Fireworks AI’s effort to “unlock” private enterprise data for LLM training/inference via autonomous intelligence—continuous, automated customization of open models per application/use case—aiming for millions of specialized models rather than one AGI model.
Guest background
Lin Qiao is CEO of Fireworks AI (Bay Area). Raised $300M+ including a $250M Series C. Holds a CS PhD from UC Santa Barbara and previously led engineering at Meta as a Senior Director (over 300 engineers), including work on PyTorch.
Key claims
Over 90% of the world’s data is private enterprise data not seen by foundation models. Fireworks automates model/inference deployment customization so app developers don’t need deep experts. Open models are cheaper post-training (no massive pretraining cost) and increasingly converge with closed models. Real-time apps need low latency; offline agents can use “slow thinking.”
Notable examples
“Small-big-small” model selection for real-time (iterate with small models, tune with largest, distill back to small). Coding agents acting like junior engineers, changing hiring/interview focus to steering/evaluating agent output. EvalProtocol open-sources a standard to connect eval systems with tuning platforms.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOLin Qiao's Background and Mission
0:45 to 1:27
Discussion of Lin Qiao's background and the mission of Fireworks AI.
“This episode of Super Data Science is made possible by Dell, Intel, Cisco, and Excel data.”
Understanding Autonomous Intelligence
1:27 to 3:36
Lin explains the concept of autonomous intelligence and its significance.
“I believe you've now raised over$300 million in venture capital, including a recent$250 million Series C, if I got that correctly.”
Impact of AI on Daily Life
3:36 to 6:20
Discussion on how autonomous intelligence will automate daily tasks and reshape professions.
“My prediction, and that's where we're betting on the future, is to be able to activate those private data and let the model absorb additional application-specific intelligence and bring the model to the next level.”
Challenges of Continuous Customization
6:20 to 8:27
Exploration of the challenges in customizing AI models continuously and the talent gap.
“can really start to behave like a junior engineer.”
Reinforcement Learning and Specialization
8:27 to 14:01
Lin discusses reinforcement learning principles and their application to model specialization.
“So we heard a lot about AI is going to free up a lot of human labor, right?”
Horizontal Scaling of AI Intelligence
14:01 to 14:58
Learn about the concept of horizontal scaling in AI and its significance.
“where nobody else built on top of an existing API would have.”
Leveraging Private Data for Competitive Advantage
14:59 to 18:31
Discover how enterprises can utilize private data to create specialized AI models.
“We've got a link to that in the show notes.”
The Shift to Mobile-First Product Development
18:32 to 21:46
Understand how the mobile-first approach is transforming product development.
“You're kind of defining a new category, which is an exciting thing to be doing.”
Navigating AI Model and Hardware Selection
21:47 to 22:46
Explore the challenges of selecting AI models and hardware amidst rapid changes.
“So we are the platform that we abstract out.”
The Role of Open Models in AI Development
22:47 to 24:46
Learn about the significance of open models and their impact on the AI landscape.
“So last year by Solve, I haven't counted.”
Show all 20 chapters
Challenges for Enterprises Adopting Open Models
24:47 to 28:00
Understand the cautious approach enterprises take towards open models and the associated risks.
“And some of them started a company with me, Fireworks.”
Adoption of Open Models in AI
28:00 to 30:05
Learn about the differences in adoption of open models between AI startups and enterprises.
“Because we as a provider, we do not have pre-training costs, billions of dollars pre-training costs to amortize over.”
The Flexibility of AI Infrastructure
30:05 to 33:50
Discover how AI is reshaping traditional business practices and product development.
“That was a very cool explanation of those kind of two customer types, the AI native startup, the enterprise.”
Understanding Reasoning Models in AI
33:50 to 40:01
Explore the differences between real-time and offline reasoning models in AI applications.
“I've learned so much from every response.”
Evaluating AI Model Performance
40:01 to 42:04
Examine the challenges of evaluating AI models and the unique approaches taken by developers.
“As you were discussing that, something that came to mind for me is how one of the most difficult things with deploying such complex models that have stochastic outputs is evals.”
Understanding Vibe Testing and Eval Protocol
42:04 to 44:00
Learn about the challenges of vibe testing and how EvalProtocol aims to standardize the evaluation process in tuning AI models.
“But those kinds of vibe testing is really hard to harden and drive further model customization because it's very subjective.”
3D Fire Optimizer: Customizing AI Models
44:00 to 48:28
Explore the 3D Fire Optimizer approach to customizing AI models across quality, speed, and cost dimensions.
“So I'll be sure to include that GitHub repo in the show notes.”
Hiring and Learning in the AI Era
48:28 to 54:58
Discuss the importance of adapting hiring practices and learning methods in light of advancements in AI technology.
“fitting to our mission of autonomous intelligence.”
Closing Thoughts and Future Engagement
54:58 to 55:26
Reflect on the episode and how listeners can engage with the guest on social media for further discussions.
“to kind of exchange notes and go from there.”
AI's Evolving Role in Technical Interviews
56:00 to 56:30
Learn how AI is changing the evaluation criteria in technical interviews.
“to iterate on data quality, moving to the largest model for best quality tuning, then distilling back down to a small model for fast real-time inference.”
Transcript
Automatic transcript. May contain errors.0:00Jon Krohn:Over 90 % of the world's intelligence is locked inside private enterprise data that no foundation model has ever seen. Today's guest is on a mission to unlock it. Welcome to episode number 971 of the Super Data Science Podcast. I'm your host, Jon Krohn. Today's guest, Lin Qiao, is the CEO of Fireworks AI, a Bay Area startup that has raised over$300 million to unlock the world's vast quantities of enterprise data for LLM training and inference, revolutionizing capabilities and performance. With a PhD in computer science from UC Santa Barbara and years of experience as a director of engineering at Meta, Lin is now a highly successful technical founder with a rich perspective on AI today and what the future holds for all of us.
0:45Jon Krohn:Enjoy this one. This episode of Super Data Science is made possible by Dell, Intel, Cisco, and Excel data.
0:55Jon Krohn:Lin, welcome to the Super Data Science Podcast. It's an honor to have you take time out of your busy schedule to be on the show. How are you doing today? I'm doing great. Thanks for having me, John. Of course. And where are you calling in from? I'm calling from Portola Valley. That's where I live. Nice. It's part of the Bay Area. Yeah, yeah, yeah. Palo Alto. Yeah, very close to Stanford. Nice, nice, nice. So we're here to talk about Fireworks AI, your business, which has done incredibly well. I mean, you've just grown so quickly. I believe you've now raised over$300 million in venture capital, including a recent$250 million Series C, if I got that correctly.
1:38Right.
1:39Jon Krohn:And so it's a platform built around open source model deployment at scale. And the Fireworks AI platform is built around open source model deployment at scale and this idea of autonomous intelligence. Tell us what that means, Lin. Yeah, sure. You're right. We raised our last round last year and we are growing really fast. So our mission is autonomous intelligence. This mission is very complementary to AGI, where the direction of AGI focuses on investing in directing a lot of intelligence into this one model and have this model be able to solve various different kinds of tasks in a great way. So the idea is you just build your application on top of the AGI model as a utility.
2:32So this is a great, AGI is a great interaction. It's very scalable if it's successful. But the reality is only a very small fraction of data goes into the foundation models for AGI. If you look at the worst data, majority of the data, by majority, I really mean like more than 90 % of the data is actually not in the public domain. It's not in public internet. It's not labeled by labeling companies, which goes into the foundation models. And majority of those data are private data locked inside applications and enterprises. And data, we all know data is intelligence. Data is knowledge. And those application-specific, enterprise-specific data is not accessible by the AGI labs.
3:32And we just leave a lot of intelligence on the table. My prediction, and that's where we're betting on the future, is to be able to activate those private data and let the model absorb additional application-specific intelligence and bring the model to the next level. And this kind of motion is more like customization, right? It is the model and the inference deployment were customized towards applications in their specific pattern. And this customization should not be just one time, right? our application enterprise product, it keeps evolving. So this customization should be continuous. And ideally, this continuous customization should be fully automated.
4:23That's autonomous intelligence. We are making great progress towards that direction. And we believe the future is not one model result. It's going to be millions of models, one per application per use case. All right.
4:39Jon Krohn:So the term autonomous intelligence, it sounds kind of vaguely to me like the buzzword of 2025 in our field, probably the buzzword for 2026, which is agentic AI. But it sounds like it's quite different from agentic AI, this idea of autonomous intelligence. It's heavily connected with agentic AI. So think about agent as a way to automate many of our day-to-day tasks. So we have been living the world that many expert intense tasks have been gradually automated so we can free up our time. Eventually, some of the even professions will be redefined. So, for example, I think there are interview agents or hiring agents where you give a job listing, it will source the candidates and even do the first rounds of filtering and interview for you.
5:37And there's marketing agent, you give your ICP list and will source the right company, right stakeholders, and start to drive customized outbound emails and riches. and their customer service agent just give the human agent some really good assistance to kind of be smart. And there are so many like agents who are doctors and so on. So this is happening, transforming our day-to-day life. But similarly, another big transformation that's happening is in my domain, software development is being disrupted. And today, a coding agent can really start to behave like a junior engineer. And I'm not kidding.
6:29This is kind of really happening. And it's actually changed our interview process. And the fundamental question we're asking ourselves is coding skills Coding interview is important anymore. So coding interview in the past is going to be replaced by how good you are at using coding agents. So it is actually happening across our day to day life. Now, let's go back to this autonomous intelligence era. So without that, currently, this work of continuously adapting the model and changing the customizing inference setup is done by a very, very small set of experts. So those experts are like they have been doing AI system for a long time.
7:30They have been researchers for a long time, accumulated their knowledge over years. So only a few companies who have those strong density talent pool are able to do that. And the question is, can that part be automated? Right? And similar to others that has been disrupted and has been reshaped, and can this part of doing product model co-design and infusing more intelligence to the model and making the inference serving tier much faster and much more efficient, can that part be accessible by a wide range of application developers without them putting in a lot of work and carrying the burden of learning all the deep knowledges.
8:27So that's what that means. So we heard a lot about AI is going to free up a lot of human labor, right? So, and this wave is interesting because it will start from a different angle. So it will actually free up a human from the high intelligence level, not from the physical level, right? The robotics is going to disrupt the physical level of engagement, but AI is going to free up a lot of kind of high intelligence level of the tasks and work. So that's kind of interesting change. And we are also innovating and disrupting in that space from a baseline platform space.
9:16Jon Krohn:Really cool. So it sounds like, yeah, autonomous intelligence builds on agentic things, but also involves lots of things not associated with agentic, like systems, like models automatically retraining and the whole kind of system humming along nicely. it seems like a key part of that working for you, especially in an earlier answer, you mentioned how you see the future as millions of different models. These could be like LLMs, but millions of different AI models that are tuned to specific tasks within specific enterprises. And it sounds like Fireworks offers a reinforcement fine tuning product that allows your customers to beat frontier closed models in under a month on specific tasks that that relatively small, tuned, open model is fine-tuned for?
10:13Right. So we are very bullish in this direction. So think about AI has been following a lot of footsteps of human intelligence, right? Artificial intelligence has been following that full step. Even the model architecture is called neural networks. It's kind of emulating like the human's brain, right? So for reinforced learning, it is actually very similar to we, how we human learn knowledges. So we learn by various different angles. So one of the angles is we get positive feedback that we know, oh, this is the correct thing to do by learning the principles. Or we get negative feedback and it's kind of we get penalized for doing something bad from our behavior point of view.
11:06Then we learn, no, don't do it anymore. And we kind of change to a different direction. So this also happened to how we do agriculture, for example. As, you know, I'm a fruit lover and, you know, today we all like to eat like sweet fruit. The fruit also grows much bigger. But this is not how originally it became, right? It's gone through multiple generations of selection process where we collect the seeds of the sweeter and bigger fruits in the planet and among those and collect another. So this is kind of another way of reinforcement learning, reinforcement seed collection selection process. So a lot of things we do, whether for ourselves, our own learning process, or we have applied in other domains, is following the same principle.
12:08And this is also similar to how model learn. Is you teach the model what is a positive reward. You teach the model what is a negative reward. And the model is going to automatically scan through a search space of possibilities and do these feedbacks and find a path to specialize self in solving certain kind of problems really, really well. With that said, it's not all like only benefits, right? It's a trade-off. Because if you let the model specialize in a certain direction in the area, it's going to be less specialized in other areas. So it's like us exactly. As John, you're specializing in driving this podcast and you're a great host.
12:53You're very knowledgeable of how to engage with guests. And I specialize in building the best AI platform which can customize towards application-specific patterns and so on. But I'm not a good chef or cook, and I don't know actually how to do gardening very well. So that's kind of a natural selection we have to direct our attention to similar to the models. So that's kind of where we are batting on this direction to make a reinforcement learning for models very, very accessible to all application developers. So they can basically have a model in tune to their product all the time. And imagine that.
13:53Imagine a model is just constant learning the intelligence from your application. And then you have your private model. And then you have your mode. where nobody else built on top of an existing API would have. So this is a special thing that we want to kind of have every application developer to get a hold of.
14:33Jon Krohn:has focused on scaling AI vertically, bigger models, more compute. Those breakthroughs matter, but intelligence also scales horizontally. Agents sharing knowledge across a network, coordinating on common intent, reasoning together. The infrastructure for that second horizontal axis doesn't exist yet. Outshift by Cisco is formalizing it. They call it the internet of cognition. They're publishing the architecture and building reference implementations. Read Scaling Out Superintelligence. We've got a link to that in the show notes. Then check out episode number 961. In it, Dr. Vijoy Pandey, the head of Outshift by Cisco, walks through how horizontal scaling of intelligence works and why it matters.
15:13Jon Krohn:Right, I see. So all of that private data that you mentioned at the outset of the episode, that accounts for most of the data in the world, enterprises can be using their particular private data to be creating a moat by not only having that private data, but also having these fine-tuned models specifically specializing in particular aspects of their data applied to specific tasks. Yep. And this is another new thing, actually. For example, doing the mobile first move. So before AI, the biggest shift is mobile. So the application moves from desktop to mobile devices, and that actually opens up a whole new domain of doing product development, which is similar to autonomous intelligence we're heading towards in terms of thinking.
16:13On mobile first, I think one thing that opens up is the access of end consumers to an application. Because the people who own a desktop versus people who own a phone, it is orders of magnitude different. And that just means now your product could have reached to orders of magnitude higher group of people. And it will go global much quickly and you will be able to access various different demographics of the cohort much wider. It changed how app developers think about product evaluation because now it's much broader and you cannot just deploy a group of PMs to understand what the customer wants.
17:04And the product no longer become monopoly, just one design, because for different cohort coverage, you may want to highlight one feature versus the other. So then it becomes, oh, how do we even do product development? How do we incorporate that feedback? It has to be customized. And that customization is being done through a new technology called AP testing. So the idea here is to use the insights from your product to compare A versus B of the product feature and make statistical decision. Based on statistical result, make the decision where your product should look like within certain cohort of groups.
17:51So the idea is It's basically similar. Like product has a lot of intelligence. We need to leverage that intelligence to make the product better. And here similarly, product has a lot of intelligence. We need to leverage intelligence to make the model better for your product. Therefore, your product built on top of a specialized model will be better. So right now it's a completely open space. It is a vacuum and there's no existing solution that has been solving this problem really well. I'm pretty sure industry will move towards looking at this space very closely. And I firmly believe there's a lot of value in there.
18:37Jon Krohn:Awesome. I love it. Yeah, really exciting. You're kind of defining a new category, which is an exciting thing to be doing. Um, so when people are trying to figure out these, you know, these relatively small, say LLMs for a particular task, does model selection really matter? I mean, do your, are your clients like picking, oh, okay, I'm going to use this llama model of this size or Quinn, or do they make those decisions or is this something that's kind of like handled automatically by the autonomous intelligence system? So model selection matters, but it's also exhausting. So, funny thing, there are two interesting phenomena that are so unique to AI.
19:28One is the model depreciation is very fast. As you can observe, every couple of weeks, there's a new model launch. While this code is open, someone topped the leaderboard and then a couple of weeks later, someone else top the leaderboard. And we often, they all are strong in different ways. So we start to see the researcher focus start to diverge, right? It's clear some labs are really good at chat-based models. Some labs are really good at coding to use agentic models. Some labs are really good at multimodality models. Some are really good at long contests. They all start to diverge into focus on and specialize in different areas.
20:20Speaking about specialization. So different use case will suit different kind of model specialty very well. But it's really hard for people to figure out which one is the best for my use case. and my use case will also evolve over time. And the public benchmark result is fully saturated and we need to figure out how to pick and choose continuously, which is exhausting. So we do help our customer figure that out. But at the same time, there's another level of complexity. It's hardware depreciation is also very fast. This is something new. Before this wave, usually every three years, there's a new hardware skill.
21:11And now, last year, myself, NVIDIA launches three skills, three new skills. And in 2026, there will be a lot more new skills from all different vendors, whether it's GPU or customer ASIC. So then how to manage hardware becomes a very hard problem. And these two combined is causing so much headache to the application developer who want to stay on top of all the different kind of wave. It's just too much deep knowledge and expertise to gain in order to kind of pick the best for them. So we are the platform that we abstract out. We eventually want to abstract out hardware. We want to abstract like hardware selection.
22:03and we want to kind of provide the best model for various different use cases. And so basically there are various different kind of mapping from specific use case, specific workload patterns to the model, to the hardware, to the inference setup, to the fleet design. So all of that requires a big team in-house where deep experts, which is really hard to find. And we want to kind of make it super easy for our customers.
22:40Jon Krohn:Right. So instead of needing to find the deep experts, they can just come to Fireworks and work with you guys, work with your solutions. Absolutely. And interestingly, there are... So last year by Solve, I haven't counted. Every month, there are a few new models get deployed and launched. and it has been very exciting. But also, I would say early on, we bet on open models. That was a big debate in the company that we debate on many what-ifs, right? But because of our background in PyTorch, which is an open source project, PyTorch is now, we built PyTorch from the ground up. It is now the dominant AI framework.
23:27we firmly believe in the power of open science.
23:36And in the long run, we believe that is going to really shape the industry in an unprecedented way. So that's why we bet on open models from very early on. And it's great to see that open model performance is converging with closed model. So I think 2026 is the year that that convergence will become more prominent. And super excited about that.
24:03Jon Krohn:Yeah, something that I haven't mentioned on air yet is that for seven years before founding and being CEO of Fireworks AI, you were at Meta as a Senior Director of Engineering, where you led over 300 engineers. And a big part of what you were doing there, based on what I could find online, is developing PyTorch. So thank you for that. This is definitely a very big team effort. And I have a very talented team at that time. And yeah, so I'm very thrilled about the industry-wide impact of PyTorch. There are many great researchers and leaders, engineers I'm lucky to work very closely with. And some of them started a company with me, Fireworks.
24:51So really great crew. And we continue to going down the same path of democratized AI to the whole entire industry.
25:02Jon Krohn:Really cool. All right. So back to Fireworks and what you were saying about open source, not just PyTorch, but open source models and how you at Fireworks have made this bet to embrace open source AI models. Do you think that a lot of organizations, a lot of enterprises underestimate what's possible with open models? I don't think it's about underestimation that much. It's just about not being familiar with open models. I think especially, I think I will put, simplify this answer. There's two groups. There's AI native startups. There is enterprises. Obviously, enterprise has digital natives and traditional enterprise in various different categories.
25:53But let's just take a look at these two groups. They are reactions different. Air native startups, they just want to try the best option, and they do not have any baggage. They're very, very brave in testing open models. They're always very curious. And at the same time, there's a very interesting phenomenon that's also unique to AI. So as you know, there are many companies, whether startups or incumbents, they're trying and experimenting with new user experiences that our day-to-day is interacting with. and those kind of experiments have seen a lot of success in terms of product market fit but product market fit at AI time doesn't mean a viable business there could be a big 10 that people are willing to pay the product, pay for the product but the cost of running the business could be much higher than the revenue you get And then people just cannot scale their business after they hit the power market bit.
27:11And it's funny, literally, they're going to scale into bankruptcy. So obviously, startups have limited funding. And even including incumbents, they cannot, they have the distribution. And they have to be so careful. Their CFOs have to be so careful in controlling the cost. So the kind of budget and this ROI totally makes sense. And because of that, they have to hold a lot of people on the waiting list and cannot open up the gate. So while the startup, going back to AI native startups, they are very brave in embracing open models. But part of motivation is open models, they have a lot more control.
27:57They can customize however they want. But also at the same time, open models are much more economical. Because we as a provider, we do not have pre-training costs, billions of dollars pre-training costs to amortize over. And we only focus on bringing the best quality speed and cost from post-training, which is much cheaper, and inference customization of the deployment. So the unit of economics of fireworks operating open model is very different from Frontier Labs. So that's kind of where we see the AI native startups are very brave in adopting and going down the path of customization with open models.
28:42The second group of enterprises, obviously, they're much more careful because many of those are public companies and they are under certain kind of obligation, especially their legal team to understand what are the implications of using open models, license, what are license limitations, just kind of trying to understand what is this beast. But from my observation, the enterprise started to open up, and they started to understand much better, oh, this model actually is not going to send data to other countries. It's just a model. If the model provider is hosting a model in the U.S., and then they can provide privacy and security around the ins and outs.
29:40So, but anyway, there are other concerns as in the kind of content generated, does that follow certain kind of guardrail and so on. But overall, I've seen the whole entire industry start to understand better and start to kind of understand the bounds of operating open models. And I only see an upper trend of embracing open models. Nice.
Read the full transcript
30:05Jon Krohn:That was a very cool explanation of those kind of two customer types, the AI native startup, the enterprise. It sounds like regardless of which category a customer of yours falls into, a huge advantage of Fireworks must be that they can not only are running open source models cheaper, like you said, but by running them through you, it means that they don't need to themselves, right? Be buying the GPUs, be managing all the ops around that, physical ops, software ops. They just can rely on fireworks to get things up and running. And then having the right GPUs running behind the scenes and controlling those kinds of costs is handled automatically.
30:52So at the beginning of this conversation, I mentioned AI is changing a lot of things in our day-to-day. I think one fundamental thing it's changing is the velocity of product development on AI is extremely fast, including the enterprises. So because of that, it's really hard for enterprise internally to forecast how much traffic this product is going to generate. it may generate no traffic because the experiment doesn't go to the phase they can go into production, or it can generate massive traffic. It's kind of the variation is very broad in this fast evolution of product development. So then it caused a big problem because without a steady prediction of how to do capacity like from a finance point of view.
31:57And then what does that even mean to procure hardware? And you either over-provision or you under-provision. It's kind of the range is very broad. So that's where we can help because we're aggregator. We aggregate demand across all different customers and we take that risk off the table from being, you know, you want to worry about that and we're very flexible in accommodating varying requests because we can kind of add them together and drive the adoption. So I think that's kind of a unique nature of AI, but it's not that unique, but think about that, right? So in the early days of cloud first, before cloud first, everyone, every big enterprise, they have mainframe.
32:50So data center locked in with keys in a cage of machines, right? This is the most rigid way of doing capacity planning. And when cloud comes into the picture, initially people think you're crazy. Why would you move from the most secure deployment of your infrastructure into renting someone else's infrastructure and you run your most important application there. But guess what? That cloud infrastructure provides so much flexibility that over the long term, it's much more efficient to run. So similar things is happening in the AI space. And that's why it's fascinating. The velocity of AI development is shifting and changing and even disrupting how we do business in the traditional way.
33:47Jon Krohn:Really cool answer again from you there. Thank you so much. I've learned so much from every response. And so it seems like you might have an interesting and informative perspective on reasoning models or slow thinking models. So that's been a big thing last year. OpenAI, Anthropic, Google, there's a big fuss around the slow thinking models that they released, and particularly their performance on complex tasks like math olympiads and writing academic papers and these kinds of things that require more processing before just outputting some kind of train of thought. I noticed on the Fireworks website that you mentioned very fast latency, like sub two second or even sub 500 millisecond latency in a lot of real world deployments.
34:41Jon Krohn:Do you think that the slower reasoning models that take time before they output something, do you think that there's a lot of enterprise need for that? So I will classify, again, extremely simplified way of looking at application. So one type of application is real-time response. So it's either human-facing interactive response latency, or it's kind of fast transaction-facing, for example, fraud detection, right? So in aggregate way, it's kind of very fast going. The other type of application is offline. So, for example, I have a case study, legal case study. I need my legal assistant to go off and analyze similar cases and come back in three days, give me a report.
35:33Now, an agent, legal agent can go off and do this study for a couple hours and come back. So those are two very big categories of agents or agent applications that has different latency requirements. So typically, for the real-time, extremely low latency requirement, you cannot afford to think. You have to basically react. And usually, the customization goes back to customization. Usually the customization process is very interesting also. You go small, big, and then small for real-time use cases. What does that mean? When you customize, a lot of time it's about your data. Is your data high quality?
36:28Keep in mind, it's kind of garbage and garbage. Same apply here. If your data is not high quality, then your end model is not high quality. So for you to test data quality and the fixed quality issue, you start with small model to see if the model is going in the right trajectory. And then if it is not going in the right, you go back to fix our data and then small model help you iterate fast. But that small model is not important. The trajectory of the model quality change is important. when you like the trajectory, and then you move to the biggest model. And usually those are biggest MOE model, really hard to get it running.
37:07So obviously that's where we will definitely help you. And use the clean data to tune the model, get the best specialized model to solve your real-time problem. But usually those largest model is not as fast, right? Because they're very big. They have trillions of parameters. Very big, but very good in terms of quality. And then you go small again, but you still find the biggest model you have tuned into a small model because you need a fast response. And then you launch the smallest model. You can get to high quality, find that process, and go from there. So small, big, small, right? Now we move to offline.
37:55are asynchronous agents. And usually those agents are doing some deep research in certain kind of domain, whether it's about legal, whether it's about finance, whether it's about software engineering, or whether it's about something else, right? But usually it takes time to think like us human beings, we're going to, hey, do some homework, go hide in a cave, figure it out. So similarly, and those models require the highest level of intelligence, highest level. So then the kind of slow thinking mode becomes extremely important. But not just that. Those big models for offline research also requires a lot more context.
38:43a lot more context to make the right thinking process, like get the right design of the flow. But at the same time, oftentimes those models also need to get specialized in solving certain kind of problem really well. The legal process, the finance process, the process of working, building PowerPoint, the process of building an executive overview, a pitch to investors, and the process of doing customer service, those are all very, very different. So we have also seen a lot of application agent developers, they start to customize offline agent where that just means they need to figure out how to tune with the thinking tokens.
39:44So that's kind of a very different classes of equation I've oversimplified here.
39:50Jon Krohn:Oh, I mean oversimplifying, but making it very easy to understand as you have with all of your responses in this episode. I love particularly the small, big, small approach to finding the right model for your use case and having that be performant in real time. As you were discussing that, something that came to mind for me is how one of the most difficult things with deploying such complex models that have stochastic outputs is evals. And so even like when you're doing that small, big, small, like how do you ensure that when you go say from the big to the final distilled small model, you're still getting the same kind of performance.
40:28Jon Krohn:Evals can be so difficult. Do you thoughts on that or does fireworks have any tooling that helps out yeah so for evals we do not offer evil product we instead partner with our um with the companies whose specialty is doing eval so here again we really respect specialty and the customization and the focus um so you're right Eval is where people get started, right? But it's interesting. If you talk to the startups or people building AI native agents, there are so many different ways to do evals. First and foremost, people do vibe eval-ing. If you build applications, the first thing you do, It's very different from software development practices.
41:23Usually when we write a piece of software, we start from unit test first. And then we build different kind of guardrail to make sure the quality, we have guarantees of quality. Find various different unit tests to integration tests, to then testing production, you collect some metrics and so on. but agent development people usually start with some hypothesis and just look at the result and do vibe testing it's very interesting because they don't they are they do not want to bound their creativity in the imagination by forming small test to big test that test big ideas.
42:08But those kinds of vibe testing is really hard to harden and drive further model customization because it's very subjective. And then the question is how you turn vibe testing into something more concrete that a model can understand or a tuning platform can understand. So I think it is a very important area. And where we have been contributing is to solve another complexity. So today, there are various different tuning platforms and the various different evaluation platform. As you know, the evals are directly feeding into tuning to drive the tuning process. Without eval, you basically don't have a guide.
43:06Hey, this direction is correct or not, right? But there are so many different eval systems, there are different tuning systems, and we're part of the tuning systems. And there's no standard across them. The integration of cross-product complexity is very high. So we open-sourced a project called EvalProtocol. The idea here is to standardize the format so any Eval system can talk with any tuning system. So for our developers, they can pick any combination as they want, but because of EvalProtocol, it kind of bridged the gap across and they have optionalities on both sides. So that's our intention to help bring more commonality across these two sides because these two sides are deeply integrated.
43:59Cool.
43:59Jon Krohn:So as you were talking there, eval protocol from Fireworks hadn't showed up actually in our research, but I quickly found the GitHub repo for it. So I'll be sure to include that GitHub repo in the show notes. a different offering of yours that did show up in our research was something called 3d fire optimizer and um so this seems like it's in a way it's a way it goes over a hundred thousand possible ways of optimizing an llm stack do you want to tell us about that yeah definitely so when it comes to customization goes back to our mission of autonomous intelligence we are customizing across three dimensions across model quality model speed and the model efficiency which is cost so when we started our journey people were sure hey they're really concerned about cost or hey they're really concerned about speed because they are very interactive or hey they're very concerned about quality wherever they start their pain points.
45:09At the end, they are concerned about all three dimensions. There's no time where like, hey, the cost is really good here. We are like 10 times cheaper compared to alternatives. And they're like, go for it. They're always like, oh, we also need a better quality to be able to launch this. Oh, we need to be faster to launch this. It's always, always all three dimensions has to be much better. So we're like, hey, this is actually not a linear problem. It's a complex problem. It's with exponential search space because the three dimensions are evolved, but it's actually more than three dimensions. If we decompose this problem, then it becomes so many different building blocks to stack onto each other and each building block has five to ten different options to pick and choose from.
46:06So that's where in combination there are more than 100 ,000 options in this search space and it becomes a search problem. Again, that's not a new problem across our system research. For example, databases. Databases has query optimizer where a data engineer or analyst to write a SQL query. But the execution, so SQL query is a semantic description of what they want to achieve. But in reality, based on whether you have an index, whether how the data is being laid out, how the data is being sorted, or how data is being partitioned, there are different ways to retrieve that data and process data in the most efficient manner.
46:59And the query optimizer basically converts this query, preserving the same semantic level of meaning, and turns that into a specific query plan customized towards the current state of a database. And query optimizer basically is the brain of a database engine. and it's massive there has been massive innovation in kind of the first decade of this Manila I was part of the research group building query optimizers but this in the space we are operating we are optimizing and customizing across quality, speed and cost And that is much more complicated than a query optimizer, which only optimizes for efficiency.
48:00Right? So that's kind of what we build, is to free up our customers, those app developers, from carrying the burden of learning how to do this three-dimensional optimization, and to find the sweet spot, the best spot among all these candidates, the best suited for their requirements. So yeah, so that's kind of one of our innovations as we build our platform, but it is directly fitting to our mission of autonomous intelligence.
48:37Jon Krohn:Very cool, Lin. Love it. That's the end of my technical questions for today. But I suspect that we have a lot of listeners out there who, having heard what an amazing speaker and thought leader you are, with all the capital that Fireworks has raised, with the impact that they're making, the challenging problems that they're solving, I bet we have a lot of listeners that would love to work for you. Do you have any open roles and what do you look for in people that you hire? Absolutely. The reason we raised CERC last year is to massively accelerate our growth. So we're actually hiring across the board, all the way from the GTM side.
49:17We have a lot of openings from sales, marketing, as well as product engineering, and the finance business operation across the board. I really mean it. Because we just have so much demand we cannot manage. and we love to work with very creative and high aptitude people who are willing to join us, take a big chunk of ownership and drive a lot of impact. So yeah, if you're interested, I'd love to please send, please contact us and send your application.
50:01Jon Krohn:Wonderful. All right, so now we're just to my last two questions that I ask every guest, but I think I'm going to get an interesting answer based on our conversation before we started recording. I always ask my guests for a book recommendation and I think you have something else. Yeah. So I think this is an interesting time of AI and everything has changed and how we learn has changed. In the past, I've been reading books from mostly around, I love to read books around business and leadership. And I really like to read books about certain kind of individuals I'm very curious about. But nowadays, I listen to a lot of podcasts.
50:55I think, John, your podcast will be a very impactful one as well. Particularly, a fundamental reason is the following, right? So this AI innovation is changing how we do business, how we create technology. All these agents that is emerging is going to change our day to day. But at the same time, fundamentally, it's also changing how we learn. How we learn what's happening in a fast-moving world. particularly I think we have a lot of people including my funding team, including my top-tier engineers, they learn a lot from a lot of posts. We were always on the cutting edge. We read a lot of paper and kind of the paper reading velocity is very high and paper generating velocity is very high.
51:56But we read a lot of best practices from social media. And there are a lot of creative ideas, a different way of thinking, approaching problems. And you have to think differently to be very innovative in the AI time, because a lot of fundamental assumptions have been disrupted. For example, we're rethinking our interview process. because now coding agents are very good at general code almost as a fresh graduate, junior engineer. So then the question is, hey, is writing code quickly and correctly an important area to test or not? So we're questioning that because with the assistant of a coding agent, And that's no longer a problem.
52:51And it's more a problem to have the eyes and the mental framework to judge how good is the code. And to have a strong way to steer a coding agent to design very well, like coming from the source of design and system architecting to drive the implementation. So it starts to shift the focus and bottleneck of the software development. Just give you an example of recruiting in the hiring process that has changed. So I will say the whole community is a book. The whole community is writing a book of AI. I think that's fascinating to me. And that's kind of, I just feel like I'm very lucky to live in this time of this world to fully embrace fast velocity of changes.
53:48And there's so much to learn from each other.
53:51Jon Krohn:Cool. What an answer. Yeah, certainly we've never had an answer like that before on the show, but I love it. You're a really progressive thinker. And I have personally learned tons from this episode. I'm sure a lot of our listeners have as well. Lin, how can people follow you? We're on social media or how can people be getting your thoughts after this episode? Yeah, I'm learning also to transition my interaction more towards social media. So you can follow me on LQIAO, my handler on Twitter, x.com. but also you can follow my LinkedIn. I share my thoughts. I will do more in the future, but also I talk about our product and the product launches and direction we're hiding towards.
54:48I love engagement. So if you have for all the audience here, if you have thoughts, feedbacks, ideas, I would love to talk with you to kind of exchange notes and go from there. Awesome.
55:04Jon Krohn:Thank you so much for opening up your inbox to our listeners, Lynn. Really appreciate it. And yeah, that's the end of the episode. Thank you so much for joining us. I can only imagine how crazy your schedule is. And so to take this time out and speak with me and share your thoughts with our audience. We really appreciate it. Thank you, Lynn. Glad to be here. Thanks, John. super episode today with the exceptional engineer and entrepreneur lin chow in it she covered how over 90 of the world's data live in private enterprise systems and never make it into foundation models representing a massive untapped source of intelligence how autonomous intelligence is about continuously and automatically customizing models with private enterprise data resulting in millions of specialized models rather than one agi to rule them all she talked about her her small, big, small approach, which means starting with a small model to iterate on data quality, moving to the largest model for best quality tuning, then distilling back down to a small model for fast real-time inference.
56:09Jon Krohn:And she talked about how coding agents now perform at the level of junior engineers, fundamentally changing what matters in technical interviews from writing code quickly to having the judgment to steer and evaluate AI-generated code. As always, you can get all the show notes, including the transcript for this episode, the video recording, any materials mentioned on the show, the URLs for Lynn's social media profiles, as well as my own at superdatascience.com slash 971. Thanks to everyone on the Super Data Science podcast team, our podcast manager, Sonja Breivich, media editor, Mario Pombo, partnerships manager, Natalie Zajski, researcher, Serge Massis, writer, Dr.
56:45Jon Krohn:Zarikar Shea and founder, Kirill Arimenko. Thanks to all of them for producing another fantastic episode for us today for enabling that super team to create this free super data science podcast for you. We are deeply grateful to our sponsors. You can support the show by checking out our sponsors links, which are in the show notes. And if you'd ever like to sponsor an episode, you can get the details on how to do that by making your way to johncrone.com slash podcast. Otherwise, help us out by sharing this episode with anyone that would like to listen to it. Review the show on your favorite podcasting app or on YouTube.
57:21Jon Krohn:subscribe, but most importantly, just keep on tuning in. I'm so grateful to have you listening and I hope I can continue to make episodes you love for years and years to come. Until next time, keep on rocking it out there and I'm looking forward to enjoying another round of the Super Data Science Podcast with you very soon.
From the publisher
Lin Qiao, CEO of Fireworks AI, talks to Jon Krohn about how she builds effective models quickly, why coding agents can perform at the level of a junior engineer, and what she attributes to the success of Fireworks AI: True to its name, the company exploded into the AI industry with over $300 million secured in venture capital, as well as netting a further $250 million Series C funding. For Lin, many enterprises miss out by not being familiar with open models. Open models give a lot of control to the user, offering customizability and at a much lower price point. Listen to hear how Fireworks AI helps companies continue to save money through AI.
This episode is brought to you by the Dell, by Intel, by Cisco and by Acceldata.
Additional materials: www.superdatascience.com/971
Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.
In this episode you will learn:
(01:19) All about Fireworks AI
(24:16) Why companies need to take notice of open models
(33:05) The commercial viability of slow-reasoning models
(38:51) Fireworks AI’s approach to model performance evaluations




