In short
Eye On A.I. Episode #281 Summary
Episode Details
- Title: Leon Song: The Research Driving Next-Gen Open-Source Models (Together AI)
- Host: Craig S. Smith
- Guest: Leon Song, VP of Research at Together AI
- Release Date: [Date not specified]
- Podcast Description: A biweekly podcast exploring the global implications of artificial intelligence.
---
Overview In this episode, Craig S. Smith interviews Leon Song, VP of Research at Together AI, discussing significant advancements in open-source AI models, their infrastructure, and how these developments are reshaping the AI landscape. The conversation covers various technical innovations, the competitive dynamics between open-source and closed-source models, and the future trajectory of AI technologies.
Key Topics Discussed
- The Shift to Open-Source AI
- Increased Acceptance: Open-source models have gained credibility and are viewed as competitive alternatives to closed-source models due to major players releasing effective open-source versions.
- Flexibility for Enterprises: Together AI’s platform allows enterprises to customize and scale their usage of open-source models.
- Technological Innovations at Together AI
- Speculative Decoding:
- A method to improve inference speed without compromising accuracy by using a smaller model (speculator) to verify outputs from a larger model (verifier).
- The success of speculative decoding depends heavily on achieving a high acceptance rate.
- Recent Releases:
- Introduction of new models like the 14B coder model performing at GPT-3 levels.
- Utilization of technologies like FlashAttention and RedPajama to enhance model performance.
- Leon Song's Background and Experience
- Previous Roles: Prior to Together AI, Song worked at Microsoft, contributing to projects like DeepSpeed and AI for Science, focusing on machine learning system optimizations.
- AI for Science Initiative: Addressed scientific challenges using AI, including protein structure prediction and climate modeling.
- Open vs. Closed-Source Models
- Competitive Landscape: Open-source models such as DeepSeq R1 and Llama 4 are now rivaling proprietary models in terms of capabilities.
- Future of Closed-Source Models: Speculation that closed-source companies will need to adapt their strategies, potentially releasing some models as open-source to remain competitive.
- Role of Hardware in AI Performance
- GPUs vs. ASICs:
- Discussion on the benefits of using GPUs for their flexibility compared to ASICs, which are customized for specific tasks but may lack adaptability.
- Exploration of partnerships with NVIDIA for enhanced performance and future hardware integration.
- Together AI's Offerings
- Platform Features:
- A range of services for enterprises, including dedicated instances for custom model training and deployment.
- Emphasis on data sovereignty, ensuring customers retain ownership of their models and data.
- Future Plans:
- Expanding hardware diversity and enhancing software solutions for better inference capabilities.
- The introduction of Together Chat for users to interact with various models in a unified format.
Conclusion The conversation with Leon Song emphasizes the transformative potential of open-source models in AI and highlights Together AI’s efforts to empower enterprises with advanced infrastructure, innovative technology, and a commitment to data ownership. The future appears promising for open-source solutions as they continue to evolve and challenge established proprietary systems.
---
Additional Resources
- Together AI Website: [Together AI](https://together.ai)
- Follow Craig Smith on X: [Craig Smith](https://x.com/craigss)
- Eye on A.I. on X: [Eye on A.I.](https://x.com/EyeOn_AI)
Feel free to explore the discussed topics further and engage with the technologies shaping the next generation of AI solutions!
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Our mission of supporting open source models gets so much attention. And I think today, this mission is no longer something that people will feel like, oh, you know, this may not be that great compared to closed source model. So this really changed when major players started releasing open source models. The platform currently is being used widely across different enterprises. Many companies may have their own models and as closed source models or whatever for their internal business. But we have customers, they use our platform to do other things, right? So the key here is the flexibility that we provide.
0:40Build the future of multi-agent software with agency. That's A-G-N-T-C-Y. The agency is an open source collective building the internet of agents. It's a collaborative layer where agents can discover, connect, and work across frameworks. For developers, this means standardized agent discovery tools, seamless protocols for interagent communication, and modular components to compose and scale multi-agent workflows. Join Crew AI, Langchain, Llama Index, Browserbase, Cisco, and dozens more. The agency is dropping code, specs, and services. No strings attached. Build with other engineers who care about high-quality multi-agent software.
1:39Visit agency.org and add your support. That's A-G-M-T-C-Y dot O-R-G. So my name is Leon Song. I'm the VP of research at Together AI. I lead the research work that doing the R &D elements of the company that we push out a lot of wonderful innovative research. you're seeing power in our platform today, you know, ranging from flash attention, like kernel works like flash attention, you know, data works like Red Pajama, version one, version two, we have a lot of, you know, research open source works on speculative decoding, a lot of the work that's been used by industry widely like Medusa, and others.
2:33And we also, So, you know, have done a lot of the work that providing modeling data optimizations to the open source community. We recently released a small 14B coder model, which performs at the level of Omini 3 level. So we just released that last week with Berkeley. So there's a lot of research in our work that's are going, but it's like very, it's a product-centric, research-driven products-centric work that I'm leading. Yeah. So prior to Together, I was at Microsoft. I was very fortunate to experience the whole, you know, the GPT, chat GPT era at Microsoft with OpenAI. I like works in training system projects in deep speed.
3:29And then I created one of the first initiative on AI for Science. Initiatives is centered around Microsoft partners. So using AI to address scientific challenges at large scale. and then later on at Microsoft I was working on you know different system machine learning system optimization techniques for inference so that we can serve you know the GPT series of models on Microsoft platforms. Yeah okay can we go back and talk a little bit to begin about uh ai for science uh because there's been a lot do you still follow i would imagine what's happening uh generally um anthropics done some amazing stuff yeah can you talk about where you guys what you guys were focused on at microsoft and where you see the field having moved since you left Microsoft and what you see in the future for that.
4:40Yeah, I'll talk a little bit about that. I think this is one of the things that the field is quite excited about for a long time until we get into the generative AI era, like early 2023 when ChatGPT came out. And then a lot of these models used to be supported by by less generative AI models, based models, where now like used to be a lot of these kind of diffusion models, where they're trying to figure out how to put different models together to address the complex biology problems or simulating weather, climate changes and things like that. So when I was there, we had a great, I think it was a great initiative.
5:27So we started with, I think, two major projects with internal folks from MSR and also external folks. So one project we started is to provide this system support for scaling protein structure prediction training. So you probably have heard of all the folks from Google. So back then, 2023, there was an open sourced community work led by Columbia that called OpenFold, which is using PyTorch lightning with deep speed backend to train these types of models. And then so trying to give the community a way of studying, using open source models to study protein structure predictions. So in the past, the problem was that model itself requires large activation during training.
6:26So the training cannot really scale well. They couldn't do really long context window and they couldn't do a larger parameter size. It's well-bounded by the system itself. System, I mean, like, you know, the hardware devices you're running, the GPUs, the CPUs, the large cloud. So we stepped in and we provide a lot of important kernels so that project can now scale and can train. I remember three times longer of sequences. We reduced the memory consumption during training by like 75 % or something like that, a significant amount. And then on that project took off. Now you guys can use it. It's an open source project.
7:12people are using it in academia. And I think there are startups using that as well. Now they formed, I think, I think, a consortium actually focusing on the effort. There were other very important Microsoft, more internal efforts, that's using very creative ways to do protein structure prediction in a different way. And the alpha fold and then the open fold way of doing it. You know, the open sort of the publications you can find is related to how they're using a probability prediction way of like looking at equilibrium in terms of a more complex, you know, protein structure prediction rather than a more straightforward way that the other models are using.
8:00So those were very successful projects. The other aspect we're looking at is weather climate models, weather models. that was there was a very significant impact at that time. I'm pretty sure now is still the case to like the Bing service because we have a lot of users are using that through Bing. So yeah. Yeah. On the open fold, did you follow the same architecture as AlphaFold? I think it's not strictly follow not strictly following that architecture. I think there were a lot of changes even in 2023 when we were working on it. I haven't really followed very closely after my time at Microsoft, but my understanding is they're trying to pull the community, the scientists into this project together and trying to ask them to contribute different model components that you can do mix of models, which is not just a one transformer architecture, you know, dominating model.
9:11You can do different things in that community. I think that's what they're trying to do. Yeah. And one thing that's, and I've spoken to DeepMind about AlphaFold, but one thing that I've never, because at the time that I spoke to them, It was still a research project early in their research. And these predict protein structures, but they're not necessarily 100 % accurate. Right. There's, I think there's always that kind of verification, validation process needs to happen post post training. You can fine tune the model to be better these days with DeepSeq being very DeepSeq models being very creative and using RL type of training strategies to learn, learn its mistakes and, and then do better with the reward signals.
10:25So all of these, if you think about protein structure prediction, is it really, some people believe it's very sort of like a restrictive structured, very like a limited in terms of scope, what the protein can, the amino acids can be structured in certain ways and other ways are not just garbage. bridge. But a lot of people feel like if you if you think it's a very restrictive fashion in terms of the scope, then how about make an analogy to coding, right? Or math, right? If that's the case, then today's model does that really, really well, using our training fashion, which which you know, can have a much more successful rate of the prediction.
11:15So back then it was not RL based. It was not reasoning model based model. It was quite simple. What they studied and investigated from off and forward. Yeah. And then because the remarkable thing is you can use these models to generate novel proteins. Right. Right. And when you do that, then, and maybe this is beyond certainly training a model, but how then do you synthesize a novel protein? I mean, is it, and then do you test how that novel protein behaves to validate the prediction? Yeah, I think that was a question a little bit out of my system researcher, computational scientist. That's sort of outside of my domain.
12:19But what I can say is that people do use these predicted structures and then they do post modeling work like validation and verification. it requires the lab validated, you know, to be valid, to be validating the labs, whether or not this legitimate or not, or they're just random compounds. But, but I do sometimes see news and read some papers and they're saying like, oh, this, you know, this wonderful creation of the new, new structure that can mean something. And it's really, really not something we have seen. Yeah. That's amazing. amazing. Okay. And then jumping to what you're doing at Together.
13:05Together was started when? Um, I believe officially was 20, uh, late 2022. Yeah. And, and it was started by Percy, right? Is, um, I think there's several, I mean, since Chris Ray, Percy, and we have, uh, and Whipple, our CEO. I think there are four co-founders and Percy is one of them. Yeah. Okay. And this latest model that you've come out with, can you talk about that? From us? Yeah. Oh yeah. So I want to say something like, we're not a modeling company. Yeah, I know. So we work with other modeling companies to come up with innovative models, sometimes with alternative architectures, not transformers, so that that and then can serve on actual hardware with certain efficiency, control, and also
14:10economics. So the recent one is a lab that in Berkeley, where we work with them, and then train this 14b coder model. The coder model performs really well and it's based on this kind of mixture of agents type of strategy that, you know, or like, and then they train this model with their own data sets and the different training strategies. And now the, I think the The coder performs really well. It's at the only three level. And what was Together's contribution to that? And in the lab, the Berkeley lab, what lab is this? I'm not very familiar with this particular project because it's external. We have one, two people, we're working with them primarily on data side.
15:06I see. Yeah. Okay, well, tell me about the speculative decoder work. Is that something you're more directly involved in? Yeah. I thought this is about a coverage of Together Innovations on our platform, but I think you would like to more understand what I do in terms of involvement. But that's okay. We can go through those topics as well. Yeah, and then we'll cover it together in a minute. Sure, sure. No problem. Speculative decoding is something that Together is really good at. a lot of the companies in the inference provider realm that provide speculative decoding. Speculative decoding is a way that is a non-lossy way of doing inference that you don't lose accuracy but you can boost inference performance.
16:02So the basic concept is you have a bigger original model as a verifier, they're verifying a very, very tiny model that's a speculator model, where if the outcome of the verifier and the speculator agrees, then the inference time will be dramatically reduced. In other ways, you would use speculator to perform the inference. So you are basically treating off a lot of the efficiencies in that realm. But the difficulty of building speculators is you want a very good acceptance rate. If your acceptance rate is very low, then your overall performance will still be hindered by the larger model. So you will not really get to the peak of the performance that you can gain.
16:48So Together has been working in this prior to, I think prior to my time, when I joined early 2023. Prior to that, Together already started working on specular decoding and then produced several open source work to the community. like there's a work called Medusa. And we and I together, we have our own way of training speculator. We have automatic pipeline we built that for internally and customers that they can provide user data for customers and provide user data. We have an automatic way of training that speculator and performing different kind of model optimizations. Okay, so I'm going to slow it down a little bit.
17:35what is a speculator? It's a smaller model. It's a smaller model distilled from... It can be, it can be, I'll give you a couple examples to clarify. So let's say you have a 670B DeepSeq R1 model. It's a large MOE reasoning model. Right, large mixture of experts. If you serve this model and you run on multiple nodes, it's going to be really slow for you depends on your decoding depths how many tokens are you generating the decoding process is going to be really yeah and just again for listeners the decoding process what are you decoding so it's um autoregressive transformer is the autoregressive way of generating tokens so for instance you have the pre-fill phase where you take the prompts into Prefill and they will generate a token where you start with that token feeding into your decoder.
18:40So anything logically coming after that token will be generated attached back to that token. And you just keep generating the context with the previous tokens that you're inputting. So that's sort of a very easy way of explaining. So what makes Transformer so powerful is this autoregressive decoding process that can actually generate in the context by itself. So that's the wonderfulness of the transformers and all the models that we're using so powerful today are based on that particular model architecture. Yeah. All the companies are using this model architecture, OpenAI, all the other modeling companies.
19:29Right. And then so you have a smaller model. And then what is the speculator? Oh, so I'm trying to get to that. So let's say you have a large model. Of course, you don't want to perform decoding on the 670B model, right? So now I can train a 1B model, model with 1 billion parameters, and you can use it to pair up with your verifier, which is the larger model, to perform speculative decoding. So the smaller model inference speed is going to be really fast because you literally reduce the parameters size by orders of magnitude. Now you have 1B model. Now how to make this 1B model that in terms of prediction capability or acceptance rate get very close to the larger model.
20:21For that smaller model, you can distill from the larger model or distill from, let's say these model providers provide several sizes of a model. You can distill from the medium size or smaller size of model to the size that, you know, 1B, 2B, 3B that you're looking for to support your model, right? That's what the speculator model means, which is as much as when you run inference, your compute is spent on the smaller model rather than the verifier. I see. and okay and these these are products that you're making available to the open source community I mean how does Together AI work? Together AI is a AI accelerated cloud company that for enterprise so that's sort of the main thread of business so we provide a platform like Together their platform that people, there are different tiers of services.
21:29One of them, the one that's open to the public, whoever has a credit card, doesn't matter if your startup company or individuals that you can use, is to go to our serverless endpoint. Then you can use all the models that we're offering. Open source models. There are now up to like 200 models that served on the platform. This has the models from Meta, from DeepSeq, from Alibaba, like all these models from different open source providers. And you can use them through the API where you can just type through the UI or you can use the APIs or other types of toolings in your own work. You can build applications based on the together service.
22:18Now, that's just one tier of the service that we're providing. Another tiers are including dedicated instances where, let's say, smaller, medium-sized companies, they want to reserve GPU resources for doing different things. Our goal is trying to provide an entire AI lifecycle from pre-training, fine-tuning to the point you want to make different service of your applications on our platform. A lot of companies are doing dedicated instances. They will buy on-demand dedicated instances. When they need to use it, they can use our software system for whatever business usage they're looking for. The key highlight or the key distinguisher of our business, open source, supporting open source model versus the cold source companies, is that the customers own their own model.
23:16They own their data, right? Whatever fantastic technologies will provide you to make your model, making your speculator, making your serving strategy, training strategy wonderful, you own that piece of the model, the data sovereignty, also the security of the deployment. So we have to go through a lot of the compliances so that we can serve customers in North America and some of the other countries. And the models, the instance of the model that you own is on your cloud or where does it exist? Yes, these are open source models on our cloud. On your cloud. Other people can use them too. They can directly download them from HackingFace.
24:05Right. But I mean, an enterprise customer can have a private instance of an open source model on your cloud. Absolutely. Fine tune it for themselves. Yeah. Yeah. And what if build the future of multi-agent software with agency? That's A-G-N-T-C-Y. The agency is an open source collective building the internet of agents. It's a collaborative layer where agents can discover, connect, and work across frameworks. For developers, this means standardized agent discovery tools, seamless protocols for interagent communication, and modular components to compose and scale multi-agent workflows. Join Crew AI, Langchain, Llama Index, Browserbase, Cisco, and dozens more.
25:09The agency is dropping code, specs, and services. No strings attached. Build with other engineers who care about high-quality multi-agent software. Visit agency.org and add your support. That's A-G-M-T-C-Y dot O-R-G. Different things with us. For instance, you say, okay, I'm looking at, I want this model to be better, right? But I don't know how to do it. I don't know how to do it in the sort of the frontier way. So we'll have a team and helping you to get the strategies up from the pre-training data all the way to post-training data pipeline, make your model look better, performance better. And now you all the way coming to the inference side, we can also help you to customize the inference solutions so that you reduce your cost and you make your performance and also the accuracy to be in the very satisfactory range that the customers demand.
26:16And that's what we do. And the open source models, I mean, typically have lagged the proprietary models.
26:29Can you bring open source models up to compete with, you know, I think you mentioned 03 mini or? Absolutely. Yeah, absolutely. let me talk about the field and why our mission of supporting open source models gets so much attention. And I think today, you know, this mission is no longer, you know, something that people feel like, oh, you know, this may not be that great compared to closed source model. So this really changed when major players are started releasing open source models, which is super aligned with our original vision of creating the company. So models today, like DeepSeq V3 model, non-reasoning model, and DeepSeq R1 model, Lama 4 models that just released last week, multi-modality, state-of-art multi-modality models in different sizes, right?
27:26A super long context window. And, you know, all these models are very, very competitive. when R1 was released from DeepSeq, it was the biggest news for weeks because the model is so good, super competitive to open, to closed source models, like open API models, right? In many aspects it's better. So open source community finally can use these models for their own way of creating their own enterprise or businesses opportunities, right? Open source model is really head-to-head with cold source modeling companies. And really, I think really making a lot of them very nervous because the DeepSeq, for example, DeepSeq, they sort of not only open weights, but they also tell you exactly how they using created way of training the models through reinforcement learning, how they created a new attention kernel that's different than everyone else was using, you know, when everyone else was using multi-head attention, they created something called multi-head latent attention, which does different things, compress KVCache really, really well to a very small size.
28:41So all of these creations, open source, the open source, the libraries, the communication libraries and all of that really strengthened and the community's mission to one day we'll have the best models from open source community. And these days where really customers, users are really enjoying using this open source model because they're either better or the same level of the closed source models. And what do you think is gonna happen to the closed source models? I mean, these companies spend an enormous amount money on research right um they have to be used through an api which they lose a huge chunk of the enterprise market anybody who who needs to keep their data on uh premise yeah and and now Now, if these cheaper open source models can do nearly the same thing, it just intuitively seems that the proprietary system will eventually disappear.
29:56Because you can't keep investing tremendous amounts of money. I mean, particularly if it's primarily just in scaling. It came, you know, let's first not deny the amazing, you know, work that in the Transformer era that OpenAN and all the other pioneer companies that delivered since 2023. That was really the biggest innovations, I think, you know, for this decade or more. Now, in January, it was when DeepSeq models released, it was really the moment that the entire community realized that, okay, I can do this in the same level. Yeah, it's a good question. So my prediction is that these companies will have to figure out their model releasing strategies.
30:52strategies. I think more and more of them, including open AI, will likely to release their open source models. Now, what likely happened is you will have open source model that has a somewhat a gap of a call source model. Now, this open source model they're releasing will be more for, let's say, smaller devices, something more commodity. So they're still maintaining the sort of the cloud level, large system level reasoning models at trying to do that as much as they can, which is very competitive, very hard these days, right? Now, but I think their business model has to transform. But my understanding or my feeling is, they have to compete in the open source market.
31:44They have to maintain that top position. They have to open source model and beats everyone else so that the community understanding is okay, these companies still can do the best model. But that game just getting harder and harder and then distinguishing products is very difficult. Yeah. Yeah. Together is you're in, you're a, I guess, ML ops or a gen AI ops. ML system. Yeah. ML system. company. Do any of the proprietary model makers use Together AI's solutions or are you really focused on the open source community? I think that's related to the business NDA, but I think the understanding is the platform currently is being used widely across different enterprises.
32:52Some companies, many companies may have their own models and as closed source models or whatever, whatever for their internal business. But we have customers, they use our platform to do other things. Right. So I think, you know, the key here is the flexibility that we provide, the innovations we provide, as well as the data sovereignty, security, model ownership that we provide on our platform. It's just a very different business model for these companies that, unless, let's exclude one or two of these top open source companies, but everyone else, my understanding, can use our platform and can do, you know, work on that.
33:38Yeah. Yeah. I mean, it seems to me that space that you're in is increasingly crowded. I mean, I talk to a lot of companies that have, you know, that accelerate training, that make training cheaper, that can elevate open source to compete with the best reasoning models. How do you operate in that market? It is competitive, but we have the greatest reputations in this realm that provide the open source community the key technologies. They can operate these models today without flash attention and other works that we contributed. We could have just closed sourced them, right? We don't have to release them.
34:26But the way that we did that, you know, first, we have the community trust that we always produce the best solution. And the numbers show you can go to the different, you know, analytical sites like, you know, artificial analysis and others. And you can look at our numbers and superior numbers, superior cost, economics benefits to enterprise. Right. So I see this game as, you know, you have really have to have a really talented research team behind this operation. you can sell GPU hardwares, right? If I can just go to the GPU hardware resource world with you and just selling, beating on each other's price, like some companies does that.
35:14But if you look at their members supporting these models, right? And the flexibility they can customize these enterprise, like unique enterprise solutions for different companies because different companies are having their own serving conditions. Some companies want very extreme serving conditions. They want large batch size. They want a large throughput rating. They care about latency, but they probably don't really care about tokens per second. All of these different models, if they fine-tune or they make, how does that really perform on today's hardware? And that's sort of the question that we answer and we connect.
35:56So it's both a research leading team behind this whole operation of accelerated AI cloud. It's not about just using open source framework and get it run on your platform. And that's not enough solution for our customers, at least the ones that we have seen in the market. that's not yeah it's getting uh there are more players um but we're confident that you know we've been the leader in this community for a while yeah yeah uh and i i somebody uh an enterprise who's working with an open source model how how do they decide on who to use i mean other than your prominence in the market. And do you compete with people like Sarvam AI or Aleph Alpha?
Read the full transcript
36:55Do you know these guys? No. Yeah, I haven't heard of them. Or RunPod? Or RunPod, I know, yeah. Yeah. Are they a competitor?
37:09So I think there is a overlapping business there. I think there are some portion of it, but I think we're running a very different business models in the market that we support. Now, I can answer your previous question about who they choose, right? My experience talking to customers is, especially enterprise customers, their models are quite unique. Even if they're based on the open source models, when they provide their models, the process has played out where we have them to further enhance their model. Anyway, most of the time, they want these very extreme serving regimes, as I talked about.
38:02without a really sophisticated serving engine and without a team that behind it that can do, can connect modeling, gain firms, training, fine tuning together with just kind of the entire AI life cycle, you will have very hard time to produce that solutions for these enterprise customers. And that's what we have been really proud of that our team can do that. So, yeah, so I think that's what enterprise customers want. Not just take the open source model with open weights and say, okay, now I want to just use this model and then just find somewhere to serve it. Of course, our pricing are competitive too, very, very competitive.
38:50And we're trying to give our customers not only the ownership, but also the best economics. Every day, our team are working very hard to make that happen. yeah yeah um and and then um uh on the on the cloud you you operate a cloud is that right and uh is it primarily a gpu cloud i'm i'm very curious about these new accelerators built for inference sure i can i have more knowledge in that realm yeah sure I think, yes, for today, Together is primarily operating on GPU cloud. We are recently become this big partnership with NVIDIA for the NVIDIA universe of, you know, one so important partnership for cloud serving.
39:48So I want to talk about the technology because the technology that we create is we have this in mind that the future, you know, it's of course, Nvidia GPUs, other GPUs are fantastic. But as a technology company, we're also looking at other possibility to mix them into a serving solutions for our customers in the near future. What I mean is in this serving situation, you will not only have GPUs, but you will have other kind of ASICs. They perform different roles in terms of how you do the economics, what is the band about the customer's demand for latency and throughput. All of these, I think the complete solution would be a cloud that utilizes different kinds of a group of hardwares like what being provided on the market today.
40:45But as a company, we're very carefully evaluating because we're very experienced about user workload, very experienced about serverless endpoint use cases. So we carefully evaluated all these hardwares and then eventually our hope is to integrate this into the Together universe, to gather serving platforms. And we have a couple ideas that we can, you know, really reduce the cost and getting the throughput, you know, significantly improved with this hardware being part of the solution. Yeah, because, you know, I've had Andrew Feldman from Cerebris on a couple of times, and I've had Rodrigo Liang from SambaNova, and I haven't talked to Grok yet, Grok with a Q, but their argument is that their inference speed is light years ahead of NVIDIA and consequently cheaper.
42:01but it doesn't seem like there's a lot of uptake. Everyone's still depending on NVIDIA and I'm kind of waiting for that point at which people sort of shift the inference to another. My personal opinion, and this doesn't represent together's opinions, purely my personal opinion on this, is that ASICs are great but they have to overcome a lot of hurdles of being the primary and only hardware serving in a cloud setup for a large enterprise. So you can just look at the decoding number. So ASICs, when they make ASICs, the pure... Explain what ASICs is for listeners. It's like a customized hardware that you understand a generic version of this model.
42:56Let's say transformer model. it's a dream you understand it like a dramatic pattern you map that into these um uh very fine grained uh circuits and uh have electrons and all this gating and you make it the you know either data flow or whatever you make it mapped that in there that does something that um customized for this generic version of the model right then you can enhance things like i can give more shared memory or on chip memory so I know I'm getting much way bigger and I can if I know the current most popular models are like this that customize a pattern or hardware pattern for it now nvidia is a ai supercomputer but it's not just ai computer it's a general purpose with something called a tensor core which tensor core does the ai matrix multiplications in different kind of precision.
43:54So ASIC's advantage is that because you purely customize for a pattern, your likelihood of doing decoding, which is we talked about autoregressive generation of tokens, that process can be really fast because you customize the hardware and you control the latency, you control the throughput, you control the size of memory for it. But there are other factors for service, You have to consider the flexibility of this ASICs. If today we're serving at a range of dense model, let's say 400 billion parameter dense model. I'm only a mix of expert models we're serving, let's say 1 trillion, 2 trillion range.
44:42Okay, today you can spend a lot of time customizing this hardware. Now, in about a month or two, our field move on to even a different realm. Let's say I go beyond$2 trillion or I did an entirely new transformer engine or transformer architecture. Then these ASIC vendors, they have to adjust it really quickly, which is not, you know, you can say it's very possible, right? You know, because their expertise, but sometimes it's also less flexible. So GPUs benefits are very flexible because they're general purpose computing plus a AI core. Does that make sense? Yeah, absolutely. Yeah. Yeah. And just ASICs stands for what?
45:31That's a good question. I've been using this forever. I think it's, I don't want to say it wrong, but I want to say it accurately. It stands for, sorry, I don't want to say this wrong. Sorry, I'm just getting it. Is it a s?
46:05A6. A-A-C-Trips.
46:19Oh, it's application-specific integrated circuits. Oh, right. OK. I remember the first part. I don't remember. Yeah. Yeah. And is SambaNova an ASIC? Yes. yeah all of these all of these companies um uh you know grog some nova uh cerebus uh these are you know asic windows grok yeah yeah mentioned that okay um okay and then um what what's next for uh together ai where are you going with us yeah so um our goal is We're trying to provide this AI accelerators cloud. We want to expand the hardware capacity or diversity that we can serve our customers. We want software solution, provide better software solutions, especially focusing on inference.
47:25If you go visit our website, now you can see tons of applications have been using open source models that built through Together Platform and Together Solutions. So we want to kind of encompass all of these capabilities of our research and translating that research into our platform. So not only enterprise users that we really put a lot of efforts for them, but also the broader community that they can enjoy and use our applications and serverless endpoints. For instance, recently we released something called Together Chat. If you go to, I was playing around it this morning, right? If you go to Google Together Chat, you can register an account and then you have a bunch of models that you can select.
48:24And from reasoning models all the way to multi-modality models, multilingual for coding tasks, for image generation. So we have all these models serving within this OneChat format through power through open source models. And then we also going to inject a lot of the things like what we create called deep research, right? Which is combining search functions with mix of agents into enhancing the power of the together chat. So users will all have all of these packaged in this, in these applications that they can use. Of course, these technologies can be taken apart and then put into our software stack that use for enterprise customers as well.
49:11And this is just one of the examples that we're trying to do. We're trying to do the cloud really well, you know, through AI. And then we're trying to provide this super fast, highly efficient, both in performance control and economics. and also putting a lot of these AI research in the modeling and data into this platform and transform them into products for our customers and be competitive against the closed source modelings through open source, yeah. Okay, wow, that's fascinating. And I think I'm gonna end it there, but this Together Chat, is there a free layer? I think you can use it for, I'm not really sure up to what point, but right now I think it's a freely offer that you can use and test.
50:09Yes. Okay. Well, I'm going to go play with it. Yeah, you're going to go play with it. It's quite good. Yeah. Tourism đanged inutter rather than an x-tur so. Self-statement, that means body patience from its name. Thank you.
From the publisher
AGNTCY - Unlock agents at scale with an open Internet of Agents. Visit https://agntcy.org/ and add your support.
In this episode of Eye on AI, we sit down with Leon Song, VP of Research at Together AI, to explore how open-source models and cutting-edge infrastructure are reshaping the AI landscape.
From speculative decoding to FlashAttention and RedPajama, Leon shares how Together AI is building one of the fastest, most cost-efficient AI clouds—helping enterprises fine-tune, deploy, and scale open-source models at the level of GPT-4 and beyond.
We dive into Leon’s journey from leading DeepSpeed and AI for Science at Microsoft to driving system-level innovation at Together AI.
Topics include:
-
The future of open-source vs. closed-source AI models
-
Breakthroughs in speculative decoding for faster inference
-
How Together AI’s cloud platform empowers enterprises with data sovereignty and model ownership
-
Why open-source models like DeepSeek R1 and Llama 4 are now rivaling proprietary systems
-
The role of GPUs vs. ASIC accelerators in scaling AI infrastructure
Whether you’re an AI researcher, enterprise leader, or curious about where generative AI is heading, this conversation reveals the technology and strategy behind one of the most important players in the open-source AI movement.
Stay Updated:
Craig Smith on X:https://x.com/craigss
Eye on A.I. on X: https://x.com/EyeOn_AI




