In short
Eye On A.I. Podcast Episode Summary
Episode Details
- Title: #186 Ronen Dar: Maximizing GPU Utilization for AI with Run:ai
- Host: Craig S. Smith
- Description: This episode explores GPU optimization in AI infrastructure with Ronen Dar, the CTO and co-founder of Run:ai. The discussion focuses on maximizing GPU utilization amidst a shortage of these resources and the technological advancements that can drive AI development forward.
Key Themes and Topics Discussed
Introduction
- Introduction of Guest: Ronen Dar, CTO and co-founder of Run:ai.
- Background: Ronen's academic and industry experience, particularly in hardware, software, and AI infrastructure.
The State of GPU Utilization
- GPU Shortage: Discussion on the current GPU shortage affecting AI development and the increasing demand for efficient GPU utilization.
- Run:ai's Value Proposition:
- Focus: Enhancing GPU availability and utilization for machine learning teams.
- Technology: Run:ai's software enables better scheduling and orchestration of GPU resources, allowing multiple workloads to share GPU resources effectively.
Technical Challenges and Solutions
- Scaling AI Models: Challenges in scaling AI models, particularly those exceeding 70 billion parameters, and the complexity of coordinating large GPU clusters.
- AI Workload Management:
- Run:ai's platform integrates with various AI tools and frameworks, allowing for an open ecosystem where machine learning teams can operate efficiently.
- Significantly increases GPU utilization, with reported improvements of 2-3 times.
Future of AI and GPU Demand
- Growing Demand: Predictions of an exponential increase in demand for GPUs, driven by advancements in AI and the need for more computational power.
- Innovations in AI: Expectations of continued innovations in AI models, requiring more compute power and optimized GPU usage.
Challenges in Inference and Cost Reduction
- Latency and Throughput: Addressing challenges in achieving efficient inference, particularly for large language models (LLMs) and associated costs.
- Reducing Inference Costs: Techniques to reduce inference costs using GPU virtualization and better workload management.
Market Dynamics and Competition
- NVIDIA's Market Dominance: Discussion about NVIDIA’s significant market share and the competition presented by companies like Cerebras Systems.
- Global GPU Market Demand: Insights into how various regions, including the US and emerging markets, are approaching GPU utilization and the need for efficient management.
Environmental Considerations
- Power Consumption and Carbon Footprint: Concerns about the environmental impact of expanding AI infrastructure and the efforts to improve energy efficiency in data centers.
Conclusion
- Future Outlook: The episode wraps up with thoughts on the future developments in AI, GPU technology, and the challenges that lie ahead in balancing technological advancements with environmental sustainability.
- Call to Action: Reminder to stay informed about the rapid developments in AI and GPU technology.
---
Key Takeaways
- Run:ai is making significant strides in optimizing GPU utilization, which is critical in the current landscape of AI development.
- The demand for GPUs is expected to continue growing, driven by advances in AI and an increasing number of applications requiring high computational power.
- Efficiency in GPU management is becoming increasingly important, not only for cost-effectiveness but also for addressing broader environmental concerns.
Additional Resources
- Follow Craig Smith on Twitter: [@craigss](https://twitter.com/craigss)
- Follow Eye On A.I. on Twitter: [@EyeOn_AI](https://twitter.com/EyeOn_AI)
- Run:ai Website: [Run:ai](https://run.ai)
Note For a complete transcript of the conversation, listeners can visit the Eye on A.I. website at [Eye On A.I.](https://eye-on.ai).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Enterprises are securing access to GPUs and they are buying and reserving those GPUs. So there is this shift in how cloud resources are being consumed from on-demand fashion, just reserving blocks of GPUs. Research institutes are also doing that. Gen AI startups are also doing that. OpenAI, the most successful AI company right now, raised$10 billion two years ago. $10 billion mostly goes to GPUs. Passwords are the bane of my existence, and I'm sure they're the bane of yours. How do you make a password that's strong enough so no one will guess it and impossible to forget? Think of a password. Now add a number and a special character.
0:45No, not that one, a different one. And make it eight, no, 12 characters. Great. Now remember it forever and do it a hundred more times. Don't reuse it and make it so everybody in your company can do the same without ever needing to reset them. Sounds impossible. Unless you have one password. That's the number one, P-A-S-S-W-O-R-D, 1Password. 1Password combines industry-leading security with award-winning design to bring private, secure, and user-friendly password management to everyone. 1Password makes strong security easy for your people and gives you the visibility you need to take action when you need to.
1:31Companies lose hours every day just from employees forgetting and resetting passwords. A single data breach can cost millions of dollars. 1Password secures every sign-in to save you time and money. 1Password lets you securely switch between iPhone, Android, Mac, and PC with convenient features like autofill for quick sign-ins. All you have to remember is one strong account password that protects everything else. your logins, your credit cards, secure notes, or the office Wi-Fi password. Protect your global workforce with simple security, easy secret sharing, and actionable insight reports. I use 1Password myself, and I encourage my family members to use it also.
2:19I don't know how I would get along without it. 1Password's award-winning password manager is trusted by millions of users and over 100 ,000 businesses from IBM to Slack. It beat out 40 other options to become Wirecutter's top pick for password managers. Plus, regular third-party audits and the industry's largest bug bounty keep 1Password at the forefront of security. Right now, my listeners get a two-week free trial at 1Password.com slash IonAI. That's E-Y-E-O-N-A-I all run together. That's two free weeks at one, number one, onepassword.com slash IonAI. Don't let security slow your business down.
3:07Go to onepassword.com slash IonAI to get a two-week free trial. Hi, I'm Craig Smith, and this is IonAI. In this episode, we dive back into the world of AI infrastructure with Ronan Dar, CTO and co-founder of RunAI. Ronan talks about the challenges of efficient GPU utilization and inference in the midst of the ongoing GPU shortage. He explains how RunAI's innovative software platform optimizes GPU clusters, empowers machine learning teams, and shapes the future of AI deployment and scalability. I hope you find the conversation as useful as I did. Happy to be here, first of all, so thank you for inviting me.
3:57I'm Ronan, I'm the CTO and co-founder of RunAI. and RunAI is an AI infrastructure company. We started six years ago and we're an Israeli-born company. And so all of our engineering and product teams are here in Israel. But all the go-to-market teams, most of them are distributed globally, right, remote companies. So as time goes by, the company becomes half Israeli-based and have remote and distributed globally. So that's fun to see how the company evolves. Before running, I've been mostly in academia. So many years in academia, I did my bachelor's degree, my master's, my PhD in Tel Aviv University in information theory.
4:49So I'm an engineer, electrical engineer. So a lot of years in academia. In Israel, I did my postdoc afterwards. In the States, I've been there for a few years. I really enjoy the academia and doing research. That's fun. And in parallel, I've been also in the industry. So I worked for chip companies, and there are quite a few of them, quite a lot in Israel. So I worked for Intel and other chip companies here in Israel. So I've been an algorithm engineer in those companies. And so I gained a lot of experience around how chips are being designed, how chips are being manufactured and all of that, those things.
5:29And Run.AI is a software company. So, right, of course, experience also in software. So I guess academia and hardware companies, software companies, that's my background. And now in Run.AI. And so tell us about Run.AI. I mean, in our email exchange or my exchange with Ali, we were talking about the GPU shortage and ways to address that. But what is Run.ai at its core? Okay, so Run.ai is an AI infrastructure company, as I said. When Omri and I started, it was six years ago. So Omri is my co-founder, right? He's the CEO. And we met in the academia, both of us. We had the same supervisor. We worked quite a lot in the academia.
6:22It was really fun back then. Still very fun to work with Omri. He's an amazing guy. And we started back then, and both of us knew that we want to start a company. We want to be in the entrepreneur in some days. And the first decision that we got is that we want to be a team. We want to start a company together. The second decision was around AI. We want to be in the AI space. back in 2017, 2018, we saw that there is this new technology, AI. Back then, it wasn't that clear that AI is going to revolutionize the world. A lot of people spoke about it, but people also spoke about AI winter and whether AI will succeed and bring value to the world or not.
7:10And I think we saw that there is a big revolution coming in, and we wanted to be part of that. We wanted to contribute our part. and we wanted to lead our space right there when it comes to infrastructure so both of us as i said we were very close to hardware we know a lot of on hardware and we saw that compute is going to be a critical component in a app and a lot of gpus are needed to train AI models. And we saw the trend that started around 2012, 2013 of more and more compute, more and more GPUs being used to train state-of-the-art models and new innovation coming from the fact that having a huge model is being trained on huge amounts of data using huge amounts of compute.
8:07So compute is there together with data and the algorithms and compute is critical. And we saw that GPUs are relatively new in the data center. So GPUs were known in the gaming industry. They were very successful in that industry, but in the mainstream data center, they were quite new. And the entire software stack that runs clouds and data centers was built for CPUs and traditional applications, traditional microservices, and not for AI workloads, which are really compute intensive. They consume a lot of compute. They may run on distributed servers, distributed computing, and they run on GPUs, which is a new beast in the data center.
8:57GPUs, if you compare it to CPUs, they are much more powerful. They are much more expensive. They are much scarce also. So there are a lot of challenges around how to schedule workloads on GPUs, how to manage them, how to utilize them. And we saw that the software stack is not the right fit when it comes to GPUs and AI. And so we took a chance of rebuilding that stack. And we're still using cloud-native technologies like Kubernetes and containers. and we brought a lot of technology when it comes to just scheduling and orchestrating and fractionalizing GPUs. So we built this engine that runs within GPUs, within GPU clusters, that provides much more advanced scheduling capabilities and ability to fractionalize GPUs and allow more workloads to share the same GPU.
9:54and essentially we brought that technology that allows machine learning teams to get much more out of their GPUs and to get much more utilization out of their GPUs, much more availability after their GPUs. And we're seeing from our customers like 2x, 3x more GPU utilization, more GPU availability after they start to work with our software and they start to use it correctly. And you're getting huge, huge benefits on that aspect. GPU utilization and GPU availability. And we've built this platform on top of that engine, which simplifies for machine learning teams the way they train models and deploy models in production.
10:43So we built tools, and we've integrated with tools in the ecosystem, with open source tools out there. And we built these abstraction layers that allows machine learning teams to much easier way, much simpler way to train their models and take their models to production, deploy them and do inference. And so we've built all of that. We're working with a lot of great companies, great machine learning teams in enterprises and in startups. And we power the most advanced research labs today in the AI space. and it was fun to see that growth of R &EI and we're always thinking about the next step and continue growing it always becomes harder and harder to continue to grow but it was amazing six years and we're looking in 2024 I think it's going to be an amazing year because there are so many GPUs out there we started to speak about the GPU shortage So a lot of GPUs are going to be required in the market right now.
11:54And we're expecting some amazing growth this year. Yeah. And so I've been talking, as I said, to people about this issue from various angles. I just had Cerebris Systems on Andy Hawk, and I've had Andrew Feldman on one of the founders previously. And then I've talked to people about different strategies for chunking prompts and different ways to get through the rate limits that large model, foundation model companies put on because of the GPU shortage. you guys are focused on increasing utilization of each GPU. That's kind of the core value proposition, isn't it? Exactly. So the core value is around more performance from GPU clusters.
13:05Clusters, yeah. Think of AI, OpenAI, they have a lot of GPUs clustered in some way. And researchers in OpenAI use those clusters, those GPUs to train their motors. and OpenAI also running chat GPT on those GPU clusters. So the way all of those workloads are running on GPU clusters and sharing those GPUs and being scheduled on these GPUs and how they utilize those GPUs is what our software brings to the table, a much better way to use and utilize those GPU scheduler. So use clusters, yeah. And is your customer base cloud providers? I mean, does this sit between the cloud and the customer, and it's employed by the cloud providers to increase utilization of their GPUs, or is it employed by the customers on their end in managing their clusters in the cloud?
14:15Okay, amazing, amazing question. So we have, well, our software is being divided into two. So one software, the entire engine runs on GPU clusters, on GPU infrastructure that is owned by the customer. A customer can run their GPUs on premises or in the cloud, on Amazon or on Google or on Azure. We support all of them. We have customers all over. And then there is a control plane of us, which runs either as a SaaS component, as a software, as a service on our machine, and connects to the different GPU clusters that are out there and controls them and runs workloads on those GPU clusters. Or that control plane can run on the customer environment as well, as a self-hosted way or as an air-gapped environment.
15:14So we have quite a few defense companies, which are defense organizations who are our customers. And so we support that as well, the most secure, no internet environment as well. Yeah. One of the things that struck me in talking to Cerebris, their thesis is that you etch a silicon wafer and then you cut it up in the little chips. and then you when it comes time to to train a foundation model or some heavy compute exercise you stitch all these little chips together again so that they can work in parallel and their vision is you know why do you cut them up in the first place if you're going to have to run them and clusters uh and my understanding that that the the the the trick to a working uh with gpus and clusters is is the cabling the hardware cabling as much as it is the software can can you talk about that that's something i don't quite understand how can you address this purely from the software end?
16:35I mean, isn't there also a sort of proximity in cabling and a hardware side to it all, a physical aspect? Right. So on the infrastructure level, you have machines, racks, you might say, with GPU servers within them. and that can be GPUs and can be any other accelerator. There are other accelerators that are not NVIDIA GPUs. And those accelerators are essentially connected with communication cables and with usually very fast communication cables. and the idea behind it is that when you're training models and usually the models are so big when it comes to the state of the art models the models are so big that they run on a lot of gpus a lot of accelerators that are not just physically within one server or within one rack right so So GPT-4, for example, according to us, was trained on more than 20 ,000 GPUs.
18:01So usually in one machine, one GPU machine, you have eight GPUs connected very fast interconnects. And then you connect multiple servers, multiple machines like that. And also with very high interconnects, very high speed interconnects. and all of those gpus need to process together in a synchronized way a lot of data and essentially train a model on that data now models became so big in the last decade that you need to train GPT-4 more than 25 ,000 GPUs. 10 years ago, when the big breakthrough in deep learning happened around 2012, 2013, AlexNet was the model that triggered that big, high, big revolution in deep learning.
19:05AlexNet was trained using two GPUs. That's it. So GPT-4 was trained on, according to Moos, on more than 20 ,000 GPUs. So much more GPUs in 10 years. Now, GPUs also became much better. And also researchers started to train those models with much more data for longer time. So research has showed that the increase in compute power required to train state-of-the-art models grew 100 million times in 10 years. 100 million times. That's mind-blowing. And so we saw that trend, and let's see what the next decade will bring, right? If we'll see more models, bigger models that consume more data, that are trained on more compute power, more GPUs.
20:04And I believe that that trend will continue. I heard, I was listening to a panel on which a deep mind person was speaking, and she was saying that a sort of master's level computer science researcher can train, can on their own work with GPUs to train a model up to about 70 billion parameters. And she was saying that after anything above 70 billion, it becomes a really challenging engineering problem to get clusters to work together efficiently and that you end up with very low utilization. So you have to employ a lot of tricks as far as she went to get the utilization up and to get all of the GPUs working together to train above 70 billion parameters.
21:20And I assume that's where run AI comes in. I mean, maybe at a university, they do it in their lab, they do it on their own. But is that right? Is that what Run.ai essentially is doing? You're like scheduling an orchestration software that optimizes the parallelization between GPUs. And can you talk about how you do that? I mean, presumably there's some AI involved in doing that. How is that done? Yeah, so we're doing also that.
22:06As you say, to fine tune a model with 70 billion parameters, taking an open source model like LAMA2 and fine tune it with available data, You might need a few machines, GPU machines to do that, like four GPU machines. So one engineer or one researcher can handle that. But scaling up beyond that and training maybe bigger models or training from scratch using much more data, then that needs much more careful engineering and fine-tuning and configurations and making sure that GPU servers don't fail and things like that. So a lot of optimizations at that level. We do also that. We have customers who are training models on hundreds of GPUs, if not thousands of GPUs.
23:06So we have that. The thing that is unique to run AI is that we give a software platform that allows a team or a few teams to do the entire scaling regime in terms of training models. So you might have a researcher just opening a notebook, Jupyter notebook on a fraction of a GPU. You might have a few of them doing that. You might have some researchers in parallel training, fine-tuning LAMA 2 with 70 billion parameters on a few machines. And you might have more advanced users training on hundreds, if not thousands, of GPUs just running one job. And you might have also inference workloads running and serving actual requests coming from actual users.
24:00So all of those workloads are running on the same clusters, on the same platform, right? And all of them need to be scheduled and orchestrated and managed and utilize the GPUs efficiently. So that's where we shine OneAI. We allow all of that regime to happen on just one cluster or maybe multiple clusters and allow all of them to get access to those GPUs and utilize those GPUs in the most efficient way and in the most simple way. I'm curious how you do that. I mean, what are the algorithms involved and is there dynamic switching between GPUs within the cluster? And also, does RunAI, does a user access the cluster through RunAI?
24:57And so RunAI is the interface for the user, or is it kind of a bolt-on where the user is still working directly in the cloud with the cluster, but there's some side optimization going on? Okay, great question. So we're bringing a few algorithms, you might say, into the engine that runs those clusters. We're bringing scheduling algorithms. That's one thing. So a lot of scheduling decisions on how to dynamically allocate the GPUs to the different workloads and making sure that all the GPUs are being utilized and no GPU is sitting idle and making sure that all of that switching, that dynamic switching between workloads is happening in a fair way between all the users according to business goals, according to priorities of the organization.
26:05And making sure that always researchers can submit a lot of work, just fire up a lot of workloads to our system, to the cluster, and all of those workloads will be queued and be allocated and run on GPUs in a systematic, automatic way. So I think as a researcher, they can just launch so many jobs. Those jobs can be of any scale and everything will run in the most optimized way. The second thing that we're bringing is that software layer that sits on top of GPUs. And we do what we call an API level virtualization. We sit at the CUDA level and we intercept CUDA calls and we manage the access to the GPUs.
26:53So with that layer, we can allow multiple workloads to get access to the same GPU without colliding and without memories being used by the two workloads in a colliding way. So we prevent those workloads to collide. we bring this isolation level and each workload sees like a virtualized smaller gpu and so that is very significant when it comes to inference not all of the models when they run in production are very big and they require you know a whole gpu or a lot of gpus any times you have small models that needs just a fraction of the gpu so we bring that a CUDA level interception layer that controls the access to the GPUs and bring more efficiency in that sense.
27:45And when it comes to launching jobs, so we got a decision early on to build a product which is open and have a platform where users can launch workloads and deploy models using run AI tools, but also use any other tool that can integrate with a cloud-native platform. Our platform runs on top of Kubernetes. So any open source or third-party tool that runs on Kubernetes can integrate with our AI. And so we have a lot of customers who are actually not running with our tools but utilizing third-party tools or just launching jobs with a command line interface and things like that. So that was a decision that we got very early to integrate with the ecosystem because the AI ecosystem is moving so fast.
28:46Tools are being built every day and the advancement is so fast and we see different machine learning teams using a variety of tools. So supporting that and supporting the advancement as time goes by was crucial as we saw it. I think now with LLMs, we see it back again, right? It was all about deep learning a few years ago. And then JITGPT came and everything now is about LLMs and so many new tools are being built right now. And so many new tools are being used by machine learning teams. So staying open, that's a decision I'm glad we took early on. And then on the inference side, you know, the foundation model providers have currently have rate limits.
29:44I mean, they keep changing, but you can only put so many requests through per minute or so many. you can only call on the APIs so many times a day or various limits like that. Does this help a user of a foundation model increase their rate limits so that there is more throughput? or does that require the foundation model provider to be running AI or some optimization software? I mean, if I'm a big company, I want to put a product that calls GPT-4 into production, but I can't get either. there's a latency issue because I can't get my inferences through and back quickly enough. Does RunAI help solve that problem?
30:52As you said, latency and throughput are big challenges when it comes to inference, and it impacts the cost of AI, and the cost of AI right now is very high. and we see companies like OpenAI deploying rate limits and trying to make sure that the compute and the models are being used in the most efficient way. So I think there is a lot of work going on in the ecosystem, in the industry right now on reducing the cost of those LLMs and in different aspects, right? You have an amazing advancement happening at the serving layer, just on how LLMs are actually running on GPUs and how the architecture of the model is being run and using the compute power of the GPU, using the memory of the GPU.
31:52So a lot of optimization at that level are happening and amazing new things are coming out. there is also a lot of optimizations that can be done at the cluster orchestration level, at the compute management layer, as we do. We're not powering OpenAI data center. If we would, then they would get a lot more efficiency coming from all the better scheduling, better management. I'm sure they're working on that. I'm sure they're working hard on their orchestration layer. And that will help to reduce the cost. So when it comes to running AI, so we help companies who are deploying their own models or third-party models on their infrastructure in their cloud environments, in their virtual cloud environment.
32:50And we help them to host those LLMs in the most efficient way, right? Yeah.
32:59There seems, I mean, you were talking about building an abstraction layer or abstracting this problem into an orchestration layer, however you want to say it. I spoke with a company called Lightning AI yesterday or the day before, And they're building something called Lightning Studio that is an interface that pulls together all these different tools and services into one environment. And so you can very easily sort of click through different. Are you integrated into Lightning Studios, or are you another sort of a competing layer? So, okay, so Lightning AI, a great company. We're integrated with PyTorch Lightning.
33:58So the builders of Lightning AI are the creators of PyTorch Lightning, which is a great tool to launch PyTorch jobs. So we're integrated with that. So we have a lot of, of course, PyTorch users, and all of them are using different tools from PyTorch Lightning to maybe using Ray and running PyTorch on top of Ray. So there are several frameworks out there, open source frameworks, which we integrate with all of them. I mean, I get pitched a lot as a journalist, and I see these all-in-one platforms coming through. I don't have time to look at each one, but I guess what I'm asking is Run.ai is one of those tools that will be integrated into whichever of these higher level abstractions sort of wins the market?
34:58Or is Run.ai itself going to become one of those higher level abstractions that runs various tools? Passwords are the bane of my existence, and I'm sure they're the bane of yours. How do you make a password that's strong enough so no one will guess it and impossible to forget? Think of a password. Now add a number and a special character. No, not that one, a different one. And make it eight, no, 12 characters. Great. Now remember it forever and do it a hundred more times. Don't reuse it and make it so everybody in your company can do the same without ever needing to reset them. Sounds impossible.
35:40Unless you have 1Password. That's the number one, P-A-S-S-W-O-R-D, 1Password. 1Password combines industry-leading security with award-winning design to bring private, secure, and user-friendly password management to everyone. 1Password makes strong security, easy for your people, and gives you the visibility you need to take action when you need to. Companies lose hours every day just from employees forgetting and resetting passwords. A single data breach can cost millions of dollars. 1Password secures every sign-in to save you time and money. 1Password lets you securely switch between iPhone, Android, Mac, and PC with convenient features like autofill for quick sign-ins.
36:32All you have to remember is one strong account password that protects everything else. Your logins, your credit cards, secure notes, or the office Wi-Fi password. Protect your global workforce with simple security, easy secret sharing, and actionable insight reports. I use 1Password myself, and I encourage my family members to use it also. I don't know how I would get along without it. 1Password's award-winning password manager is trusted by millions of users and over 100 ,000 businesses from IBM to Slack. It beat out 40 other options to become Wirecutter's top pick for password managers. Plus, regular third-party audits and the industry's largest bug bounty, keep 1Password at the forefront of security.
Read the full transcript
37:23Right now, my listeners get a two-week free trial at 1Password.com slash IonAI. That's E-Y-E-O-N-A-I, all run together. That's two free weeks at 1, number 1, 1Password.com slash IonAI. Don't let security slow your business down. Go to 1password.com slash INAI to get a two-week free trial. Okay, so I'll try to help with that. So, yeah, companies who are training models, right, deploying models using AI, they have a choice when they build their AI platform or they build their AI tools. They have a choice whether to go with a managed and serverless solution, like maybe SageMaker. Probably you heard about Amazon SageMaker, their machine.
38:20So that's a managed solution by Amazon. And some companies go with such a solution, right? They don't control the infrastructure. They let SageMaker control the infrastructure and spin up the jobs. And it's all managed by SageMaker. So there are companies who are going that way. There are other companies who are controlling their infrastructure. They want to own their GPUs. They want to own their infrastructure and control and prioritize how it's being accessed and how it's being managed. And then we come in. So if there is a company owning their own GPUs, then we come in and we put that software layer that runs on top of those GPUs.
39:18Do you charge by usage? Are you a flat subscription model in their different tiers? How do you charge for this? And how does it impact the cost of inference? I'm most interested in the user. So we charge by the GPU. That's our business model.
39:41What we see, we work closely with NVIDIA. NVIDIA is our best partner. We're working very closely with their field teams, different levels. And there is a bundle that you can buy from NVIDIA people with a SKU that lets you buy GPUs with RunAI on top of that. And it's charged by the GPU. And then from what we see is that when companies are using RunAI, they're becoming much more effective in how they use the GPUs. Suddenly it becomes a no issue to get access to GPUs. It can be a big issue to get access to GPUs. So GPUs become a non-issue. It becomes so easy to spin up jobs, run workloads, and also deploy workloads much more efficiently.
40:46Deploy your models, right? Much more efficiently because in computer vision, models are relatively small and they don't take a whole GPU. So we have quite a lot of companies, customers who are deploying their computer vision models on run ai with llms it's also going that way models are becoming smaller and smaller when it comes to inference right when if you if you have just one task that your application is doing and you don't need a whole big general model very generic model, then you can fine-tune a small model and just use that and you'll get amazing results. So think about it, a LAMA 2 model with 7 billion parameters.
41:39So with 7 billion parameters, you need around 14 gigabytes of GPU memory. so now the newest GPUs are coming with 80 gigabytes of GPU memory. So you can store several Lama 2 with 7 billion models on just one GPU. So that brings massive cost reduction. You can deploy 10x more models on the same GPU. That's amazing, right? so those factors we have customers using in inference and just reducing the cost of inference yeah there are various strategies that people if they're not using run.ai use to bring down the cost of inference or to improve latency and relevancy I mean, you know, to get around hallucinations.
42:43And it seems that the, I mean, I'm trying to think of some of these strategies. I know that one of the things people are doing are splitting queries up into, you know, different pieces and then routing different queries to different models because you don't need a GPT-4 to do something simple like summarization. You could use LANA-2, but you may need GPT-4 to use something more challenging. So that's one strategy. There are other strategies. I mean, does Run.ai use those same strategies and you just take it off the shoulders of the user? Yeah, I mean, you talked about scheduling algorithms and things, but are there strategies like that that you're using?
43:46So a good question, first of all, because we're actually working on those things. And we're continuing to innovate. And now the last year, LLMs came out, right? LLMs, like a big breakthrough. And there are new needs of companies running LLMs. There are new pains compared to a few years ago, right? Just with deep learning. So some of those pains are around, as we said, throughput, latency, and cost, but also around how requests are being scheduled into the different servers, and how load balancing is being done, and how rate limits are being implemented, because rate limits can really hurt the utilization of the GPUs and how the compute is actually being used.
44:42So there are a lot of pain points around requests coming in. And those pain points grow and become more and more difficult to companies as they scale. And they have more and more design productions and more requests coming in, more users. But certainly there are a lot of problems there. And we're working now on solving them. There are also a lot of problems when it comes to starting and launching those LLMs and do autoscale. Maybe you heard about that. So autoscaling is a common practice in traditional applications, in traditional microservices. You run those microservices on CPUs, and when the load of requests increases, then you spin up more replicas of those microservices of those applications to support the increase of the load.
45:46And the other option is to run a lot of replicas always, and always be ready for your peak load, but then it will cost more. So the best practice is to do this auto-scaling, which adds more replicas and shrinks the number of replicas with the load of your users. So same thing in LLM can also be applied. Do this auto-scaling. You spin up more replicas of your LLMs as you have more users, as you have more demand. And that would bring down the cost because you don't need those replicas to always be on and ready for users and holding very expensive GPUs. But autoscaling in LLMs is very hard because a few reasons.
46:33First of all, it's not CPUs, it's GPUs. Spinning up GPUs takes much more time than with CPUs because you have much more drivers to install. GPUs are much heavier, much more time. LLMs are much bigger. It's not very thin, low-sized applications. We're speaking about models that can weight more than 100 gigabytes. LAMA2 with 70 billion parameters. And just the weight themselves is a file of more than 100 gigabytes. So just downloading that to your new replica can take minutes. So autoscaling is a big, big problem that has a big impact on the cost of inference. And it's a pain point that different companies are working on to solve.
47:33We also, we're working on solving that and bringing much faster autoscaling. it becomes a big pain when companies scale and they have more and more users and more and more LLMs in production. Yeah. Where do you think this is going? And what do you think of Cerebris Systems' wafer scale chip? And do you think NVIDIA and other chip makers will start going that direction? I mean, already they're integrating more and more GPUs into a system. But that just makes a lot of sense to me. Why split the processors up into separate packages and then link them together? Why not leave them all on a single wafer where the communication is much faster and cleaner and then, you know, build a package around that wafer scale.
48:47And not only that, I mean, not only that wafer scale idea, but do you see this, where do you see the GPU market going? you know there are all kinds of new fabs in the works that presumably will increase the number of silicon starts available and for chip makers and then that'll accelerate the production of whatever processor is required. and on top of that how does somebody like run ai keep up with what's happening i mean it's a moving target so you you create a solution for for today uh but but it's constantly evolving okay and big question so yeah the gpu market the gpu market we mentioned it before grew tremendously fast in the last decade.
49:53We spoke about 100 million X improvement, increase in the demand for compute power, compute operations, flops essentially to train state of the art models. And I think that demand for GPUs, we saw it in the last year, just jumping so high, right? Because of the chat GPT hype, LLMs, generative AI, everyone wants to get that. There's a new technology that's changing the world, changing a lot of industries. Everyone wants to get into that. And you need GPUs for that. So big jump in the last year in the demand for GPUs. And it took time to supply that demand. And so a shortage was created, right? because to supply that big jump in demand takes time.
50:45We were speaking about manufacturing fabs that need to adjust themselves to produce more GPUs. So that takes time. Usually you plan those things a few years in advance. So a sudden increase in the demand for GPUs will take some time to satisfy and to take time until those manufacturing fabs will adjust themselves to the new demand. But it happened right now. So, but in terms of demand, I think the demand for GPUs and accelerators in general will just continue to grow very fast. I think it's continued to go maybe in the same pace. And because we'll see more and more AI companies, we'll see more and more companies becoming AI driven companies.
51:37and AI companies, they rely on GPUs. They need a lot of compute. OpenAI, the most successful AI company right now, raised$10 billion two years ago. $10 billion mostly goes to GPUs. It's a 500-people company. Most of that money goes to GPUs. Mark Zuckerberg just a few days ago has said that Meta is buying 350 ,000 new H100 GPUs. And, right, so that's that amount of GPU worth, I think, more than$1 billion. $10 billion, sorry. More than$10 billion. That's a lot of money, right? So AI companies are spending a lot on GPUs and we see more and more companies like that. We also will see more and more applications, more and more new products being based on AI.
52:41And those AI will run on GPUs. Those products will have more and more users. Those products will scale and those products will need more GPUs. So more GPUs will be required. And also more innovation will come. Innovation in AI will come for sure. I think that we'll see bigger models coming out. More and more GPUs will be used just to create state-of-the-art models. The new breakthroughs will come from just using more compute power. So I think the demand for GPUs will go very fast in the next decade. With no doubt, I really believe in NVIDIA. We're speaking about Cerebras. I like Cerebras. They have an amazing offering.
53:29I think NVIDIA is an amazing company. I think NVIDIA, they know what the market will need in three years for now. They know they're already building what the market will need and what the next innovation will need in a few years. They're controlling the market right now. I think people speaking about 87 % market share of NVIDIA. I think it's much more than that. I heard numbers of like 95 % market share of NVIDIA. So crazy numbers. I think NVIDIA is an amazing company. We'll see, but more companies, more companies like Cerebras that will take market share. And that market is going to be so huge that second place will be a big company.
54:14Third place will be a big company. And there's a lot of market for everyone. Is your market primarily, I mean, presumably it's global, but there are a lot of countries that have limited access. I was just in Armenia, not just, this was four months ago, and they have a shortage of GPUs, and there are a lot of people in line ahead of them. And certainly China, I spent much of my adult life in China, I'm very close to what's going on there, And there's a real panic there over limited access to GPUs. So I would think that run AI would be an important technology for those markets, because if you're not going to get GPUs, the next thing you do is increase your utilization of those chips that you have.
55:18So, yeah, can you talk about where the demand for run AI is coming from? Yeah, the demand for run AI comes from enterprises who are getting decisions to secure access to GPUs. You cannot trust the cloud that your engineers or data scientists will try to spin up a GPU when they want it on demand and they will actually get a GPU. My engineers tried it. You can wait for days to get one of the newest GPUs on every cloud right now. So enterprises are securing access to GPUs and they're buying and reserving those GPUs. So there is this shift in how cloud resources are being consumed from on-demand fashion to just reserving blocks of GPUs.
56:28Research institutes are also doing that. Gen AI startups are also doing that. And in those use cases, we're shining. We allow them to manage all of those reserved GPUs and utilize them in the most performant way. And then we're shining. So, yeah. Yeah. So it's enterprises, but it's global. Does the demand correlate with the number of GPUs available in that market? I mean, certainly as much as there's a constraint in the U.S., that's where most of the GPUs produced by NVIDIA are sold. But there are markets that, you know, they want much more. So do you see high demand from those more constrained markets?
57:31Yes. So our biggest market is in the U.S. Most of our customers are coming from U.S. And we also have customers in Yemen. We also have customers in Israel, of course. We also recently have customers in the APAC region. We saw that growing in the last year. That was so fun. In Singapore, a lot of activities are happening around AI. There are some really cool companies starting there. And we're getting a lot of demand around enterprises, research institute, and Gen AI startups and governments, different organizations as well. Yeah, that's interesting on the government side.
58:20is the U.S. government using RunAI? Because they certainly have the same problem with everybody else. They certainly should. We have several federal research institutes that are using RunAI. You were saying NVIDIA is, you know, working three years or so ahead of the markets, preparing for what the market's going to need. Do you think that the whole, I mean, NVIDIA was kind of came out of the woodwork, right? It was really Jeff Hinton and Alex Kurzewski and Ilya Svitskoper that said, hey, let's use a GPU. And I'm sure it was being played with, but that's what gave AlexNet its edge in that and the data set that they were trained on.
59:30But people in the AI world were still working with CPUs and then GPUs came along. And that's what interests me about Cerebris. But I don't know that Cerebris, their architecture is the answer. But do you think that there will be a new architecture either coming out of NVIDIA or some other startup that will overtake? I mean, again, going back to this idea of why are we chopping silicon up into little pieces and then wiring them back together? I mean, it just seems to me that there must be intuitively an architecture out there that's more efficient. And are you guys watching for that if you think that there will be that shift?
1:00:25Okay. So you said a few interesting things.
1:00:36NVIDIA actually, they created CUDA. And you heard about CUDA, right? They created CUDA already in 2007. And I think around that, they started to work on that, I think. But it was early then, and they started to build CUDA because they saw that GPUs bring a lot of value to graphics applications. But they understood that essentially any linear, processing can be accelerated a lot with GPUs, any linear workload, any linear processing. And so they created CUDA for those applications. Back then, it was really hard to use GPUs and actually run workloads on a GPU. So they had this vision that any linear processing algorithm could run on a GPU.
1:01:41And for that, they created CUDA just to simplify that for a lot of users, right? So they won't need to engineer kernels that run on GPUs. It would be much easier. So when Jeff Hinton and his students came in 2012, and they had this amazing idea to use GPUs because deep learning essentially has a lot of linear processing algorithms. It's mostly linear algorithms. So they had this amazing idea to use GPUs, but they could do that because CUDA was there. Because NVIDIA already enabled that big vision of linear applications running on GPUs. So AlexNet was enabled a lot because of the GPUs and the software that runs on top of the GPU.
1:02:32I think when it comes to Cerebras and stacking up more GPUs, as you said, Now the problem is about scaling up or scaling out. How are you looking? Just shoving more GPUs into your workload, running the workload on much more GPUs and connecting all of those GPUs at very high speeds. Right now, I think the bottlenecks are not on the connectors themselves. There are ways to train those models on thousands of GPUs and make the communication not an issue, right? Not the bottleneck. But I think more this becomes bigger and bigger, that becomes a bottleneck and the communication will need to improve and improve faster.
1:03:22Cerebras are taking this approach of stacking more chips, right? And do it in a smarter way. Let's see. Let's see what... Let's see how it goes. Future applications. Yeah. Yeah, the metaverse, if it is fully realized, but, you know, there's been progress, is going to require just a mind-boggling amount of compute to render these environments in real time. How do you think that's going to play out?
1:04:02I mean, there are many issues with the metaverse. One is just the compression required. But from the hardware side, and is RunAI working at all with companies that are building environments for the metaverse, for Meta, for example? Yeah. The Metaverse is a big vision. And both Meta, Facebook, and NVIDIA are pushing that vision forward. They pushed it very strongly before LLMs and Generative AI came out, right? And now it's all about Generative AI. And in any case, I think we're going to have a future full of GPUs and a lot of compute power to run the metaverse or to run any other application that is going to be based on AI algorithms on the backend.
1:05:04So we'll see a lot of AI and we'll see a lot of GPUs that are being utilized to run those models in the future. Yeah, I mean, I have a lot of questions about the metaverse and how that could be realized. I mean, one of the things I think about... What's that? I have to say that I'm not an expert in metaverse. Yeah, sure.
1:05:36But beyond, I mean, not even talking about the metaverse, just the power consumption and the carbon footprint And at what point is the public going to start pushing back on these larger and larger data centers that are consuming more and more energy? Do you have any thoughts about that, about where that might go? Because certainly run AI, your technology, the more efficient you can be in using these things, in using the hardware you have, the less you need to add hardware and consume more power. What's your thoughts there? My thoughts, that's a big issue. the trend of having more data centers out there, bigger data centers with more carbon footprint and a lot of compute power being used right now.
1:06:51And it's only increasing. It's a big problem. I think a lot of smart people are working on that problem, always trying to improve the efficiency of the data centers out there and the compute consumption.
1:07:10But I think, unfortunately, that it will only increase. We'll see more data centers, bigger data centers, more compute being consumed. And I don't know where it's going, actually. It's a big problem, but it's just going to be worse. Maybe it'll go to the moon. That's what I heard some people talking about. The moon will become just a big data center that's streaming data up and down. It's possible. It's possible. Humanity was always creative. Let's see what smart people will do. That's it for this episode. I want to thank Ronan for his time. If you want to read a transcript of the conversation today, you can find one, as always, on our website, IonAI.
1:08:00That's E-Y-E hyphen O-N dot A-I. In the meantime, remember, the singularity may not be near, but A-I is changing our world, so pay attention. Passwords are the bane of my existence, and I'm sure they're the bane of yours. How do you make a password that's strong enough so no one will guess it and impossible to forget? Think of a password. Now add a number and a special character. No, not that one, a different one. And make it eight, no, 12 characters. Great. Now remember it forever and do it a hundred more times. Don't reuse it and make it so everybody in your company can do the same without ever needing to reset them.
1:08:42Sounds impossible. Unless you have 1Password. That's the number one, P-A-S-S-W-O-R-D, 1Password. 1Password combines industry-leading security with award-winning design to bring private, secure, and user-friendly password management to everyone. 1Password makes strong security, easy for your people, and gives you the visibility you need to take action when you need to. Companies lose hours every day just from employees forgetting and resetting passwords. A single data breach can cost millions of dollars. 1Password secures every sign-in to save you time and money. 1Password lets you securely switch between iPhone, Android, Mac, and PC with convenient features like autofill for quick sign-ins.
1:09:35All you have to remember is one strong account password that protects everything else. Your logins, your credit cards, secure notes, or the office Wi-Fi password. Protect your global workforce with simple security, easy secret sharing, and actionable insight reports. I use 1Password myself, and I encourage my family members to use it also. I don't know how I would get along without it. 1Password's award-winning password manager is trusted by millions of users and over 100 ,000 businesses from IBM to Slack. It beat out 40 other options to become Wirecutter's top pick for password managers. Plus, regular third-party audits and the industry's largest bug bounty, keep 1Password at the forefront of security.
1:10:27Right now, my listeners get a two-week free trial at 1Password.com slash IonAI. That's E-Y-E-O-N-A-I, all run together. That's two free weeks at 1, number 1, 1Password.com slash IonAI. Don't let security slow your business down. Go to 1Password.com slash INAI to get a two-week free trial.
From the publisher
This episode is sponsored by 1Password. 1Password combines industry-leading security with award-winning design to bring private, secure, and user-friendly password management to everyone. Companies lose hours every day just from employees forgetting and resetting passwords. A single data breach costs millions of dollars. 1Password secures every sign-in to save you time and money.
Right now, my listeners get a free 2-week trial at: https://www.1password.com/eyeonai
In this episode of Eye on AI, join us as we explore the cutting-edge world of GPU optimization with Ronen Dar, CTO and co-founder of Run:ai.
Delve into the intricacies of managing and maximizing GPU utilization in an era marked by a severe GPU shortage. Ronen shares his insights on how Run:ai's innovative software is revolutionizing AI infrastructure, making GPU resources more efficient and accessible.
The conversation spans the technical challenges of scaling AI models, the evolution of GPU demands from basic algorithms to complex systems like GPT-4, and the strategic innovations helping enterprises overcome these hurdles. Ronen also reflects on the future of AI development, predicting an exponential increase in demand for computational power and the innovative solutions poised to meet these needs.
Tune in to uncover the technological advancements that are propelling AI capabilities forward and shaping the future of AI deployment across industries.
Don't forget to like, subscribe, and hit the notification bell for more deep dives into the technologies that are transforming our digital landscape.
Stay Updated:
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI
(00:00) Preview
(02:46) Introducing Ronen Dar
(04:00) Ronen's background and RunAI's origins
(09:13) The need for efficient GPU utilization in AI
(13:14) RunAI's core value proposition
(15:33 RunAI's deployment model
(18:55) The growing demand for compute power and GPUs
(22:08) Challenges in scaling models beyond 70 billion parameters
(27:52) RunAI's open platform approach
(31:00) Addressing latency and throughput challenges in inference
(34:36) RunAI's integration with AI tools and frameworks
(39:37) Reducing the cost of inference with GPU virtualization
(43:54) Challenges in auto-scaling for large language models
(47:06) The future of the GPU market and demand
(50:49) NVIDIA's dominance and the role of competitors like Cerebras
(54:20) RunAI's global customer base and demand patterns
(57:52) NVIDIA's vision and the evolution of GPU architectures
(01:01:25) Compute requirements for the metaverse and future AI applications
(01:03:56) Concerns about power consumption and carbon footprint




