Powering AI with the World's Largest Computer Chip with Joel Hestness - #684

13 May 2024 · 55 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The TWIML AI Podcast Episode #684: Powering AI with the World's Largest Computer Chip with Joel Hestness

Podcast Overview Title: The TWIML AI Podcast Host: Sam Charrington Guest: Joel Hestness, Principal Research Scientist at Cerebras Episode Title: Powering AI with the World's Largest Computer Chip Description: Discussion on Cerebras’ custom silicon for machine learning, the Wafer Scale Engine 3 (WSE3), and its evolution to support large language models.

Key Themes and Discussions

  1. Guest Background
  2. Joel Hestness has a PhD in heterogeneous processor design and experience in large-scale language and speech recognition models at Baidu Research.
  3. He emphasizes the need for hardware to scale up machine learning applications, particularly in natural language processing.
  1. Introduction to Cerebras
  2. Cerebras is known for its unique hardware architecture, particularly the Wafer Scale Engine (WSE), which is designed for large-scale training of language applications.
  3. The WSE is a single large device that simplifies model training by avoiding the complexity of traditional GPU setups.
  1. The Wafer Scale Engine (WSE3)
  2. Unlike traditional chips that are cut from a wafer, the WSE retains the wafer structure, allowing it to function like a single large device.
  3. The WSE architecture includes:
  4. Homogeneous Design: Unlike GPUs or TPUs, which require complex coordination, the WSE operates as one unit.
  5. Memory Architecture: 40 GB of SRAM memory, vastly exceeding traditional GPU memory capabilities, enabling high-performance training of large models.
  1. Software and Model Support
  2. Supports open-source ML frameworks like PyTorch, and additional support for TensorFlow.
  3. Weight streaming stack allows the WSE to handle large models efficiently by separating weights from activations.
  1. Innovations in Training Techniques
  2. Weight-Sparse Training: Streaming non-zero weights leads to significant bandwidth and computational savings.
  3. Advanced Optimizers: Research aims to implement second-order statistics optimizers like KFAC and Shampoo to improve training efficiency.
  4. Activation Sparsity: Techniques to reduce computation during training by ignoring zero values in activations.
  1. Comparison with Other AI Hardware
  2. The WSE offers unique advantages over TPUs and AWS Inferentia, particularly in scaling and simplicity of use.
  3. The architecture allows for efficient execution of matrix operations without the need for complex parallelization strategies.
  1. Applications and Use Cases
  2. Partners include Core 42 for language modeling in the Arabic-speaking world and collaborations with medical institutions like the Mayo Clinic for drug discovery and patient outcome analysis.
  3. The WSE's capability in handling large language models opens doors for fine-tuning and alignment of pre-trained models.
  1. Future Directions
  2. Cerebras is working on expanding the capabilities of their hardware to support more innovative model architectures and optimizing existing models for deployment.
  3. The potential for the WSE to be used in high-performance computing tasks and other AI applications beyond language models is being explored.

Key Takeaways

  • Cerebras is pioneering a new approach to AI hardware with its Wafer Scale Engine, designed specifically for large-scale machine learning applications.
  • The architectural innovations of the WSE allow for efficient training and deployment of complex models, helping academic and commercial organizations improve their AI capabilities.
  • The discussions reveal a strong emphasis on research and open-source contributions, aiming to advance the field of machine learning while providing practical solutions for users.

Conclusion This episode provides a deep dive into the innovations at Cerebras and how its hardware is positioned to tackle the challenges of modern AI workloads, particularly in natural language processing. Joel Hestness shares valuable insights on the intersection of hardware design, algorithm optimization, and practical applications in the industry.

For complete show notes, visit [twimlai.com/go/684](https://twimlai.com/go/684).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:03All right, everyone. Welcome to another episode of the TwiML AI podcast. I am your host Sam Charrington. And today I'm joined by Joel Hesnes. Joel is the principal research scientist and lead of the core machine learning team at Cerebris. Joel, welcome to the podcast. Thanks for having me, Sam. Pleasure to be here. Pleasure to meet you. And I'm looking forward to digging into our conversation. We'll be talking a bit about Cerebris and the ways that hardware innovation is driving AI forward. To get us started, I'd love to have you share a little bit about your background. So my background started in heterogeneous processor design.

0:43I did a PhD 2010 to 2016. The good old days? The good old days before this crazy ML effort that's going on now. We were looking at heterogeneous systems that included CPU and GPU processors and sort of looking at how those applications could coordinate or how we could develop applications that could coordinate between those two kinds of devices. It turned out that machine learning was one of those good applications that makes a lot of sense there, where you have phases of computation that are very parallel and phases that are somewhat serial, and you want to get through them very quickly. And so that makes a nice dichotomy with the parallelism of GPUs and sort of single threaded nature of CPU cores.

1:35After my PhD, I went to Baidu Research. I was there for just about three years, and we were looking at very large scale language and speech recognition models. One of the things that we were studying while I was there was scaling laws for these models. It was one of the first deep learning scaling laws papers that we put out, essentially showing that we could predict how much compute in terms of model size and data set size that we needed to get to particular accuracy levels for different kinds of applications. And then sort of based on that study, I was kind of looking around. I knew hardware was going to be the big thing.

2:21I knew that language was going to be the hard application for us to solve. and so I use that as motivation to jump over to Cerebris where we're building hardware that's specifically targeting large-scale training for language applications we've broadened that out a bit at Cerebris where we're now doing multimodal but still large-scale and training applications I think I'd like to start by trying to position Cerebris from an approach perspective um the ai hardware space is exploded uh over the past few years you know along with ai in general i guess and so you've got all of the major hyperscalers producing their own ai chips you've got standalone folks like cerebris and salminova and others that have you know each kind of have a unique vantage point to trying to innovate around hardware, not to mention, of course, NVIDIA with GPUs.

3:28How do you think about all of the different offerings and the different ways that folks come about trying to accelerate machine learning? So my take on it was coming from a background where we were looking at the kinds of things you could do to get improvements in deep learning applications. so it was pretty clear that we needed scale so um in 2017 we predicted you'd need something like 10 000 times more compute than we were using at the time i think we're about uh maybe 2 000 x into that today uh so um we i definitely had the view that uh it would take large scale And another unique opportunity that I had at Baidu is we were working with some of the first data centers that used many GPUs.

4:23So at the time, we were running jobs on hundreds of GPUs. And trying to scale that out and run across those different GPUs was a very complicated distributed programming situation. So recognize the need for the hardware to be simplified or abstracted at least, and to have sort of software stack come in and make it easier for us to program such a large number of devices or a large amount of compute. Coincidentally, at Baidu in their research group, they were doing diligence on startup companies that Baidu Ventures was investing in and Cerebrus was one of those. So I got to kind of got an overview of Cerebris before I started looking around.

5:14And the solution to me came across as very unique because it programs like a single device, a single very large device. So that takes out the complexity of splitting up your models and sort of trying to distribute the chunks of the model and communicating, coordinating between them in order to do computation for these large language models. The approach is fundamentally one of trying to look like a GPU, but abstracting a distributed collection of otherwise semantically traditional GPUs behind a single device abstraction. so uh the nice thing is it's not actually an abstraction that the device is actually the scale of tens of gpus so um our our solution is the cerebus wafer scale engine which is a full wafer instead of cutting the wafer apart into chips and then repackaging it onto separate devices and then stitching it all back together with with software we just leave it together as single large-scale device.

6:39It takes a lot of, it's fairly complex to package this into a system. We got to figure out how to power it. We power it from one side, we cool it from the other side. And it's, we actually have - When you say sides, like sides of the horizontal wafer, the top is power and the bottom is cooling? It's actually parallel to the plane of the wafer. So the power is coming in from one side and the cooling is mated up against the other side. It's a pretty unique packaging compared to what we're used to, where we package something on a piece of a board, I guess, and package it into a system. So there's an intricate mating process where we sort of fuse these pieces together into a block and then we connect it up to cooling and, you know, a power array in the box.

7:37And so, and the result is that that appears like a single large device, like a single GPU. And then, so in our history at Cerebrus, back when I started, we had been building our software stack to program this kind of like FPGAs, where we would lay out the whole computation of a model on the wafer. That stack is very interesting because it gives you ultra low latency for inference, for instance. But we were targeting training first as a company. So we wanted to support large language models as those were coming up. So we actually built out another software stack, which we call our weight streaming stack.

8:25And the idea there is the weights actually sit off on a separate system, like a parameter server system. And we stream those weights into the wafer for each training step. Then on the wafer, we're storing the activations for the model. And when you do it that way, it ends up looking a lot like a GPU, that you launch a kernel, We load in the memory for the computation that you want to do. We do the computation, and then that kernel finishes, and you go on to the next one. So the whole wafer behaves like a single large device in that setting.

9:08That abstraction has helped a lot, I think, because it helped us separate the, you know, in a pipelined execution mode, we have to lay the whole model on the wafer. And so we're limited by the amount of memory that we have on the wafer, which our current generation is about 40 gigabytes of memory. So that's, you know, 100 times larger than what you'd see on a single device for something like a GPU. When we switch it over to the wait streaming mode, we get this 40 gigabytes that can be used to store all the activations. and that's a huge amount of activation memory for these models so that's it's really nice it gives us a lot of flexibility for how to how to operate is that 40 gigs comparable to um the gpu memory when we conventionally talk about the gpu memory or something else in the gpu the interesting thing about it is it's sram it's like a cache on the wafer so it behaves like cache in a GPU, which the caches in a GPU are on, you know, on the order of about 50 megabytes total.

10:19We have 40 gigabytes. So actually maybe about a thousand times larger. The so the distinction there is instead of going out to a memory, which a memory has a fairly high capacity on a GPU like HBM memory, I think on H100s has about 80 gigabytes of memory there. It's sort of bandwidth limited, getting that data to the chip. On ours, it's directly on the chip. And so then we only have to consider the bandwidth coming from our wait servers into the wafer. What I'm hearing you say is that it's kind of very similar to GPU architecturally, but on a chip as opposed to more discrete like a traditional GPU.

11:10Yeah, we definitely previously viewed ComputeGraph as like the whole computation. We want to get the whole computation on the wafer, and that was our pipelined mode. But now that we've sort of switched it around to this more time multiplexed kernel-centric mode, we can disaggregate the graph much further and sort of break it down into chunks. um so yeah it behaves a little bit more like running kernels on a gpu now there are some things in there that are um there are certainly interesting ways to optimize that computation given the the wafer architecture that we have uh so for instance in our um in our attention mechanisms we have uh a way of streaming the so if you're familiar with these they get broken down where you calculate some keys that are associated with the tokens that you're operating on.

12:11And then you compute some queries and you have to calculate the cosine simile, the dot product between each key and each query roughly. So we have a mechanism for doing that, which is going to stream one of that set of activations across the other set of activations. And that's sort of different than our weight streaming mode where we're streaming the weights. But essentially those kernels have the same underlying software architecture. So you can make optimizations that are like that. And we can do things like fusions and whatnot. So our stack compiles all this down to run on the wafer. We can sort of find places in the compute graph where it makes sense to fuse operations and then fuse those in the compute graph to make sure we get high performance or high utilization.

13:04I think that's actually a very, that's very analogous to what you'd see in many different inference settings where at inference time, you only care about latency. And so then you end up looking at big chunks of the graph and trying to sort of flatten them or fuse all the components together to run through them very quickly. It sounds like you're saying the graph stuff that I remember hearing, like that was one of the things that Cerebrus was talking about at the time. Did the hardware change at all? Yeah, that's a good question. We just announced our third generation wafer scale engine. And we've gone through a few different sort of minor adjustments to the architecture through our generation.

13:53So our first wafer had about 400 ,000 cores. And it was equivalent to about 25 V100 GPUs in terms of performance. That had a sort of a narrow vector width and could do floating point 16 computation. our second generation architecture went to 850 ,000 cores. So we doubled the number of cores. We didn't change the core architecture at the time. So it was still a vector width of four. This recent announcement of our CS3, our third generation, that's just a small increase in number of cores to about 900 ,000 cores. But we increased the vector width to eight. and so that effectively doubles the throughput of each core.

14:45So we got just about a 2x performance lift with this latest generation that we're building now. You know, at the time that you were talking about graph, deep learning use cases were a lot broader maybe. Scale was lower and use cases were broader, but we've kind of centralized in on, you know, transformers and language models and is the change that the software has been kind of tuned to those kinds of workloads definitely yes uh i think the there was a recognition in our pipelined execution effort um roughly 2019 sort of shortly after i joined that it was going to be hard for us to do very large models actually some of the first projects that i worked on were figuring out whether we could get the same benefits of large models, but still run in our pipelined mode.

15:40And so we decided to make that switch over to a wait streaming execution mode. And that was essentially born out of the recognition that if you plot over time the scale of models, you can see it's an exponential increase. You have to actually look at the vertical axis on log scale, and it's linear, right? And when we saw that, we said, hey, we know that we need to be able to train larger models. So we need to shift our software architecture so that it would support that more easily. Our third generation hardware actually adds the ability to hold our parameter servers can hold up to 1200 terabytes of memory capacity.

16:29So that'd be NVMe, kind of RAM memory. That would allow us to train models that are tens of trillions of parameters if we needed to. In practice, we have customers that are interested in training models that are hundreds of billions or trillions of parameters now. So we're working up towards that in the process. So definitely we're following the trend in the field on that. I do think over time, what's going to be interesting here is we've also been expanding out our software offering in terms of the kinds of models that we can support in this new execution mode. So the way that you design kernels ends up being different also, where you want to make sure that you can saturate and fully use the wafer.

17:25And so for something like a matrix multiply, if you're doing a reduction, you might have to do a reduction across the full dimension of the wafer. We have to sort of take that into account, make sure that we're still getting good utilization, even when you're doing communication that would cross the full wafer. So there are some design aspects that are going into the software stack. And then some of those experiences that we have are going to be informing our future generations of the hardware also. Can you give us a sense for the ways in which your approach differs from what we see in TPU or in the case of AWS, Inferentia, these kinds of chips?

18:11Sure. I'm familiar with the TPU from back at our time at Baidu. The core architecture of the TPU is a systolic array. So it's set up to do matrix-matrix multiplication. on the whole our wafer is essentially built to do the same thing except it's broken down into what i what i mentioned eight wide vector execution so um in that structure we can actually do a scalar times vector multiplications at throughput so at peak performance so it's a little bit more fine grained in terms of what we calculate but then when you take it in aggregate across the wafer it behaves a lot like a systolic array execution and then we have execution units in there that do transcendental functions and other things for things like non-linearities in these models.

19:21When you combine all that together it probably looks a lot like individual kernels that you would run on a TPU or a GPU. I believe Inferentia is similar. So I think if I recall, their architecture might have extra execution units, like there's a systolic array and then there are executions for nonlinearity, execution units for nonlinearities. So that's a little bit different than what we have, But it's taken on the whole, I think, fairly similar. But then sort of because we have this large scale wafer where you would train a super large scale model on these other devices like GPUs or TPUs, you would use, you would probably break up the calculation that you're doing and split it across devices.

20:19So for a large scale matrix multiply like GPT-3, you would probably shard the matrix multiply so that, you know, a set of columns of the weights are on one device, the next columns on another device. You run those executions, those kernels in parallel, and then you'd combine the results afterward. All of that happens on our wafer, and it's just sort of part of our kernel architecture. And so it can be easier to execute something like that conceptually. You mentioned that there's specific models that you support. And so it sounds like that remains an issue. Like you can't just deploy arbitrary models to the wafer.

21:08It's got to be a pre-supported model. Am I hearing that correctly? It's not exactly the model architecture that's supported, but sort of the basic building blocks of large language models were the first things that we built out in this new stack. So we support essentially all the important large language models today. And that's because we support matrix multiplies, the nonlinearities, the different position embeddings and attention architectures. So basic transformer backbones are all supported. And we just released in our latest software release support for multimodal applications. So now we're doing transformer based vision models also.

21:57And we can combine those with language models to do multimodal training and inference. And so from a developer user experience, am I working with my existing tools, PyTorch or whatever? And your libraries are handling compilation to the wafer transparent to me? Or do I need to be aware of the fact that I'm working with this specialized hardware? Yeah, so we do our best to make it transparent that you're using specialized hardware. We do support PyTorch. Basically, all the ops in PyTorch are things that we can consume into our stack, and we can compile most of those down to the wafer at this point.

22:50We also previously supported TensorFlow, so we still have some support for that if people are still using it. That sounds like a test. Yeah, well, we've seen a trend that most publications and most researchers prefer PyTorch. And even, you know, we've been hearing the cool kids prefer Jacks these days. So we're setting up our stack in such a way that we can sort of abstract. So our compiler stack is based on LLVM. And so we use MLIR intermediate representations to compile. So if you can extract MLIR from one of these frameworks, it is possible to bulk it up to our stack. And we, you know, we go try to do the last mile on every bit of software.

23:44So PyTorch, we have a package that you can pip install that gives you Cerebrus capabilities. We have our model zoo now, which can be pip installed. So you'd be able to even run our models. Some of our models even will run on GPUs or CPUs. Like if you just want to test out what our software stack looks like. And we also try to release the implementations that are extremely simple. So, we sort of took a cue from Andre Karpathy last year in his nano GPT release. 600 lines of code can run up to something like a billion parameter model on a GPU. We have a similar release, which we call Giga GPT. It's also 600 lines of code.

24:36But on our hardware, you can run models that are up to about a trillion parameters. so um and it's essentially the same code um we we wanted to sort of emphasize that that's uh that's what our capability is you get to as a researcher you get to work in a code base that's very simple to make improvements very quickly you can sort of focus on the ml instead of focusing on how do i I distribute my computation across many devices. Is the core idea that you're changing the price performance that you're able to reach, or is it you're trying to push absolute performance or something different than those two?

25:24Yeah, that's a good question. I think in order to be competitive as a startup, we've got to push on all dimensions. So right now, I would say we are price performance competitive with contemporary hardware. So we consider NVIDIA and maybe AMD to be the direct competitors there. In terms of there are a lot of other dimensions that people take into account when they're trying to consider whether they should buy certain hardware. So things like total cost of ownership, the footprint in your data center, the power footprint. So we are probably better on performance per watt than most device-based solutions just because we don't have communication between devices.

26:23we can train models with data parallel instead of orchestrating a lot of communication for large-scale training. And so that just means that we eliminate a lot of that extra hardware that costs power. So we're trying to be competitive on all of those fronts. The absolute performance, though, is something that we're pushing. like this release of our latest generation of hardware is equivalent to about maybe 20 H100 GPUs in a single box. So it's definitely a very interesting offering if you're daily trying to figure out how to run on many GPUs that are something like H100s. Is a wafer delivered as a card or as a machine?

27:20We package it as an appliance, so it's a full machine. As an appliance? Yep. And is there a notion of kind of networking multiple machines to create a cluster or a supercomputer of sorts? Definitely, yes. So we have two large-scale clusters currently that are 64 systems, and we've just announced a third. These are the Condor Galaxy cluster series that we've been bringing up. The third one is in conjunction with our partner Core42 that's being brought up in Dallas, Texas. And that one is going to be based on our new generation of hardware, the CS3. So when we network these together, we have some specialized components that are specifically the parameter servers, which are called MemX nodes, MemoryX.

28:18Those are large memory CPU-based systems, and we have to network those with our devices to send the weights through. In order to send the weights through, we have a separate set of gear called our SwarmX nodes, which are networking cards. And those can do a lot of specialized primitives for that sort of computation. So we can distribute the weights. So they broadcast the weights out to the CS2s or CS3s. And then when they're sort of aggregating the results back together, these would be like gradients coming back to the parameter server. Those can be added along the way. And so we get a tree-based architecture for doing the collectives between the two.

29:08it adds some kind of interesting opportunities which which my team tends to look at which are things like we can because the parameter server is often a separate set of nodes and the computation there doesn't interfere with what's going on on the wafers you can do non-trivial optimizers for these models if you're training them so you could consider things that'd be like second order approximations to different optimization techniques. Some prior works have done things like the add a factor optimizer, which will do approximations to Hessian calculations. That'll be sort of at sort of narrow spots in the model, like for a particular layer, it will do approximations to second order optimization statistics.

30:01Those can just run off on the side off of our wafers, where if you're running that same thing on a TPU or a GPU, you might end up running that directly on the device. And if you do that, you don't want to spend much compute to do that. Alternatively, if you want to try to do it off on a separate set of devices like CPUs, you got to do this very careful orchestration between the two, which is sort of taken care of by our software stack. And so are those optimizations part of the standard pipeline for LLM training? Or are these things that, you know, that are being explored in research as? Yeah, these are in certain organizations, these are the standard.

30:55So for instance, I think at Google, their Gemini and Gemma works might be based on some of their prior work, which does use more advanced optimizers like Adifactor. um so maybe for them it's considered the state of the art um a lot of the organizations we're working with are are maybe more um they're they're not considered considering state-of-the-art techniques like that but they are still using um the sort of bread and butter large language model optimizers like Atom and AtomW. For someone adopting these clustered systems, you now introduce kind of two layers of software. The first is the compiler that compiles code to run on the wafer.

31:51And the second is some kind of distributed parameter server type of scheme. makes me think of like a Horovod or something like that. Is it those kind of the two main pieces? Yeah, that's a good way of sort of breaking it down. Something like Horovod when it came out was basically just focused on data parallel computation. That was sort of maybe an inspiration even for our approach that it can be that simple. Like we just used distributed data parallel execution currently. And we're able to scale that out to, so we're easily scaled out to our 64 systems in our current clusters. And we believe we'd be able to scale it out even to something like 200 total systems.

32:43So that'd be on the order of something like a few thousand GPUs in terms of performance. And so that would be just data parallel rather than needing to do things like model parallel, you know, pipeline or tensor model parallel execution. And then, yeah, the stack does also the lowering for kernel selection. And so we have a kernel library that would be equivalent to something like maybe KubeLaz or KudNN from the GPU side of things. Are you able to talk at all about who's using this and what are some of the things they're doing and what it enables them to do that is differentiated? Yeah, sure. So we have a couple partners that we've announced and have done a lot of work that's publicly visible.

33:38So Group 42 in the United Arab Emirates, they're an umbrella organization that has this internal group called Core 42 now. So Core 42's aim is to be the artificial intelligence powerhouse of the Middle East. And so their effort is to build out infrastructure that would allow the Arabic speaking world to have the same kinds of applications that you would see we're using a lot of large language models for. So they're focused on sort of the basic language modeling applications, the sort of foundational capabilities. That's directly at Core 42. And then they have a few different partner projects internally inside of Group 42 that are focused on things like finance and medical and their sort of civil support services in Middle East countries.

34:48um so they're they're essentially covering all of the the kinds of workloads that you might be hearing about um in with with using large language models um we're we're also uh finding uh strong collaborations with medical domain and and pharma organizations uh in particular they're They're using large language models for things like drug discovery and genomics. So we just announced a partnership with the Mayo Clinic, and we're looking into some of those directions for them. They're interested in electronic health records and trying to improve patient outcomes with those. do some analysis, see if we can sort of get predictive capabilities about what would be assistive to a patient or diagnosis.

35:51And then we're also working with a few pharma companies like GlaxoSmithKline. And there it's like drug discovery using these for analysis like protein folding and genomics. Is there a way of using the wafer that is kind of leveraging its unique characteristics to do totally different things, or has that been kind of obviated by the focus on LLMs? That's a good question. In the space of LLMs, I think the capability of doing large scale has been very helpful. In particular, we have a number of organizations that are doing things like fine-tuning and alignment of models that they've picked up from the open source.

Read the full transcript

36:44So things that are pre-trained models. So a couple important ones would be LAMA 2 and LAMA 3, For instance, when you want to do fine tuning of a 70 billion parameter model, it's pretty difficult to do that on even the latest GPU hardware just because of the like memory and scale requirements. So our hardware makes that fairly simple. Like it's that's like a everyday scale. scale.

37:19So that's maybe the kind of thing that we target in large language models, I guess, in addition to pre-training. We've pre-trained a number of large language models from scratch, and we sort of focus on doing that in a compute efficient way. It's easy to distribute across many devices. So it's easy to train those at scale. You don't have to depend on sort of open source code bases and trying to sort of understand the intricacies of the parallelism strategies and, you know, sweeping a bunch of different parameters to find good utilization. um the our hardware is also um being used in a few high performance computing settings where there may be smaller markets but the potential advantage there is substantially larger so it'd be applications like computational fluid dynamics um in those applications the in order to do so the computation is structured in in a very local way but you're trying to do like very large grids or over meshes that are maybe irregular.

38:37Trying to lay that out on CPUs or GPUs means that you're constantly hitting the memory. You're doing access to memory all the time because we can store the full set of activation, this full set of data on the wafer in the memory. And it's ultra low latency to access it. we don't have this uh this issue of going back and forth to a memory uh and so the the performance can be 10 or you know 10 to 100 times faster than if you're running on a corresponding uh on a set of cpus or gpus with a corresponding amount of um like flop throughput Okay. In talking about the use cases, you mentioned folks that were on the LLM side, you mentioned fine tuning and pre -training.

39:38But I thought earlier you mentioned inference as a focus area. Did I get that right? Yeah, so I think there are two different angles on that. We're certainly thinking about our original implementation of our software stack, the pipeline execution mode would work really well for inference. That's been sort of on hold while we've brought up the training side for these large language models. That's something I think we'll revisit in the future here. the hardware though as an appliance probably makes more sense for training and so because of that we want to make sure that anybody that would be training on our hardware is able to use their devices and deploy them on the inference side so we've partnered with a couple different organizations that are kind of looking into ways to make that very efficient and picking particular hardware architectures that would be good for deploying the models that we've trained.

40:43So I mentioned a lot of architectures beyond the ones that you're providing. Yes. Right. So this would be partnered hardware vendors. So one of those vendors specifically is uh is qualcomm with their um with their uh aspire i don't recall the name their uh processor uh specifically targeted at ai um they had like a cloud ai 100 is there like enterprise box there you go i think i think ai 100 is the is the code name yeah um so the the those uh pieces of hardware are targeted at doing inference for things like large-scale language models. So it's sort of a natural platform. But then those architectures also have capability to leverage techniques like quantization and sparsity.

41:48And so a while back, I mentioned that our hardware architecture allows us to do scalar times vector multiplications. A unique opportunity with that kind of hardware is that we can do sparse computation that's essentially unstructured. And we can do that at very high performance. So we study things like weight sparse training. We put out a few papers from my research group on this. essentially the we're streaming the weights into the wafer we can sparsify them uh sparsify those weights and so then we only have to send the weights that are non-zero and we ignore all the computation that would be multiplications with zero so we get improvements in the bandwidth because of the the uh movement of fewer weights and then we get a reduced amount of compute that we do.

42:46So on the whole, that gets us significant acceleration when we're running sparse. So we've done a fair bit of work in sparse pre-training. And I think there's still a lot to be done there that in that field, the research community is maybe just getting its hands around a holistic understanding of how that works. We want to continue. Tom's also done some research around like thresholding weight values to further sparsify them. I forget the specific name of the paper, but we've talked about it on the podcast here. Yeah, they've done a lot of great work that focuses both on sparsification and pruning and on quantization.

43:30and we can do the combination of those two things for inference so that when you deploy this thing it's very low latency and still very high accuracy so we've been partnering with them on the kind of the deployment the implementation side of how you deploy these models and we've also partnered with another organization called neural magic they're a startup company that was founded by some some faculty members that have been focused on sparsity in the past, like Dan Alistar from IST Austria. And their organization is focused on the techniques that get you from a pre-trained model to something that's very high quality, but very sparse.

44:17And then they also have their own software stack they can deploy on CPUs. So if you already have a bunch of CPUs in your data center, you're considering Cerebrus hardware for the training, you'd be able to deploy those directly on CPU systems for inference. um those techniques uh that that that group has studied have focused on um essentially looking at things like the hessian or second order uh optimization statistics to decide is this weight something you should prune or is it a very sensitive weight that even if it has a small value and you prune it it's going to have a huge effect on the loss They have sort of ways to decide carefully which weights should be pruned, and then they can prune to a larger degree than you would during pre-training, for instance.

45:13The pre-training problem is a little bit harder, where the weights that you have are not just transmitting information through the model, but they also might need to be around so that your model can sort of carefully navigate the lost surface. in your optimization trajectory. So there are opportunities to improve sparsification and pruning for inference time in deployment. Can you also deploy it on traditional NVIDIA-style GPUs? Like, is the idea that you train the model, you take advantage of the scale of the wafer, but at the end of the day, you just end up with a weights file and you can put that weights file wherever you want?

45:57or do you have to then further like manipulate it to get it to run on specific hardware yes uh and that's that was uh kind of goes back to what we were talking about with our our software offering that the at the end of the day we have checkpoint converters that pull these models out to sort of common formats that people use typically we target hugging face because hugging face has a really strong software architecture ecosystem that just allows, if you have the right organization of the weights, you can just run it there. Like safe dentures format or something like that? Yeah, exactly. And that allows you to deploy pretty directly to CPUs or GPUs if that's where you'd prefer to run.

46:43So we use that for things like evaluations internally even and with the research organizations that we're working with externally. You mentioned from a research perspective, one of the efforts that you've worked on kind of taking advantage of the hardware, in particular, the weight sparse training. What are some other ways that your group is trying to take advantage of the unique opportunities afforded by the architecture? Sure. Yeah. So in addition to weight sparsity, I did also mention the idea of looking at optimizers. Certainly, if there are ways to do what we'd maybe consider to be non-trivial optimizers.

47:36So techniques like KFAC or shampoo optimizers, which are using second order statistics, We can do some sort of non-trivial compute associated with that and potentially speed up training by improving each training step. We can do fewer training steps. Can you briefly summarize KFAC and Shampoo? They're factorized implementations that can approximate some second order statistics about optimization. and uh so you get you get approximations to the hessian uh through um like low rank projections and uh because they're low rank you can actually store a big chunk of the hessian in an approximation that maybe be the the high level view um those are uh those are quite interesting to us i think Because the, you know, we still have a limited amount of memory, even though it's a much larger memory, but you'd be able to apply those techniques to extremely large models on our system.

48:49And we're looking at those because we think that there are interesting opportunities in the context of like weight sparse training. So like I said, in weight sparse training, there are some weights that are maybe not just responsible for transmitting information through the model, but they're also responsible for transmitting gradient back through the training step. So if you could do second order approximations to the gradient, the Hessian, you'd be able to say whether a weight is very important for that purpose. And so I think they're sort of nice ways to coordinate those things. Another direction we've been thinking about is in the space of activation sparsity.

49:39So in addition to weight sparsity, where you'd maybe get the benefit of reducing the number of weights you stream to the wafer, we can still take advantage of activation sparsity directly on the wafer. and that's in context where you'd have high computational intensity anyway so you could trim out some of the computation so like an obvious one is relu sparsity and there's been some really interesting recent work from groups like apple that essentially show if you layer in relu's throughout throughout large language models they still behave pretty well and you can cut down the amount of computation that's needed per sample or per token by a substantial amount.

50:28So something like you could get maybe 90 % sparsity on a per token basis. And then even like the union over many tokens, like maybe during a batch, you'd still be able to get a significant amount of sparsity, maybe in the range of 30 to 60%. In those settings, if you already have the rel use in the computation, you're actually just wasting the computation anyway, because some of those values are zero. What if you could just forget those and accelerate the computation by ignoring them? That's sort of what we're thinking about. And then like kind of a natural extension to that or a natural analog to that is the mixture of experts style models where you would do very strong gating, like you'd take a token and then make a decision, this is only going to go through this chunk of the model.

51:18We can do that sort of activation sparsity also. So we're working on those techniques and hopefully those will be available soon to our users. Do you see these efforts that your team is doing around research as being applicable only to cerebrus users or do they inform things more broadly in the industry that others can build on from a research perspective? Our effort is to certainly build things that support our hardware and can show unique advantages there. But I think a lot of these are techniques that will eventually be broadly applicable. So for example, one of the techniques that we, one of the things that we studied last year was the chinchilla scaling laws in a very open context.

52:07So in the chinchilla setting, they trained some language models where they didn't give full details on the models and they didn't release the data sets. We re-ran those experiments on our hardware, scaling up to a 13 billion parameter model and sort of showing that the scaling laws hold roughly the same as Chinchilla. 20 tokens per parameter is roughly compute efficient for pre-training. And we We released the, this is based on the pile data set. So it was an open source data set. And it was just a standard GPT-2 model architecture. So it's very reproducible. We released the checkpoints after training.

52:52That's something we want to continue doing as an organization is contributing to the open source so that people have that sort of infrastructure to build on top of. And then we'd also like to augment the results that have come out from prior works. So there have been a number of scaling laws studies, for instance, in weight sparsity. We think we have some unique observations that might push the envelope a bit there and maybe shift our way of thinking about scaling laws in the context of weight sparsity. And so that's something that we're currently working on, maybe potentially working with some partners there, too.

53:32And I think the same is true in something like the mixture of experts context, where a lot of the organizations that are building mixture of experts models are keeping that very proprietary. They might release the model that they trained, but they didn't tell you how they got there along the way. Like, what does the scaling law look like? Does it give you computational benefits? So those are things that we want to demonstrate and release also. Well, Joel, thanks so much for taking a bit of time to share what you're doing at Cerebrus and some of the research that your group is doing to kind of build on the hardware.

54:11It's been great for me to kind of catch up. It's been a while since I've looked at it, but it was super informative. Definitely. Thanks for having me. I was glad to be able to share a bit more about Cerebrus and thanks for the tough questions.

From the publisher

Today we're joined by Joel Hestness, principal research scientist and lead of the core machine learning team at Cerebras. We discuss Cerebras’ custom silicon for machine learning, Wafer Scale Engine 3, and how the latest version of the company’s single-chip platform for ML has evolved to support large language models. Joel shares how WSE3 differs from other AI hardware solutions, such as GPUs, TPUs, and AWS’ Inferentia, and talks through the homogenous design of the WSE chip and its memory architecture. We discuss software support for the platform, including support by open source ML frameworks like Pytorch, and support for different types of transformer-based models. Finally, Joel shares some of the research his team is pursuing to take advantage of the hardware's unique characteristics, including weight-sparse training, optimizers that leverage higher-order statistics, and more.

The complete show notes for this episode can be found at twimlai.com/go/684.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
Powering AI with the World's Largest Computer Chip with Joel Hestness - #684The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 55 min
Listen in VO