In short
Eye On A.I. Podcast Episode Summary
Episode Title
#165 Andy Hock: Will AI Revolutionize How We Do Business?
Host
Craig S. Smith
Guest
Andy Hock, VP of Product at Cerebras Systems
---
Episode Overview In this episode, Craig S. Smith discusses the transformative power of AI in business with Andy Hock, focusing on how advanced technologies are reshaping corporate strategies, operations, and their societal implications. Hock provides insights on integrating AI into traditional business models, ethical challenges, and the future of AI in various industries.
Key Themes
- Cerebras Systems and Its Technology
- Introduction of the Wafer Scale Engine (WSE) and its significance in AI compute.
- Overview of how Cerebras Systems is positioned to solve bottlenecks created by GPU shortages.
- AI in Business
- The changing landscape of corporate strategies due to AI advancements.
- Adaptability and continuous learning as essential traits for businesses in the AI landscape.
- Ethical Considerations
- Discussion on privacy, responsible deployment, and the ethical implications of AI technology.
- Current Trends in AI Research
- Latest developments in natural language processing and machine learning applications.
- Advice for Businesses and Individuals
- Importance of being adaptable and continually updating skills in this fast-evolving field.
---
Episode Breakdown 00:00 - Introduction
- Introduction of the podcast and guest, Andy Hock.
01:09 - Decoding Cerebras Systems' Technology
- Details on how the WSE is distinct from traditional GPUs.
03:25 - The Power of the Wafer Scale Engine
- Explanation of the WSE's capabilities, featuring 850,000 cores optimized for AI tasks.
05:24 - Training and Inference Capabilities
- Insights on training large AI models and how Cerebras addresses inefficiencies in traditional setups.
10:44 - Accessibility for Students and Researchers
- Discussion on making AI technology available for educational purposes, emphasizing collaboration with researchers.
12:08 - GPU Equivalents and WSE's Efficiency
- Comparison of WSE performance relative to conventional GPUs.
13:53 - Utilization and Power of the Chip
- Discussion on high utilization rates of the WSE and its impact on AI workloads.
19:02 - Understanding Wafer and Circuitry
- Technical overview of the wafer design and its production process.
23:27 - Computing Power of the Wafer Scale Engine
- Performance metrics and implications for large-scale AI applications.
28:59 - Challenges in Cooling the Wafer Scale Engine
- Overview of the cooling systems employed due to high energy requirements.
32:31 - Utilization of Silicon Waste
- Discussion on sustainability and recycling silicon waste from the production process.
36:21 - Cost Comparison with Other Technologies
- Overview of pricing and value proposition compared to GPU clusters.
40:01 - Customer Usage Patterns – Cloud vs On-Premise
- Insights into how businesses are utilizing Cerebras systems, either on-premises or via cloud services.
45:22 - Training Models for Edge Devices
- Exploration of capabilities to support smaller models for edge computing.
47:30 - Collaborations with OpenAI
- Overview of partnerships and collaborative projects with major AI stakeholders.
49:26 - Addressing GPU Shortages and Enterprise Needs
- How Cerebras is positioned as a solution to industry-wide GPU shortages.
51:01 - Inference Capabilities and Future Developments
- Discussion regarding future directions for inference tasks on the WSE.
57:46 - AI as a National Compute Resource
- Implications of AI technology as a critical resource for national development.
01:03:41 - Closing Remarks
- Final thoughts and encouragement for businesses and individuals to embrace AI.
---
Key Takeaways
- Cerebras Systems is revolutionizing AI compute with its Wafer Scale Engine, allowing for unprecedented training speeds and capabilities compared to traditional GPU clusters.
- There is a critical need for adaptability among businesses as AI technology continues to evolve and integrate into various sectors.
- Ethical considerations must be prioritized as AI technologies grow, particularly concerning privacy and responsible use.
- The industry's shift towards AI-driven solutions presents opportunities for innovation, efficiency, and competitive advantage, making it essential for both businesses and individuals to stay informed and agile.
Conclusion The episode reinforces the notion that AI technology will dramatically change business practices and societal structures. With pioneers like Cerebras Systems leading the way, understanding and adapting to these changes is vital for future success in the evolving landscape of artificial intelligence.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00What we're looking at here is a square piece of silicon that was cut out of a circle of silicon that was 300 millimeters in diameter. So you see data scientists and ML researchers using libraries like DeepSpeed and Megatron, specialized versions of ML frameworks like distributed PyTorch or distributed TensorFlow, tools like OpenMPI, Horovod, to effectively make the problem of spreading a large model out over many small workers make it incrementally easier. They spend tons of their time doing sort of supercomputer engineering, Not thinking as much as we might like about the AI, the application.
0:40And then once you actually get that model up and running, because you're training it then on a cluster of many small workers, there's inefficiencies. If you go from one to, say, 512 GPUs, you don't get 512 times faster because you're bottlenecked by memory bandwidth and communication bandwidth between chips. That's why we ended up with such a unique looking processor, because at the end of the day, AI work is a different kind of work than the computational workloads that preceded it. Hi, I'm Craig Smith, and this is Eye on AI. I'm at NeurIPS 2023 in New Orleans this week, and I'm thrilled to have Andy Hock, head of product for Cerebra Systems, here to discuss the company's revolutionary wafer scale chip technology.
1:26Andy talked about how this new chip architecture is being used to train foundation models and how it may help break the bottleneck in generative AI caused by the ongoing GPU shortage. I hope you enjoy the conversation as much as I did. I want to give a shout out to our sponsor this week, Babbel, the science-backed language learning app. Babbel is something I feel strongly about because language is the key to opening the world and broadening horizons. I know because as a journalist, I've reported out of more than 40 countries around the world, and it's amazing how knowing just a few words of the local language will open doors and build bonds.
2:11Be a better you in 2024 with Babbel. Don't pay hundreds of dollars for private tutors or waste hours on apps that don't really help you speak the language. Babbel's quick 10-minute lessons are designed by over 150 language experts to help you start speaking a new language in as little as three weeks. Babbel's designed by real people for real conversations. Babbel's tips and tools are approachable, accessible, rooted in real-life situations, and delivered with conversation-based teaching so you're ready to practice what you've learned in the real world. It's so easy to learn how to order food, ask for directions, speak to merchants without having to consult language apps while on vacation.
2:57The key to learning languages is not to be embarrassed for speaking poorly. Plus, Babbel's speech recognition technology helps you to improve your pronunciation and accent, something that I need a little work on. Studies from Yale, Mish State University, and others continue to prove Babbel is better. Right now, get 55 % off your Babbel subscription, but only for our listeners, at babbel.com slash ionai. This is an actual production wafer scale engine from Cerebrus. This processor is the heart of the system that we build and deliver to customers and make available remotely and through cloud. It is a beast.
3:44850 ,000 cores in this generation. and 40 gigabytes of on-chip memory, and all those cores are directly connected over silicon. So you can think of this in a sense as a cluster worth of AI compute all on one device, and you can also serve breakfast on it. So that would be the equivalent. Yeah, I had a couple of questions about that. People are using it now for training large models. Is that right? Absolutely, yeah. And what is the difference between that and an NVIDIA GPU? Sure. So this is a fundamentally different device, right, from the architectural level of the cores all the way up to the wafer.
4:34So this isn't just a bunch of GPUs together. This is a processor that is built from the ground up for AI, not for graphics, not for database management, but really built to accelerate large-scale AI. And so there's a couple of fundamental differences. First and foremost is at the processor level, all those cores, those 850 ,000 cores, they're designed to handle sparse, tensor-based, linear algebra operations from the get-go. So there's not different functional units on that chip that are, say, designed for ray tracing or, say, 64-bit computations for physics-based simulations, right? The cores on this thing are built to accelerate the sparse linear algebra ops that are fundamental to all of AI compute.
5:25Both inference and training. Both inference and training. And what we've focused on in terms of our software and our go-to-market has been on accelerating training in the data center. But this thing is built from the ground up for that work. And what that really means is that, look, what we observed is that AI compute isn't just, in particular, data center training, it isn't just a problem of bringing a lot of compute to bear. It's also a problem of bringing that compute to bear with high memory bandwidth and high communication bandwidth. If we look today at problems like training GPT models or chat GPT size models, right?
6:03These are models that don't fit on a single GPU anymore, and they would be slow if you did, but they don't fit on a single GPU anymore, So it requires you to train these large models, if you're using GPU, on large clusters of, say, hundreds to thousands of tiny general purpose chips. And look, that's gotten the industry a long way. It is an extraordinary machine, and it's a suitable engine for large-scale AI. But when you're training a large model on a cluster of many small chips, you run into a couple problems. The first is that, first of all, it's just hard to distribute that model, to program the model to run on so many small workers.
6:49So you see data scientists and ML researchers using libraries like DeepSpeed and Megatron and specialized versions of ML frameworks like distributed PyTorch or distributed TensorFlow, tools like OpenMPI and Horovod to effectively make the problem of spreading a large model out over many small workers make it incrementally easier. But that still requires thousands or tens of thousands of lines of code. It takes days or weeks or even months sometimes of software engineering just to get the model set up to run. And then, well, good luck if you change something. If you change the model architecture, you change something about the data, or you change the size of your cluster, often have to go back to the beginning and redo the distribution problem.
7:44So what ends up happening in that case is that your data scientist and your ML researcher is often the most valuable people in an organization that's pursuing AI. They spend tons of their time doing what you think of as sort of supercomputer engineering, right? Parallel programming. Not thinking as much as we might like about the AI, the application. And then once you actually get that model up and running, because you're training it then on a cluster of many small workers, there's inefficiencies. If you go from one to, say, 512 GPUs, you don't get 512 times faster because you're bottlenecked by memory bandwidth and communication bandwidth between ships.
8:30Okay, so then let's come back to our machine, right? The wafer and clusters of our machines. Our machine, because of its physical scale and because of our cluster architecture, we can run even the largest models on the planet today. That is 10 billion, 100 billion, even trillion parameter larger models on a single machine. and then if you want to scale out that is if you want to add more machines to go faster because we can run the whole model on just one machine then we can run the whole model on every machine in the cluster and just ask every machine in the cluster to work on a different part of the data set and that's called data parallel scaling right and it's it's fundamentally different than the distribution techniques that are required for a cluster of gpus where you can't do just data parallel, you also have to do model parallel, tensor parallel.
9:27And what this being able to run the largest models on one machine and scale with simple data parallelism only on Cerebrus, what that really means to an end user is I can program a cluster of Cerebrus machines that give me the horsepower of say thousands of GPUs with the same code that I would use to train my model on a single desktop machine. And that takes that days or weeks of months of software engineering and all the required expertise of parallel programming and supercomputer architecture. And it just makes it all go away and makes it far simpler for users to get up and running. And then when we run on, say, 16 or 32 or 64 of our machines, like our latest clusters, you actually get 16 and 32 and 64 times faster than one.
10:13And at the end of the day, especially for someone like me coming from a research background, what that means for our end users that are researchers themselves or researchers in an enterprise organization building new applications, it just means that they can go that much faster. right they can ask and answer questions more quickly they can set up and train these billion 10 billion 100 billion parameter models in days or weeks rather than weeks or months and bring new applications to life that much sooner yeah it's interesting i was talking to somebody from deep mind yesterday or i was at a a talk at which somebody from deep mind was speaking along with somebody from meta and the question was how difficult is would it be for you know a master's level student computer science student to train an llm on uh on one gpu could they do it with all the tools available and they were saying yeah uh up to about um up to about 70 billion parameters is what they were claiming uh but once you get beyond that once you need more than one gpu it just becomes this massive problem exactly what you're talking about and that how much time they spend and how much expertise is required to coordinate multiple multiple GPUs so two questions sure how many GPUs the GPU equivalents does one wafer sure represent Sure.
12:09So the, you know, you see a lot in marketing materials about flops, about floating point operations per second. But what really matters to an end user in this business is time to solution. That is, how quickly can I train a model to state of the air accuracy? So sometimes it's not even just throughput that is training samples per second. Sorry about that. It's training throughput that yields state-of-the-art accuracy. So no tricks, just tell me how fast I can train a model to state-of-the-art accuracy. And in some recent work, actually one of our customers published an independent study of novel AI accelerators for large AI model training.
13:01And this is some work from Argonne National Labs, Department of Energy, U.S. Research Laboratory. And they're looking at big AA models for science, right? Some really, really interesting stuff under the hood. But the result that they found when comparing our CS2 machine to A100 was that for training, particularly large GPT models, they found the CS2 was 152 times faster than an A100, right? And that's a meaningful difference when you couple that together with a simpler programming model. It just lets you iterate that much faster. So long story short, each one of our machines typically delivers the compute equivalent of several tens to, in some cases, hundreds or more GPUs for the same AI training task.
13:53Yeah, and you mentioned efficiency or utilization. That was one of the things that the DeepMind person was saying, that one thing is getting them all to work together. The other is you're running it, I can't remember the technical term, but basically at 50 % capacity. and to get above that then you need to do a lot of other uh stuff so do you have that problem on cerebris to that that you're actually using the power of the chip we're actually using the power of the chip so our our software takes care of that challenge also under the hood yeah um by virtue of architecture and the way we execute a model we can achieve very high utilization is is the term of our very high utilization for this chip and i would actually say that while it's true that some of the some of the the leading ai research labs in the world that is the you know the cream of the crop while it's it's it's true that some of those researchers might be able to get something like 40 to 50 percent utilization out of a gpu it's not that common and it's very difficult in order to achieve that kind of utilization on a on a general purpose processor like a like a gpu for an ai workload you end up tweaking sort of lower level microcode kernels and tweaking communication patterns between devices i think that's probably what they were alluding to that It is possible to squeeze that much performance juice from a GPU, but it's very difficult and it requires highly specialized expertise that really fundamentally only lives in a few organizations on the planet.
15:55And I know this is a related point, a little bit of a jog, but I think one of the things that we are really trying to do at Cerebris is not just build the most powerful AI processors and systems on the planet to accelerate AI research, but also put that power into more users' hands. That is, democratize access to the highest performance AI compute. And I think part of that is actually making the programming easier and giving users who might not know how to architect their own microcode kernels for compute operations or communication patterns given the tools that they need to harness that power from a processor on their own with the tools that they use today so the short answer is although our chip looks you know different than a than a regular processor you program our machines with standard ml frameworks like pytorch and as we said before you don't need any specialized libraries um you don't need to go underneath the hood and optimize kernels all of that power gets put in your fingertips through our compiler and our software stack and you get to use the the tools that you're comfortable with yeah actually can you hold that up to the camera again and sure just uh explain what we're looking at.
17:30I mean, it's a piece of silicone. Right, right, right. And what are the individual rectangles? Sure. Yeah. Yeah, so as I mentioned, Cerebrus was born out of an interest to really fundamentally transform the AI compute landscape. And as a startup, up we got our start back in 2016 we had the opportunity to approach the processor design problem from first principles that is we could ask what is the right processor shape size architecture for ai full stop whereas i i think if you're if you're already in the industry and already making computer chips then you can make small changes to the architecture to move it towards the AI workload, but it's much harder to take a clean slate approach.
18:30So we had this opportunity in front of us to take a clean slate approach. And that's why we ended up with such a unique looking processor, because at the end of the day, AI work is a different kind of work than the computational workloads that preceded it. So what we're looking at here is a square piece of silicon that was cut out of a circle of silicon that was 300 millimeters in diameter that's the largest and that that's the 330 millimeter diameter that's what do you call that it's a loaf it's a it's a wafer yeah that's a way cut into a wafer the oh oh oh the yeah you might be right yeah you might be right yeah it's it's a long yeah uh it's like a long cylinder so that's right That's right.
19:21And it gets sliced into very thin wafers. So you probably can't resolve this in the camera, right? But what we're actually looking at here is basically as thin as a pane of glass, right? If I tried to hold this in my fingers, you might cut yourself on the edges. So yeah, we start with a blank circular wafer. and then our fabrication partner uses optical lithography basically to to etch the circuitry onto the silicon wafer and the way they do that optical lithography is is in these individual steps so each one of these tiny rectangles that you might be able to see in the camera each one of these tiny red tiny rectangles and there's 84 of them on on my wafer each one of those tiny rectangles is called a reticle.
20:10And so that's the physical unit of optical lithography exposure. It's like a camera shot, right? And so we take one and then step it to the right, take another, take another, take another. On our wafer, all of these are identical. And for other processors, they might be identical too. The big difference is that once we're done writing all the circuitry on this wafer, we don't cut it any further. If I was making other chips, like a GPU or a CPU, I'd start with something that looks like this, but then I would dice it up into many tiny chips. Sort of an interesting way to think about it, right? Because then what they end up doing is they end up mounting each one of those little chips on a motherboard with memory and interconnect.
20:55And then we spend a lot of time as a community figuring out how to cluster them and get them to work together again. And we just said, well, let's not do that. Let's not cut it apart. Let's keep it all together so that we can have all this compute close to memory and able to talk to one another without having to put it into individual chips and then string them all back together with a bunch of cabling and other machines. So you're looking at 84 reticles. And this wafer, our wafers are fabricated by TSMC. there's only a few fabs in the world that can produce chips and circuitry at this resolution in this generation we're fabricating at the seven nanometer process known and that seven nanometer just uh that that's that's kind of a reference term it doesn't mean that the circuits are literally seven nanometers wide that's right that's right yeah at at one point in the in in the history of optical orthography, they were actually very close to that.
21:58They still are, but it's not exactly seven nanometers right there. Particularly as we get the future higher resolution, it's finer scale nodes like five nanometer, three nanometer, two nanometer, they're not exactly that either, but they're very close. That fine pitch just means I can put more transistors per unit area. right and as a as a chip designer our primary resource just like an artist working with a canvas right our primary resource is how much silicon do we have right and what do i want to do with that area so if i can put more transistors per unit area that means more cores or more memory and therefore more capability yeah so coming back to this guy we have 84 identical radicals 850 ,000 identical cores all interconnected all with their own memory one clock cycle away is a think of it as the fundamental unit of compute processing or a single processing element right so it includes a data path with transistors that when data comes in it executes the the computation right and it also includes uh interconnect to adjacent cores and interconnect to memory so it's sort of the the fundamental cellular unit of computation yeah on a machine it's yeah okay because uh when they talk about dual core right you're talking about one chip but on that there are two processing units many tiny computers exactly exactly yeah and right If you're like me, you remember maybe when the first dual-core CPUs came out or four-core CPUs.
23:49And even state-of-the-art CPUs today maybe have 24 or 48 cores. We're talking about 850 ,000 on this device. And we were talking yesterday. So that's a very thin slice of this loaf, this cylinder of silicon that's etched. it's got a copper color is there uh is it then covered with something yeah so there's there are many layers right so the actual fabrication process will be laying down many layers one on top of another of uh of circuitry and then a final upper layer that will connect all the cores and maybe connect the chip itself to other I.O. systems. So what you're looking at here on our device is we have a uniform 2D array of those 850 ,000 identical cores.
24:52So it's a very simple design. And then on the edges, we have I.O. circuitry so that when we hook this into a machine, we can get data from the outside world. Say in AI, we can get the weights of the model and the input data batches actually onto the wafer. And then the wafer, those 850 ,000 cores on the wafer all execute the multiply and accumulate operations that constitute the AI training problem. Just sort of a fun fact. I know we're talking a lot about the wafer today, But once we landed on this architecture because we thought it was the right chip architecture for AI and solved this decades-long challenge of integrating a single device at wafer scale, that was really just the beginning of the engineering challenges that we faced.
25:52because then we had to figure out how to package this thing and put it into a system that could go into a data center and connect with standard power, standard interconnect to other machines. So that meant we had to figure out how to power this thing, how to keep this thing cool. This is consuming about 17 to 18 kilowatts worth of power at about 20 ,000 amps when it's running in a machine. So how do I keep all that cool and uniform? So building the wafer itself was, in some sense, just the beginning. Then we had to figure out how to package it, how to power it, how to cool it, how to deliver data so that it could live and breathe in today's standard data center environments.
26:46And so we couldn't be just a chip company. we had to be a full systems company and then write the software on top of it to make that whole thing programmable by data scientist or ml researcher who's probably miles away from the data center um so i think one of the things that i that i enjoy about cerebrus as a company and a team is that there's this common thread amongst our engineers both on the hardware and the software and the ml research side that they uh in some sense they are unafraid of of these big challenges, right? Whether it's architecting a wafer scale chip or bending metal and, you know, cutting gaskets to cool this beast in a data center machine or building the software on top of it or training new world leading models, right?
27:35We get excited by those kinds of challenges that I think might throw other people for a loop. And this is the second version, right? Yeah, this one right here is a second generation wafer scale engine. We introduced this part and the associated system called our Cerebra CS2 in 2022. And we introduced the first generation machine, which we call the CS1. You can sense a theme here, CS1, CS2. We introduced the first machine in 2020. and we've got a roadmap into the future of multiple wafers and multiple machines that'll be coming in in future years right now it's one way for per system per machine right now it's one way for per machine and the the work that we've done recently has allowed us to cluster these things together very very efficiently like i said we've we're building clusters right now of 16, 32, 64 systems and larger.
28:41And we're seeing really good scaling across those when we train big AI models. We're getting 64 times faster than one in a cluster of 64. So right now, that approach is serving us really well. That is one wafer in one machine. Yeah. Yeah. And the cooling thing, after I spoke to Andrew, I got a lot of emails and comments saying, yeah, that's great, but how do you cool it? And so that, I would imagine, was one of the biggest challenges. How do you cool it? It was one of the biggest challenges, yeah. So unfortunately, I don't have a picture today, but you can imagine and your listeners can imagine, this thing's about the size of a dinner plate.
29:32The system that it goes into, that is the physical box that it goes into, is about the size of a dorm room refrigerator or a kitchen mini fridge. And it weighs in at about 600 pounds. and about the top third and quarter of the machine is power io and the wafer package which we call engine block itself is about the top third quarter of that tiny refrigerator size device everything else in the box is cooling so the way we cool it is we actually have a closed internal water loop that is taking cold water pumping it towards the back of the chassis then we have a series of custom built manifolds that take that cold water flow and distribute it uniformly across the back of the wafer where there's a copper cold plate so you have a thermal contact to the back of the wafer that's got basically a series of micro etched grooves that are a manifold for the cold water to spread out across the back of the wafer so that we're not only keeping the wafer cool but we're keeping the temperature of the wafer across the face uniform then we have warm water and we have to we move that warm water by pumps down to the bottom half of the machine which is basically a big heat exchanger andrew often makes a good joke about this heat exchanger that we're effectively taking the 1970s radiator technology and building it into a 21st century machine.
31:20It sounds like a radiator. Yeah. Even 19, did you say 70s, even 20s? Earlier. Yeah. And so if your listeners are gamers or PC builders, if you imagine the machine in an X-ray view, it looks a little bit like a gaming PC on steroids, right? Big engine block and then pumps with an internal closed water loop and a big heat exchanger, which is effectively a radiator at the bottom. And then the heat exchanger itself can be cooled either by facility air in a data center or by facility water. so most data centers that we deploy into today whether they're our own laboratory data centers or customers or cloud partners most of those data centers have water cooling and they'll end up you know bringing effectively a feed of cold water to the machines that then cools the internal closed water loop of the cs2 so closed loop cooling with a heat exchanger that's then either cooled by facility water or facility air yeah that's fascinating and and um just a fun fact that we talked about last night this starts out as a disc you cut it into a square and the uh the ends are discarded and we were talking about all the ways you could use that uh that discarded silicone and you've got to send me one.
32:57I'll figure out a business for using. I love it. I love it. Cut me in. These things would make great jewelry or artwork. And actually they are. I don't know, Craig, if you've ever been to our office, but if you come to our office, I mean, first of all, this thing looks pretty cool and it's in its final state like this. But we also have in the office large prints of the microscopic view of each one of these reticles and all the different layers of circuitry. And those are quite beautiful also. I mean, it's a very regular geometric pattern. But when you sort of imagine what all these literally trillions of transistors map out to, it looks like a city grid in a way.
33:46And it's just, it's a really cool piece of visual art also. So yeah, we got a couple of side hustles that we can work on. Chip parts are art. One of the big topics globally right now as these models proliferate and they get bigger is the shortage of GPUs or the shortage of compute. You were saying when we spoke earlier that these wafers are on a separate production line or in a separate process at the fab. So you're not constrained in the way that other chip makers are constrained. Is that right? Can you talk a little bit about that? Yeah, so we still end up getting in line with our state-of-the-art fab partners like other chip builders do.
34:53But we have a really well-controlled and redundantly sourced supply chain after that. So we don't have the same supply chain complexities or risks or sort of single points of failure in vendors that other manufacturers might have. And so we today are still in a position where we can deliver systems in a reasonable timeframe, something like 90 days from order, whereas other system vendors might be quoting something like six months or nine months or even a year. And so, yeah, we're in a relatively good position from that standpoint. And as a proof point to that, over the past six months or so only, we've built a system with our strategic partner and customer, G42, that we are calling Condor Galaxy 1.
35:53and condor galaxy one is a cluster of 64 of our machines so four exaflops of sparse ai compute which is and sounds tremendous but really at the end of the day with this system like this we've been able to train 10 to 30 billion and larger models in a matter of days or weeks so it's a massive machine and it's built just in in the past six months only um yeah yeah how does the cost compare because uh it you know not being in the industry uh as a journalist you wonder why isn't everybody training on cerebrous chips sure so the short answer on cost is that we are going to be price performance competitive or better than a cluster of gpus and that's just sort of the dollar capex outlay right after that we're going to deliver more performance to your users with that simpler programming model so your users are going to be able to do more in less time and therefore your operational cost is also better from the ai development cycle and we also because we didn't break apart our chip and reconnect it with literally hundreds of yards of cabling, we also have inherent power advantages from a power efficiency standpoint.
37:27So we're going to be at time of purchase, we're going to be price performance equal or better. And then we're going to pile on advantages for you as a user afterwards. right so that gets at your question of of price and cost and then i think the question about adoption is is really um less a matter of of price performance and more a matter of the fact that we're just bringing a fundamentally new tool into the marketplace right and in my head it's it's a little bit like inventing the wheel for the first time and bringing it into market right There is a massive incumbent system that's out there, and we're an alternative to that.
38:11And I think we have to prove to our customers and to our users and to the market that you can use this machine to train state-of-the-art models. And we've introduced a library of GPT models this year called Cerebrus GPT that prove that. We've released multiple models with customers like JACE, 13 and 30 billion, which are Arabic, English, state of the art GPT models, BTLM, which is a world leading small language model that can run on devices. And we recently released a model called Crystal Coder in partnership with MBZ UAI and Patum that speaks not only English language, but also speaks code. So we've illustrated that proof point through our customers and our own work that you can use these machines to build state-of-the-art models and you can see these performance and operational cost advantages.
39:09and I think that's the beginning right that's the foot our foothold in the market to then expand from there so my hope is honestly when we talk you know maybe next year at the next Nureps you'll see even more customers even more models and we also hope to make these machines available to more developers in the open source community so that they can start building their own right and at that point, I think we start to see the dominoes fall and we start to see a bigger and bigger chunk of the AI data center compute market falling in our direction. But we, you know, we really have had to prove ourselves and show the world the value of a fundamentally new platform for this work.
40:01Yeah. Are most of the people, most of your customers using it through the cloud or buying it and installing it on-prem? Great question. This has changed over time. The short answer today is that most of our users and customers are using these systems remotely, either directly through us or through our cloud partner, Cirascale. In the beginning, it wasn't that way. Some of the earliest adopters of our machines were traditional supercomputing centers and research laboratories, like the AI research group at GlaxoSmithKline and Department of Energy labs like Argonne National Laboratory and Lawrence Livermore National Laboratory.
Read the full transcript
40:43These folks have data centers. They know how to build supercomputers. They were very comfortable with the idea of bringing a new refrigerator-sized box into their data center and working with us to figure out how to power and cool that with their infrastructure. So they're comfortable with that. Part of their charter is to adopt these new classes of machines, see what they're good at. So our earliest customers, we were deploying direct on-prem. And we still do that, right? If you're a customer that happens to have a couple of megawatts worth of power and wants to build a big cluster, we will do that.
41:18um alternatively if you're a customer that is say a new ai startup um or just a software oriented enterprise organization that doesn't own or operate your own data center and is used to consuming compute through the cloud we can roll with that too and so we've built that out over the past year basically just letting users be able to to log into to to our machines just like any other virtual machine, bring their data to the machines and get to work right away without ever having to set foot in the data center. Yeah. And who are the cloud partners? Or do you have your own cloud? Yeah. So right now we can make our systems available directly.
42:02So we could work directly together and get you remote access to machines in our data center. We also have a cloud partnership with a company called Cirascale. So they are offering pre-trained models as a service and training compute time as a service through their cloud. And we've also, of course, been talking to some other cloud partners as we expand our reach. I should say also, additionally, some of our supercomputing partners, like I mentioned Argon earlier, but we also have systems at Pittsburgh Supercomputing Center funded by the National Science Foundation and in Europe at EPCC which used to be Edinburgh Parallel Computing Center in Scotland and LRZ in Germany and those centers in particular ANL, PSC, EPCC and LRZ also offer time on their machines to local and regional researchers right so you can actually apply for grants of time if you're a researcher.
43:07And so while that's not a traditional commercial cloud kind of engagement, they are offering the machines up for researchers in that cloud kind of consumption model. Yeah. If a couple of questions, I can see the value for training foundation models on a Cerebris machine. if if uh are there use cases uh for where you don't need the entire wafer can you partition the wafer and run multiple workloads on it yeah great question um the short answer is yes so for for big foundation models and gpts the the way we train is we load a mini batch of data onto the wafer and then compute one big model layer at a time right so you can imagine these big rectangles basically large matrices of compute coming down and landing on the wafer where there's already data and the compute happens there so that's that's how we execute training for for large models today but if if the individual layer isn't that big we can run multiple layers and if the whole model fits on the device we can actually run the whole model on the device at one time and to your point if your workload is smaller still that is it doesn't require even a whole wafer then we have the capability to do that too all of that is in software effectively describing to the machine how to compile the compute job onto that array of cores and then how to move data to or through that compute job.
44:57The only reason I say that is because what we've really optimized the software for recently, particularly the past 12 to 18 months, is for these large language model training tasks, these sort of foundation model training tasks. and so we can run smaller models but it's it's it's not the most of our users are are focused on a lot bigger things sure and you also mentioned along that same idea that I don't remember exactly what you said but that people can train models that run on the edge? Did you say that? Oh, yeah. Yeah, so maybe there's two related points, right? One is that one of our design principles from the beginning, from a product standpoint, was that you should be able to take code that, say, runs on a GPU, and you should be able to run that on our machine.
45:59That is, if you train a model on a GPU, you should be able to train it on us. And that's true for PyTorch transformer models today, right? You should also be able to take a model that was trained on our machine and run it elsewhere for inference. So if you want to train a model on our machine and then run inference in the data center on CPU, or if you want to run inference on that model in an edge device like a phone or a MacBook, you can do that too. so yeah I mentioned one of our particular models is a model called that we developed with a customer it's called BTLM it stands for BitTensor Language Model it's a 3 billion parameter language model it's a large language model just not as big as the others and that 3 billion parameter model can because of its size once it's trained it can run say on a MacBook Pro or another capable edge device And so for, you know, for academic developers or people that are the sort of hobbyist developers or just researchers interested in highly efficient and capable small models, that's been a real hit in the market.
47:18It's one of the top, if not the leading model for its size posted out on Hugging Face in the open source community because of that. Yeah. Your partners, I don't know how freely you can talk about them. Are you working with OpenAI? So we're working with a large number of partners, both in research and enterprise. We have lots of friends over at OpenAI right now. So I would say the biggest public partners that we've announced are, once again, the U.S. Department of Energy labs like Argonne and Livermore, Pittsburgh Supercomputing, GlaxoSmithKline, OpenTensor, Together. we're also have recently announced just earlier this year a massive strategic partnership with a commercial company called g42 based in the uae they're the uh they're a private company that is sort of the national champion of artificial intelligence in the emirates and they're working with us to build large state-of-the-art language models for arabic as well as large foundation models for medical clinical assistance we worked with a part of g42 called m42 recently to to fine-tune a open source model that happened to be built by meta on gpus one of their llama models And we continuously pre-trained and fine-tuned that model with some medical knowledge and open-sourced with M42, the world's first open-source model that was able to pass the U.S.
49:07medical licensing exam. So our partner G42 is working with us not just to build big compute clusters, but also to build state-of-the-art models for things like Arabic language, medicine, coding, science, and climate. yeah yeah the the reason I mentioned open a is is I've been talking to a lot of people in industry and what I'm hearing is that and it backs all the way up to the fabs but because of the limited GPU supply. OpenAI, you know, limits the number of tokens per minute or requests per minute that you can push through the model on the API. And that is limiting or constraining enterprise scale production applications because people can't you know go into production on something if they're gonna run into this bottleneck issue right so it would this seems to me particularly if you don't have a supply problem that it would be the answer to that.
50:33So are you talking to any of the really big, you know, open AI or Google or Amazon about, I guess Amazon's got their own solution, but about using these as an alternative to smaller chips, just simply to train and run inference on very large models. Right. A couple of points here. So first and foremost, as a novel AI accelerator builder and a relatively new company in this ecosystem, yeah, we're talking to all of those organizations because I think at some level, we all as a community recognize the need for or the value of alternative compute architectures for this really economically and maybe even societally important workload.
51:32You see Google working on TPU. You see AWS working on Tranium. You hear about other hyperscale and research organizations thinking about rolling their own silicon. Because AI compute is a different workload than what we've seen before. And in some sense, that's just a point that substantiates It's our founding hypothesis, which is that building and running these models in production on existing infrastructure is okay. It got us off the ground, but it's not an optimal solution. And so we are talking to all of those folks about the value of cerebrous systems in their infrastructure. So that's one point.
52:15And I think there's, you know, as I look into my own crystal ball for the future, I think it's inevitable in some sense that future data centers that are doing AI work or that are built for AI work are going to have different kinds of accelerators. That is, these are going to be heterogeneous clusters where you can imagine a world where you have many users that are bringing to your infrastructure AI problems of different shapes and sizes. Big training jobs, small inference jobs, dense jobs, sparse jobs, even jobs that maybe border up against or also do HPC. Things that couple, say, physics-based simulation and AI.
53:00So we're going to see a broad landscape of pure AI or AI augmented workloads coming into the compute infrastructure of the future. And I think it's therefore very likely that on the backend, you end up with a heterogeneous compute infrastructure that's composed of AI and HPC accelerators that are each best athletes in some sense in their own domain. Right. And, and that in between you have a software layer that can interpret those inbound jobs and sort of decompose them into the requisite compute and communication and memory and storage attributes of the engine, and then compile down to, you know, maybe a combination of training jobs on a, on a big CS cluster and then inference maybe on some purpose built inference device.
53:51Yeah. So, we are having those conversations and I think that is sort of the direction that I see the industry going.
54:03Then I wanted to make one comment also on business model, right? Some organizations today that are doing say enterprise GPT training for customers end up keeping the model, right? Like you asked about inference I know and so I don't want to dodge that question. I'll come back to it, but I also wanted to make mention of something that I think is sort of unique about our business model. That is, when you look out at, say, existing businesses that do AI model training, often as a customer, you don't get to keep the model. You end up, the organization that built the model keeps the model, and then you have to run inference there.
54:41And I think that business model has clearly worked for those organizations. But when you work with Cerebrus, not only is your data safe, so when you bring your data to, say, our cloud or our machines, it's kept secure by all standards. But also, we're not going to build other models with that data. And when you have a fully trained model from our system, we give it back to you. So you get to own the weights. and therefore if you have your own say infrastructure for model serving or if you want to fine-tune that model a little bit further for another purpose you can do that you have full control and ownership and i think customers have come back to us and told us that they like this because it it lets them in some sense own their own destiny in their ai work right their data is safe when they train and then when the model is done they get the the weights and they can they can hold on to those and of course we hope that they come back to us to to train again or update those models but but they get to own the outcome and this leads back to your jury inference question because it means that they can run inference say in a data center with us um or they can they can run it say in their their own cloud or in their own on-premise compute still doesn't answer the question as to whether or not we do inference so i'll come to that now the the short answer is this This is back to the discussion that we were having earlier about large models and small models.
56:10This chip doesn't care if we're running training or inference. The chip can run inference and in fact has some big architectural advantages in terms of being able to compute at high throughput with low latency even at very small batch sizes. We can run inference on the machine. Today we haven't written all the software to do that. So you can run eval, which is one form of inference on our machine today. We also have some really exciting early work in inference for computer vision for some of our government customers who are interested in processing lots of imagery and video feeds. And last but not least, our team today is working on some architecture studies to be able to support large language model inference in the future.
57:03So I think it's unlikely in a sense that we will be a fully general purpose inference engine. But we're taking the same approach as we did to training to inference that is really taking a very close look at the variety of inference workloads, what each one of them is asking of the underlying machine, everything from computer vision inference on say, tiny images to inference on large language data sets. And from a product standpoint, then we're making decisions about where to focus first, where our architecture can give the biggest performance value back to users. So long story short, yes, and stay tuned.
57:46Yeah. And you mentioned Dubai and we spoke yesterday about, you know, this evolving or developing idea of a national compute resource for AI in the United States. and whether that's cloud credits or physical data centers. And I'm a believer that this generative AI and behind that other large model technologies are going to fundamentally change economies and even geopolitics. Are you talking to nation states about building national compute resources that then would power their economies, the AI economy, including the United States? Yeah, this is a great question. I think following our chat last night, it's an area that I'm really passionate about.
59:06The short answer to the question is yes. And I think like you, we at Cerebris see massive and positive potential for this technology to fundamentally change the way that we interact with data, make decisions on a daily basis. Everything from retail applications and things like recommendation engines that find good cat videos on YouTube to models that can help predict the next pandemic outbreak or help us develop smarter vaccines faster or help us figure out pathways to more efficient and cleaner, renewable energy generation and environment. So I think there's tremendous and yet fundamentally untapped potential for these large AI models, both for economic and good.
1:00:08And I think the way that we discover these positive applications as a society is by accelerating our research and making the tools for that research more broadly available. And it comes back to my remark about how do we make systems like these accessible to more users? And that's the broad base, right? But in the meantime, back to your question, we're also seeing around the world big enterprise companies and big enterprise nation-state governments sort of having that aha moment and realizing that these technologies could be transformative for economy and environment and provide tremendous return.
1:00:53turn on investment for their people and for all of us. And so, yeah, we are seeing a lot of interest in our systems from both large commercial enterprise and from governments to help them think about how to build that right compute infrastructure and even help building models on that infrastructure as part of their AI initiatives. and we are directly talking to the US. As you're probably familiar, I think we chatted a little bit about it last night. The National Science Foundation has announced a program called the National AI Research Resource or NAR. That was even called out in the administration's recent executive order as a priority for the nation.
1:01:41So we're engaged and talking to our friends at NSF and NAR about that. And also in the US, the Department of Energy is formulating its own position. These are the folks that have built generations of world-leading supercomputers, historically, primarily for physics-based modeling and simulation, but know how to build these sort of massive world-class computational infrastructure. Department of Energy is also looking at how to build that next generation of supercomputers to power the US AI research initiative, both for science and for security. And so we're deeply engaged with both of those communities.
1:02:29And at a personal level, I think it's something that can really, really benefit the nation. So much of this technology like ours and GPUs have been developed here. And so I think we're in an extraordinary position in the country and have an extraordinary opportunity, therefore, to lead the world in this with our partners. and co-develop the next generation of capable and efficient, safe and responsible AI methods. And I think the way that we do that is by starting at the infrastructure level and then putting that infrastructure into the hands of users that can ask and answer questions and build those solutions.
1:03:23I'm stoked. I'm thrilled for what the next several years and the next decades in this field is going to bring. And I couldn't be more proud to be a part of a team that's bringing potentially one of the platforms that powers this into the future. It's cool stuff. It's exciting. Yeah. I'm very interested in longevity research because I want to be around as soon as someone is. We got to build you a cluster, Craig. Let's get to it. I want to give a shout out to our sponsor this week, Babbel, the science-backed language learning app. Babbel is something I feel strongly about because language is the key to opening the world and broadening horizons.
1:04:09I know because as a journalist, I've reported out of more than 40 countries around the world, and it's amazing how knowing just a few words of the local language will open doors and build bonds. Be a better you in 2024 with Babbel. Don't pay hundreds of dollars for private tutors or waste hours on apps that don't really help you speak the language. Babbel's quick 10-minute lessons are designed by over 150 language experts to help you start speaking a new language in as little as three weeks. Babbel's designed by real people for real conversations. Babbel's tips and tools are approachable, accessible, rooted in real life situations, and delivered with conversation-based teaching so you're ready to practice what you've learned in the real world.
1:05:00It's so easy to learn how to order food, ask for directions, speak to merchants without having to consult language apps while on vacation. The key to learning languages is not to be embarrassed for speaking poorly. Plus, Babbel's speech recognition technology helps you to improve your pronunciation and accent, something that I need a little work on. Studies from Yale, Michigan State University, and others continue to prove Babbel is better. Right now, get 55 % off your Babbel subscription, but only for our listeners, at babbel.com slash IonAI. That's it for this episode. I want to thank Andy for his time.
1:05:43If you want to learn more about the conversation today, you can read a transcript on our website, IonAI. That's E-Y-E hyphen O-N dot A-I. In the meantime, remember, the singularity may not be near, but A-I is fast changing our world. So pay attention.
From the publisher
This episode is sponsored by Babbel. Babbel is conversation-based learning built with science-backed cognitive tools like spaced repetition and interactive lessons created by real language teachers and voiced by real native speakers.
Get 55% off your Babbel subscription 👉 https://babbel.com/eyeonai
Unveil the transformative power of AI in business and society.
Join us on #165 of the Eye on AI podcast as we sit down with Andy Hock, VP of Product at Cerebras Systems, an AI hardware startup out to accelerate deep learning and change computing forever.
Focusing on the intersection of AI and business, Andy discusses how advanced technologies are reshaping corporate strategies and operations. He offers insights into integrating AI into traditional business models, emphasizing the transformative role of technology in innovation and growth.
We also explore the societal implications of AI. Andy delves into the ethical considerations and challenges posed by AI, including privacy and responsible deployment. He sheds light on the latest trends in AI research, particularly in areas like natural language processing and machine learning, and their diverse applications across industries.
Andy concludes with advice for businesses and individuals navigating the AI landscape, underscoring the importance of adaptability and continuous learning in this rapidly evolving field. Make sure you watch till the end!
Stay Updated:
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI
(00:00) Preview and Introduction
(01:09) Decoding Cerebras Systems' Technology.
(03:25) The Power of the Wafer Scale Engine
(05:24) Training and Inference Capabilities
(10:44) Accessibility for Students and Researchers
(12:08) GPU Equivalents and Wafer Scale Engine's Efficiency
(13:53) Utilization and Power of the Chip
(19:02) Understanding Wafer and Circuitry
(23:27) Computing Power of the Wafer Scale Engine
(28:59) Challenges in Cooling the Wafer Scale Engine
(32:31) Utilization of Silicon Waste
(36:21) Cost Comparison with Other Technologies
(40:01) Customer Usage Patterns – Cloud vs On-Premise
(45:22) Training Models for Edge Devices
(47:30) Collaborations with OpenAI
(49:26) Addressing GPU Shortages and Enterprise Needs
(51:01) Inference Capabilities and Future Developments
(57:46) AI as a National Compute Resource
(01:03:41) Closing Remarks




