#222 Andrew Feldman: How Cerebras Systems Is Disrupting AI Inference

28 Nov 2024 · 42 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Eye On A.I. Podcast Episode #222: Andrew Feldman: How Cerebras Systems Is Disrupting AI Inference

Episode Overview In this episode, Craig S. Smith interviews Andrew D. Feldman, Co-Founder and CEO of Cerebras Systems, to explore how the company is transforming AI inference and high-performance computing. Feldman discusses Cerebras’ innovative wafer-scale engine, record-breaking inference speeds, custom architectures for AI, and the broader implications for enterprise workflows.

Key Topics Discussed

  • Cerebras Systems and AI Inference: Introduction of Cerebras' technology and its focus on AI inference.
  • Wafer-Scale Engine: Discussion of the architecture that underpins Cerebras' innovative approach.
  • Performance Metrics: Fast inference speeds and record-breaking capabilities compared to traditional GPU architectures.
  • Deployment Strategy: The role of a cloud-based service in facilitating AI deployment.
  • Open-Source vs. Proprietary Models: Debates around the benefits and challenges of both approaches in AI development.
  • AI in Enterprise Workflows: How AI is transforming business processes and decision-making.
  • International Collaboration: Partnerships with global supercomputing centers and industry leaders.

Detailed Summary

Introduction to Cerebras Systems

  • Andrew Feldman provides insights into Cerebras Systems, focusing on its groundbreaking work in AI inference and high-performance computing.
  • The episode highlights Cerebras' wafer-scale engine, which integrates vast amounts of memory and computing power onto a single chip.

The Rise of AI Inference

  • Feldman explains the transition from AI training (creating AI models) to inference (using AI models) and its implications for businesses.
  • He notes that the period from 2014 to 2023 was primarily about AI development, but starting in 2024, there is a significant shift towards practical applications of AI.

Performance Advantages

  • The Cerebras wafer-scale engine offers 7,000 times more memory bandwidth than traditional GPUs, enabling significantly faster inference speeds.
  • Feldman emphasizes that speed and capacity are crucial in redefining what is possible in AI computation.

Cloud-Based Services

  • Cerebras has developed an API-driven cloud service that allows easy and rapid deployment of its inference capabilities.
  • Feldman mentions the rapid uptake of their service, with tens of thousands of developers signing up shortly after the launch.

Competing with Established Players

  • The discussion includes insights on the competitive landscape, particularly how Cerebras intends to challenge NVIDIA's dominance in the inference market.
  • Feldman argues that the long-standing reliance on NVIDIA's CUDA is being challenged by the rapid adoption of alternative architectures and platforms, such as Llama and other large language models (LLMs).

Open-Source vs. Proprietary Models

  • Feldman shares his views on the ongoing battle between open-source and proprietary models in AI.
  • He discusses Meta's approach to open-source AI and its implications for developers and enterprises.

AI's Role in Enterprise Transformation

  • The podcast delves into how AI is increasingly being integrated into enterprise workflows and decision-making processes.
  • Feldman predicts that AI will become foundational to many corporate operations.

Edge Computing vs. Cloud AI

  • The conversation addresses the role of edge computing in delivering AI applications, especially for consumer-facing technologies.
  • Feldman explains that while some applications will operate at the edge, many enterprise applications will still rely on more robust centralized computing resources.

Addressing Uncertainty in AI Models

  • The discussion transitions to the importance of managing uncertainty in AI models and the innovative approaches being adopted by Cerebras.
  • Feldman mentions the use of multiple models to enhance certainty in outputs and reduce potential errors.

Global Partnerships and Climate Solutions

  • Feldman discusses Cerebras' collaborations with international supercomputing centers and its involvement in climate mitigation efforts.
  • He emphasizes the company's commitment to contributing to environmental sustainability through AI innovations.

The Future of AI Training and Inference

  • The episode concludes with a discussion on the relationship between AI training and inference.
  • Feldman argues that the demand for training will only grow as more organizations adopt AI, creating a virtuous cycle of development and application.

Key Takeaways

  • Cerebras Systems is leading the charge in AI inference technology, with a unique architecture that offers unmatched performance.
  • The shift towards practical AI applications is reshaping how businesses integrate these technologies into their workflows.
  • The competition landscape is evolving, with increasing opportunities for new players as they challenge established giants like NVIDIA.
  • AI's future will revolve around the interplay between training, inference, and the collaborative efforts of the global tech community.

Final Thoughts This episode of Eye On A.I. provides a thorough exploration of how Cerebras Systems is not only innovating in AI inference but also shaping the future of AI in enterprise environments. It highlights the broader implications of these advancements for technology and society.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Training is how you make AI and inference is how you use AI. And for between 2014 and the end of 23, we were making AI. It was a novelty. It wasn't important in any action. The type of AI we're talking about now. But starting in 2024, it left novelty, left cool, interesting to move into productive. And people began to use AI. And you saw this extraordinary boom in inference. And I think that's why companies like us and others are now announcing leading products. That's why we invested behind this idea of why we announced the fastest AI inference in the world. We're in the dial-up era of inference.

0:46It's slow. You can see the spinning ball of death we used to get. You could almost hear that modem noise going back and forth. Now, we have a near memory compute architecture. It's a data flow architecture on a wafer scale chip. So all our memory is on the chip and it's S-firm. And it's close, right? Each compute core has its own dedicated memory. And what this means is we have 7 ,000 times more memory bandwidth, right? The path that data moves to get to compute. It's closer and we have more bandwidth. And the result is we are vastly faster at inference. and uh that's a a sort of un over overcomable advantage for the gpu that we have there this is gpu impossible speeds because we're running faster than their memory bandwidth can allow yeah so the hbm connected to a gpu has a certain upper bound that you can move data across yeah and we run faster than that right And that's constrained by the length of the connection, the material of the connection.

2:18And the technology and the memory. Yeah. And we are faster than that. And so what we announced was the fastest inference in the world bar none for LAMA 8B 3.1 and LAMA 70B. and so we're really proud of these results they are blisteringly fast as one of our customers said speed and capacity changes everything yeah now the question is in order to to to get people to use cerebrous inference and this is a cloud-based service you're providing. You have to convince the clouds to install Cerebris wafer engines, wafer scale engines. You're starting with your own cloud. Correct. Correct. I think waiting for other people to deploy your equipment may not be the best strategy.

3:25So we stood up our own cloud and we have exaflops of capacity available and we made it dead simple to move it's less than 30 seconds it's API based there's a chat engine you can move by a single code snippet just directing to the Cerebris API it is really extraordinarily simple and we had tens of thousands of developers sign up in the first 10 days. Yeah but for this really to go to dominate the market which presumably at some point you would like to do. I would like that yes. It has to move into the big cloud providers and replace GPUs at least for for certain kinds of inference. How big a barrier is that?

4:25I know that NVIDIA had this moat for developers because of CUDA. People were comfortable with CUDA, with Cerebrus. In this case, it's an API call for people building applications. How do you get the big cloud providers to provide it at scale? So I think that's really a good question. There are a couple points there. We can provide today hundreds of millions of dollars a year of revenue for us in our cloud. But we want more, obviously. I think the dynamics of the hyperscale world has changed with the rise of people like CoreWeave and the rise of Lambda and the rise of Sovereign Clouds, all looking to do training and inference.

5:28So the relative importance of the traditional three or four US-based hyperscalers has reduced. We would still love to be a part, but there are other vehicles and paths to market as well. I think the CUDA moat today is vastly overestimated. CUDA, over the past 18 years, has had approximately 40 million downloads. In the last 18 months, LAMA has had 350 million downloads. All right, what's happened is 15 years ago, CUDA was it. Then people moved up into TensorFlow, and then they moved up into PyTorch. And right now, they've moved up into Lama. And application developers don't want to bother with C-like languages.

6:33they want to bolt on to uh to llms yeah and so um you know when you have uh you know 9x more downloads in the 12th the time right something's happening there yeah and you have to take that very very seriously and so that that's what we see and that we our view is that the building block of the future is going to be an LLM, not some C-level language that nobody likes to program in. Right. LAMA has done great. Mistral has done great. Done great. But OpenAI still dominates. They do. And why is OpenAI so married to NVIDIA chips, GPUs? Have you talked to them? It seems to me this would be natural to move some of their big models, foundation models, onto a cerebrous wafer.

7:40I think OpenAI has a complex, large relationship with AWS. Excuse me, with Azure. Sorry about that. With Azure. and that included for a period of time delivering all their compute and so I think the compute that that Azure had or GPUs as you saw in Nvidia's earning announcement they have four customers that are approximately half of their business but extraordinarily consolidation for customer consolidation for a company that big. So I think OpenAI has done extraordinary work with what was available to them. And I think based on what Sam has said publicly, they would very much like other hardware options.

8:40Now at the same time, they've built optimizations and around what they had. And so I think that's the race. But the WSE, the wafer scale engine that you're putting LAMA 3.1 on, and I don't know if you have 405B on it yet. Soon to come. You could put GPT-4.0. GPT-4, GPT-5. We can put any model on it. And so I think they are now so big that it involves a wide ranging partnership. Right. And they're not easy to do. Yeah. Have you done any pilots with them? Look, let's just see. We don't generally talk about people we're doing pilots with. We will be announcing a whole range of really interesting customers.

9:45in the weeks and months ahead. Yeah. The, what's happened generally, I mentioned to you that I interviewed Rodrigo from Sambanova, a company, frankly, that I really wasn't aware of. And they're hitting benchmarks. They claim - They do some good work. Yeah. They do some good work. What's happening with inference? I mean, there was the in the hardware market was really focused on training and NVIDIA had that locked up. And there were some other players. As you said, these these foundation models now people are building on top of. They don't need to build their own foundation model. So is the market shifting toward inference?

10:41Look, the way I think about it is training is how you make AI and inference is how you use AI. And for, you know, between 2014 and the end of 23, we were making AI. Right. It was a novelty. It wasn't important in any action. Right. The type of AI we were talking about now. But starting in 2024, it left novelty, left cool, interesting to move into productive. And people began to use AI. And you saw this extraordinary boom in inference. And I think that's why companies like us and others are now announcing leading products. That's why we invested behind this idea. why we announced the fastest AI inference in the world.

11:47Because what we saw was this movement from making it to using it. Right. And what we know about hardware is speed makes multibillion-dollar markets. Yeah. And we know this, right? If you remember, and you've got a little bit of gray hair like I do, right? You remember when Netflix was sending you DVDs and envelopes. Yes. The guy started in the town that I live in, Chappaqua. Yeah. Yeah. Right. That's what happens when your internet is slow. Yeah. Right. And when your internet is fast, right, you have a multi-billion dollar streaming business. And when your internet is really fast, my 11-year-old granddaughter is a creator.

12:32Yeah. Right. She's not just the recipient of video. she's a creator of video and sending it to her friends. It's become a two-way medium. Each one of those, each step as the internet got faster created enormous new markets. We believe the same happens with inference. By making it instant, yeah, all right, we make whole new markets possible. Again on OpenAI, they're about to come out with, I don't know what they're going to call it, GPT-5. Cool new stuff. Yeah, something new. It is reasoning agents. Right. Do you think meta will follow quickly after? I mean, this goes to the proprietary open source debate.

13:23Is open source going to be able to keep up? Is it going to be able to take the lead? I think that's an extraordinarily interesting battle. Yeah. Right. That is unfolding before us. I think the strategy that Meta has undertaken is interesting and that they undertook a similar strategy in hardware. Yeah. Right. If you remember, they pioneered OCP, the open compute projects where they open source model, open source chassis and various hardware building blocks for the community. and that was very successful for them. I think they've been a company that has pioneered openness. I think it is TBD, whether it will win in the big model space.

14:21We're rooting for it. We think openness is extremely valuable. I think their licenses are fairly permissive so that many large enterprises can build on their open models. You're talking about Meta's. I'm talking about Meta's model. Yeah, Meta's license for LAMA. I think they have seen tremendous benefit to the company for their open source work. And so we hope that there continues to be a flourishing open model creation ecosystem. Yeah. And you mentioned that inference were kind of at the beginning or its beginning. I mean, generative AI, there's a lot of pilot programs. There's very little in enterprise production yet, other than maybe customer services starting.

15:19Correct. So how big is this market going to get, do you think? More than half a computer. Yeah. Yeah. I think this will be, there will be two or three ways that compute is done. And one of them will be the way we do it for AI. I think AI will be foundational in most, or maybe not more, many of corporate workflows. And I think if your organization isn't thinking hard about these right now, you're going to be left in the dust. Yeah. And the advantage here, again, is that it's an API call. It's an API call. Yeah. That's right. I think, you know, what Python brought to the world was an interpreted, simplified way to get code written.

16:27And I think what large language models will become are these building blocks that are called upon to do tasks. And they will be underneath many, many of your corporate workflows yeah the um uh the you know 5g technology 6g i mean the the transfer of data uh wireless transfer of data is is getting better is i'm trying to ask about the edge is how does this can this reach the edge or are we really going to talk about which is what it is today a different class of chips a different class of models that can operate at the edge or or do you think that the day will come where those transmission speeds and the bandwidth is high enough that you can reach the edge from the cloud?

17:42I think many, if not most, consumer applications will eventually be delivered by the edge. And I think that is our history. It is that consumer applications get pushed to the edge. I think enterprise applications are much heavier and generally benefit from a different class of compute. And so as we see models improve algorithmically, as our ability to prune them, to use techniques like sparsity to make them smaller, as our ability to build better edge chips. I think you'll see a great deal of work done at the edge. But it's a fallacy to think that the edge subtracts from the core. I think the more work done on the edge, the more work the core needs to do as well.

18:53I think the more AI diffuses through our ecosystem, the more AI is done, the more training that needs to be done, the more value accrued to advantages in training, the more heavyweight inference you need to... Is Cerebus working on any hardware for the edge? We are not. Yeah, it's just not your business. Yeah, that's not... That's a different business. It's differently challenging. and I think the returns to Apple and others who build the edge are different. It's been a place that's very challenging for startups to enter. But with the speed of inference from the cerebrus in the cloud and with wide enough bandwidth, how much compute actually has to take place at the edge other than safety critical applications?

20:00I think it's TBD. I think you certainly want your car doing inference. You don't want that going back to cloud. I think you'd like your phone doing some basic inference, but I think just like search on your phone, there's some some hooks that use the phone CPU and some work done in the cloud. And I think that's sort of a model you'll see for inference. And when you're as fast as we are, what's left really is transport latency. And that matters a lot. Your car. You definitely don't want to do it. But I think a lot of inference, and this is, I think, a mistake people make right now, is thinking of inference as human readable inference.

21:05So it generates a stream of text or images. I think increasingly that the customer for most inferences is other machines. And so the consumer for the results of inference is going to be another chunk of software. And you're going to have collections of these where the input of one produces an output, which becomes the input of the next. And so blistering speed matters an enormous amount. Otherwise, you have a stacking of latencies. Yeah. And so I think that's the way enterprise workflows are going to unfold. You have sort of collections of inference. You've seen this already with agentic workflows.

21:59You see this already with a train of thought. So I think that's more complicated inference. And I think that's a direction we're going. Did you ever think why some businesses are all over the place and you're familiar with their brands and you interact with them seamlessly? When you think about those businesses whose sales are rocketing, like Feastables or Mr. Beast or Thrive Cosmetics, you think about an innovative product, a progressive brand and button-down marketing. But an overlooked secret is actually the business behind the business, making, selling, and for shoppers buying, simple. For millions of businesses, that business is Shopify.

22:50Nobody does selling better than Shopify. And the not-so-secret secret with ShopPay that boosts conversions up to 50%, meaning way less virtual shopping carts going abandoned and way more sales going. So if you're growing your business, your commerce platform better be ready to sell whatever your customers are scrolling or strolling on the web in your store, in their feed, and everywhere in between. Businesses that sell more sell on Shopify. Upgrade your business and get the same checkout that Feastables uses. Sign up for your dollar-per-month trial at shopify.com slash ionai. That's all lowercase.

23:44It's Shopify, S-H-O-P-I-F-Y.com slash IonAI, E-Y-E-O-N-A-I, all run together. Go to Shopify.com slash IonAI to upgrade your selling today. That's Shopify.com slash IonAI. Yeah, and make sure it's an ensemble of models where you... That's right. You want an average of several models. Yeah, we have a very interesting case study where a customer has multiple different models stacked on the same wafer. and the same stimulus, the same signal comes in. And the models are completely different classes of models. And they use the results, the correlation between the results as a measure of certainty. So we have four different types of models agree on their result, result, then they are more certain of the result.

25:08That's an extremely interesting approach to manage uncertainty. So we've seen some very, very creative work. Yeah, and that'll eventually reach consumer products. Of course. And solve the hallucination problem. That's right. That's exactly right. Yeah. How many players do you think there will be in this inference hardware market? I mean, right now it's, I don't know, four or five? Yeah, I think there are very few healthy markets that have only two players. Yeah. Right. So my expectation is they will be a handful, three to five. Yeah. I think, you know, we're proud to be among the leaders in the challenger category.

26:08and so I we intend to be here and you have the the production pipeline to to meet capacity which is always an issue it's an issue with with nvidia well and it is we increased manufacturing capacity 5x last year this year and we intend to do something similar next year so capacity is a challenge but we're up to the task yeah and this is coming out of TSMC TSMC makes the way for but remember we build a system and so and in fact that was a good decision right that when we made that decision, NVIDIA hadn't doubled down on building boxes. AMD hadn't acquired a company to help them build boxes. And several of our chip competitors were thinking they could build PCI cards.

27:13And we bet on a system because that was clearly, in our view, the way to deliver this extraordinary technology. and that's proven out to play in this market. You have to build a system. You guys are restricted from doing business in China with the current export controls. What do you think is going to happen? I mean, I have a conversation tomorrow with, I'm not going to remember his name, but he wrote the book on the chip wars. Chip wars. Yeah. Chris. Chris. Yeah, yeah. Yeah, it's a good book. Read the book. Extremely about it. Where do you think that's going to go? I mean, my view, I spent a lot of my life in China.

28:03I see where things have gone with Xi Jinping and the rapprochement with Russia. I just don't see us coming back from that. and it's to China's detriment because they're 10 years behind in chip production and chip manufacturing. Do you think that this is a fork in the road? First, I think because you are a CEO doesn't give you great insight into two political matters. No, I think it's somehow you're good at building teams and product and people ask you geopolitical questions. And I think most CEOs don't, including myself, aren't experts in that. I think what we can say is the following. We have, over the course of a 30-year career, built product together with Chinese companies and found really good people.

29:09like-minded engineers that governments are different and the behavior of governments are difficult to faddle. I think we are I think obligated as good citizens to abide by not just the letter of the law but the spirit of the rules put down by the Department of Commerce and by the executive orders. And so we've been following that diligently. I think that everybody loses if the two nations take adversarial positions. Both. Everybody loses, I think. And so I'm no expert in this domain, but we're not in a good spot, clearly. And that's sort of the observation that we have. Yeah. And China is, I mean, I frankly agree.

30:28I think that China was happy to let the Dutch do the ultraviolet etching machines and let TMC do the wafers and everybody was happy. but they're going to build their own ultraviolet etching machine. So it'll take a while, but the Chinese, they've got a lot of incredibly talented people and a lot of money. So I think the U.S. kind of misplayed that because it triggered that response in China and they'll be a competitor in the chip market. Right. As I said, this is some Sunday couch quarterbacking. You have to do quarterbacking from guys who barely played football in high school. It's easy to talk about what's going on on TV.

31:38But I think right now there are returns to industrial policy, strategic industrial policy. And I think we are in the U.S. sometimes challenged in making long-term policy. our budgeting process, the way we allocate, the way we approve funding. And I think that's a challenge for us. And large capital investments that pay dividends over decades have proven challenging in the last 10 or 15 years to get appropriated. And I think as we look forward, things like artificial intelligence, continued investment in STEM, being sure that our colleges and universities continue to produce world-class engineers, that we produce electrical engineers who, while at university, can do cool things, can build chips, can work on interesting projects, that investment needs to be made.

32:54And I think the CHIPS Act is a tiny little start. 40 billion, I mean, sounds like a big number. It's a tiny little number. NVIDIA did more than that. It's chasing that in a quarter. So this is for a nation, for economic policy and industrial policy, for a nation. And so it's a tiny little bit. It's a start. And so I think it's a good start, but I think we need to think much, much bigger. Yeah. Have you heard of anyone in China doing wafer scale? I've seen papers that they've published written about it, but we've not seen anybody who's done it. It turns out to be pretty hard. I mean, it's no joke.

33:43And we were the first to succeed at it in the 75 year history of the computer industry. Gene Amdahl failed at it and no dope was he right interesting he was lived down the street where I live in Chappaqua I always tell my kids Gene Amdahl lived in that house and they have no idea so my neighbor was William Shockley growing up the inventor of the transistor and we were young and had no idea that this was seminal in sort of invention in the industry. And all we knew was that his wife gave full-size candy bars at Halloween. So we love stopping there for Halloween. Yeah. Yeah. Okay. I'm running out of questions.

34:35Is there something I haven't touched on that you want to talk about? We didn't really talk about the inference service. The inference service is blisteringly fast. There are three tiers, and this is our cloud-based service. There's a free tier you can jump on, it's a playground, come and play. I think if you read some of the comments from early users, I mean, their eyelids were blown back at how fast it was. There's a pay-as-you-go tier where you buy a million tokens. And then there's a dedicated throughput tier, an enterprise class tier. We have seen extraordinary things made already. And as infrastructure builders, it's a pleasure when people build cool stuff on your stuff.

35:30That's one of the great joys of being an infrastructure builder. And we've seen take-up in our enterprise customers, too. and both immediately in the cloud and for deployment on their premise. And so it has just created an explosion of opportunity for us. What about the government? I mean, I have the head of the institute at Lawrence Livermore on the podcast recently. Fascinating, brilliant guy. Brawnis? No, it was... Never mind. Yeah, but it'll come to me in a minute. He was telling me about the models that they're building at the labs. Right. These are government models. I was asking him, why doesn't the government build its own foundation model?

36:30And they are. They are. You've seen large foundation models under the AI for science that's headed out of Argonne you've seen the tri-labs which is Lawrence Livermore, Sandia, Los Alamos we have a large program with them in which models are being built and our equipment is deployed you see work being done out of Oak Ridge in which foundation models are being run on frontier you're seeing work at Argonne where foundation models are being run on Aurora. I think there's a great deal of work being done both in the tri-labs and in the scientific DOE labs. And how difficult is it to get them to adopt wafer scale engines?

37:21Lawrence Livermore is a customer. Argonne is a customer. Sandi is a customer. National Center for Supercomputing is a customer. Pittsburgh Center for Supercomputing is a customer. European Parallel Computing Center is a customer, LRZ and Bavaria is a customer. We sort of ran the table on large computing sites. And these are on-prem? These are on-prem deployments. Yeah. And you mentioned Aramco when we were talking about. Well, thank you for bringing that up. We signed an MOU with Aramco here at the show. We've done a fair bit of work with Total Energies. We now have made ourselves knowledgeable on various simulation techniques and models for oil exploration and for simulation and reservoir modeling.

38:09And so we look forward to a long and successful partnership with Aramco. Yeah, yeah. Biggest company in the world. Yeah, about 10 % of the world's oil. It's an amazing organization. Yeah. On the other side of the fossil fuel coin, are you doing much for simulation and climate mitigation? We are. We actually have large models that customers have built for climate. I mean, Lawrence Livermore is doing this. But there are also models being done and used for carbon sequestration. There are models and there are techniques that are being used for carbon capture. There are all sorts of interesting companies and ideas.

39:00we've seen Total be a pioneer in this, where they inject into the empty oil well carbon gas, all sorts of other ways of sequestering carbon. Excuse me. So we've done all sorts of work in this domain. And interestingly, the nations that you might not expect it, the oil producing nations, Emirates, Saudi, have taken the leadership position in thinking about climate and environmental impact and have really been aggressive. Yeah. Yeah. Certainly the impact is most extreme in this part of the world. So they're doing really interesting work in mitigation and in clean energy. Yeah. They also have good sun, right?

40:00They have opportunity for solar. They use solar for desalination of water. They've already used clean energy. They're building clean data centers in the UAE and the UAE. and here as well. I know Neom is doing that here. So, well, I think we should talk just a sec about training. I think we are, because inference has exploded, it hasn't reduced demand for training. I think this is a myth that there's some sort of zero-sum game. I think quite the contrary. I think there's a virtuous cycle that is created, that as people use AI, they want better AI. The way you make better AI is with training. And so what we have is vast amounts of new data coming in by virtue of the inference.

Read the full transcript

40:50This puts pressure on model builders and on the computer infrastructure to produce better models, to produce more accurate models. And so we're seeing a concomitant sort of explosion or acceleration in the demand for training. I think I hear this all the time, is it going to be inference or training? And I think you've misunderstood the problem, right? You can't use which is inference unless you make which is training. And the more you use, the more you want to make. Right, except that NVIDIA really has the training market locked up. I think that's, certainly they have a great deal of market share.

41:37But I think a great deal of work is being done right now on TPUs at Google. I think real work is being done on our equipment around the world. I mean, in this region, for example, we trained on our equipment in combination with our partner, G42, the leading Arabic model. There are other models that are trying to compete with it, but for nearly a year, we have had the premier Arabic LLM. I think there are plenty of examples. I know the guys at Anthropic are doing training on non-GPUs. The guys at Stability are doing training on non-GPUs. And so I think there's, I think it's a mistake to say it's game over.

From the publisher

This episode is sponsored by Shopify. 

Shopify is a commerce platform that allows anyone to set up an online store and sell their products. Whether you’re selling online, on social media, or in person, Shopify has you covered on every base. With Shopify you can sell physical and digital products. You can sell services, memberships, ticketed events, rentals and even classes and lessons.

 

Sign up for a $1 per month trial period at http://shopify.com/eyeonai



In this episode of the Eye on AI podcast, Andrew D. Feldman, Co-Founder and CEO of Cerebras Systems, unveils how Cerebras is disrupting AI inference and high-performance computing.

 

Andrew joins Craig Smith to discuss the groundbreaking wafer-scale engine, Cerebras’ record-breaking inference speeds, and the future of AI in enterprise workflows. From designing the fastest inference platform to simplifying AI deployment with an API-driven cloud service, Cerebras is setting new standards in AI hardware innovation.

 

We explore the shift from GPUs to custom architectures, the rise of large language models like Llama and GPT, and how AI is driving enterprise transformation. Andrew also dives into the debate over open-source vs. proprietary models, AI’s role in climate mitigation, and Cerebras’ partnerships with global supercomputing centers and industry leaders.

 

Discover how Cerebras is shaping the future of AI inference and why speed and scalability are redefining what’s possible in computing.

 

Don’t miss this deep dive into AI’s next frontier with Andrew Feldman.

 

Like, subscribe, and hit the notification bell for more episodes!



Stay Updated:

Craig Smith Twitter: https://twitter.com/craigss

Eye on A.I. Twitter: https://twitter.com/EyeOn_AI



(00:00) Intro to Andrew Feldman & Cerebras Systems

(00:43) The rise of AI inference

(03:16) Cerebras’ API-powered cloud

(04:48) Competing with NVIDIA’s CUDA

(06:52) The rise of Llama and LLMs

(07:40) OpenAI's hardware strategy

(10:06) Shifting focus from training to inference

(13:28) Open-source vs proprietary AI

(15:00) AI's role in enterprise workflows

(17:42) Edge computing vs cloud AI

(19:08) Edge AI for consumer apps

(20:51) Machine-to-machine AI inference

(24:20) Managing uncertainty with models

(27:24) Impact of U.S.–China export rules

(30:29) U.S. innovation policy challenges

(33:31) Developing wafer-scale engines

(34:45) Cerebras’ fast inference service

(37:40) Global partnerships in AI

(38:14) AI in climate & energy solutions

(39:58) Training and inference cycles

(41:33) AI training market competition

 

More from Eye On A.I.

All 266 episodes
#222 Andrew Feldman: How Cerebras Systems Is Disrupting AI InferenceEye On A.I. · 42 min
Listen in VO