Gaudi processors & Intel's AI portfolio

7 Aug 2024 · 46 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Summary: Gaudi Processors & Intel's AI Portfolio

Podcast Information Title: Practical AI Description: A show focusing on making artificial intelligence practical, productive, and accessible, featuring discussions on various AI-related topics.

Episode Details Episode Title: Gaudi Processors & Intel's AI Portfolio Episode Description: Discussion with Intel representatives on their AI strategy, focusing on Gaudi processors, AI workloads on CPUs, and open-source collaborations.

Key Participants

  • Daniel Whitenack - Host, Founder & CEO at Prediction Guard
  • Chris Benson - Co-Host, Principal AI Research Engineer at Lockheed Martin
  • Benjamin Consolvo - AI Engineering Manager at Intel
  • Greg Serochi - Developer Ecosystem Manager at Intel Gaudi

Episode Highlights

Introduction

  • The episode opens with a discussion about the upcoming Intel Innovation 2024 event, emphasizing the importance of developers in tackling technological challenges.

Intel's AI Strategy

  • Benjamin Consolvo shares insights into Intel's strategy in AI:
  • Hardware Options:
  • Xeon Processors: Used for AI workloads, particularly inference.
  • Gaudi Processors: High-performance alternatives to NVIDIA GPUs, aimed at AI training and inference.
  • AI PC: Combines CPU, GPU, and NPU for optimized local workloads.
  • Future Releases:
  • Falcon Shores: A new all-purpose GPU that integrates features of both Gaudi and the existing GPU lineup.

Competitive Landscape

  • Intel aims to differentiate itself from competitors like NVIDIA through:
  • Cost-effective performance.
  • Availability: Offering alternatives when cloud resources are limited.

Software Ecosystem

  • Key software initiatives include:
  • PyTorch Support: Integration of Intel hardware into widely used frameworks.
  • Open Platform for Enterprise AI (OPEA): Encourages community contributions and collaboration for generative AI workloads.

Transitioning from NVIDIA to Intel

  • The transition to Intel's hardware is described as seamless, with existing tools requiring minimal adjustments.
  • Intel provides extensive support to facilitate this transition, including extensions for frameworks like PyTorch.

Open Source and Community Engagement

  • Greg Serochi discusses Intel’s commitment to open source:
  • Partnerships with communities like Hugging Face to ensure ease of migration and usability of AI models on Intel hardware.
  • Focus on maintaining compatibility and full documentation to support developers.

Gaudi Processors

  • Gaudi Overview:
  • A dedicated AI processor designed specifically for AI workloads, not just a GPU.
  • Offers significant memory and networking capabilities to enhance performance in training and inference tasks.

Hands-On Opportunities

  • Prospective users can explore Intel’s products via:
  • Intel Tiber Developer Cloud: Access to hardware for testing and development, with plans for free access to Gaudi in the future.

Future Vision

  • Benjamin Consolvo anticipates ongoing advancements in AI, drawing parallels to the evolution of the internet.
  • Greg Serochi emphasizes personalization and efficiency in AI applications, with Intel committed to supporting the latest technologies and models.

Key Takeaways

  • Intel is actively positioning itself as a competitive player in the AI market by diversifying its hardware offerings, emphasizing cost-efficiency, and enhancing its software ecosystem.
  • Gaudi processors provide a robust alternative to existing GPU solutions, with a focus on AI-specific performance.
  • Engagement with the developer community and open-source projects is central to ensuring the usability of Intel’s AI technologies.
  • The Intel Tiber Developer Cloud allows developers to experiment with AI capabilities without the need for personal investment in hardware.

Conclusion The episode concludes with encouragement for listeners to engage with Intel's AI ecosystem through hands-on experimentation and to stay tuned for future developments in AI technology.

For more information and to join the AI community, listeners are invited to visit the Practical AI website.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:05Welcome to Practical AI. If you work in artificial intelligence, aspire to, or are curious how AI-related tech is changing the world, this is the show for you. Thank you to our partners at Fly.io, the home of changelog.com. Fly transforms containers into micro VMs that run on their hardware in 30 plus regions on six continents. So you can launch your app near your users. Learn more at Fly.io.

0:44What's up, friends? Intel Innovation 2024 is right around the corner. Accelerate the future. Registration is now open, and it takes place September 24th and 25th in San Jose, California. This event is all about you, the developer, the community, and the critical role you play in tackling the toughest challenges across the industry. Ignite your passion for AI and beyond. grow your skills to maximize your impact, and network with your peers as they unleash the next wave of advancements in technology. Here's what you can expect. Understand the emerging innovation and trends in dev tools, languages, frameworks, and technologies in AI and beyond to empower you and the solutions you're building.

1:28Get in-depth technical experience doing hands-on workshops, labs, meetups, and hackathons to collaborate and solve problems in real time. You can explore featured partner and Intel solutions. They have partners there, startups there, customers there. And Intel is showcasing the latest in products, services, and solutions across keynotes, tech sessions, and the show floor to help you meet your development needs. Collaborate with experts, learn and have fun, engage in interactive sessions to connect, get certified, gain unique ideas and perspectives, build long-lasting networks, and of course, have fun.

2:05And get inspired, hear from leading industry experts, technologists, startup entrepreneurs, and fellow developers, along with Intel leadership, CEO Pat Gelsinger, and CTO Greg Lavender, as they take you through the latest advancements in technology. Don't miss this chance to be at the forefront of innovation. Take advantage of early bird pricing right now until August 2nd. Register using the link in our show notes. Or to learn more, go to intel.com slash innovation. Once more, that's intel.com slash innovation or go to the show notes and click that link.

2:46Welcome to another episode of Practical AI. My name is Daniel Whitenack. I am founder and CEO at Prediction Guard. I'm joined as always by my co-host, Chris Benson, who is a principal AI research engineer at Lockheed Martin. How are you doing, Chris? Doing great today, Daniel. How's it going? It's going well. I am really happy to bring a little bit of content to the show today that is from a little bit of the world that I've been operating in and some of the groups that I've been collaborating with at Intel, which I think is a really kind of cool set of stuff that maybe people are, maybe they're less aware of it, maybe they're aware of it, but really happy to have with us today Benjamin Consalvo, who is an AI engineering manager at Intel, and then also Greg Siraki, who is a developer ecosystem manager at Intel Gaudi.

3:45Welcome. Yeah, thank you. Thanks for having us. Yeah, well, like I say, I think, you know, people are, of course, aware of Intel and that, you know, Intel is sort of everywhere in one degree or another in cloud providers and PCs and all of that. But I think some people might not be aware of kind of the strategic moves that Intel is making in the AI space and what they're focusing on. Ben, I'm wondering if you could give us a little bit of a high level sense of what Intel is doing as related to AI and how that's kind of featuring in their strategy right now. Yeah, no, I think a lot of people don't know Intel in terms of our strategy with AI.

4:28And so I'm happy to talk through that a little bit. So both in terms of hardware and software, we're really aimed at AI in terms of the hardware front, which people are perhaps more familiar with. We have our Xeon product line, which is a data center CPU that we use for inference in a lot of cases for AI workloads. And then more recently, we've announced the AI PC and kind of started that whole category, which includes the machine has a CPU, a GPU, and then what's called an NPU, a neural processing unit. And so that's exciting. That's kind of to optimize workloads on your local machine. And then back to the data center side, we have the Gaudi product line, which is really a really good performance substitute for a lot of the modern GPUs that are out there.

5:29So that's really exciting as well. Really powerful data center hardware that we have. So that covers some of our hardware. I guess the other big one that I want to emphasize as well is we have Falcon Shores coming out in the future, which is an all purpose, you know, GPU, a data center GPU as well. So kind of leading into that is Gaudi, and we'll get more into that in the episode. But on the hardware front, we have multiple products for AI. So that's really exciting. And then on the software front, Intel, again, spans kind of the whole gamut of software for AI, for enabling workloads in AI. But rather than kind of going through the whole software stack that we have, I'll just talk about a couple of things that I'm excited about.

6:22So the PyTorch 2.4 release includes support for the Intel GPU. So right now is the Mac series GPU and will be Falcon Shores. So that's really exciting that the upstreamed mainstream version of PyTorch has now support for that. And then coming soon, PyTorch 2.5 will have support for the ArcGPU, which forms a part of the discrete GPU product line that we have. And that's also included in the AI PC. So those are a couple of exciting things with PyTorch that are happening. And then we also have the OPEA ecosystem. That's the open platform for enterprise AI. And it's kind of this open framework that we have that multiple people can contribute to for Gen AI workloads, such as chat Q &A, code copilot, and different other Gen AI examples.

7:25So, yeah, those are a few of the things that I'm excited about on the software front. Wondering if you could talk a little bit about how Intel is uniquely positioned in the AI landscape. if you have both some very large competitors out there, there's a bunch of smaller ones in the space. How do you guys see yourselves? And how do you approach that point of differentiation? Yeah, so I think where we can really compete well in this space is cost and performance and availability. So when you're running out of availability or don't have the ability to rent that GPU you need from the cloud, from NVIDIA and, you know, you need something cheaper that also, yeah, runs AI workloads really well, I can say with confidence there's, there are some good options coming out of Intel that I've used personally.

8:16And I, you know, I come from a background in AI engineering where I was only using NVIDIA GPUs and I have gotten to work at Intel and to, you know, to try out our, our Gaudi products for training AI, you know, deep learning AI models. And then I've also gotten to run inference on our Xeon products and even run training as well on our Xeon data center products quite a bit to do more fine tuning. So yeah, I would position Intel in terms of those things, just being able to have good options for out-of-the-box performance hardware that people might not be aware of. Yeah. Could you take just a moment, just a quick follow-up to what you were saying.

9:02You brought up the big competitor being NVIDIA out there. And a lot of folks out there, just as you were, had been on their hardware using their GPUs and such. You talked about that transition. Could you talk a little bit about what is it like coming over to Intel when historically maybe you were on NVIDIA as a platform and you guys are now coming as a real powerhouse? What does that migration and transition look like and feel like if you decide to go for that? Yeah, no, and I think I can relate in terms of my background in, you know, in deep learning as I was first working on with TensorFlow and then with PyTorch mainly.

9:44And then now since I've been at Intel, there's been a lot on, you know, Hugging Face and Transformers and those libraries as well. And what I'd say in general is that the transition is not difficult. A lot of the same tools that I'm used to using for development on the NVIDIA platforms, I can use those same tools with some slight code modifications for the Intel products. So in terms of like a software leap, it hasn't been too difficult. And Intel actually historically has had, we have both the upstream support into those frameworks where we have our own developers and the community developing for Intel in the mainstream frameworks.

10:29But we've also in the past had Intel extensions where there are gaps. So for example, Intel has had Intel extension for PyTorch where there's not yet support in the mainstream framework. You can install this extra package to get all the support that you need, again, with just a couple lines of code change. But we're constantly aiming to get our changes into the mainstream framework so that it's just easy for developers to use. But yeah, in terms of software development, I haven't had huge obstacles for transitioning over to different hardware. Yeah. And you mentioned kind of some of those important open ecosystem projects, whether that be things that Intel is maybe driving more directly like the OPEA stuff or it's kind of more community things.

11:24Greg, I know being kind of developer ecosystem manager for part, I know you're focused more on the Gaudi side, but we were talking before the show even about like our team is utilizing a lot of these great, great packages that actually aren't even in, you know, an Intel repository on GitHub. There may be a Hugging Face repository like Optimum has been a big one that we've used. But there's other frameworks like TGI I know that are important for what Intel is trying to do. So maybe we'll, of course, get into more details later. But just at that kind of open source level for those out there that are not only wanting to utilize these great packages, but also contribute to them.

12:12How has Intel engaged in that open source community? Right. The key thing here is wanting to maintain as much of that connection with the open source. And this really started with Gaudi when the project was introduced four years ago, where we started with TensorFlow and PyTorch. And now the industry has moved to PyTorch, and so have we. And as Ben said, our goal here is to have full PyTorch support in native PyTorch. So we're working towards that. the same thing with DeepSpeed and with Megatron DeepSpeed for large language models. So we engaged with the full ecosystem to support those things. So we can talk specifically about our support for PyTorch.

12:51So a customer can run their PyTorch models and migrate them directly onto Gaudi. And so, for instance, we have a tool that takes a model that maybe was running on a GPU architecture and in real time migrate some of those things and move some of that code that was for GPUs and change them over to things that Gaudi can understand. But the key thing is if you're running on PyTorch, you can bring your PyTorch models over to Gaudi. If you've been using Optimum, if you've been using Hugging Face, we've partnered very closely with Hugging Face and have a dedicated library called Optimum Havana or Optimum for Intel Gaudi.

13:29And that is a dedicated set of fully performant and fully documented examples of LAMA 2, LAMA 3, OPT, Mixtrel, Mixtrel, all the important models that people are using today for both fine-tuning and inference in Hugging Face. So if you're taking advantage of using Hugging Face, then it's really easy to bring those models over. And then we look at, again, for training, we have our partnership with, we're using DeepSpeed, specifically using Megatron DeepSpeed, which is really optimized for doing those large scale, large language model training where we're taking an advantage of the tensor parallelism, pipeline parallelism, and data parallelism that Megatron provides to be able to really get customers to scale and start using our product very quickly and very easily.

14:17I know one of the things that was really cool for me and just, I know a lot of people have been working on this and contributing a lot but i i just love the because the reality for us when we were building prediction guard is we had a bunch of transformer based code running and one of the things i liked about the examples when i was trying out this stuff was you sort of had the the example of like loading a model in with with optimum or with transformers and then kind of the after and it was just sort of like like you'd see a git diff in a repo it's like hey change this line to this line and then you're basically um pretty good to go i remember i was actually on a plane to india during like a hackathon back uh last june and had got access to some of these gaudi hardware these gaudi processors um in intel developer cloud and i remember doing that and going through those examples.

15:16And by the time I had landed in Bangalore, I had the models up and running and on the Gaudi processors for what we needed to have, which was pretty cool. So yeah, great work to you all and the whole team in terms of providing some of that sort of functionality and tying in very closely with that ecosystem. Yeah, it's really important that because Hugging face is so huge and so pervasive. We wanted to make sure it was really easy for people to migrate over and just even take advantage of the models that we have already optimized. So a lot of the work we do there is really managing at the lowest level of managing some of the static shapes and managing the bucketing and making sure that we have the most optimized models.

16:02And as you said, Daniel, they're fully documented, right? So it's really easy for people to go into the repository on GitHub and you will see examples of running something as simple as doing text generation with GPT on one card or going and running a full Llama 3 70 billion parameter model on eight cards. Or if you have access to more nodes, up to 16 or 32 cards and everything is fully documented. So like I said, you get off the plane and you're already running. And it's a great starting point for people to begin their development, either to take their existing model and fine-tuning it with Gaudi and taking advantage of that performance, or being able to run inference and applying that to their applications, just like you've done with Prediction Guard.

16:52So we've kind of dived into talking a bit about Gaudi, but I'd like to pull us back. And for those of us that are out there listening and are not familiar with it, could you possibly kind of give us a what is Gaudi and kind of introduce the whole platform in the broad and talk about kind of where it came from, how it came about, you know, what Gaudi is versus maybe some of the other things that Intel has that are not specifically Gaudi. and just kind of give us a context setting about what Gaudi looks like? Yeah, Chris, that's a great question. So let's talk a little bit about what Ben has mentioned a moment ago about the overall Intel product roadmap.

17:34And you look at sort of three key areas. You think about the PC and we have the AI PC that Benjamin was talking about. Then we have the Edge where we, you know, the Edge has huge latency requirements and performance requirements. And there's great solutions there with Xeon. to handle those low latency, you know, on-premise requirements. And then the final is really that large language model training and inference in the data center, where we're looking at fine-tuning and pre-training models, as well as running inference on large batch loads or batch sizes or large batches or dealing with users running an application where we have multiple, multiple, multiple users trying to take advantage.

18:16And the reason why we have Gaudi and the reason why Gaudi exists really is to give the ecosystem a choice. One thing we've been hearing from customers over and over is they want an alternative to the standard mainstream GPU solutions because of cost and because of availability. So Gaudi really is that low-cost alternative to the standard NVIDIA GPU solutions that are in the market today. So in a little bit of a history, the company Habana was an independent company. and Intel really saw the value in the product they were building, the performance. So Intel made an initial investment in 2016, 2017, and then fully purchased the company in 2019.

19:03So it was a small company and they really needed to scale. So Intel invested in the company and brought them inside of Intel and brought many Intel employees. I was one of them into the structure within the Gaudi team. So we started to build the product and started to ship. And so the first real milestone there was launching the first generation of Gaudi on AWS. So today that's available in the DL1 instance on AWS. It's still available today. And the next step was in Gaudi 2. And Gaudi 2 launched a year and a half ago. And one of the real key milestones with Gaudi 2 was the submission for MLPerf on training and inference.

19:45And, you know, you look at, for those that don't understand, may not know, MLPerf is a benchmark that the larger ecosystem uses today to measure these really standardized benchmarks. And so the way that you run those MLPerf benchmarks makes it easy to have a direct comparison from product to product. And one of the really key benchmarks was the large language model training benchmark for MLPerf. And Gaudi was one of the only products other than NVIDIA that submitted an actual benchmark for MLPerf, showcasing the performance. So why Gaudi? What is Gaudi? Is Gaudi a GPU? No, Gaudi is not a GPU. It is a dedicated AI processor that is used to manage and train and run inference on today's largest and most complex workloads for both training and inference.

20:36Could you differentiate a little bit for me as I'm learning as we go here, when you say a dedicated AI processor versus a GPU, could you kind of distinguish between those? Well, in some cases, a GPU still has the ability to do some other things. It's got additional programmability to do other types of workloads, whereas Gaudi as a dedicated AI accelerator is built specifically for AI training and inference. So we don't have those additions. It's specifically built for AI. So it's a product where if you're wanting to go run these workloads in the cloud, on the edge, or in an on-premise solution in your on-premise data center, Gaudi really is that low-cost solution.

21:19As an analogy with a potential competitor that people may know just to transition, Google has their TPUs, which sound sort of similar to that where it doesn't have all the extra stuff that a GPU has. is it and i and i i get gaudi as its own thing but is it uh just for people to make that connection a little bit similar to that yeah come with somewhat similar and you know that i had one more thing that really is a great differentiator in gaudi is is part of the hardware architecture we provide two things one is uh 96 gigabytes of on board on card hbm memory and for those the audience that really know about training and inference you know having that local hbm memory is critical for storing your weights or your parameters, the things that actually are stored when you're actually running the inference or running the training.

22:08So having that large HBM memory allows you to do one of two things. You can either run a larger model on a single card or you're more efficient as you scale out to more cards. And the other thing that really is a key differentiator for Gaudi is the onboard networking. So on DAI, Gaudi offers 24 100 gigabit Ethernet ports. So instead of relying on the latency of having to use a third-party network controller to scale to multiple nodes in a rack, for example, these dedicated Ethernet ports allow, in some cases, a direct all-to-all connection for eight Gaudis in a node or a single server. And then the additional Ethernet ports then can go to the other nodes in a rack and then scale out to a switch to go to multiple nodes and then a full pod.

23:03So what you end up with is a virtual alt-alt connection, which significantly improves the scalability when looking at large workloads. Ben, you were talking a little bit about your prior work in deep learning, in training models, and doing inference, and how you've transitioned a bit and also done some projects with training on Gaudi and various other hardware that we've already talked about. Could you speak a little bit more to that kind of practical side and just to give people a sense of the kinds of projects that are possible on this hardware, you know, just so they can form in their mind kind of both the scale and possibilities with what can be done?

23:46Yeah, no, and my use cases might not be the full expanse of what's possible on our hardware, but I can certainly speak to the things I've been able to work on and had some fun with over the last few years. So one of them is working with OpenAI's Whisper model, which is a translation and transcription model that works very, very well out of the box. And so I speak both English and French. And so I tested this, you know, the abilities of this model, you know, to transcribe my own voice, English to English, and then also its ability to transcribe from English to French and then going the other way as well.

24:30And yeah, first of all, the model does a really great job. I was really impressed with the fidelity of the model that's been trained that OpenAI has released in the open source. So that's great. And then what I was able to do was to run this model, you know, on our Xeon product line to run inference on Xeon and run that really, really quickly and kind of put together a notebook and a workshop around that. So that's been something that's been fun for me to work on, on this generative AI model that is, you know, used for translation and transcription. Yeah, that's one. Another one that has been exciting for me to work on is with my, I have a background in imaging and computer vision kind of prior to working at Intel, where I was applying a number of different techniques like pixel segmentation and image classification, object detection to different problems in computer vision, especially around geophysical imaging where we're imaging the subsurface.

25:33I kind of describe it like an ultrasound for the earth. So if you know what an ultrasound is, looking at imaging, you know, the subsurface of the earth to try to find, extract and find certain minerals and that kind of thing. So where I've applied some of the techniques there is to help, you know, get some of those images more quickly and both in fine tuning models and also in inference. And so I've been able to use some of our Xeon, again, product line to fine tune some of those models as well as run inference. So that's been exciting. Yeah. I'll add this too, just to add on to that. One area that's really getting a lot of focus now is the use of RAG, Retrieval Augmented Generation.

26:22We've invested a lot of effort there. You'll see that in the OPEA project where we have a lot of RAG-based examples. But not only for just general RAG usage, which is important, but also transitioning into multimedia. So we're seeing a huge request for, and we see this in the market, for both video, audio, and text all coming together. So whether it's a prompt of text to get a picture or a prompt for text to get now a video or to use a rag type of usage to be able to parse through not only text, but parsing also through video and then get a response. So those are all of the types of things that we have available for use on Gaudian and the products.

27:10Yeah, and actually, that's something I'll pick up on, too, because I've been working on some multimodal models on the AIPC, actually running some multimodal, smaller, 2 billion parameter multimodal models, and actually successfully ran one on the NPU, the neural processing unit, to run inference and to extract essentially some information about an image into text. So using that multimodal capability.

28:10From the ground up in AI, quantum technologies, cloud native, and more, their newest AI innovation, Motific, addresses a critical challenge in the rapidly advancing world of Gen AI. Bridging the gap between concept and deployment, this model and vendor agnostic solution supports the entire Gen AI journey. From assessment and experimentation, Motific accelerates deployment from months to days while safeguarding against Gen AI security. trust, compliance, and cost risks, all while empowering business function and IT teams to rapidly configure end-user assistance powered by organizational data. Motific provides advanced customizable policy controls to prevent unauthorized access to sensitive data and helps ensure compliance throughout the entire process.

29:01With deep visibility into operational and business metrics, Motific enables you to track ROI, optimize costs, and make informed decisions. By offering a centralized view, Motific deters shadow AI usage and empowers teams to innovate responsibly. So move beyond the traditional constraints of AI implementation, utilizing AI deployment that is both responsible and is revolutionary, ensuring your projects are not just quickly launched, but built on a foundation of trust and efficiency. Visit Motific.ai. That is M-O-T-I-F-I-C dot A-I.

29:55so uh ben and greg we've talked about a lot of interesting things both on the kind of software and migration side but also on the hardware and what that hardware enables people might though be out there and be wondering this is cool i'd love to experiment with this stuff, how can I get hands-on with some of this hardware that I'm hearing about? What are some of those ways that people can, they might have an Intel processor in their PC, or likely many do, but they don't have a Gaudi or a Xeon sitting around at the moment. So if they were to want to kind of explore this ecosystem and get hands-on, try some things, how might they do that?

30:38Yeah, I can start. And then Greg, you can fill in anything, Greg, that I'm missing. But I think the best way is to get onto the Intel Tiber developer cloud, which is, you know, a way it's kind of our Intel cloud where developers can come and try out both our hardware and our software that's set up right there for them. and we'll offer kind of our latest, whatever we have, you know, on that platform. So we'll offer, you know, we have our Xeon product, we have our Gaudi product. We have, we even are going to be having like a dev kit kind of for the AIPC, a simulated environment, even though it's not a, you know, a local machine, it's still a cloud.

31:16So that's, I think the best way to get started is the Intel Typer Developer Cloud. And Greg, did you have anything to add there? That's the best, that is the best way. You know, we're going to have some pretty soon, some free access to Gaudi. So right now you need a developer credit. We're working to get some nodes available for free on the developer cloud. So people will have the ability to try out our tutorials and examples and code examples and be able to see and experience the Gaudi usage and see how easy it is to run. Yeah, that's awesome. And I mentioned earlier some of the experimentation on my end, even on the plane, but a lot of that was enabled by this developer cloud.

Read the full transcript

31:55And I think there are, you all can correct me if I'm wrong, but people can sign up on the site, get access to, and there's also some like training resources. People can spin up notebooks, try a variety of things. Maybe if they're not as familiar or they're learning, get access to various things. But there's also a kind of transition to, within that environment, utilize these powerful products in a production sense or in an enterprise sense rather than just a kind of developer experimentation sense, which is definitely the transition that Prediction Guard has taken. So we've been able to operate very price performant at scale with our with our LLM engine AI platform on top of Gaudi and running that in Intel Tiber Cloud.

32:46So I don't know if there's anything you'd want to highlight, whether that be kind of success stories or just commenting on that kind of transition to to production, that there kind of are people running this stuff in production, not just in a kind of developer experimentation sense. Anything you'd want to highlight there? Yeah, Daniel, that's a great point. You know, the Tyber Developer Cloud is really meant to do two things. One is just to give people the access to our products. And so they have the ability to experience and run them and test them. But Daniel, to your point, it is also a place for people running a business to be able to have really easy access to our products as well.

33:28and specifically with Gaudi, we're enabling customers now with very large scale out to be able to do full production workloads and use the Typer Developer Cloud as a baseline for their business. So we invite those listening to be able to reach out to your sales contacts in Intel and really be able to talk about how we can help those people really using this for business to be able to scale as a real product. And also, as we look to Gaudi 3, which is our new product that we've announced at previous events, and we're going to make a very large announcement at Intel Innovation in September, that's also going to be a place where we'll see Gaudi 3 also begin to scale.

34:13So it's definitely a place where it makes it easy to partner directly with Intel or with partners in the future. Yeah, I'm curious on whether it is existing or maybe a roadmap item. Could you talk a little bit about Gaudi at the edge? And when you're not in the cloud and you're out maybe wanting to use Gaudi in devices that are out there or platforms that are out moving about, what's the roadmap look like on that? Right. So going forward, we're going to, so today it is a, you know, OCP compliant part. So it's on a mezzanine card. So it's meant for data center, right? So today the Gaudi platform is in a 6U or 8U rack mount server.

34:59That's the form factor it has today. And that OCP form factor of spec is eight Gaudis on a single baseboard. And as you may have noticed at Intel Vision a few months ago, we announced full packages that you can buy. So you can buy a Gaudi 2 baseboard or a Gaudi 3 baseboard that has the full baseboard that's OCP compliant. So you can drop that into a chassis from Supermicro or WeWin or other products. But in the future, we're also going to have a standalone PCIe card that will be available. And that's going to be exactly for those type of more on-prem sensitive solutions. Chris, like you said, at the edge where people can take advantage of a single Gaudi 3 and its capability on the edge.

35:42So you'll see that PCIe card coming soon. Yeah. And maybe that's a good transition to talk a little bit, like we've talked a lot about the things that are kind of the now of what's available and kind of tooling or hardware wise with Intel, but both of you have alluded to kind of the future to one degree or another. So maybe, maybe Ben, I'll start with you. But I know you mentioned certain things, whether it be the Falcon Shores or some cool things that are happening with AI PCs and that sort of thing. But what kind of strategy wise and kind of positioning wise is Intel really thinking about and investing in moving kind of into into the next phase of what AI is becoming and where Intel thinks the market is going, I guess.

36:31Yeah, no, thanks. Probably the best place to start on this is, and what I'm excited about is Falcon Shores, like you pointed out. So Falcon Shores will be kind of the culmination of combining the Gaudi product line with the GPU, with our current Max series GPU. And it will be a GPU, graphics processing unit. So it will be a full GPU capable of not only the AI workloads, but also graphics and other applications that people want to use GPUs for. And so I think that's probably the most exciting thing that we're kind of aiming at, that we know the market needs. And then, yeah, iterations on the, as you mentioned, the AI PC.

37:19So we'll be, in the future, will also have a new, more powerful AI PC with the Lunar Lake chip coming out. And again, it will include the CPU, GPU, NPU, but it'll be a much more powerful, more memory form factor that where developers will get even more out of their local machines. So yeah, those are a couple things. And then the other thing that I'm excited about just as an AI software developer myself is the integrations with, like I pointed out at the beginning, the integrations with PyTorch. I think it's huge that we're aiming at getting all of our optimizations and everything we can into this framework that is by far one of the most popular deep learning frameworks and one that I use on a regular basis.

38:14So that's really exciting to me as well that just natively I'll be able to, you know, work with PyTorch and say, hey, I want to use the XPU, which is, you know, for the Intel GPU and same with the AI PC, just have that direct integration. So yeah, those are a couple of things that are exciting for me coming out soon. And I'll add to that to say, you know, the key thing here is forward compatibility, right? So people that are using our products today saw specifically speak to Gaudi. If you're running workloads on Gaudi 2, you'll be able to run those workloads directly on Gaudi 3. And that same architecture will move forward into Falcon Shores.

38:55So people that make their technological investments now, those will remain viable and relevant far into the future. Yeah. And I guess one other piece of this, which, you know, Greg, you mentioned kind of in passing, one of the things that people are really interested in with Gaudi, of course, is the fact that there's some diversity in the market and there's another choice, right, for hardware out there. I think one of the other interesting things that I don't know if I know you two are kind of only in pieces of Intel and focused on certain things, but I found it really interesting how Intel is very much investing in chip production kind of diversity as well outside in various geographies around the world as a key part of their business.

39:47And I don't know if you have any comment on that or thoughts on how that influences the market as a whole and availability or supply chain sorts of robustness that that can build in. But I know we've seen, you know, we've seen some interesting things over the years, both in terms of availability and supply chain issues with hardware. Yeah, and you could talk about, let's talk about this at a macro level and maybe at an AI level, right? At a macro level, you look at what our CEO, Pat Gelsinger, has talked about when we promoted the CHIPS Act, for example, that we need, Daniel, to your point, we need to be able to build that infrastructure worldwide.

40:26So you can see from an Intel perspective, our investments in our fab in Ohio and our fabs now in Germany, as well as just our fab worldwide, we have that worldwide capability to support significant growth and expansion as the world continues to need more and more silicon. From an AI perspective, the growth and the need for AI compute is insatiable and will continue to be that way for the foreseeable future. So again, this goes back to why we've invested in Gaudi and brought that as a product as part of Intel and as part of just growing the broader AI portfolio. We really want to be able to give the ecosystem an alternative to getting access to AI compute as they need it today.

41:11So as we start closing up, I would really love to hear from each of you, and you guys can decide who wants to go first, but where do you see it going? Where Gaudi is going and where the overall ecosystem and these technologies are going? And this is a moment where you can take a little bit of poetic license and speculate a bit. I would love to see what you think will unfold and happen in the times ahead and how each of you may see it a little bit differently as individuals. Sure. Yeah, I can. Yeah, I can start. So I think, yeah, just looking at kind of AI broadly and, you know, what's happening, it seems like things progress with these incremental, you know, these incremental changes.

42:02And sometimes there's a leap, but there's often just these incremental changes. And one of the questions I often get from friends, maybe you guys do too, as you're working in AI is like, you know, is AI going to take over? And, you know, are we going to have kind of robots controlling everything we do? And, you know, so that seems to come up a lot as I say that I work in AI. Sometimes I regret saying I work in AI and just say I should just say software engineering. I mean, but no, it's always an interesting conversation that I have with different friends. And, you know, my perspective is that like the internet, the coming of the internet, AI has come along and has changed the way we work and changed the way we operate.

42:45But just like the internet, we have people behind building these things. And as these technologies evolve, we will have safeguards and we will have things in place to help regulate the different technologies of AI that come out. And lately with the AI agents, I think is one of the most recent things where you have the agent kind of do more things for you than maybe previously where you had to ask it to do more things. So that's been a really interesting part of AI that's kind of come out and that I think is going to see a lot more adoption in the future to answer your question about the future.

43:27I think just getting the AI tools to build more things and to be able to kind of do more complex tasks in sequence is something that's evolving and happening. Yeah, Greg, did you have some more to add? I love, I'm so excited about the personalization of AI. You know, I see cases where now things we didn't have when we went to college, but now, you know, you can have your phone or your AI PC open in your college classroom. And, you know, the AI will summarize and create notes for a lecture or quiz you on a lecture and create all that content for you automatically. I love that. I want to see AI do better with my email and be able to organize my email better, do searches better, make my life better.

44:17Those are things I'm really excited to see from a general perspective. From an Intel perspective, I think you're going to continue to see us lean in on giving customers what they need, which is being able to have more and more compute for fine-tuning and for inference, either on-prem or on the edge, supporting the world's largest models. And as we see innovation happening on a monthly basis that used to take a year, now we're on a monthly basis. We will keep up with that innovation. Just as an example, Meta just launched their Lama 3 400 billion parameter model in the market last week. So we already been running that model and we've support that model.

45:00And we wrote a blog on it a couple of days ago. So we're going to continue to support the most bleeding edge, latest and greatest technology that's coming out again on a monthly basis. Thank you both for taking time out of a lot of things going on in a fast moving ecosystem to come and chat with us. And I would definitely recommend to all the listeners to check out some of the show notes and the links and go try some things hands on and have some fun and start building. Thank you both, Greg and Ben. I appreciate you taking time. Yeah, thank you. Thank you, Daniel and Chris. Thank you.

45:40All right, that is Practical AI for this week. Subscribe now. If you haven't already, head to practicalai.fm for all the ways. And join our free Slack team where you can hang out with Daniel, Chris, and the entire ChangeLog community. Sign up today at practicalai.fm slash community. Thanks again to our partners at fly.io, to our Beat Freakin' Residence, Breakmaster Cylinder, and to you for listening. We appreciate you spending time with us. That's all for now. We'll talk to you again next time.

46:25Game on!

From the publisher

There is an increasing desire for and effort towards GPU alternatives for AI workloads and an ability to run GenAI models on CPUs. Ben and Greg from Intel join us in this episode to help us understand Intel’s strategy as it related to AI along with related projects, hardware, and developer communities. We dig into Intel’s Gaudi processors, open source collaborations with Hugging Face, and AI on CPU/Xeon processors.

Join the discussion

Changelog++ members save 5 minutes on this episode because they made the ads disappear. Join today!

Sponsors:

Featuring:

Show Notes:

Something missing or broken? PRs welcome!

More from Practical AI

All 157 episodes
Gaudi processors & Intel's AI portfolioPractical AI · 46 min
Listen in VO