#308 Christopher Bergey: How ARM Enables AI to Run Directly on Devices

19 Dec 2025 · 54 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Eye On A.I. Episode #308 - Christopher Bergey: How ARM Enables AI to Run Directly on Devices

Podcast Title: Eye On A.I. Host: Craig S. Smith Guest: Christopher Bergey, Executive Vice President of Arm's Edge AI Business Unit Episode Date: [Insert Date Here] Sponsored by: Oracle Cloud Infrastructure (OCI)

---

Episode Overview In this episode, Craig Smith interviews Christopher Bergey, focusing on how ARM's technologies are enabling artificial intelligence (AI) to function directly on various devices instead of relying solely on cloud computing. They discuss the implications of edge AI across multiple platforms, including smartphones, PCs, and wearables.

Key Concepts Discussed

  1. Movement from Cloud to Edge
  2. Rationale: There is a significant shift from AI processing in the cloud to devices, termed "on-device intelligence," due to improved processing capabilities and the need for low latency and trust in user experiences.
  3. Example Applications: Smart cameras, hearing aids, robotics, and automotive systems.
  1. ARM Architecture and V9
  2. History: ARM architecture, pioneered by companies like Apple and Nintendo, has been foundational for mobile computing for nearly 30 years.
  3. ARM V9: The latest architecture focuses on security, performance, and AI capabilities, enabling AI inference at the edge.
  4. Matrix Extensions (SME): Enhancements that facilitate AI processing, improving efficiency without burdening the developer with new programming languages.
  1. Heterogeneous Computing
  2. Definition: The integration of CPUs, GPUs, and NPUs to perform various tasks efficiently.
  3. Programming Challenges: The need for a user-friendly programming model that allows developers to harness the power of heterogeneous computing without the complexity of learning new languages (such as CUDA).
  1. Memory Bandwidth Challenges
  2. The most pressing issue for AI workloads is memory bandwidth, which affects performance significantly. ARM is working on scalable solutions to improve memory efficiency.
  1. Real-World Applications of Edge AI
  2. Use Cases:
  3. Hearing Aids: AI processing to filter and enhance sounds.
  4. Smart Cameras: Facial recognition and event detection without cloud delays.
  5. XR Devices: Enhanced user interfaces that respond to gestures and commands.

Future of AI and Technology Interaction

  • User Experience Evolution: AI is becoming an essential interface; users will expect intelligent interactions with their devices.
  • Agentic AI: Future devices will communicate and perform tasks seamlessly, anticipating user needs much like how modern touchscreens are expected to function intuitively.

Challenges in AI Development

  • Power Management: Balancing power consumption with performance in edge devices.
  • Model Size and Efficiency: The need to optimize AI models to fit within the constraints of edge devices while maintaining performance.

ARM's Role in the Semiconductor Ecosystem

  • ARM provides intellectual property (IP) for chip design, enabling other companies to build upon their architecture. This model allows for widespread innovation across a variety of devices and sectors globally.

Closing Remarks

  • The episode emphasizes the importance of developing AI technology that is practical, efficient, and user-friendly, pointing towards a future where devices are expected to be inherently intelligent.

Key Takeaways

  • The transition to edge AI represents a substantial advancement in how users will interact with technology.
  • ARM's V9 architecture and innovative approaches to memory and processing are pivotal for the development of devices that incorporate AI capabilities.
  • The future interactions with technology will increasingly involve seamless, intelligent devices that can respond to user needs without latency or reliance on cloud services.

---

Stay Updated

  • Follow Craig Smith on X: [@craigss](https://x.com/craigss)
  • Follow Eye on A.I. on X: [@EyeOn_AI](https://x.com/EyeOn_AI)
  • Learn more at ARM's Developer Portal: [developer.arm.com](https://developer.arm.com)

---

For more insights on AI and its applications, tune into the Eye On A.I. podcast biweekly!

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00The ARM architecture has existed now for almost 30 years. And that started from early investments from companies like Apple, early adopters of the ARM architecture that made it the stalwart that it is, were companies like Nintendo, companies like Nokia back in the 80s and 90s. These smartphone revolution, all that kind of stuff really started around ARM. And that's what's driven us to be where we are today. We have big CPUs and little CPUs, and we're actually moving the workloads back and forth because certain times you need the performance, certain times you don't. And so that's really the way these devices work, where to your doorbell example, you're looking for maybe some motion, you're looking for something.

0:39And then once you trigger that event, OK, now let's fire up some more computing elements. In business, they say you can have better, cheaper or faster, but you only get to pick two. What if you could have all three at the same time? That's exactly what Cohare, Thomson Reuters and Specialized Bikes have. since they upgraded to the next generation of the cloud, Oracle Cloud Infrastructure. OCI is the blazing fast platform for your infrastructure, database, application development, and AI needs, where you can run any workload in a high availability, consistently high performance environment, and spend less than you would with other clouds.

1:27How is it faster? OCI's block storage gives you more operations per second. Cheaper? OCI costs up to 50 % less for compute, 70 % less for storage, and 80 % less for networking. Better? In test after test, OCI customers report lower latency and higher bandwidth versus other clouds. This is a cloud built for AI and all your biggest workloads. Right now, with zero commitment, try OCI for free. Head to oracle.com slash IonAI. IonAI, all run together, E-Y-E-O-N-A-I. That's oracle.com slash IonAI. So it's great to be here, Craig. Thanks for inviting me. So my name is Chris Berge. I'm a Senior Vice President and General Manager of the client line of business at ARM.

2:32And that means I very much focus on all of the rich edge devices that ARM is so prominent in. So things like smartphones, but also we are obviously making quite a bit of inroads into things like PCs. We participate all parts of your house, whether that's your TVs, your smart speakers, all kind of rich end points that that you have powered and are quickly becoming AI enabled. And I think that's what we're going to talk about today, Craig. Just a little bit about my background. I've spent almost 30 years in semiconductors. various big companies started out of school at AMD and moved to Broadcom for almost a decade and also did some startups in between.

3:15So along, I can't believe it's gone by this quickly, but I guess, you know, semiconductors have never been this cool as it seems like governments and everyone really cares about semiconductors. So I guess I was very fortunate in my career choice. Yeah, that's kind of funny, isn't it? How things happen that way. Yeah, and particularly semiconductors on the edge. That's like where everything's at right now. So ARM, and we're talking about chips for AI inference. So ARM's ARM version 9 Edge AI platform has a couple of new processors designed to enable complex AI models. Can you talk about those, or is there something even more recent that you want to talk about?

4:13Well, I think that's a good starting point, Craig. And, you know, and we've actually been on quite a journey, you know, in the edge evolution, not just actually at the, sorry, in the AI evolution, not just at the edge, but also we're actually a big part of the infrastructure build out as well. And data centers that, you know, really leverages a lot of our business model and capabilities. And so we're doing quite well there as well. But I'd like to spoke probably mostly about the edge today. And so we did a Lumix platform, which is really around kind of high-end smartphones, high-end multimedia experiences at the edge.

4:54We launched that back in September. And already there's several products, chipsets that have been launched and now phones that are being launched by different leading phone manufacturers based on Lumix. And what's cool about it is it is the first, it's addition to the platforms of V9, as you mentioned, that starts rolling out SME, which is the matrix extensions to the V9 architecture. and starts pulling that into the ecosystem to allow people to really take advantage of AI at the edge in these devices and do so with the traditional CPU programming model. So obviously, there's a lot of discussions around accelerators.

5:44Accelerators are great and are quite dominant in data center and are also finding their home in the edge. But it is always this balance of accelerators have great metrics around maybe tops per watt, but they definitely are a bit more challenging to program versus, let's say, a CPU. And so you really got to get – the reality is we talk about heterogeneous computing because I think the answer is everything, right? I think, you know, in the AI world, it's about CPU capabilities. It's about GPU capabilities. It may be about dedicated accelerator capabilities. And then it's about memory bandwidth, you know, quite frankly.

6:28And really, AI is putting stress on that whole system. Okay. And let's back up a little bit for the listeners that are not deep in the chip space. and I'd like you to talk a little bit about V9 and what that is. But when you talk about heterogeneous computing and the programming language, NVIDIA, one of the reasons it has such a strong position is very early on it developed a very user-friendly programming language called CUDA and and everyone has adopted that for programming AI on to GPU processors and that has become the standard and so a lot of new chips are arriving but you need to learn a new programming language and that's a barrier but you in this heterogeneous particularly at the edge edge uh what i understand is that you still use a traditional processor with a programming language that you're familiar with and then there's some sort of a conversion that sends some of those workloads to the edge uh chip uh and it could all be packaged together but is is that right yeah i think so you you you got a lot wrapped up into there so let me just first start with v9 so um The ARM architecture has existed now for almost 30 years.

8:11And that started from early investments from companies like Apple, early adopters of the ARM architecture that made it the stalwart that it is. where companies like Nintendo, companies like Nokia back in the 80s and 90s as kind of the smartphone revolution, all that kind of stuff really started around ARM. And that's what's driven us to be where we are today. So V9 is an architecture that we launched about five or six years ago now, which was the next generation of architecture, which was focused on a couple different things. One was security, another was performance, and third was AI. And, you know, we are very much at the forefront of looking at what these next generation systems are going to require.

9:02And so that's really what V9 is about. And at this point in time, you know, a large percentage of both iOS and Android handsets ship with V9 CPUs. And you'll see that continue to accelerate over the next couple of years. And so and now V9 also is getting perforated across the other markets that ARM participates in, whether that's data center, whether that's automotive and then kind of AI, IoT as well. So that's that's that's V9. So let's talk about the programming comment that you made there. So you're right. CUDA is an amazing language. I've actually had the pleasure of working very closely with Ian Buck, who is the, I guess, you know, who basically was the Stanford student that really kind of started this and is obviously now, you know, has been a very important part of the NVIDIA story and where they're at today and still is there today.

10:02You know, but and it is a great program for, you know, basically programming GPUs, especially taking care, taking advantage of some of the capabilities in the accelerators there. But, you know, it is a accelerator language, right, versus a CPU language. And, you know, what people do is that, you know, you basically have to start making things like driver calls and you have to move your workload off of the CPU to that. Now, it makes a ton of sense if you're going to get, you know, I think, you know, Ian and I actually used to work on quite a bit of high performance HPC stuff for high performance computing, big government labs.

10:40And we would have these kind of rules where it would almost want to have 100x uplift when you kind of move that workload over. Maybe at least you'd have to have 10x, but you'd really want to have that versus the convenience and keeping it on the CPU because you really need to load that and then you stream that workload on an accelerator. And so that's kind of as you think about this heterogeneity of you're really moving these workloads around based upon, you know, what's required. Is it latency? Is it, you know, is it high performance? Is it lower power? All those kinds of things. And so this is, you know, what I would say the ecosystem is doing.

11:17What ARM's focused on is we're really focused on making all of those pieces, the CPU be as performing as possible on AI workloads and be super developer friendly, which it is. Also moving some of those workloads to GPUs. ARM is actually, many people don't know this, but ARM is actually the highest volume GPU out there. We've shipped over 9 billion GPU cores because of the mobile handsets and so much of the mobile industry that's based on ARM GPUs. And then also in support of accelerators through ARM's extension of our architecture. And so many of those accelerators hang off of things like CHI buses that are kind of part of the ARM architecture.

12:00Yeah. And when you say the ARM ships GPUs, where does the NPU fit into all that? Yeah. So in a mobile handset, let's talk about that since that's when we started. Traditionally, there was two large computing elements or there was the CPU and the GPU, right? And their main function, CPU, was to run the OS and then eventually apps. The GPU function was to display and play games and do all those kinds of things. Well, as these other workloads become important, whether it was, you know, camera imaging has obviously become super important and the amount of video that you take. Well, then we start putting little accelerators that are maybe doing some of the compression for the different video codecs and those kinds of things.

12:56And so MPUs have kind of evolved as a way to efficiently do matrix multiplications or CNNs and different kinds of models as an accelerator. And so the system can choose to send that. But what really happens in the real world, and this is the heterogeneous, is you actually end up kicking off the job, usually on the CPU. You may actually run it to the GPU. You may send it to the NPU. And then it comes back to the CPU usually to kind of conclude. And so that tends to be how these workloads actually work in a system. It's all obviously not seen by the user, but it is the developer has to make some of those decisions as a software developers thinking about, one, what kind of capability is in the handset.

13:44They're also making the decision of do they want to do it on the cloud or the edge, and we should have that conversation as well today. So these workloads move around, and there's a whole set of reasons that they move around, but we're making the CPU at the edge be as AI-friendly as AI-performant and power with a good power envelope as possible. Yeah, and are these all packaged together, the CPU, GPU, and NPU? if you're using the NPU? Or are they, you know? It depends on the system. In today's cell phone chips, yes, they are largely all in a single SOC. So all these computing elements exist in SOC.

14:26As you get into, you know, things that have, let's say, larger, you know, larger battery windows, you know, in PCs, many of the NPUs are integrated in the SOC. We are seeing trends around people adding accelerators. Obviously, people are also using the GPU that's in a laptop or something like that. And again, that may be integrated. It may be discrete. I would say there's a trend towards integration in many of these markets. And that is because of the memory pressure and the memory size requirements for AI. And so I think if you look at some of the latest, I'm going to say put it on your desk, kind of computing platforms.

15:12We're especially proud of our partnership with NVIDIA on the GB10, the product they've actually just started shipping last week. It's puts, you know, I think Jensen likes to say it puts a supercomputer on your desk, right, where they're actually providing, I think it's a petaflop of AI performance on your desk. And it is because of, You know, you've got these ARM CPUs, again, you know, 20 of them coupled tightly to this accelerator with a single memory system that offers up to 128 gigabytes of DRAM and very high bandwidth. And so you've got this whole system that you can put together or that developers can use.

15:55And we're seeing this as a trend, whether you look at, you know, the latest V9 M5 that Apple just announced this week as well, or I think it was last week. Again, you see amazing AI performance, leveraging the ARM CPUs, GPUs and other accelerators, and then a tremendous memory system for that. So that's really what we're seeing happening. And those are integrated, but some of them, as they get a little bit bigger, they get discrete. The problem with discrete is then you have to split the memory system and it gets it. Yeah. AI is so memory heavy that it just it's a whole balancing act. Yeah. And again, I'm mindful of listeners who are not deep in the chip space.

16:39SOC is system on a chip. And that's where you combine different kinds of chips into one. From the consumer's point of view, it's one looks like one chip, right? It's all compressed in there with a cover. and the reason this is important is that increasingly, I mean, right now I have an app that I built for myself. When I drive around, it talks to me about the history of the places I'm in and the voice drops out a lot of time because it's got to, you know, read my GPS and send the data to the cloud and then, you know, do a lookup in the model of whatever model I'm using, you know, get the text from the model, convert it to speech, send it back to the phone.

17:43Maybe the text of speech is on the phone. I'm not sure. But that whole chain gets, there's a million ways that it can get interrupted. And if you have that all happening on device, you don't have those problems with connectivity and things like that. But can you talk about how you see, you know, V9 or these SOCs changing the way we interact with AI? Yeah, absolutely, Craig. And you're at the forefront here. It, you know, it's my job, super fun because I get to talk to many of the industry leaders and visionaries, whether that's both on, you know, companies like Google that we work so closely with and all the chipset companies and then all the the CE companies, the consumer electronics companies that are actually putting these things in your hand.

18:45And, you know, it's fun because, you know, I always I'm an I'm an edge guy. I think, you know, I want edge computing. That's kind of that's the business I run. And so, you know, of course, I want edge computing to happen. And so but I do often go to these partners and say, hey, you know, OK, why can't you run it in the cloud? You know, why can't we do it? You know, it seems to work just fine. People love ChatGPT. They love Gemini. You know, they're starting to really utilize these services. And, you know, they basically said, you know, what gives me reassurance is all these companies are like, no, no, we need to put these in devices.

19:22And one of the reasons is what you just said, Craig, which is, you know, the goal of AI is that, you know, we're seeing some of the early cool use cases, real time translation, you know, some smart, some agents that are starting to do some things. But we're in the early, early innings here. And one of the analogies that I like to use is touch. and if you take a child, let's say less than 10 years old, and you give them a screen, they just start touching it because they don't know anything that's not a touchscreen. And, you know, whereas, you know, you and I, you know, like the mouse was a big thing.

20:08Once we thought that was cool. But, and I use that analogy because that is how AI is going to be. If something doesn't have AI and you can't interact with it and it can't start figuring out what you're trying to do, it's going to be like that child that says, this thing's not going to have a touchscreen. I don't know how to use it. I don't want to use it. Right. And so so first off, you need to believe that that is that is how essential this is going to be. It's it's the you know, it's how annoying it is when it takes a long time for the app to boot up. And it's how annoying it is if you have an app die or, you know, those kinds of things that we work out the edges and they don't really happen as much anymore.

20:49But it's technology is new. And so it's the same thing with AI where, you know, in talking to one of these partners, you know, they said, hey, Chris, we, you know, OK, yes, maybe the AI in the cloud works great, you know, 90 percent of the time. time, but you're driving up 101. I'm here in Silicon Valley in San Jose. You're driving up the highway, there's a dead spot. And basically, you're going to get a bunch of latency that it's not going to be conversational. You're going to be waiting. And the reality is that people aren't going to complain about to their carrier saying, hey, you got this dead spot on 101.

21:30When are you going to fix it? They're going to say, hey, Craig's app, I don't like that experience. It's frustrating. It doesn't work three times a day. I'm trying to use it. It's like doing these video calls, right? When it doesn't take long for somebody to have a shaky connection, you're kind of like, hey, let's just talk next week or let's talk here in the office, right? It's just too frustrating. That's just one example. There's privacy, there's performance, there's all kinds of other things that really is going to drive this to the edge. The counterside is the models have to get smaller. There is a real cost because of the computing and the memory computing pressure it puts on it.

22:11And so it's a balance, but it is happening. And again, training, and there'll still be quite a bit of inference that will happen in the cloud. But as much as we can move to the edge, I think it's pretty unilateral, whether you're a hyperscaler, whether you're an app developer, whether you're a device manufacturer. all are quite incentivized to try to make it run very very well on the edge in the device but there are challenges and there are trade-offs as you said uh one is uh that uh there's a power issue uh and uh a heat issue as a result of that uh how do you manage that because and that was i mentioned this company brain chip i had a conversation with that's using neuromorphic chips that only uh fire when the when there's enough activity to wake them up so for example in a doorbell camera you don't want the doorbell absorbing the the the video sending it to the cloud computing and sending it back when there's nothing happening in the scene, but that's in fact what happens.

23:32How do you manage that? Yeah, so I think that, I mean, we have these management techniques that get used in, you know, in almost every aspect of a computing device, right? I mean, we get, we have very clever engineers who figure out how to kind of fool you that the device is running full power yet we're obviously very aggressively changing things in the background. And that was actually one of the ARM innovations over 10 years ago. We came up with this big little concept around we have big CPUs and little CPUs, and we're actually moving the workloads back and forth because certain times you need the performance, certain times you don't.

24:11And so that's really the way these devices work, where to your doorbell example, you're looking for maybe some motion, you're looking for something. And then once you trigger that event, okay, now let's fire up some more computing elements. Let's go figure out, oh, it's a face. Okay. Is that a face we know? If we don't, let's trigger the cloud or whatever. And that's the way these systems work. So, you know, I think it is about really these more intelligent computing platforms and getting smarter on how you do these accelerators. And that's, you know, that's, for example, are the SME I mentioned that's part of this SME2 that's part of V9.

24:51It is this Kitely coupled accelerator, but we do it in a very, very efficient manner where it's actually shared across multiple CPU cores. So you're also being very area efficient, which really drives cost. So, I mean, those are things, but the other thing is there's tons of innovation. I, again, don't want to get too technical, but your listeners maybe have heard about HBM, right? HBM memory is stacked DRAM with a very high bandwidth interface. And it's become a huge enabler for AI in the cloud, right? Because we have to have all this memory bandwidth and we put these HBM stacks next to these accelerators.

25:33we're looking at very similar things at the edge relative to, you know, if where is the power, right? Is it in the computing element? Is it in the joules per bit of memory transfer? All of these things are opportunities for innovation. And there's a ton of research and a ton of investments going into them. So I don't worry about the power. I actually think the power is fairly manageable relative to inference. I think the bigger pressure we have right now is actually memory, memory size. So the model size and how big, how do we make it small enough? Because at the end of the day, that drives, it drives costs, it drives power and other kinds of things.

26:19So it's happening. And the good news is we're in this innovation cycle where, you know, there's two things you can do. You can shrink the model and get the same performance. And that seems to be shrinking, you know, almost, you know, 50 % a year, if not faster than that. And on the other side, you can say, hey, I want a three gig, you know, gig, 3 billion parameter model. And it's just getting more and more intelligent every six months, every year, right? And so there's different ways to solve the problem, but this is, that's really the enabler, I think. Yeah. And again, i'm just thinking about listeners so that so we don't lose them sme is scalable matrix extension and that's a way to optimize the the math operation correct that's used and are you you know on the models uh are you guys exploring state space models uh that optimize memory usage?

27:20I mean, that have a different way of handling memory? So I'm not as familiar personally with State Spade. I think I understand the general concept of it. I would say that, you know, we provide this platform that many of these innovations sit on top. So, you know, many of the, what we've basically done with, for example, our our matrix extension engine that you just talked about uh we basically built what we call clidy framework on top of it which is these libraries that now developers can basically leverage so it just does the right thing in the hardware relative to taking advantage of which you know do you have the latest rmv9 etc so a lot of those how the model works and states states and and what you know are they using a kbcash or how are they doing the different updates those actually sit at a a little bit higher level in the stack where we're kind of more of a, I guess a plumbing enabler is kind of the way that I would say it.

28:24So we support all of that and we're making it super easy for developers to explore that and to build their innovations on top. But there's nothing unique, at least at this point in time that we're actually doing from a state state point of view in our CPUs. And can you talk about some of the current applications and where you see this going, this move to the edge. I mean, obviously, you know, self-driving cars is one where you can't afford to risk connectivity and latency by sending data to the cloud. It's got to be computed on the platform or in the car. uh what are some of the other applications that you're currently uh putting chips into and or current devices i should say and and where do you see that going i mean how another the form factor of these chips how small can they be uh so that they're they're sitting i mean when i was talking to a guy the i wear hearing aids uh that uh you know you'll be able to have these chips in the hearing aid, you know, filtering or isolating sound and that sort of thing.

29:52So can you talk about that a little bit? Yeah, I mean, that's a great question. I mean, I think the scalability is almost everywhere, right? I mean, your hearing aid example is a good example. And this is one of the reasons why ARM builds the huge set of portfolio that we do, because we are in many of those hearing aids and those kinds of things. And we've actually announced AI across, I mentioned Lumix, but we have our Edge AI platforms that we've also announced back in February. So we are enabling that. And it just comes down to if you know the task you are trying to do, you can make things quite small and efficient, right?

30:37So in a hearing aid, right? In a hearing aid, you're probably not trying to run a large language model, not yet at least. You're doing things like trying to reduce noise or picking out, amplifying only the interesting, what are you trying to hear? Now, obviously adding translation and those kinds of things are probably going to be possible soon. Now, if I know what language I'm translating to, then that's going to make my model smaller. but yeah everything you know one of the things that I like to think about is how much you know we use apps today to configure things like a good example is like a security camera right so today you probably you know when you try to install your security camera in the past you would have went to your laptop and connected it now probably use your smartphone to do that well you know that is a use model we've gotten used to it but but why can't you just do it with your voice?

31:35Like, why can't it talk to you and have the camera have a large language model? And you say, hey, it asks for what is your SSID? Here are the ones I see. And, you know, now again, that may or may not be a better user experience, but we just see the applicability across the board. Again, I use that touch analogy because touch has become so prevalent. AI is going to be way more prevalent than even touches today in changing the use model, how you interact with these things. You know, a great example I'll give you, Craig, is we work very closely with Meta, who is really doing a great job, I think, with some of the advances in their glasses, right?

32:17So the XR glasses and the product they just announced two weeks ago, three weeks ago, they now have a wristband that goes with that product. And that actually uses one of our ethos NPUs, which is, again, super small, super low power that you can have this wristband that basically, you know, has a huge battery life. I forget how many days or weeks you can wear it. And it literally just, you know, you manipulate things just by moving your fingers like this, and it senses in your wrist the changes in what's happening, you know, below your skin using AI to basically figure out that, oh, you just did your second finger, so that's going to do this, or you just did this, so that's, or, you know, you're doing this.

33:06So basically, that is the new UI of a wristband that has AI in it that has, you know, a long, long battery life and has a tiny, tiny battery. That's how small we can make AI for very specific use cases. yeah uh and and what's the focus uh with with arm right now is it a reducing size i mean obviously it's going to be all these things but but where do you see the next breakthrough is it reducing size, reducing power consumption, increasing the model size that can be on these chips. Yeah. Where do you see it going? It's a great question, Craig. And it is a very open design space right now, to be honest with you.

Read the full transcript

34:03Definitely, there has been a focus on power in the past. And that's kind of where our ethos product line comes from. But a lot of that was around CNN networks and some of the early stuff. Now you're starting to see the move to transformer networks. What's next, right? I don't think that transformers are actually the end goal. And so we also need to allow flexibility as these models change because, you know, back to kind of educating your users. I mean, it takes two years almost at a minimum to design silicon and get it in a shipping product and you can see how fast AI is moving. So there's also this flexibility piece, but I think most of it is performance.

34:46It's really memory performance and it's tops and CPU performance and trying to shrink that down as the best we can for whatever the power envelope is to how much computing you can get. And I think that's obviously happened in the data center as well, right? Where it's okay, we're trying to get, we're now building gigawatt data centers, but how much, how, you know, how many tokens can you create? How much performance can you get in that? So almost everywhere in the space, it's like, okay, tell me what your power envelope is. It's a gigawatt here. It's, you know, seven, six Watts in a phone. It's, you know, in my wrist example, it's literally, you know, a couple hundred milliwatts you know how much ai performance can you get and so it it really is spanning a large um space yeah and you guys don't uh you you're producing ip right you're not uh building uh you don't have a fab you're not uh sending your designs out to a fab to make your chips, you're licensing this to other chip makers.

36:03That's correct. Yes. Our business model is we provide IP to a large portion of the semiconductor ecosystem. And then our partners build on top of it and create silicon solutions that can span, like we talked about, anything from a wristband or your hearing aid to you know, your next self-driving car. Yeah. Yeah. And, and where is most of this? I mean, obviously most of it's going into the U S but do you have partners that are, are licensing your IP? Where, what are the other markets that you're operating in? Well, you know, I would say that traditionally semiconductors have been quite global, right?

36:57You know, that the cost of developing a semiconductor is tremendous, right? Obviously, the foundry costs, and I'm not even saying like put aside the foundry costs, but literally building a chip, you know, if your example, if you're using one of the latest nodes, just the mass costs alone of when you're done, your design and getting that manufacturer is in the tens of millions of dollars. The development costs up front can often be in the$100 million plus range. So you're talking about tremendous amount of cost to get to the first unit. But what's made semiconductor so great is that you can scale that.

37:39And because we have these amazing manufacturing capabilities, et cetera. And so they tend to be very large markets. And so generally it's been a global market for us. And yes, obviously, there's many great IC or semiconductor companies in the US that do designs. We work very closely with European semiconductor companies in Taiwan, in Korea as well. And we also have had an ARM China, a JV, where semiconductors get built in China based on RIP. And so it's a pretty global market. Yeah. I would imagine the Chinese relationship has been strained by the sanctions. How does that affect you guys? Well, yeah, I mean, we always make sure we're following the local laws and obviously where we're located.

38:37Most of the constraints have been around the manufacturing process and not necessarily some of the IP. There is some around IP enablement, but, you know, it's something that we always are very careful about. Make sure we're complying to what is required and we'll see as things evolve. Yeah, yeah.

39:02And, okay, what am I missing here? We're closing in on your hour. is there something that I haven't talked about that you guys want people to hear? Melissa, anything tops to your mind? I don't know if you're still on. I'll put Melissa up here. Okay, sorry.

39:27Melissa, anything that comes to mind?

39:47okay yeah and and and what again you know for for the less technical savvy technically savvy uh listeners i'd like to talk about uh so this is uh ip for chip design uh it's sort of looking forward where you know i'm not sure how to even ask the question whether it's you know you mentioned the wearables uh you know what what should the public be imagining is we're going to see as a result of of these ai enabled chips uh but i'll i'll let you uh okay and i'll cut all of this stuff out okay that's fine okay you want to ask a question then start as a question or uh yeah so uh uh and i forgot melissa what was your your point

40:58right right right right yeah so so in the that you guys are uh are primarily selling ip uh do you work with a large developer community and what's that relationship like uh is how much of this is open source for example and uh and then if you could talk about where you see i mean you mentioned wearables where wearables are going to be a hot thing where you see the edge really exploding the public uh interaction or experience with ai okay yeah so i mean let's let's start about developers so you know arm is um you know developers are super important to arm and in fact we now have uh the world's largest developer ecosystem we believe we have over 22 million software developers that are developing software on arm um and that's just because of our our footprint um that spans all the way from you know ios to now windows on arm and then of course things like Chrome, Android, and Linux.

42:12You know, we've been a longtime supporter of. So, you know, really the world's largest software ecosystems are quite, you know, work closely with ARM and we support them. And we've really focused on that developer experience and how do we improve that developer experience. I mentioned Clyde before in this conversation. that's really about, again, trying to make it seamless for developers to use AI. Because, you know, you talked about CUDA. AI is not simple from many of the programming elements. And I would say we're still seeing not necessarily the hardware abstraction that we've gotten other software ecosystems to.

42:57We still see a fairly tight coupling to today's models to actually the hardware that they run on, whether that's the operators they support, whether that's, you know, the way they kind of assume the model is going to propagate and those kinds of things. So, you know, but we are we're clearly focused on supporting those developers and and seeing what they can build. Right. And I think that kind of goes to your second question of, you know, what's kind of on the forefront. And that's why it would makes my job so fun because I get to interact with so many of these super innovative folks, both startups and established companies on kind of what is next or what's the art of possible.

43:41And if I was just to use a general concept, it would just be intelligence. The idea of something is capable of interacting with you in a way that exceeds your current expectations. Right. And maybe I'll use another product example. I'm a huge Amazon fan. But one of the things I've noticed is that before I upgraded to Alexa Plus, you know, because of the way I was using a chat bot or chat GPT or Gemini, I started talking to it in a more conversational manner because that's what I could do. But then when I would come home and use Alexa, it wouldn't necessarily, it was expecting more of kind of a search kind of give me, you know, really break it down.

44:31I'm going to give you three words and try to contextualize those three words. You know, and so I think, by the way, Amazon's done a great job with Alexa Plus. I think they're on the right path here. But that's just a change in expectations. And I think that's going to only go, you know, just through the roof relative to, you know, the idea of, you know, the things that we think are how you interact with a device for it just to be smart. You know, one of the things people ask me about is like, what's agentic AI or how's that going to feel? Right. And one of the things I like to use, and I think I think Microsoft has done an amazing job of really integrating AI into their products as large as a company they are, you know, the way they push copilot, all those kind of things.

45:18But what I like is a is a silly example of settings. So when you use Windows and Windows settings, we've all done it and you have to figure out, OK, I want to add a screen or I don't like this resolution screen. How many like windows do you have to click in to kind of say, figure out how to make this change, right? Well, they've changed it to be more of an agentic experience in Windows 11 where you just say, what do you want to do? I want to add a screen. I want to reduce this. And it just says, OK, well, here's how you do it or I'll do that for you. And so, you know, to me, that's a touchy feely thing of how many times do we, you know, click through three apps.

46:02oh, I'm going to cut this. Now I have to go to my flight app. Then I have to go to my rental car app. And I like, you know, so many of these things that we're just trained to think like, oh, yeah, it works good enough. I, you know, but like this idea of things are just so smart that, you know, when you call an Uber, it of course knows where you're going to go. And it's already told people when you're going to show up and all these kinds of things. So, you know, those are, and of course, those are consumery examples. There's huge amount of enterprise examples of, teams that are not sharing data or people don't understand or they're looking for information they can't find.

46:36And so I think it's just going to be the general idea of just how do we make things just super smart and really kind of exceed expectations of what technology limitations are today. Yeah. Yeah. And robotics is another one. I'm not a great believer in humanoid robotics. I uh what you're describing i mean it's um you know the day will come when uh our children our grandchildren will will look back and can you believe that you know any object that you use daily uh was a dumb object back then instead of being able to talk to it and and have it configure itself or whatever it is. And that's pretty exciting.

47:30What about robotics? I mean, obviously, you guys are involved in that. Yeah, I mean, it's pretty amazing what's happening there and the ability to train the robots in virtual worlds as well as real worlds. And we are a huge believer in that. And we're quite involved with many of the ecosystems. systems, you know, I think it is really taking the advantage of, you know, the camera sensor has been an amazing enabler, right? If you just look at what our cell phones are able to do in capturing images, but really the camera sensor has now unleashed this idea of physical AI and how do you, you know, because we can use vision systems and AI is so good with vision systems, Right.

48:17Whether that's looking at an MRI and or, you know, maybe in augmenting a doctor's reading of that or in a system and trying to understand what's going on. So it is it is a huge area. I think we're going to see a lot of innovation. I think, you know, to your robotics example, you know, I think I don't know, go back to the Jetsons or whatever. I mean, people said they're going to have flying cars and they were going to be able to do all these autonomy things. it's taken a long time, but, you know, another great arm powered device, my Tesla, you know, it is gotten pretty darn good relative to self-driving.

48:51And, and, and I really enjoy having that augment me in, in many, many scenarios. And so it is, like you said, it's going to just be an expectation and you're just going to be disappointed when it doesn't do something that it clearly can't do today. But these things are hard. I mean, you know, I think that that blend of physical um and and you know is going to be a hard thing but like anything it's probably you know i go back to i was involved with bluetooth in the early days and and you know oh it's going to be amazing you have all these bluetooth devices and just and and it it didn't work great in the beginning the experience wasn't great but just look at how fundamental bluetooth is today again when i walk up to my tesla it just unlocks it just like it locks and it's all over Bluetooth and augmented by UWB.

49:41Or, you know, so it's just these technologies, you know, sometimes the hype cycle gets a little bit ahead of expectations and that's our fault as technologists. But at the end of the day, the technology delivers and it's gonna be an exciting future for our children and grandchildren. Yeah, absolutely. And yeah, with moving into physical, the challenges aren't really the AI, It's the mechanics, the actuators, and the wear and tear, and the dust, and the grease, and all that stuff that people kind of forget about. It's still got a long ways to go. Yeah, okay. Okay, so is ARM on the verge of announcing anything new, or is it really the V9 that you're focused on?

50:43We've got quite a broad set of products, right? So, yes, I have my responsibilities, but we have a whole automotive and robotics group. We have a whole group that focuses on data center. Same thing with Edge AI and additional intelligence. So I think you'll keep seeing quite a bit from us. Lots of great stuff going on. And again, we're just super excited about the partnerships that we get to partner with and how much developers and hopefully some of your listeners are inspired to build things on top of the ARM architecture and help make the future show us what's possible. Yeah. And on that note, we'll end on that.

51:26Developers, if they're interested, do you have a portal, developer portal? We do. Developer.arm.com. Great, great, great resource to to find out more, figure out how to, you know, build and and and explore many of these new technologies. Obviously, many of the makers are familiar with Raspberry Pi. That is also arm-powered and something that many kids and educators start their journey. And those things are getting super excited now with robotics and beyond. In business, they say you can have better, cheaper, or faster. But you only get to pick two. What if you could have all three at the same time?

52:13That's exactly what Cohare, Thomson Reuters, and Specialized Bikes have, since they upgraded to the next generation of the cloud, Oracle Cloud Infrastructure. OCI is the blazing fast platform for your infrastructure, database, application development, and AI needs, where you can run any workload in a high availability, consistently high performance environment, and spend less than you would with other clouds. How is it faster? OCI's block storage gives you more operations per second. Cheaper? OCI costs up to 50 % less for compute, 70 % less for storage, and 80 % less for networking. Better? In test after test, OCI customers report lower latency and higher bandwidth versus other clouds.

53:15This is a cloud built for AI and all your biggest workloads. Right now, with zero commitment, try OCI for free. Head to oracle.com slash IonAI. IonAI, all run together, E-Y-E-O-N-A-I. That's oracle.com slash IonAI.

From the publisher

Try OCI for free at http://oracle.com/eyeonai 

This episode is sponsored by Oracle. OCI is the next-generation cloud designed for every workload – where you can run any application, including any AI projects, faster and more securely for less. On average, OCI costs 50% less for compute, 70% less for storage, and 80% less for networking. 

Join Modal, Skydance Animation, and today's innovative AI tech companies who upgraded to OCI…and saved.


Why is AI moving from the cloud to our devices, and what makes on device intelligence finally practical at scale?

In this episode of Eye on AI, host Craig Smith speaks with Christopher Bergey, Executive Vice President of Arm's Edge AI Business Unit, about how edge AI is reshaping computing across smartphones, PCs, wearables, cars, and everyday devices.

We explore how ARM v9 enables AI inference at the edge, why heterogeneous computing across CPUs, GPUs, and NPUs matters, and how developers can balance performance, power, memory, and latency. Learn why memory bandwidth has become the biggest bottleneck for AI, how ARM approaches scalable matrix extensions, and what trade offs exist between accelerators and traditional CPU based AI workloads.

You will also hear real world examples of edge AI in action, from smart cameras and hearing aids to XR devices, robotics, and in car systems. The conversation looks ahead to a future where intelligence is embedded into everything you use, where AI becomes the default interface, and why reliable, low latency, on device AI is essential for creating experiences users actually trust.


Stay Updated:
Craig Smith on X: https://x.com/craigss    
Eye on A.I. on X: https://x.com/EyeOn_AI 

More from Eye On A.I.

All 266 episodes
#308 Christopher Bergey: How ARM Enables AI to Run Directly on DevicesEye On A.I. · 54 min
Listen in VO