In short
Edge AI (“AI at the edge”) in 2026: what “edge” means (not in the cloud), why it’s constrained (size/power/connectivity/cost/reliability/latency/privacy), and how modern AI architectures use smaller models and cascades/pipelines to act in real time. It also covers edge vs physical AI, tooling/ML ops for distributed devices, and how to start experimenting.
Guests
Brandon Shibley, Edge AI Solutions Engineering Lead at Edge Impulse (a Qualcomm company). Hosts: Daniel Leitnack (CEO, Prediction Guard) and Chris Benson (principal AI and autonomy research engineer).
Key claims
Smaller LLM/SLM models (single-digit to tens of billions of parameters) can run on edge hardware with NPUs/GPUs; large models stay in data centers. Edge systems often combine lean models in cascades (e.g., detector → VLM → optional RAG → LLM) to save power/latency. Edge ML needs continuous updates due to drift; devices are centrally managed when connected (OTA updates).
Notable examples
YOLO-style object detection to drop 99% of frames; cropping detections for VLM analysis; vehicle detection then license-plate detection; RAG from documentation; maker projects like leak detectors and cat-feeder detection using Arduino.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOMeet the Hosts and Guest
0:45 to 2:12
Daniel and Chris introduce the episode's guest, Brandon Shibley, and discuss the topic of edge AI.
“I am CEO at Prediction Guard, and I'm joined as always by my co-host, Chris Benson, who is a principal AI and autonomy research engineer.”
Understanding Edge AI in 2026
2:12 to 4:23
Brandon discusses the current state of edge AI and its evolving definition.
“maybe ways, if there are different ways in which AI is being applied at the edge, then it has maybe traditionally been applied in previous years or eras, if you will.”
Trends in AI Models
4:23 to 8:20
The conversation shifts to the evolution of AI models, particularly generative models and their implications at the edge.
“companies on is really understanding what it means to achieve a positive outcome for them.”
Characteristics of Edge Operating Environments
8:20 to 11:42
Brandon explains the unique constraints and characteristics of operating AI at the edge compared to the cloud.
“Because I think most people probably listening out there that have done stuff have been operating cloud environments instead of this.”
Understanding Physical AI
11:42 to 13:11
The discussion covers the concept of physical AI and its relationship with edge AI.
“And how does this, we've talked about sort of real people using this technology in the kind of physical real world, maybe not at their computer screen.”
The Importance of Latency in AI
13:11 to 14:00
Brandon elaborates on the criticality of latency in different applications and how it affects the use of AI at the edge.
“sensing and making predictions of the data that's out there in the real world.”
Understanding Real-Time Performance at the Edge
14:00 to 15:01
Learn how latency requirements vary by application in edge computing.
“connectivity, and then maybe tie that to what might need to be run at the edge in order to not operate at that model or in that kind of API endpoint model.”
Cascading Models in Edge AI
15:01 to 18:04
Discover the significance of cascading models for efficient data processing at the edge.
“And so it all comes down to what is the requirement for the type of behavior we're trying to get out of the system.”
Tooling Advancements for Edge AI
18:04 to 21:46
Explore the latest tools and frameworks that enhance edge AI development.
“So what we'll do in many cases is we have this pipeline or cascade where on the front end is some kind of very initial detection that can be done very efficiently.”
Workflow Shifts and Agency in Edge AI
21:46 to 24:16
Understand the transition in workflows and the concept of agency in edge AI.
“So it is a way of being able to easily work with data, train models, tune and optimize them for target devices, and then generate a deployment that's easy to run on a device.”
Show all 17 chapters
Governance and Management of Edge Devices
24:16 to 28:06
Learn about best practices for managing and governing distributed edge devices.
“Yeah, in a lot of ways, machine learning is math and it's statistics.”
The Evolution of Smaller AI Models
28:06 to 29:14
Learn about the advancements in smaller AI models and their effectiveness at the Edge.
“We can now apply also to the models that we're deploying to the Edge.”
Edge Impulse's Unique Approach
29:14 to 32:45
Discover Edge Impulse's strategies for optimizing Edge AI workflows.
“I think like the what's happening with the state of the art is kind of overshadowing some of the, you know, also advancements that are happening at the at the edge and with small models.”
Hardware Challenges and Innovations at the Edge
32:45 to 37:45
Explore how advancements in hardware are enabling new capabilities for Edge AI devices.
“So Edge Impulse is the leading edge AI platform.”
Getting Started with Edge AI Projects
37:45 to 42:05
Learn how to initiate your own Edge AI projects using accessible tools and platforms.
“And, you know, there's always an opportunity to do more.”
The Future of Edge Computing
42:18 to 45:41
Explore the exciting possibilities of edge computing and AI advancements.
“I guess as we are winding up and we're kind of looking at the future, you know, you have Edge Impulse and Qualcomm and the kinds of work you're doing.”
Encouragement to Experiment
45:41 to 46:15
Get inspired to experiment with AI technologies in your own projects.
“Yeah, well, I'm certainly excited to see some of those things.”
Transcript
Automatic transcript. May contain errors.0:02Welcome to the Practical AI Podcast, where we break down the real-world applications of artificial intelligence and how it's shaping the way we live, work, and create. Our goal is to help make AI technology practical, productive, and accessible to everyone. Whether you're a developer, business leader, or just curious about the tech behind the buzz, you're in the right place. Be sure to connect with us on LinkedIn, X, or Blue Sky to stay up to date with episode drops, behind-the-scenes content, and AI insights. You can learn more at practicalai.fm. Now, on to the show.
0:41Welcome to another episode of the Practical AI Podcast. This is Daniel Leitnack. I am CEO at Prediction Guard, and I'm joined as always by my co-host, Chris Benson, who is a principal AI and autonomy research engineer. How are you doing, Chris? Hey, doing great. You know, looking forward to another show. And like always, we're getting really edgy out there in the AI topics, aren't you? I'm definitely on the edge of my seat for this discussion. I've been thinking about it a lot because today we have with us Brandon Shibley, who who is the Edge AI Solutions Engineering Lead at Edge Impulse, which is a Qualcomm company.
1:22Welcome, Brandon. How are you doing? Doing great. It's an honor to be here. I've been a fan of the podcast, so it's great to join. Oh, that's great to hear. It's always a good connection to make. Thanks for putting up with our terrible puns here as we start the show off. We're famous for terrible puns. Yeah. I'm here for it. Nice, nice. Yes. Well, it's been a while since we've had a full episode talking about edge AI or AI at the edge or machine learning at the edge or however, whatever combination of things you want to make. I'm wondering if you could just give us a little bit of an update or a kind of state of edge AI or AI at the edge in 2026, maybe highlighting first, what does the edge mean in 2026?
2:11And then maybe ways, if there are different ways in which AI is being applied at the edge, then it has maybe traditionally been applied in previous years or eras, if you will. Sure. So allow me to start with the definition of edge. I take a pretty broad view of the edge. And practically speaking, in my mind, it's anything that is not in the cloud. Depending on who you ask, they have far more specific definitions and we get into like far edge, near edge, edge of network and all of these things. in my world we can we deal with all of it so you know edge just means we're taking ai we're going to embed it somewhere that's it's not in a data center not in the cloud but usually close to the real world where real data is captured and where the sensors are to tech you know where is it going at the edge.
3:13The good news is with everything that's going on with AI, we're seeing a lot of innovation around silicon. That's enabling us to embed models at the edge with greater efficiency, more capability. And so we're seeing that the industry is adapting to the needs of AI as we're going to bring it into the real world. And so that's very exciting. We're seeing also some other, you know, we call it pressures or trends, you know, economically speaking, tons of money has been going into AI research. At the same time, you know, the economy is also putting pressure to, you know, to achieve productive outcomes, you know, an ROI on that investment.
4:04That pressure has always been there at the edge, by the way. So I think what that means is there is a rationalization that's actually pretty healthy, ensuring that when we apply AI that, you know, it is doing something productive and ultimately achieving some kind of return on investment. So that's a lot of what I end up collaborating with companies on is really understanding what it means to achieve a positive outcome for them. And then we can discuss like the technical methods on which we're going to get there. And would you characterize, I mean, I know the last few, three or four years have been dominated by certain types of models, specifically generative AI models.
4:50And many people thinking, of course, of these large models. I mean, it's even in the name, large, large language model. and those, you know, I guess those models people might not think of as living kind of in the physical world or at the edge. Is that a fair assumption? Or I know we've seen kind of people talking about SLM, small language models. How has that shifted over time? I guess, you know, If we look back five to 10 years, the types of models that were being run maybe in disconnected environments or on a factory floor or at the edge in some sense versus now, has there also been that shift as the market has moved to kind of these Gen.AI tools?
5:45Yeah, absolutely. I mean, language models are still relatively new phenomenon, right? But they, in the last couple of years, we've seen them kind of explode in different directions. They're getting bigger in the cloud, they're getting smaller at the edge. And that's a good thing. It means that, you know, there's a broader range of possibilities to solve problems with. So we're going to have trillion plus parameter models that you can never practically fit at the edge that are going to be in a data center somewhere. And then we're going to have much smaller versions of LLMs, even SLMs that we can embed into devices.
6:22So, you know, edge devices are growing to accommodate those small to midsize LLMs. LLMs, we're talking on the order of single digit to tens of billions of parameters can be accommodated in some form of edge hardware, even edge AI appliances. These are things that have, let's say 64, 128 gigabytes of memory, for example. They have powerful NPUs or GPUs to be able to do the inference and they can be embedded into a premise or even into vehicles or things like that to accommodate these kinds of models. Now, the smaller size of those models makes them, you know, the implications of them being small mean that they don't have quite the same kind of like knowledge capacity, as we call it.
7:15It's kind of a rough term, but the idea, I mean, is they're not going to be necessarily great for retaining tons of like real world knowledge, but where they shine is where they're specialized and fine-tuned for specific specialized, you know, data. And so there, I think this is what the industry is beginning to become more effective at is achieving that. That means doing more with less essentially. And it's not just SLMs, it's all kinds of AI models where we see this. I personally work with a lot of other kinds of neural networks as well. And there, it's always been about, you know, curating data sets, training those specialized models for specialized needs.
8:03And if anything, what we're seeing now is a lot more of combining these models into really interesting, you know, cascades or ensembles of models in order to leverage sort of the best of all of them. And at the edge, we really have to remain pretty lean. So it means in many cases, what we're doing is a combination of different lean models to get exactly the characteristics we need to solve problems at the edge. I guess one of the things that I'd like to take a moment and maybe back up a little bit, and as we've kind of dived into smaller models at the edge, but maybe for listeners who have not had experience themselves at operating at the edge, could we talk a little bit about or could you kind of explain a little bit about some of the characteristics that you find at the edge that make it kind of a distinct operating environment that you're having to cater to, you know, in terms of security, latency, comms between things, but just kind of like the whole set of characteristics that makes it very distinct from the cloud environment.
9:10Because I think most people probably listening out there that have done stuff have been operating cloud environments instead of this. Absolutely. I mean, this is a key point. I'm glad you brought it up because these constraints are what we have to live and die by at the edge. So what are those constraints? Size, power, connectivity, which may or may not be there or be reliable. We're dealing also with cost constraints at the edge. As I mentioned, many of these products have to be sold into very cost-sensitive markets. That plays a huge factor. Reliability may be key. Latency in the case where we're dealing with, let's say, robots or anything that's got to take immediate action in the real world based on the data that it's, you know, collecting.
10:00And then also privacy. So, you know, users are going to be, you know, in many cases, we're talking about systems used by people and the kind of data that's being captured with cameras and microphones and other sensors is sensitive data that should be kept private. And so that's, you know, another element we're often dealing with. In fact, a lot of these are, you know, you can think of it almost like double-edge issues, both the challenge that you face at the edge, but it's also the opportunity. Privacy is a good example of that. Edge is an opportunity to keep that private data at the edge and not proliferate it out onto the internet and into the cloud and in places where, you know, users would prefer it not go.
10:41And so, yeah, to contrast that with the cloud, obviously we've got far fewer constraints around power, around compute. you know usually things are compute in the cloud where latency is less of an issue although it still may be important but yeah the pressure to bring things to the edge is often driven by things like latency privacy and you know there's in general the economics as well because you're thinking about where is it most efficient to do computation. You'd want to do it near the data. I mean, otherwise we need connectivity. We're going to be paying for cloud services. A lot of these systems already have some compute at the edge.
11:29And a lot of times it's underutilized, right? It's being used to do certain things. But if you've already got compute there, you can also use it to do a lot of your data computation and AI. You know, there's a lot of economical, you know, economic benefits of leveraging the compute at the edge rather than have to, you know, pay a lot to compute at scale in the cloud. And how does this, we've talked about sort of real people using this technology in the kind of physical real world, maybe not at their computer screen. I know one of the big topics kind of coming into this year that I even, I just saw LinkedIn post about is physical AI.
12:10How does that jargon kind of overlap with edge AI or relate to it? Maybe for people that are kind of trying to parse through some of the hype and the jargon. Yes, it's difficult because there is some jargon and buzzwords. And I think in some way, edge AI, physical AI can be a buzzword. But it's also referring to a real, you know, use case and phenomenon, which is that we can put AI out in the real world. And in the case of physical AI, I would say it sometimes is distinguished from edge AI in that it really relates to taking physical action in the real world. Think about robotics or self-driving vehicles.
12:54Not only are they sensing the world and making predictions about it, but then they're also translating those predictions into taking action at the edge. So I would say, you know, if there is a distinction, that's generally what the distinction is between edge and physical. but there's also a ton of overlap. Obviously, any physical AI is essentially also about sensing and making predictions of the data that's out there in the real world. And could you describe a little bit, I know there's so many people and many developers that are probably listening to this show that have primarily interacted with AI through API endpoints, you know, over the internet.
13:36Those seem fairly fast in many cases. And there might be some people thinking out there, you know, oh, well, now we have Starlink and we have these endpoints. What are you talking about with latency or these sorts of things? Could you just kind of drive home on that point and maybe with, you know, theoretical examples or something illustrating how sometimes that's not an assumption that can be made, that kind of connectivity, and then maybe tie that to what might need to be run at the edge in order to not operate at that model or in that kind of API endpoint model. Yeah, absolutely. When we talk about real-time performance, it means that we need some kind of response or output of a system within a certain timeframe.
14:32Now, what that is, it depends very much on the application. So, if we're talking about a high-speed manufacturing line, that may be on the order of microseconds. Take a self-driving car. Again, maybe it's microseconds or single-digit milliseconds. If it's a chat app where I'm chatting to an agent and I need a response, it might be on the order of many microseconds or even seconds is acceptable for latency. So the application really drives home the requirement. And so it all comes down to what is the requirement for the type of behavior we're trying to get out of the system. Based on that, we can make a decision about, you know, where should the computing be done?
15:16Where should the models live? Is it acceptable to send that data over the internet or not? Do we need to do it right at the sensor? Can we do it somewhere, maybe on premise, but somewhere else on the network? Those are the kinds of things that we can determine based on those latency requirements. And, yeah, again, there's a wide range of different possibilities. You know, the great thing, I think of AI, Edge AI, and, you know, even cloud, it's, we have many tools in the tool chest. You know, we need to kind of approach this from first principles design thinking, which is what are we trying to accomplish at the end of the day?
15:52And then that will inform us about what tools we can use to get there. You made a comment earlier as we were getting into the description of that edge environment. And you talked about, I think I can quote you as cascades of models at the edge. And as we're, I think, you know, with most people, even outside of the industry itself, just people using, you know, Gemini, ChatGBT, Claude, and they're kind of used to thinking of, I'm going to go to the AI, you know, that's the large language model that is going to solve whatever it is that I want to solve. And yet on the edge, as you've just kind of described all these characteristics that are very common that teams have to address, and, you know, you have lots of potentially different models coming into bear.
16:47And some of those are LLMs and some of them are small language. They're kind of moving from large language to small language, and some of them may have nothing to do with generative. It may be reinforcement learning in a lot of cases or other types of models that we've talked about on the show. Can you talk a little bit about the relationships of having that cascade of models to the types of actions that you need to take, you know, the sensing and the actions that you need to take on platform when you're at the edge to kind of give a sense of, you know, the different architectural thinking that goes into these edge environments that way?
17:26Yeah. Yeah. So let me start by giving you an example, which I think kind of makes clear why combining models and cascading them, or you can think of it as like a processing pipeline, is a common pattern that you'll see here. So, you know, the thing about the edge environment is we're often compute constrained. And we're also trying to minimize power in many cases, which means we don't want to just use the most powerful processing technique we have at all times. If you were using a large language model or maybe a vision language model on camera data and running it continuously on every frame that came through, it's a very quick way to burn through a lot of power.
18:11So what we'll do in many cases is we have this pipeline or cascade where on the front end is some kind of very initial detection that can be done very efficiently. So maybe it's an object detector. Maybe listeners are familiar with YOLO is a common form of an object detector. It can be used to detect objects in the frame. and maybe we throw away, you know, 99 % of the frames that ever come through based on this initial object detection. But then when we see an object that looks of interest, we can maybe use the bounding box that we've predicted around this object, crop the image out, and then cascade it into something like a VLM where we can do maybe much deeper or more dynamic analysis on the image.
18:59And it can give us like much more detailed, you know, metadata about what's there. That's an example of where these cascades are useful. And we don't just use them for image processing. They get used for audio. Sometimes we're doing multi-stage detection. Sometimes we'll do initial dissection, then detect, have other object detectors that can detect different features. So maybe you detect a vehicle. And then what you want to do is, once you've detected a vehicle, now I want to detect the license plate. Maybe I want to detect certain features of the vehicle. And then based on that information, maybe I need to perform some retrieval augmented generation.
19:37For example, I'm looking up information from a database of documentation and then combining all that information to request a response from an LLM, which will craft like a, you know, textual reply to a user, for example. So, you know, these are all the tools that we think about using when we're going to solve a real world problem and, you know, trying to get the best possible performance. I mean, balancing many different constraints and also traits that we're trying to get in the solution that we built. And I'm wondering, I'm having flashbacks to maybe, I don't know, like earlier on in my career where a lot of what I was doing was running kind of models next to data.
20:27A lot of that due to just the size of the models and how we wanted to deploy them and that sort of thing. And I remember part of the trauma of, not to people experience, obviously, flashbacks and trauma, Dan, oh boy. But I'm thinking of trying to get all the right dependencies to get TensorFlow to run with this particular model and debugging all of that kind of chain of things. what is maybe from another perspective from the developer perspective what taking a look at the kind of state of tooling around edge ai now and like the ability you mentioned the advance and hardware which we can talk on here in a second i'd love to kind of hear some of that in a little bit more detail but just in terms of the tooling what is the state of that i'm guessing things have advanced and changed maybe.
21:26But yeah, how is that advanced? How is the kind of tool set and frameworks kind of advanced to support these kinds of pipelines? Yeah, the good news is the industry has responded with options for tooling. And so, you know, I personally work for a company that builds a platform with this kind of tooling called Edge Impulse. So it is a way of being able to easily work with data, train models, tune and optimize them for target devices, and then generate a deployment that's easy to run on a device. That's the kind of state of the art in terms of simplifying this development. Of course, there are frameworks below that, things like TensorFlow, as you mentioned, and others.
22:14Those are also, you know, many machine learning developers work directly with these frameworks. But I would say the difference is, you know, people that are, and you've seen this in software forever, right? Abstraction layers. There are people that specialize at different layers of this stack. And to reach the general developer, somebody who's not necessarily an expert in TensorFlow or these frameworks, they can leverage, you know, easy to use tools. They're out there. And Edge Impulse is a great example of one that's specifically designed for the edge and the fragmented hardware ecosystem that's out there.
22:55The advantage of the cloud is really that there's kind of been largely some, what I want to say, almost unification around some common hardware, right? NVIDIA is obviously very dominant there. It means that most developers are using very similar tooling, targeting a very, you know, similar hardware target. And at the edge, things are still very fragmented. So, you know, this is where using tools like Edge Impulse really does help developers and helps them make developed models that are highly portable, can still also be optimized for the specific features of the hardware as well. As you're talking about your kind of the tooling there and and recognizing that, you know, as we've moved from cloud to edge and that maybe the workflow is a little bit different, you're trying to develop systems that are planning and executing, you know, multiple tasks with some level of autonomy, you know, and and the various support framework that has to go around that.
24:00Could you talk a little bit about what, you know, most people, you know, we're so used to hearing, you know, about inference in the cloud and stuff, and you still have that at the edge, but you hear the word agency a lot more when we get to the edge. And can you talk a little bit about kind of what that workflow shift and that objective shift is like and how the tooling impacts that? Yeah, in a lot of ways, machine learning is math and it's statistics. And at that level, it's very similar, like cloud edge, the similar concepts apply. The difference comes in generating efficient runtimes that are going to work on a processor at the edge versus, you know, a GPU and server in the cloud.
24:49and you know also there are also other differences i think are pretty important like how do you get data from the edge how do you continuously deploy newer and greater models we talk about ml ops is like a best practice here which is uh just because you've deployed a model out into the real world doesn't mean that um it's going to be you know um it's always going to be good enough uh the world changes, right? And sometimes we're also deploying things into new environments. Those models will need to be adapted and improved. And the way we do that is collect newer data from, you know, over time, we talked about this concept called drift, where the world changes for whatever reason.
25:32And so the model will perform less well in these new environments. So we'll have to get new data, train a new version of a model and redeploy it out. And so that can be challenging in the physical world. These devices, they live out in a world where connectivity may be an issue. The environments vary vastly, unlike the cloud, where you have a very uniform environment, centrally managed. It's highly distributed and chaotic out in the real world. So that is also one of the major factors that comes into play here. And when you're thinking about, I guess, that distributed nature of the environments that you're working with, immediately my mind goes to sort of like complication and control.
26:24Like, how do you govern and manage both the operational component of that and the governance component of that? What would have what's been learned, I guess, as some of the best practices and thought process that goes into making sure that you as you have more and more of a distributed set of things out there in the world, you have some concept of kind of control or governance or however you put put that. Yeah, where possible, where these devices are connected to the Internet, we still leverage that connectivity in order to manage the devices. That means that we're still centrally managing a lot of this in the cloud.
27:07We're obviously aggregating a lot of data to do training in the cloud, be able to generate models from data that's been captured from many different devices. It helps us train more generalized models than if we were to try to train a model on a per device basis, right? Because each device has only got a small sliver of the total universe of data. So by bringing all the data together, we can train models that are really more generalized to work broadly throughout the whole world. And the same goes for how we're going to manage deployments as well. So if we can bring that connectivity of those devices centrally, it means that we can also roll out new versions of the model in a controlled way, Often using something like an over-the-air update framework as a way of helping manage not just the software on the device, but the models as well.
27:59So revision control, all these best practices that we have from software, and we at the Edge have been dealing with that for quite some time. We can now apply also to the models that we're deploying to the Edge. You know, as we're talking about models at the Edge, one of the things that has definitely been very pronounced. It's been this, as we've moved to smaller models in terms of number of parameters over time, and you're kind of comparing like where we're at today with that and the advances there, which maybe a lot of folks aren't, you know, the general public still focusing very much on the frontier, you know, large language models out there.
28:41That's what they read about most of the time and in the news. And maybe this is one of those topics that kind of gets missed is the advances in smaller models. Can you talk a little bit about the fact of like, what can you do now that we, as we are doing this in 2026, discussing this and, you know, you have some incredibly capable models that are small that may have 3 billion parameters, You know, instead of many times that number of some of yesterday's large language models. Can you talk a little bit about why those smaller models have gotten so effective? And what are the decisions that you have to make when you're using these small models, both for their strengths and their weaknesses, so that you can kind of put them in an architecture that makes it work for the mission that you're trying to address in that architecture?
29:37Sure. You're correct. I think like the what's happening with the state of the art is kind of overshadowing some of the, you know, also advancements that are happening at the at the edge and with small models. The good news is also a lot of that is applicable to what we're doing at the edge as well. And so one of those techniques that we use is knowledge distillation. So a way of leveraging big, powerful models and being able to distill out the knowledge into a small model. And this is one of these techniques that allows us to achieve this. We don't need like the whole universe of knowledge into a small model that's meant to do something very specialized.
30:16We only need the knowledge that's relevant to that specialized thing. And so these knowledge distillation techniques mean that we can use big, large models, extract basically their knowledge through a lot of, you know, if it's a language model, then we use a lot of like, you know, hit it with a bunch of queries where we can get a response, train a simpler model based on that. And there are other techniques we use as well. Fine-tuning as well. So, you know, taking a model like this, fine-tuning it specifically on the data that it's going to work on in its specialized task. There's a lot of those techniques.
30:52And then, of course, there's non-generative models too. So these classically have always been pretty purpose-built on data sets that are targeting specialized use cases and enabling us to generate very small models. At Jimples, we've been able to, we work with a lot of wearable devices. So this is like the smallest of edge devices you can think of. Wearable rings, for example, pointing to microcontrollers. We've always been able to do that using, you know, what's been coined a tiny EML, right? Small machine learning models. So, yeah, there's a whole spectrum of possibilities there, many techniques that are applied.
Read the full transcript
31:35And yeah, I think it's great that we've been able to leverage the advances that continue to come in the frontier of AI. And I know, Brandon, that you mentioned that Edge Impulse has their kind of own take on some of the framework and the tooling used to enable some edge AI. I'm also curious, you know, Edge Impulse now being a part of Qualcomm or a Qualcomm company, there's kind of a vertically integrated, I guess, component to that. I'm not going to, you know, put you on the spot to talk through why, you know, Qualcomm would want to acquire Edge Impulse or something like that. But could you talk a little bit about maybe first kind of Edge Impulse's unique take on or opinionated take on how the tool set should look for enabling these kinds of workflows?
32:32And then maybe also if there's anything relevant to that kind of vertically integrated take on Edge AI that kind of makes a vertically integrated approach maybe appealing in certain ways. Absolutely. Yeah. So Edge Impulse is the leading edge AI platform. It's also the reason why. And really the goal for it, I mean, for it to be the leading platform really had to deal with the diversity and the fragmentation of the silicon in this space and continues to do so.
33:07So our opinionated take on how to serve that space has really been to try to, I think of it as kind of a duality when it comes to hardware, right? We're trying to, in some ways, abstract away all the hardware differences. The machine learning is essentially math and statistics. And so we, on some level, want to treat it that way. Then when it comes time to deploy, we do target aware optimization and generation and conversion of the models for those targets. So by kind of thinking about it in those two different terms, we have this flexibility to go and serve the broad market. Now, how we bring in and empower the processors and platforms from Qualcomm is we make sure that Qualcomm is, of course, supported best in class with this optimization and tuning and leveraging all the competencies that Qualcomm has.
34:08So that means, you know, extreme power efficiency and leveraging their accelerators. You know, we're talking about like the Hexagon MPU, for example, that's in Dragonwing processors used in many different use cases, industrial and also things like automotive. you know everything from like very low power infrastructure out in the world up to very powerful like I mentioned AI appliances these are like basically AI servers that go on-prem so it's a broad portfolio there that they you know awesome range of different silicon with and it's not just the MPUs as well. It's DSPs, it's ISPs. It's a lot of specialized processing.
34:59So it's a lot that we can tap into and leverage in order to bring, you know, the most efficient models out to the edge device. Yeah, so I think in a lot of ways, Edge Impulse hasn't had to change its opinion about the world. It's like we understand how we need to be able to go out and bring ML into the edge space. And, you know, it means also being able to accentuate all the different silicon that we can serve with our platform. I'm curious kind of to dive into the hardware again a little bit more, because again, I think this is a little bit of a new topic for folks that are used to, you know, big servers that you're plugging in in a data center and, you know, or a cloud environment.
35:44You know, all of these things are battery driven out at the edge. you know, kind of by definition, especially if they're a moving platform, you know, we're talking, you know, autonomous vehicles and stuff like that. When you're looking at trying to do that computation out there and, you know, you have neural processing units that have become quite, you know, quite advanced, to your point about Qualcomm, and you see that the number of operations per second per watt are really have gotten pretty amazing in terms of what they can do. How has that, uh, how has that changed the math, uh, of, or, or kind of changed the way you think about, uh, operations at the edge when you're talking about different platforms that don't have traditional power available, um, that efficiency that you're having there, how is that, how has that kind of yielded new capability at the edge for battery-powered devices?
36:46Yeah, it certainly means we can do more. And so just the amount of ML that we can bring or the size of the models obviously allows us to scale out. And, you know, that's very, it's just, it means that what we've previously been able to do, we can just keep building on in ways that allow us to bring more intelligence, more processing. So I don't think it's anything more than that necessarily. It's just that extreme efficiency means like when developers are building a product, they're trying to differentiate, they're trying to do better than the last iteration of the product. They're able to get processors that are both very cost efficient, very power efficient.
37:33They have the compute so that they can go and deploy models or their software in a way that's going to give them best in class performance. And that translates to being able to market their end product, you know, competitively or best in class relative to all their competitors. So that's the way I look at it. And, you know, there's always an opportunity to do more. I mean, I think that's the way when once you have sort of AI in the tool chest, there's, you kind of just broadens the perspective of like, what could the world be like if we put intelligence right where the data is at? Suddenly, the possibilities start to explode.
38:17And that question is like, well, what's actually feasible with the devices that we have there or that we're going to put there in the next generation and so on? And it's usually power and cost constraint. So, you know, that's how the calculation usually works out. I hear from folks all the time, like have really interesting, sometimes crazy ideas about what they want to achieve. And that's exciting. But we also, again, are forced to rationalize a bit. What brings real value to the end users of these products? What can like bring, you know, let's face it, revenue for the companies building these products?
38:55and that helps you know it's a forcing function for making sure that what we're building is actually valuable ultimately and some of that's really i mean i i get excited thinking about some of those possibilities also i i love the idea of kind of creativity that comes when you're working in constraints and you know trying to to work through some of those things as uh develop if If you think of developers or kind of AI practitioners out there, do you have any recommendations for the person who is maybe inspired by this conversation and says, hey, I want to try an AI at the edge thing? What might be a way that they could, you know, not necessarily there's all sorts of use cases, as you mentioned, but, you know, create some type of lab environment or maybe it's Arduino or whatever that is.
39:49do something, you know, create a kind of minimal setup that would help them explore and experiment with some of these edge AI things? Where should they start both on the, maybe on the hardware side and the software or kind of use case tooling side? What's a good starting point and how can they get going? Yeah. What I think is awesome about edge AI is it's, you can honestly think about any real world problem out there and start to think about how can I go and like solve it with AI. And we can put these processors anywhere now. So that's the first place to start. Is there something interesting or a pain that, you know, somebody is dealing with?
40:31There's so many cool projects that people have built around their homes because they're like, there's not something that does this for me. And so I can take a simple board, maybe it's from Arduino. You know, I can create a simple model in Edge Impulse. By the way, it's free to sign up. So in terms of tooling, you know, that's a great place to start. And then go and solve your problem, right? Like, does your basement leak and you want to know when it leaks? Create a leak detector, super easy. Do you want to detect when your cat walks by so you can dispense some food from your cat feeder? You can do that too.
41:07It's amazing that so many of these things are readily achievable with commodity like maker hardware that's out there. That's a great place to start. And we see this even in enterprises like, you know, I've been a developer in enterprise and I've known many of them. They also use a lot of this like stuff to get started. And it's an easy way to generate a proof of concept. And once you've got a working example, then of course you go get some real enterprise hardware. Qualcomm's got a lot of it, so definitely check it out. and you know you've got tools like edge impulse which also scale into production and can support you when you go up to you know serving models at scale with like true ml ops continuous deployments and all of that as well so you know long story short i think there are some great examples i'll stick with arduino here great options for um getting people started on these projects for very inexpensive amount of money, check out edgeimpulse.com.
42:09You can sign up for free, start using it. And there's also great content to help you get started using these tools. That's a great answer. And I point out a very fun answer to go implement when you're actually bringing these capabilities into your own life, into your own world, not just through your work or whatever, or through your phone on an app, but actually having things happen that you said I want to go do. So a great answer. I guess as we are winding up and we're kind of looking at the future, you know, you have Edge Impulse and Qualcomm and the kinds of work you're doing. And you also have the larger Edge space.
42:53And where is one of the things we like to ask, and you may have heard this on other episodes, it's just like, where do you think things are going. And, you know, the nature of the question is a little bit less structured in the sense of like, kind of when you're not trying to solve a problem and you're just kind of letting your mind wander, what kinds of things do you think of that might come to pass? They might not, but they might come to pass as you're looking at this industry you're in that excites you and you kind of go, that's where I really want to go. Like, and I know that there are other people that would probably want that too.
43:28What are those kinds of thoughts that you have about where edge compute may be heading over the next few years? Yeah, I think, so the way I think about it is like, what if power and cost and compute, they basically kind of go to almost zero or like the cost of these things, right? It means that we could put intelligence literally anywhere right at the edge. And where we're at today with a lot of intelligence is it's kind of in the cloud. So it's gated by connectivity, you know, the cost of using the cloud and things like that. I think when you're at the edge, and I think it's also important to think about like biological intelligence, right?
44:09We've had these incredible organisms, right, that have sensors, they have intelligence directly where the sensors are. And like we see the world around us that we've managed to create with that. It's incredible. That's like just the most amazing inspiration. What if we can get closer to that with AI? And so just the realm of possibility is enormous. What I see is that we're going to continue bringing models to the edge, more of them. We talked about cascades and things like that, I think is how we, one of the techniques we go. And then there's also, you know, world models, VLAs and stuff on that spectrum as well, where we're talking about, you know, very large models.
44:58And if those become more economical to bring the edge, it means that they bring real more like true intelligence about the broader world and the ability to act in the world. So, you know, I also think that these action models are going to become more prevalent. We're seeing that robotics, self-driving, but also many other places they could be applied. So, you know, that's what I see. I think we're going to see a lot more robotics, which is going to be exciting and interesting. And hopefully, like, even, you know, maybe they're like robotic-like systems, but things that just live around us and can help take action in the world using intelligence.
45:40It's awesome. Yeah, well, I'm certainly excited to see some of those things. And also, I really encourage folks out there, you have no excuse to not go experiment and try some things with all the great hardware and tooling available to, you know, experiment with some of those things in your own, that fit your own passions and your home environment or wherever that is. So I really appreciate you coming on the show to inspire those things, Brandon, and the work that you're doing with Edge Impulse. Appreciate that and hope to have you back on. It's been a real pleasure. Thank you for having me on.
46:22All right, that's our show for this week. If you haven't checked out our website, head to practicalai.fm and be sure to connect with us on LinkedIn, X, or Blue Sky. You'll see us posting insights related to the latest AI developments, and we would love for you to join the conversation. Thanks to our partner, Prediction Guard, for providing operational support for the show. Check them out at predictionguard.com. Also, thanks to Breakmaster Cylinder for the beats, and to you for listening. That's all for now, but you'll hear from us again next week.
From the publisher
What does “AI at the edge” really mean in 2026, and why does it matter now more than ever before? In this episode, we’re joined by Brandon Shibley, Edge AI Solutions Engineering Lead at Qualcomm’s Edge Impulse, to discuss the current state and future of Edge AI in 2026. We discuss Gen AI, Small Models, and Cascades of Models, along with real-world constraints like latency, power, and privacy. We also dive into the role of MLOps, evolving hardware, and how developers can start building practical edge AI systems today.
Featuring:
- Brandon Shibley – LinkedIn
- Chris Benson – Website, LinkedIn, Bluesky, GitHub, X
- Daniel Whitenack – Website, GitHub, X
Links:
- Read our Ultimate Guide to Edge AI
- Download your copy of O'Reilly's AI at the Edge
- Check out the Edge Impulse blog
- Sign-up for an expert led trial of Edge Impulse
Upcoming Events:
- Register for upcoming webinars here!




